Unstructured data processing system based on large language model and vector database

By using multiple large language models to extract vectors and query vector comparisons of private domain data in the unstructured data in the unstructured data of large language models and vector databases, the most suitable large language models are selected for answer output, and the existing large language models have different answer quality on different types of questions are solved, and the answer accuracy and fast operation are achieved.

CN118939848BActive Publication Date: 2025-05-13BOYUN VISION (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411407256.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-05-13
Estimated Expiration
2044-10-10

AI Technical Summary

Technical Problem

The existing large language models have different answers when answering different types of questions, resulting in unsatisfactory answers on certain types of questions.

Method used

Design an unstructured data processing system based on large language model and vector database, and vector extraction and query vector comparison of private domain data through multiple large language models, and filter out the most suitable large language model for answer output.

Benefits of technology

It realizes the strengths and avoids weaknesses between different large language models, ensures the accuracy and rapid calculation of the final answer, and improves the overall performance of the system when processing unstructured data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118939848B_ABST
    Figure CN118939848B_ABST
Patent Text Reader

Abstract

The present invention provides an unstructured data processing system based on a large language model and a vector database. A plurality of large language models are pre-set in the system. Private domain data presented in the form of unstructured data is input into these large language models to form a plurality of vector matrices to vectorize and sort the private domain data in the form of vectors. The user's question statements are also input into these large language models to form query vectors respectively. The vector matrices and the query vectors are corresponded to obtain the vector distance. The judgment module selects the terminal large language model according to the vector distance, and inputs the prompt vector obtained in the vector distance calculation process, and finally generates the answer to the question. The present invention initially uses a plurality of large language models for post-screening, so as to make the best use of the strengths and avoid the weaknesses of different large language models. The prompt vector generated in this process can ensure the accuracy of the final answer. The operation of the entire system ensures both the accuracy of the result and the speed of the operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to an unstructured data processing system based on a large language model and a vector database. Background Art

[0002] In recent years, large language models have been widely used. For example, Wenxin Yiyan launched by Baidu and chatGPT launched by OpenAI have been widely used in people's daily lives and are responsible for answering user inquiries.

[0003] However, these large language models have their own strengths and weaknesses in practical applications. A single large language model system is often good at answering one type of question, but its answers are not satisfactory when answering other types of questions.

[0004] Technically speaking, this is because such large language model systems have different logics for unstructured data. Their different underlying logics cause the database to operate in different ways when performing retrieval, calculation, and induction, resulting in answers that may be very different from each other.

[0005] Therefore, how to give full play to the strengths of large language models while avoiding the shortcomings of specific types of large language models as much as possible has become a problem that needs to be solved urgently. Summary of the invention

[0006] The present invention provides an unstructured data processing system based on a large language model and a vector database, which effectively solves the technical problems existing in the prior art.

[0007] Specifically, the present invention provides an unstructured data processing system based on a large language model and a vector database, the system comprising N large language models S1, S2, ..., S N , vector database, judgment module, execution module, N is an integer and N ≥ 2, wherein the private domain data are respectively input into the N large language models S1, S2, ..., S N Thus, N vector matrices A1, A2, ...A are generated one by one. N , these N vector matrices A1, A2, ...A N are input into the vector database; the question statements from the user are respectively input into the N large language models S1, S2, ..., S N Thus, N query vectors B1, B2, ...B are output one by one. N , these N query vectors are input into the vector database; the integer i takes values ​​from 1 to N in sequence, and each vector matrix A i All column vectors in B that correspond to the query vector i The column vector with the shortest vector distance is set as the prompt vector a*i , the hint vector is consistent with the query vector B i The vector distance between them is set as the vector matrix A i With the query vector B i The vector distance L between i As the integer i changes from 1 to N, N prompt vectors a*1, a*2, ..., a* N and N vector distances L1, L2, ..., L N , the N prompt vectors a*1, a*2, ..., a* N Take the average value to get the comprehensive prompt vector a*=(a*1+a*2+…+a* N ) / N, the judging module is at the N vector distances L1, L2, ..., L N Find the kth vector distance L k is the minimum vector distance among the N vector distances, wherein k is an integer and 1≤k≤N, and the N large language models S1, S2, ..., S are determined according to the k value. N The large language model S in k As a terminal large language model S k , the execution module inputs the comprehensive prompt vector a* into the terminal large language model S k Output the answer to the question.

[0008] Preferably, the vector matrix A i With the query vector B i The vector distance is calculated as: vector matrix A i is an m*j matrix with m rows and j columns, where m and j are integers, and m ≥ 2, and j ≥ 2, then the matrix forms j column vectors: a1, a2, ..., a j , each column vector contains m numbers, query vector B i It also contains m numbers, so we need to find the j column vectors: a1, a2, ..., a j Each column vector in is related to the query vector B i The vector distances between are l 1. l 2. … l j .

[0009] Preferably, the vector matrix A i With the query vector B i The vector distance L between i = Min{ l 1. l 2. … l j}, that is, the vector distance l 1. l2. … l j The minimum value in .

[0010] Preferably, the j column vectors a1, a2, ..., a j and the query vector B i The vector distance is the vector distance L i The column vector of is set to the prompt vector a* i , as i changes from 1 to N, N prompt vectors a*1, a*2, …, a* N .

[0011] Optionally, the private domain data includes user basic information data, user purchase behavior data, and user interaction data.

[0012] Preferably, in the process of obtaining the comprehensive hint vector, the N hint vectors a*1, a*2, ..., a* N Combined into matrix A*, each prompt vector is a column of matrix A*, and the covariance matrix D = A*ⅹA* is calculated T , where the matrix A* T is the transposed matrix of the matrix A*, performs eigenvalue decomposition on the covariance matrix D, and obtains N eigenvalues, which are combined into the comprehensive prompt vector a*.

[0013] In summary, the present invention provides an unstructured data processing system based on a large language model and a vector database. A plurality of large language models are pre-set in the system. On the one hand, private domain data presented in the form of unstructured data is input into the plurality of large language models to form a plurality of vector matrices so as to vectorize and sort the private domain data in the form of vectors. In parallel, the user's question statements are also input into the plurality of large language models to form query vectors respectively. The vector matrices are corresponded with the query vectors to obtain the vector distance. The judgment module selects the terminal large language model according to the vector distance, and the prompt vector obtained in the process of calculating the vector distance is input into the terminal large language model to finally generate the answer to the question. In the system provided by the present invention, a plurality of large language models are initially used for post-screening, so as to make the best use of the strengths and avoid the weaknesses of different large language models. The prompt vector generated in this process can ensure the accuracy of the final answer. The operation of the entire system ensures both the accuracy of the result and the speed of the operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be discussed below. Obviously, the technical solutions described in conjunction with the drawings are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments and their drawings can be obtained based on the embodiments shown in these drawings without paying creative work.

[0015] Figure 1 The operating flow chart of the unstructured data processing system based on a large language model and a vector database according to the present invention is shown. DETAILED DESCRIPTION

[0016] The technical solutions of various embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments described in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0017] The present invention provides an unstructured data processing system based on a large language model and a vector database, which combines multiple large language models with a vector database, extracts vectors from private domain data using multiple large language models, and selects one of the large language models for answer output in a preferential manner after comparing it with the query vector, completes the screening of the large language model in the entire feedback process, and thus selects the most suitable large language model for answer output. The specific composition and operation of the unstructured data processing system based on a large language model and a vector database according to the present invention will be described in detail below.

[0018] Figure 1 The operating flow chart of the unstructured data processing system based on a large language model and a vector database according to the present invention is shown.

[0019] like Figure 1 As shown in the figure, the private domain data is first input into N large language models S1, S2, ..., S N , where N is an integer and N ≥ 2. Each large language model outputs a vector matrix based on the private domain data through vector representation. Therefore, the N large language models output a total of N vector matrices, counted as vector matrices A1, A2, ...A N In other words, the private domain data is input into the large language model S1 to output the vector matrix A1, and is also input into the large language model S2 in parallel to output the vector matrix A2, and so on, until the large language model S is also input into the large language model S2 in parallel to output the vector matrix A2. N Output vector matrix A N .

[0020] It should be noted that private domain data refers to user behaviors and data collected by enterprises during private domain operations, including basic user information, purchasing behaviors, interaction status, etc. These data constitute the most original data in the data processing of the big language model. This type of private domain data is typical unstructured data, so it needs to be processed by the big language model into a quantifiable vector matrix for subsequent data analysis.

[0021] For example, private domain data contains five sentences (this is just for the convenience of description. In reality, private domain data often contains a huge amount of information, which may contain tens of thousands of sentences). After these five sentences are input into the large language model, each sentence is quantized by the large language model to form five sentence vectors. These five sentence vectors are combined to form the vector matrix output by the large language model. Therefore, the private domain data are input into N large language models in parallel, and each of them will output a total of N vector matrices.

[0022] The N vector matrices A1, A2, ...A N It is input into the vector database. It should be noted that the vector database here does not only play the storage function of the ordinary vector database, but also adds a judgment module outside the database. The role and function of the judgment module will be described in detail below.

[0023] like Figure 1 As shown, in parallel, the question sentences input by the user are also input into the N large language models respectively, and the N large language models output their own query vectors, thereby outputting a total of N query vectors, which are counted as B1, B2, ...B N These N query vectors are also input into the vector database.

[0024] The N vector matrices A1, A2, ...A N and the N query vectors B1, B2, ...B N The vector distance is calculated one by one. In other words, the vector distance between vector matrix A1 and query vector B1 is calculated, and the vector distance between vector matrix A2 and query vector B2 is calculated, and so on, until the vector matrix A is calculated. N With the query vector B N The vector distance of .

[0025] For example, take an integer i, where 1≤i≤N, and the integer i is taken from 1 to N in sequence, the vector matrix A i With the query vector B i The vector distance can be calculated as, the vector matrix A i is an m*j matrix, then the matrix forms j column vectors: a1, a2, ..., a j , each column vector contains m numbers, query vector B iIt also contains m numbers, so find the difference between each column vector and the query vector B i The vector distances between are l 1. l 2. … l j .

[0026] Any of the column vectors and the query vector B i The vector distance between them can be calculated by the Euclidean distance, for example. As mentioned above, each column vector contains m numbers, so the column vector is set to [at1, at2, ...at m ], and at the same time, query vector B i Also contains m numbers, then the query vector is set to [b1, b2, …b m ]. Thus, the Euclidean distance between the column vector and the query vector l It can be calculated as:

[0027]

[0028] Furthermore, the vector matrix A i With the query vector B i The vector distance L between i Values ​​from vector distance l 1. l 2. … l j The minimum value in L i =Min{ l 1. l 2. … l j The j column vectors a1, a2, ..., a j The distance between the vector and L i The corresponding column vector is set to the prompt vector a* i .

[0029] As the integer i takes values ​​from 1 to N, the N vector matrices A1, A2, ...A N and the N query vectors B1, B2, ...B N One-to-one correspondence is used to obtain the vector distances L1, L2, ..., L N , accordingly, these N vector distances L1, L2, ..., L N Each of them corresponds to the prompt vector a*1, a*2, ..., a* N .

[0030] The N prompt vectors a*1, a*2, …, a* N Take the average value to get the comprehensive prompt vector a*=(a*1+a*2+…+a* N ) / N.

[0031] The comprehensive prompt vector can be obtained by superimposing the different calculation processes of the above-mentioned multiple large language models, thereby ensuring that the prompt vector is more accurate and closer to the correct answer when the subsequent large language model outputs it.

[0032] Of course, the comprehensive hint vector can also be obtained in a more complex and accurate manner. For example, in the process of obtaining the comprehensive hint vector, the N hint vectors a*1, a*2, ..., a* N Combined into matrix A*, each prompt vector is a column of matrix A*, and the covariance matrix D = A*ⅹA* is calculated T , where the matrix A* T is the transposed matrix of the matrix A*, and the covariance matrix D is subjected to eigenvalue decomposition. There are N eigenvalues ​​in the covariance matrix D, and these N eigenvalues ​​are obtained. These N eigenvalues ​​are arranged into vectors and combined into the comprehensive prompt vector a*.

[0033] Next, the judgment module determines the N vector distances L1, L2, ..., L N Perform numerical comparison to find the kth vector distance L k is the minimum vector distance among the N vector distances, where k is an integer and 1≤k≤N. Then the vector distance L k The corresponding prompt vector is a* k , and in the N large language models S1, S2, ..., S N Filter out the large language model S k As a terminal large language model S k .

[0034] In this step, by comparing the vector distances, the large language model that is closest to the input question sentence in context and is the most practical among the N large language models so far is selected, that is, the terminal large language model.

[0035] So far, the system provided by the present invention has generated both the most accurate comprehensive prompt vector and the large language model that is closest to the question sentence. Next, the system provided by the present invention will start the corresponding execution module.

[0036] The execution module inputs the comprehensive prompt vector a* into the terminal large language model S k In the initial step, as described above, the question sentence is input into the large language model to generate the corresponding vector matrix. In the execution module, the reverse operation is performed, and the comprehensive prompt vector a* is input into the terminal large language model S k , terminal large language model S k The answer to the question will be finally output under the guidance of the comprehensive prompt vector a*.

[0037] So far, the unstructured data processing system based on a large language model and a vector database provided by the present invention has been basically introduced. In summary, the present invention provides an unstructured data processing system based on a large language model and a vector database. Multiple large language models are pre-set in the system. On the one hand, private domain data presented in the form of unstructured data is input into the multiple large language models to form multiple vector matrices so as to vectorize and sort the private domain data in the form of vectors. In parallel, the user's question statements are also input into multiple large language models to form query vectors respectively. The vector matrices and query vectors are correspondingly calculated to obtain the vector distance. The judgment module selects the terminal large language model according to the vector distance, and the prompt vector obtained in the process of vector distance calculation is input into the terminal large language model to finally generate the answer to the question. In the system provided by the present invention, multiple large language models are initially used for post-screening, so as to achieve the advantages and disadvantages of different large language models. The prompt vector generated in this process can ensure the accuracy of the final answer. The operation of the entire system ensures both the accuracy of the result and the speed of the operation.

[0038] The above description is only an exemplary embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An unstructured data processing system based on a large language model and a vector database, characterized in that: The system includes N large language models S1, S2, ..., S N , vector database, judgment module, execution module, N is an integer and N ≥ 2, where The private domain data are respectively input into the N large language models S1, S2, ..., S N Thus, N vector matrices A1, A2, ...A are generated one by one. N , these N vector matrices A1, A2, ...A N is entered into a vector database; The question sentences from the user are input into the N large language models S1, S2, ..., S N Thus, N query vectors B1, B2, ...B are output one by one. N , these N query vectors are input into the vector database; In the vector database, the integer i ranges from 1 to N, and each vector matrix A i All column vectors in B that correspond to the query vector i The column vector with the shortest vector distance is set as the prompt vector a* i , the hint vector a* i With the query vector B i The vector distance between them is set as vector matrix A i With the query vector B i The vector distance L between i , As the integer i changes from 1 to N, N prompt vectors a*1, a*2, ..., a* N and N vector distances L1, L2, ..., L N , Based on the N hint vectors a*1, a*2, ..., a* N Get the comprehensive hint vector a*, The judging module determines the distances L1, L2, ..., L N Find the kth vector distance L k is the minimum vector distance among the N vector distances, wherein k is an integer and 1≤k≤N, and the N large language models S1, S2, ..., S are determined according to the k value. N The large language model S in k As a terminal large language model S k , The execution module inputs the comprehensive prompt vector a* into the terminal large language model S k Output the answer to the question.

2. The system according to claim 1, characterized in that Vector Matrix A i With the query vector B i The vector distance is calculated as: vector matrix A i is an m*j matrix with m rows and j columns, where m and j are integers, m ≥ 2, j ≥ 2, then the matrix forms j column vectors: a1, a2, ..., a j , each column vector contains m numbers, query vector B i It also contains m numbers, so we need to find the j column vectors a1, a2, ..., a j Each column vector in is related to the query vector B i The vector distances between are l 1. l 2. … l j .

3. The system according to claim 2, characterized in that Vector Matrix A i With the query vector B i The vector distance L between i = Min{ l 1. l 2. … l j }, that is, the vector distance l 1. l 2. … l j The minimum value in .

4. The system according to claim 3, characterized in that The j column vectors a1, a2, ..., a j and the query vector B i The vector distance is the vector distance L i The column vector of is set to the prompt vector a* i As i changes from 1 to N, N prompt vectors a*1, a*2, …, a* N .

5. The system according to claim 1, characterized in that The private domain data includes the user's basic information data, the user's purchase behavior data, and the user's interaction data.

6. The system according to claim 2, characterized in that The j column vectors a1, a2, ..., a j Each column vector in is related to the query vector B i The vector distance between them is calculated using the Euclidean distance. Each column vector contains m numbers, so the column vector is set to [at1, at2, ...at m ], query vector B i Also contains m numbers, then the query vector is set to [b1, b2, …b m ], each column vector is related to the query vector B i The Euclidean distance between l Calculated as: 。 7. The system according to claim 1, characterized in that In the process of obtaining the comprehensive hint vector, the N hint vectors a*1, a*2, ..., a* N Combined into matrix A*, each prompt vector is a column of matrix A*, and the covariance matrix D = A*ⅹA* is calculated T , where the matrix A* T is the transposed matrix of the matrix A*, performs eigenvalue decomposition on the covariance matrix D, and obtains N eigenvalues, which are combined into the comprehensive prompt vector a*.

8. The system according to claim 1, characterized in that In the process of obtaining the comprehensive hint vector, the N hint vectors a*1, a*2, ..., a* N Take the average value to get the comprehensive prompt vector a*=(a*1+a*2+…+a* N ) / N.

Citation Information

Patent Citations

  • Dialogue processing method and system, electronic equipment and computer readable storage medium

    CN118296126A