Database retrieval method and device, storage medium and electronic device
By vectorizing the target problem and determining the nearest cluster center vector and knowledge block vector in the vector database, the problem of low data retrieval efficiency in the prior art is solved, and the fast and accurate data retrieval effect is achieved.
Patent Information
- Application Number
- CN202510142101.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-03
AI Technical Summary
The data retrieval efficiency in the prior art is inefficient and has failed to effectively solve the problem of determining the similarity between the problems in knowledge search and database knowledge.
By vectorizing the received target problem, the problem vector is obtained, and the cluster center vector closest to the problem vector is determined in the vector database. Then, the closest knowledge block vector is searched for the candidate knowledge block vector corresponding to the cluster center vector as the search result.
The data retrieval efficiency is improved. By calculating the distance between the problem vector and the cluster center vector and the residual vector of the knowledge block vector, the calculation amount is significantly reduced, and fast and accurate data retrieval is achieved.
Smart Images

Figure CN120086251A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computers, and in particular, to a method, apparatus, storage medium, and electronic device for retrieving a database. Background Art
[0002] In the related art, when performing knowledge search, it is usually necessary to determine the similarity between the problem to be determined and all the knowledge included in the database to determine the answer to the searched problem.
[0003] It can be seen that there is a technical problem of low data retrieval efficiency in the related art.
[0004] For the above problems existing in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present invention provide a method, apparatus, storage medium, and electronic device for retrieving a database, so as to at least solve the problem of low data retrieval efficiency existing in the related art.
[0006] According to an embodiment of the present invention, there is provided a method for retrieving a database, including: vectorizing a received target problem to obtain a problem vector; determining a first target cluster center vector included in a vector database, where the target distance between the first target cluster center vector and the problem vector satisfies a first predetermined condition, where the vectors included in the vector database are knowledge block vectors obtained by vectorizing knowledge blocks, the first target cluster center vector is a vector included in multiple cluster center vectors obtained by clustering the knowledge block vectors, and the first predetermined condition includes that the target distance is less than the distance between the problem vector and other cluster center vectors included in the cluster center vectors except the first target cluster center vector; searching for a first knowledge block vector among candidate knowledge block vectors belonging to the same category as the first target cluster center vector, where the distance between the first knowledge block vector and the problem vector is less than the distance between the problem vector and other vectors, and the other vectors are vectors included in the candidate knowledge block vectors except the first knowledge block vector; and outputting a first knowledge block corresponding to the first knowledge block vector.
[0007] In an exemplary embodiment, searching for a first knowledge block vector from candidate knowledge block vectors belonging to the first target cluster center vector includes: determining first residual vectors corresponding to each of the candidate knowledge block vectors included in the vector database, where the first residual vector is a vector obtained by taking the residual between the candidate knowledge block vector and the cluster center vector; determining a second residual vector between the problem vector and the first target cluster center vector; determining a first distance between the first residual vector and the second residual vector; determining a second distance included in the first distance that satisfies a second predetermined condition, where the second predetermined condition includes that the second distance is less than other distances included in the first distance except the second distance; determining a third residual vector corresponding to the second distance; and determining the knowledge block vector corresponding to the third residual vector as the first knowledge block vector.
[0008] In an exemplary embodiment, determining a first target cluster center vector whose target distance from the problem vector in the vector database satisfies a first predetermined condition includes: obtaining a target tree included in the vector database, where the target tree is obtained by classifying vectors included in the vector database; and searching for the first target cluster center vector in the target tree whose target distance from the problem vector satisfies the first predetermined condition.
[0009] In an exemplary embodiment, before vectorizing a received target problem to obtain a problem vector, the method further includes: partitioning a obtained knowledge base document according to a document structure to obtain a plurality of the knowledge blocks; vectorizing each of the knowledge blocks to obtain the knowledge block vectors; clustering the knowledge block vectors to obtain a plurality of the cluster center vectors; and constructing the vector database based on the cluster center vectors.
[0010] In an exemplary embodiment, constructing the vector database based on the cluster center vectors includes: for each second target cluster center vector included in the plurality of the cluster center vectors, performing the following operations to obtain target data with the second target cluster center vector as the cluster center, where the second target cluster center vector is any one of the plurality of the cluster center vectors: determining second knowledge block vectors belonging to the second target cluster center vector and determining the knowledge block identifiers of each of the second knowledge block vectors, determining a fourth residual vector between each of the second knowledge block vectors and the second target cluster center vector, associating the fourth residual vector and the knowledge block identifier to obtain an association relationship, and determining the fourth residual vector and the association relationship as the target data; and determining the database including the target data, the knowledge block vectors, and the plurality of the cluster center vectors as the vector database.
[0011] In an exemplary embodiment, clustering the knowledge block vectors to obtain a plurality of the cluster center vectors includes: constructing a target tree from the knowledge block vectors; and performing clustering on the target tree to determine a plurality of the cluster center vectors.
[0012] In an exemplary embodiment, after outputting the first knowledge block corresponding to the first knowledge block vector, the method further includes: obtaining a feedback result of the target problem; when the feedback result indicates that the first knowledge block is valid, assembling the target problem and the first knowledge block to obtain an assembled knowledge block; vectorizing the assembled knowledge block to obtain an assembled knowledge block vector; determining, based on a target tree included in the vector database, a third target cluster center vector that is closest to the assembled knowledge block vector; determining a fourth residual vector between the assembled knowledge block vector and the third target cluster center vector; assigning a target identifier to the assembled knowledge block vector and associating the target identifier with the fourth residual vector; and saving the target identifier and the fourth residual vector to the vector database.
[0013] According to another embodiment of the present invention, there is provided a retrieval device for a database, including: a vectorizing module configured to vectorize a received target problem to obtain a problem vector; a determining module configured to determine a first target cluster center vector included in a vector database, where a target distance between the first target cluster center vector and the problem vector satisfies a first predetermined condition, where vectors included in the vector database are knowledge block vectors obtained by vectorizing knowledge blocks, the first target cluster center vector is a vector included in a plurality of cluster center vectors obtained by clustering the knowledge block vectors, and the first predetermined condition includes that the target distance is less than a distance between the problem vector and other cluster center vectors included in the cluster center vectors except the first target cluster center vector; a searching module configured to search for a first knowledge block vector from candidate knowledge block vectors belonging to the same category as the first target cluster center vector, where a distance between the first knowledge block vector and the problem vector is less than a distance between the problem vector and other vectors, and the other vectors are vectors included in the candidate knowledge block vectors except the first knowledge block vector; and an output module configured to output a first knowledge block corresponding to the first knowledge block vector.
[0014] According to still another embodiment of the present invention, there is further provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0015] According to another embodiment of the present invention, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0016] According to another embodiment of the present invention, a computer program product is further provided, including a computer program, and when the computer program is executed by a processor, the steps of the methods described in the various embodiments of the present application are implemented.
[0017] Through the present invention, the received target problem can be vectorized to obtain a problem vector. Calculate the distances between the problem vector and multiple cluster center vectors included in the vector database, and determine a first target cluster center vector that meets the first predetermined condition in the vector database. Then determine the distances between all knowledge block vectors included in the candidate knowledge blocks corresponding to the first target cluster center vector and the problem vector, and determine a knowledge block vector (the first knowledge block vector) with the closest distance. Use the first knowledge block corresponding to the determined first knowledge block vector as the retrieval result of the target problem. By only calculating the distances between the problem vector and multiple cluster centers, and then determining the corresponding knowledge block vectors in the candidate knowledge blocks corresponding to the cluster center vectors, the computational amount is much smaller than the computational amount of calculating the similarity between the problem and each knowledge base in the database. Therefore, the problem of low data retrieval efficiency in the related art can be solved, and the effect of improving data retrieval efficiency can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a hardware structure block diagram of a mobile terminal for a database retrieval method according to an embodiment of the present invention;
[0019] Figure 2 is a flowchart of a database retrieval according to an embodiment of the present invention;
[0020] Figure 3 is a flowchart of a database retrieval according to a specific embodiment of the present invention;
[0021] Figure 4 is a structure block diagram of a database retrieval device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The embodiments of the present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0024] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a database retrieval method according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0025] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the database retrieval method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0026] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0027] In this embodiment, a database retrieval method is provided. Figure 2 is a flowchart of the database retrieval according to an embodiment of the present invention. As Figure 2As shown in the figure, the process includes the following steps:
[0028] Step S202: Vectorize the received target problem to obtain a problem vector.
[0029] Step S204: Determine a first target cluster center vector in the vector database whose target distance from the problem vector satisfies a first predetermined condition. Among them, the vectors included in the vector database are knowledge block vectors obtained by vectorizing knowledge blocks, and the first target cluster center vector is a vector included in multiple cluster center vectors obtained by clustering the knowledge block vectors. The first predetermined condition includes that the target distance is less than the distance between the problem vector and other cluster center vectors except the first target cluster center vector included in the cluster center vectors.
[0030] Step S206: Search for a first knowledge block vector among candidate knowledge block vectors belonging to the same category as the first target cluster center vector. Among them, the distance between the first knowledge block vector and the problem vector is less than the distance between the problem vector and other vectors, and the other vectors are vectors other than the first knowledge block vector included in the candidate knowledge block vectors.
[0031] Step S208: Output the first knowledge block corresponding to the first knowledge block vector.
[0032] In the above embodiment, the received target problem of the user can be vectorized using the bge-large-zh embedding model to obtain a problem vector. Among them, bge-large-zh embedding can be understood as converting text into a high-dimensional vector representation, and learning the semantic representation of text through two stages of pre-training and fine-tuning. Among them, the pre-training stage uses a large-scale corpus for training, and the fine-tuning stage is optimized according to specific task requirements.
[0033] In the above embodiment, multiple target distances between the problem vector and multiple cluster center vectors obtained by clustering the knowledge block vectors in the vector database can be calculated, and the cluster center vector corresponding to the smallest target distance value among the multiple target distance values (i.e., the above first predetermined condition) is selected as the first target cluster center vector. Obtain candidate knowledge block vectors belonging to the same category in the first target cluster center vector, calculate the distances between all candidate knowledge block vectors and the problem vector, and determine the knowledge block vector corresponding to the smallest distance value as the first knowledge block vector of the problem vector. The first knowledge blocks corresponding to the retrieved first knowledge block vectors can be spliced together according to the paragraph structure and output as the retrieval result of the final problem vector. Among them, the target distance between the problem vector and the cluster center vector can be calculated by the Euclidean distance.
[0034] Through the present invention, the received target problem can be vectorized to obtain a problem vector. Calculate the distances between the problem vector and multiple cluster center vectors included in the vector database, and determine a first target cluster center vector that meets the first predetermined condition in the vector database. Then, determine the distances between all knowledge block vectors included in the candidate knowledge blocks corresponding to the first target cluster center vector and the problem vector, and determine a knowledge block vector (first knowledge block vector) with the closest distance. Use the first knowledge block corresponding to the determined first knowledge block vector as the retrieval result of the target problem. By only calculating the distances between the problem vector and multiple cluster centers, and then determining the corresponding knowledge block vectors in the candidate knowledge blocks corresponding to the cluster center vectors, the computational amount is much smaller than that of calculating the similarity between the problem and each knowledge base in the database. Therefore, the problem of low data retrieval efficiency in the related art can be solved, and the effect of improving data retrieval efficiency can be achieved.
[0035] Optionally, the execution subject of the above steps may be a background processor, or a terminal, a server, etc., but not limited thereto.
[0036] In an exemplary embodiment, searching for the first knowledge block vector among the candidate knowledge block vectors belonging to the first target cluster center vector includes: determining a first residual vector corresponding to each of the candidate knowledge block vectors included in the vector database, where the first residual vector is a vector obtained by taking the residual between the candidate knowledge block vector and the cluster center vector; determining a second residual vector between the problem vector and the first target cluster center vector; determining a first distance between the first residual vector and the second residual vector; determining a second distance that meets the second predetermined condition among the first distances, where the second predetermined condition includes that the second distance is less than the other distances except the second distance included in the first distances; determining a third residual vector corresponding to the second distance; and determining the knowledge block vector corresponding to the third residual vector as the first knowledge block vector.
[0037] In the above embodiment, the first residual vector between each candidate knowledge block vector and its corresponding cluster center vector and the second residual vector between the problem vector and the first target cluster center vector can be calculated first, and the first distance between the first residual vector and the second residual vector can be determined. Since there are multiple first residual vectors, the second distance with the smallest distance among the first distances can be determined, and the third residual vector corresponding to the second distance can be determined as the retrieval result of the problem vector, that is, the knowledge block vector corresponding to the third residual vector can be determined as the first knowledge block vector. Calculating using the residual vector can effectively reduce the order of magnitude and greatly reduce the computational workload.
[0038] In an exemplary embodiment, determining a first target cluster center vector in the vector database whose target distance from the problem vector satisfies a first predetermined condition includes: obtaining a target tree included in the vector database, where the target tree is obtained by classifying vectors included in the vector database; searching in the target tree for the first target cluster center vector whose target distance from the problem vector satisfies the first predetermined condition.
[0039] In the above embodiment, the kd-tree (i.e., the above target tree) obtained by classifying all knowledge block vectors in the vector database can be obtained first, and the first target cluster center vector whose target distance between the cluster center vector in the vector database and the problem vector satisfies the first predetermined condition can be searched through the target tree. That is, the target distance between all cluster center vectors and the problem vector can be calculated, and the cluster center vector with the closest target distance can be selected as the first target cluster center vector.
[0040] In an exemplary embodiment, before vectorizing the received target problem to obtain a problem vector, the method further includes: partitioning the obtained knowledge base document according to the document structure to obtain a plurality of the knowledge blocks; vectorizing each of the knowledge blocks to obtain the knowledge block vectors; clustering the knowledge block vectors to obtain a plurality of the cluster center vectors; and constructing the vector database based on the cluster center vectors.
[0041] In the above embodiment, the documents in the knowledge base can be segmented and numbered according to the document structure to obtain a plurality of knowledge blocks. The segmented knowledge blocks can also be vectorized using the bge-large-zh embedding Chinese vector large model embedding model to obtain a plurality of knowledge block vectors. The number of clusters of the knowledge block vectors can be set according to the size of the knowledge base document, and then the clustering center vectors (i.e., the above cluster center vectors) of all knowledge block vectors can be calculated and the vector database can be constructed.
[0042] In an exemplary embodiment, constructing the vector database based on the cluster center vectors includes: for each second target cluster center vector included in the plurality of cluster center vectors, performing the following operations to obtain target data with the second target cluster center vector as the cluster center, where the second target cluster center vector is any one of the plurality of cluster center vectors: determining the second knowledge block vectors belonging to the second target cluster center vector, and determining the knowledge block identifiers of each of the second knowledge block vectors, determining the fourth residual vector between each of the second knowledge block vectors and the second target cluster center vector, associating the fourth residual vector and the knowledge block identifier to obtain an association relationship, and determining the fourth residual vector and the association relationship as the target data; and determining the database including the target data, the knowledge block vectors, and the plurality of cluster center vectors as the vector database.
[0043] In the above embodiments, the second knowledge block vector of any one of the second target cluster center vectors in the cluster center vector, and the number corresponding to the second knowledge block vector (i.e., the above-mentioned knowledge block identifier) can be determined, and the fourth residual vector of each second knowledge block vector and its corresponding second target cluster center vector is calculated. The fourth residual vector is associated with the number corresponding to the knowledge block to obtain an association relationship. The fourth residual vector and the association relationship are determined as target data, and the target data, the knowledge block vector, and multiple cluster center vectors can be stored in the vector database together.
[0044] In an exemplary embodiment, clustering the knowledge block vectors to obtain multiple cluster center vectors includes: constructing the knowledge block vectors into a target tree; performing clustering on the target tree to determine multiple cluster center vectors.
[0045] In the above embodiments, the k-means++ algorithm can be used to calculate the cluster center vectors of all knowledge block vectors, and a data point can be randomly selected as the first cluster center. For each knowledge block vector, an operation of constructing a target tree is performed, that is: a dimension can be first selected, and the knowledge block vectors are divided into left and right parts according to the median of this dimension. The same operation is recursively performed on the left and right knowledge block vectors until a certain stop condition is reached (for example: the number of points in the left or right knowledge block vector is less than a certain threshold). Each generated node can be stored, and its left subtree and right subtree are recorded until the construction of the target tree is completed to determine multiple cluster center vectors.
[0046] In the above embodiments, the nearest cluster center can be searched through the target tree: starting from the root node of the target tree, the coordinates of the problem vector and the current node vector can be compared, and based on the splitting dimension of the current node vector, it is judged whether to continue searching in the left subtree or the right subtree. If the coordinate of the problem vector is less than the value of the corresponding dimension of the current node vector, continue searching along the left subtree; otherwise, continue searching along the right subtree. The search process will form a recursive call until the leaf node of the tree is reached.
[0047] In the above embodiments, when traversing to a leaf node, the target distance between the current node vector and the problem vector can be calculated. If the target distance is less than the currently stored best distance, it is updated as the best candidate point (i.e., the above-mentioned first target cluster center vector). Then, it can backtrack and check the other side. During the backtracking process, it can be determined whether to check another subtree (right subtree or left subtree) based on the distance between the current node vector and the problem vector and the dist of its splitting dimension, that is, if the distance of the splitting dimension of the current node vector (the absolute value of the difference between the coordinate of the problem vector on this dimension and the coordinate of the current node vector) is less than the current best distance, the opposite subtree needs to be searched, and finally the found best candidate point is returned. For each knowledge block vector x, its target distance can be inversely proportional to the selected probability, that is, the stored probability of being selected from the knowledge block vectors is: where y can be understood as the unselected knowledge block vector. The above process is repeatedly executed until k cluster center positions are selected, where k can represent the number of vector dimensions.
[0048] In the above embodiments, when aggregating the nearest cluster center through target tree search, each knowledge block vector belongs to the nearest cluster center. The coordinates of the first target cluster center vector can be updated according to the belonging knowledge block vectors, which can be calculated by computing the average value of all knowledge block vectors, that is, C(x) = ∑ y∈C(x) y / N.
[0049] In an exemplary embodiment, after outputting the first knowledge block corresponding to the first knowledge block vector, the method further includes: obtaining a feedback result of the target problem; when the feedback result indicates that the first knowledge block is valid, assembling the target problem and the first knowledge block to obtain an assembled knowledge block; vectorizing the assembled knowledge block to obtain an assembled knowledge block vector; determining a third target cluster center vector closest to the assembled knowledge block vector based on the target tree included in the vector database; determining a fourth residual vector between the assembled knowledge block vector and the third target cluster center vector; allocating a target identifier to the assembled knowledge block vector and associating the target identifier with the fourth residual vector; and saving the target identifier and the fourth residual vector to the vector database.
[0050] In the above embodiments, the vector database can also be updated through the result feedback of the user. The user will evaluate the answer to the target question (i.e., the above first knowledge block) to generate a feedback result. The target questions with valid feedback results can be assembled with the first knowledge block to form an assembled knowledge block. The spliced assembled knowledge block can be vectorized using the bge-large-zhembedding model to obtain the assembled knowledge block vector. Then, the third target cluster center vector closest to the assembled knowledge block is searched through the target tree, and the fourth residual vector between the assembled knowledge block and the third target cluster center vector is calculated. The fourth residual vector is associated with the corresponding knowledge block number (i.e., the above target identifier) and written into the vector database together, that is, only the original answer content and number are saved to the vector database.
[0051] The retrieval method of the database will be described below in conjunction with specific embodiments:
[0052] Figure 3 is a flowchart of the retrieval of the database according to a specific embodiment of the present invention, as Figure 3 shown. The process includes the following steps:
[0053] S302, obtain the knowledge base document;
[0054] S304, segment the knowledge base document to form knowledge blocks;
[0055] S306, vectorize the knowledge blocks through the embedding model to form knowledge block vectors;
[0056] S308, store the knowledge base number, cluster center vector, and knowledge block vector residual in the database;
[0057] S310, obtain the question input by the user;
[0058] S312, vectorize the question input by the user through the embedding model to form a question vector;
[0059] S314, search for the nearest cluster center of the question vector;
[0060] S316, obtain the knowledge block vector corresponding to the nearest cluster center;
[0061] S318, splice the retrieved knowledge blocks together according to the paragraph structure;
[0062] S320, retain the valid answers feedback by the user;
[0063] S322, splice the answer and the question into a knowledge block;
[0064] S324, vectorize the spliced knowledge block through the embedding model to form a knowledge base vector;
[0065] S326. Use the kd - tree to search for the centroid vector of the cluster of knowledge base documents that is closest to the problem vector.
[0066] S328. Write the residual vector and the corresponding number into the vector database.
[0067] In the above - mentioned embodiments, when constructing the knowledge base, the k - means++ algorithm can be used for index construction. Assume that the number of knowledge base blocks obtained by splitting the knowledge base is k, the dimension of the text vector is n, and the number of iterations is g. The time complexity of the initial center selection of K - means++ can be O(kn), and the time complexity of the iterative process can be O(gtn). Therefore, the time complexity of a complete K - means++ algorithm is O(gtn). Since the number of its iterations is much smaller than the number of knowledge base blocks and the dimension of the text vector, and the time complexity is much smaller than O(knn). In addition, when updating the knowledge base, only the similarity between the knowledge block and the centroid needs to be calculated, and the time complexity can be O(mn). Since the number of centroids m is much smaller than k and n, it is also much smaller than O(knn), which can greatly reduce the complexity of constructing the knowledge base and the computational workload of updating the knowledge base.
[0068] Through the description of the above - mentioned implementation manners, those skilled in the art can clearly understand that the method according to the above - mentioned embodiments can be implemented by means of software plus a necessary general - purpose hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0069] In this embodiment, a retrieval device for a database is also provided. This device is used to implement the above - mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation by hardware, or a combination of software and hardware is also possible and contemplated.
[0070] Figure 4 is a structural block diagram of the retrieval device for the database according to an embodiment of the present invention. As Figure 4 shown, this device includes:
[0071] A vectorization module 42, configured to vectorize the received target problem to obtain a problem vector.
[0072] A determination module 44, configured to determine a first target cluster center vector included in a vector database, where the target distance between the first target cluster center vector and a problem vector satisfies a first predetermined condition. The vectors included in the vector database are knowledge block vectors obtained by vectorizing knowledge blocks. The first target cluster center vector is a vector included in multiple cluster center vectors obtained by clustering the knowledge block vectors. The first predetermined condition includes that the target distance is less than the distances between the problem vector and other cluster center vectors except the first target cluster center vector included in the cluster center vectors.
[0073] A search module 46, configured to search for a first knowledge block vector from candidate knowledge block vectors belonging to the same category as the first target cluster center vector, where the distance between the first knowledge block vector and the problem vector is less than the distances between the problem vector and other vectors. The other vectors are vectors included in the candidate knowledge block vectors except the first knowledge block vector.
[0074] An output module 48, configured to output a first knowledge block corresponding to the first knowledge block vector.
[0075] In an exemplary embodiment, the search module 46 may search for a first knowledge block vector from candidate knowledge block vectors belonging to the first target cluster center vector in the following manner: determine a first residual vector corresponding to each candidate knowledge block vector included in the vector database, where the first residual vector is a vector obtained by taking the residual between the candidate knowledge block vector and the cluster center vector; determine a second residual vector between the problem vector and the first target cluster center vector; determine a first distance between the first residual vector and the second residual vector; determine a second distance included in the first distance that satisfies a second predetermined condition, where the second predetermined condition includes that the second distance is less than other distances except the second distance included in the first distance; determine a third residual vector corresponding to the second distance; and determine the knowledge block vector corresponding to the third residual vector as the first knowledge block vector.
[0076] In an exemplary embodiment, the determination module 44 may determine a first target cluster center vector included in the vector database, where the target distance between the first target cluster center vector and the problem vector satisfies a first predetermined condition, in the following manner: obtain a target tree included in the vector database, where the target tree is obtained by classifying the vectors included in the vector database; and search for the first target cluster center vector in the target tree, where the target distance between the first target cluster center vector and the problem vector satisfies the first predetermined condition.
[0077] In an exemplary embodiment, before the device is used to vectorize the received target problem to obtain a problem vector: the obtained knowledge base documents are chunked according to the document structure to obtain a plurality of the knowledge chunks; each of the knowledge chunks is vectorized to obtain the knowledge chunk vector; the knowledge chunk vectors are clustered to obtain a plurality of the cluster center vectors; and a vector database is constructed based on the cluster center vectors.
[0078] In an exemplary embodiment, the device can construct the vector database based on the cluster center vectors in the following manner: for each second target cluster center vector included in the plurality of cluster center vectors, the following operations are performed to obtain target data with the second target cluster center vector as the cluster center, where the second target cluster center vector is any one of the plurality of cluster center vectors: determining a second knowledge chunk vector belonging to the second target cluster center vector, and determining the knowledge chunk identifier of each second knowledge chunk vector, determining a fourth residual vector between each second knowledge chunk vector and the second target cluster center vector, associating the fourth residual vector and the knowledge chunk identifier to obtain an association relationship, and determining the fourth residual vector and the association relationship as the target data; and determining the database including the target data, the knowledge chunk vectors, and the plurality of cluster center vectors as the vector database.
[0079] In an exemplary embodiment, the device can cluster the knowledge chunk vectors to obtain a plurality of the cluster center vectors in the following manner: constructing the knowledge chunk vectors into a target tree; and performing clustering on the target tree to determine a plurality of the cluster center vectors.
[0080] In an exemplary embodiment, the device is further configured to, after outputting the first knowledge chunk corresponding to the first knowledge chunk vector: obtain a feedback result of the target problem; in the case where the feedback result indicates that the first knowledge chunk is valid, assemble the target problem and the first knowledge chunk to obtain an assembled knowledge chunk; vectorize the assembled knowledge chunk to obtain an assembled knowledge chunk vector; determine a third target cluster center vector closest to the assembled knowledge chunk vector based on the target tree included in the vector database; determine a fourth residual vector between the assembled knowledge chunk vector and the third target cluster center vector; assign a target identifier to the assembled knowledge chunk vector, and associate the target identifier with the fourth residual vector; and save the target identifier and the fourth residual vector into the vector database.
[0081] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0082] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0083] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks or optical disks and other various media that can store computer programs.
[0084] An embodiment of the present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0085] In an exemplary embodiment, the above-mentioned electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0086] An embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the methods in various embodiments of the present application are implemented.
[0087] The specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0088] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0089] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A database retrieval method, characterized in that: include: Vectorize the received target question to obtain a question vector; Determine a first target cluster center vector included in a vector database, whose target distance from the problem vector satisfies a first predetermined condition, wherein the vector included in the vector database is a knowledge block vector obtained by vectorizing the knowledge block, the first target cluster center vector is a vector included in a plurality of cluster center vectors obtained by clustering the knowledge block vectors, and the first predetermined condition includes that the target distance is less than a distance between the problem vector and other cluster center vectors included in the cluster center vectors except the first target cluster center vector; Searching for a first knowledge block vector from candidate knowledge block vectors that belong to the same category as the first target cluster center vector, wherein the distance between the first knowledge block vector and the question vector is less than the distance between the question vector and other vectors, and the other vectors are vectors included in the candidate knowledge block vectors except the first knowledge block vector; Output the first knowledge block corresponding to the first knowledge block vector.
2. The method according to claim 1, characterized in that Searching for a first knowledge block vector from candidate knowledge block vectors belonging to the first target cluster center vector includes: Determine a first residual vector corresponding to each of the candidate knowledge block vectors included in the vector database, wherein the first residual vector is a vector obtained by performing a residual between the candidate knowledge block vector and the cluster center vector; Determine a second residual vector between the problem vector and the first target cluster centroid vector; Determining a first distance between the first residual vector and the second residual vector; Determine a second distance included in the first distance that satisfies a second predetermined condition, wherein the second predetermined condition includes that the second distance is smaller than other distances included in the first distance except the second distance; Determine a third residual vector corresponding to the second distance; The knowledge block vector corresponding to the third residual vector is determined as the first knowledge block vector.
3. The method according to claim 1, characterized in that Determining a first target cluster centroid vector included in the vector database and having a target distance from the problem vector that satisfies a first predetermined condition comprises: Acquire a target tree included in the vector database, wherein the target tree is obtained by classifying the vectors included in the vector database; The target tree is searched for the first target cluster centroid vector whose target distance to the problem vector satisfies the first predetermined condition.
4. The method according to claim 1, characterized in that: Before vectorizing the received target question to obtain the question vector, the method further includes: Dividing the acquired knowledge base document into blocks according to the document structure to obtain a plurality of the knowledge blocks; Vectorizing each of the knowledge blocks to obtain the knowledge block vector; Clustering the knowledge block vectors to obtain a plurality of cluster center vectors; The vector database is constructed based on the cluster centroid vectors.
5. The method according to claim 4, characterized in that Constructing the vector database based on the cluster centroid vector includes: Determine each second target cluster center vector included in the multiple cluster center vectors, perform the following operations to obtain target data with the second target cluster center vector as the cluster center, wherein the second target cluster center vector is any one of the multiple cluster center vectors: determine the second knowledge block vector belonging to the second target cluster center vector, and determine the knowledge block identifier of each second knowledge block vector, determine the fourth residual vector of each second knowledge block vector and the second target cluster center vector, associate the fourth residual vector and the knowledge block identifier to obtain an association relationship, and determine the fourth residual vector and the association relationship as the target data; A database including the target data, the knowledge block vector, and a plurality of cluster center vectors is determined as the vector database.
6. The method according to claim 4, characterized in that Clustering the knowledge block vectors to obtain a plurality of cluster center vectors includes: constructing the knowledge block vectors into a target tree; Clustering is performed on the target tree to determine a plurality of cluster center vectors.
7. The method according to claim 1, characterized in that After outputting the first knowledge block corresponding to the first knowledge block vector, the method further includes: Obtaining feedback results of the target problem; When the feedback result indicates that the first knowledge block is valid, assembling the target question with the first knowledge block to obtain an assembled knowledge block; Vectorizing the assembled knowledge block to obtain an assembled knowledge block vector; Determine a third target cluster center vector which is closest to the assembled knowledge block vector based on the target tree included in the vector database; Determine a fourth residual vector between the assembled knowledge block vector and the third target cluster center vector; assigning a target identifier to the assembled knowledge block vector, and associating the target identifier with the fourth residual vector; The target identifier and the fourth residual vector are saved in the vector database.
8. A database retrieval device, characterized in that: include: A vectorization module is used to vectorize the received target question to obtain a question vector; A determination module, used to determine a first target cluster center vector included in a vector database, whose target distance from the problem vector satisfies a first predetermined condition, wherein the vector included in the vector database is a knowledge block vector obtained by vectorizing the knowledge block, the first target cluster center vector is a vector included in a plurality of cluster center vectors obtained by clustering the knowledge block vectors, and the first predetermined condition includes that the target distance is less than a distance between the problem vector and other cluster center vectors included in the cluster center vectors except the first target cluster center vector; A search module is used to search for a first knowledge block vector from candidate knowledge block vectors belonging to the same category as the first target cluster center vector, wherein the distance between the first knowledge block vector and the question vector is smaller than the distance between the question vector and other vectors, and the other vectors are vectors included in the candidate knowledge block vectors except the first knowledge block vector; An output module is used to output the first knowledge block corresponding to the first knowledge block vector.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.