Data completion method and device, electronic equipment and medium

By indexing embedded vectors that meet the similarity criteria in the vector database, the problem of inaccurate data completion caused by data hiding or abbreviation is solved, achieving efficient and accurate data completion and improving data integrity and usability.

CN120910049AActive Publication Date: 2025-11-07ZHEJIANG LAB
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511035485.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-07
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Existing technologies, when supplementing incomplete data, suffer from data hiding or the use of abbreviations, which prevent large models from accurately understanding and supplementing the data, thus reducing the integrity and usability of the data.

Method used

By acquiring the user's input request text to be completed, converting it into a query vector, and indexing the target embedding vectors that meet the preset similarity conditions in a pre-built vector database, the target metadata is extracted as completion data, and the embedding vector technology is used to perform efficient and accurate data completion.

Benefits of technology

It enables efficient and accurate data extraction and completion even when data is encrypted or abbreviations are used, improving data integrity and usability, with a matching accuracy of over 83%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910049A_ABST
    Figure CN120910049A_ABST
Patent Text Reader

Abstract

The invention discloses a data completion method and device, electronic equipment and a medium. The method comprises the steps that a to-be-completed request text input by a user is acquired; converting the request text to be complemented into a query vector; in a pre-constructed vector database, indexing a target embedded vector of which the similarity between the target embedded vector and the query vector meets a preset condition; extracting target metadata corresponding to the target embedded vector from a vector database; and returning the target metadata as the target completion data to the client. Therefore, on the basis of an embedded vector technology, the target embedded vector meeting the region can be quickly searched in the pre-constructed vector database, so that the target metadata with high reliability is extracted to serve as complementation data. Besides, even if the data are encrypted or represented by using abbreviations, the efficient and accurate extraction and complementation of the data can be met through the vector similarity index mode provided by the invention, so that the integrity and availability of the data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a data completion method and device, electronic equipment and medium. BACKGROUND

[0002] With the development of the information age, a large amount of data will be generated in various industries every day, and these data are usually stored in a structured format. Among them, some columns are used to store core data content (for example, event description), and other columns record meta information (for example, time, place, involved person, etc.) related to the data content.

[0003] In the actual data information collection, arrangement and storage process of the user, the meta information part often has a missing situation. At present, the collected data information can be completed by a large language model, so as to ensure the integrity of the data. However, in the completion process, some data information is hidden or encrypted for the purpose of secrecy, in addition, some data information is represented by using abbreviations and the like, so that the large model cannot accurately understand and complete the semantics, thereby reducing the integrity and usability of the data.

[0004] Therefore, how to efficiently and accurately complete the non-complete data information, so as to improve the integrity and usability of the data, is a problem to be solved by those skilled in the art. SUMMARY

[0005] Therefore, an aspect of the present application provides a data completion method, which comprises:

[0006] obtaining a user inputted to-be-completed request text;

[0007] converting the to-be-completed request text into a query vector;

[0008] indexing a target embedding vector between the query vector and the vector database pre-constructed, wherein the similarity between the target embedding vector and the query vector meets a preset condition;

[0009] extracting target meta data corresponding to the target embedding vector from the vector database;

[0010] returning the target meta data as target completion data to the client.

[0011] Optionally, the target embedding vector between the index and the query vector and meeting the preset condition of similarity comprises:

[0012] extracting the target similarity greater than a threshold value to obtain a candidate embedding vector corresponding to the target similarity;

[0013] when the target similarity is one, taking the candidate embedding vector as the target embedding vector;

[0014] when the target similarity is multiple, scoring candidate metadata of different data attributes corresponding to the candidate embedding vector; and taking the candidate embedding vector corresponding to the candidate metadata with the maximum score value as the target embedding vector.

[0015] Optionally, the extracting, from the vector database, target metadata corresponding to the target embedding vector comprises:

[0016] when the target similarity is one, taking all metadata corresponding to the target embedding vector as the target metadata;

[0017] when the target similarity is multiple, taking the metadata corresponding to the candidate metadata with the maximum score value as the target metadata.

[0018] Optionally, the scoring of the candidate metadata of different data attributes corresponding to the candidate embedding vector comprises:

[0019] determining, according to the target similarity, a Boltzmann score corresponding to each of the candidate embedding vectors;

[0020] determining whether there is identical candidate metadata in the same data attribute;

[0021] if not, taking the Boltzmann score as a score value of the corresponding candidate metadata;

[0022] if yes, performing the following steps:

[0023] determining that a score value of target candidate metadata is a corresponding Boltzmann score; the target candidate metadata is other metadata in the same data attribute except the identical candidate metadata;

[0024] performing an aggregate calculation on the Boltzmann score corresponding to the identical candidate metadata; and taking an aggregate calculation result as a score value of the identical candidate metadata.

[0025] Optionally, the scoring of the candidate metadata of different data attributes corresponding to the candidate embedding vector comprises:

[0026] calling a specified large model to score the relevance of the to-be-completed request text and each of the candidate metadata.

[0027] Optionally, in the absence of the target similarity, the data completion method comprises:

[0028] marking the target completion data as a NaN value; and returning the NaN value to the client;

[0029] determining whether artificial completion data input by a user is acquired within a preset time length;

[0030] if yes, completing the request text to be completed by using the artificial completion data;

[0031] if no, entering the step of indexing a target embedding vector that meets a preset condition in similarity between the query vector and the target embedding vector in the vector database constructed in advance every preset period, and performing subsequent steps.

[0032] Optionally, the step of constructing the vector database comprises:

[0033] acquiring a data set to be processed; the data set to be processed comprises core data content and metadata of different data attributes related to the core data content;

[0034] batch-converting the core data content in the data set to be processed into embedding vectors;

[0035] constructing an index for the embedding vectors based on a FAISS vector index to obtain index numbers;

[0036] establishing a mapping relationship between the embedding vectors, the index numbers and the metadata of the different data attributes to construct the vector database.

[0037] Another aspect of the present application provides a data completion device, which comprises:

[0038] a request acquisition module configured to acquire a request text to be completed input by a user;

[0039] a query vector acquisition module configured to convert the request text to be completed into a query vector;

[0040] a target embedding vector indexing module configured to index a target embedding vector that meets a preset condition in similarity between the query vector and the target embedding vector in a vector database constructed in advance;

[0041] a target metadata extraction module configured to extract target metadata corresponding to the target embedding vector from the vector database;

[0042] a data output module configured to return the target metadata as target completion data to a client.

[0043] Another aspect of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps of the data completion method when executing the computer program.

[0044] Another aspect of the present application provides a computer readable storage medium, which stores a computer program, and the program implements the steps of the data completion method when executed by a processor.

[0045] The data completion method, device, electronic device and medium provided by the present application have the following beneficial effects: based on the embedding vector technology, the target embedding vector meeting the region can be quickly searched in the pre-constructed vector database, so that the target metadata with high reliability is extracted as the completion data. In addition, even if the data is encrypted or represented by abbreviations, the vector similarity index method provided by the present application can also meet the efficient and accurate extraction and completion of the data, thereby improving the data integrity and usability. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 A flowchart of a data completion method provided by an embodiment of the present application;

[0047] Figure 2 A principle diagram of a data completion method provided by an embodiment of the present application;

[0048] Figure 3 A principle diagram of a data completion method provided by another embodiment of the present application;

[0049] Figure 4 A structure diagram of a data completion device provided by an embodiment of the present application;

[0050] Figure 5 A structure diagram of an electronic device provided by an embodiment of the present application.

[0051] The reference signs are as follows: 40 is a request acquisition module, 41 is a query vector acquisition module, 42 is a target embedding vector index module, 43 is a target metadata extraction module, 44 is a data output module, 50 is a memory, 51 is a processor, 55 is a display screen, 53 is an input / output interface, 54 is a communication interface, 55 is a power supply, 56 is a communication bus, 501 is a computer program, 505 is an operating system, and 503 is data. DETAILED DESCRIPTION

[0052] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this application and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or," as used herein, refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0053] It should be understood that although the terms first, second, third, etc. can be employed in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish one piece of information from another. For example, without departing from the scope of this application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining."

[0054] Figure 1 A flowchart of a data completion method provided by an embodiment of the application is shown in FIG. 1, which includes the following steps. Figure 1

[0055] S10: obtaining a user-inputted to-be-completed request text;

[0056] S11: converting the to-be-completed request text into a query vector;

[0057] Figure 2 A principle diagram of a data completion method provided by an embodiment of the application is shown in FIG. 2. In a specific embodiment, when a user has a data completion requirement, the user inputs a to-be-completed request text from a client. In order to achieve fast data completion, the server converts the request text into a query vector upon receiving the to-be-completed request text. Figure 2

[0058] It should be noted that in a specific embodiment, the to-be-completed request text refers to an object that needs to be completed by the user, and is a descriptive text of a data core content. The query vector refers to a vector obtained after encoding, which is used to initiate retrieval.

[0059] In an optional embodiment, in order to achieve fast conversion between text and vector, the to-be-completed request text can be converted into a query vector through a pre-trained target large language model, which includes but is not limited to ChatGPT, GPT-4, Claude, and Tongyi Qianwen, without limitation.

[0060] ​​In addition, the conversion can be implemented by using an encoder pre-trained to realize encoding, distributed static word vectors, and vector models, etc. The specific manner of converting the to-be-completed request text into a query vector is not limited in the present application.

[0061] S12: In the pre-constructed vector database, a target embedding vector between the index and the query vector satisfies a preset condition of similarity;

[0062] It can be understood that, in the case that the to-be-completed request text is encrypted or uses abbreviations, if the to-be-completed request text is subjected to semantic analysis by using a large model to generate target completion data, the large model may not accurately understand the semantics due to encryption and abbreviations, and the data completion accuracy is low.

[0063] Therefore, in order to solve the technical problem, in the data completion method provided in the present application, as shown in Figure 2 It should be noted that the vector database is used to store complete structured data, and specifically includes embedding vectors and metadata of various data attributes, wherein the embedding vector refers to a fixed-dimension vector obtained by encoding an indexed object.

[0064] It can be understood that, in the specific embodiment, the indexing in the vector database means matching between the query vector and the embedding vector, and the embedding vector between which the similarity satisfies a preset condition is taken as the target embedding vector after indexing. In an optional embodiment, since the request input by the user is a text completion request, the embedding vector refers to a vector obtained by encoding core data, and the metadata refers to relevant data information about the core data. In an optional embodiment, the query vector and the embedding vector can be understood as event descriptions, and the metadata refers to relevant information (for example, time, place, etc.) of the event.

[0065] It should be noted that the storage manner of the vector database is not limited in the present application. In an optional embodiment, the vector database can be stored by using a row-column structure, wherein a row is a complete data record, and specifically includes an embedding vector, metadata under different data attributes, and an index number of the embedding vector. The first row of the row-column structure stores a descriptive text of a data type, and each column in the list structure is specific metadata corresponding to a data attribute. Of course, in the row-column storage, the column can be used to store complete data, and the row can be used to store specific metadata under different data attributes, which is not limited in the present application.

[0066] Table 1 is a data storage schematic table of a vector database provided by an embodiment of the present application. In order to facilitate understanding, the following will be described in combination with Table 1 and taking an application scenario of intelligence analysis as an example.

[0067] Table 1 is a data storage schematic table of a vector database provided by an embodiment of the present application. In order to facilitate understanding, the following will be described in combination with Table 1 and taking an application scenario of intelligence analysis as an example.

[0068] Index number Embedding vector Subject side Object side Time …… Place 001 E1 A1 B1 T1 …… D1 002 E2 A2 B2 T2 …… D2 003 E3 A2 B3 T3 …… D3 004 E4 A3 B4 T3 …… D4 005 E5 A4 B5 T3 …… D5 …… …… …… …… …… …… …… 00n En Ai Bj Tx …… Dt

[0069] As shown in Table 1, in an optional embodiment, the vector database stores complete intelligence information, wherein the first row is used to store data attributes, which can also be understood as data types. In addition to the first row, each of the following rows stores a complete intelligence data. Each column stores specific metadata information under different data attributes.

[0070] It should be noted that the embedded vector refers to the vector of the intelligence description text coding in each complete intelligence data, and the subject party, the object party, the time and the place refer to the related metadata information of the complete intelligence data, that is, the metadata information of the intelligence description.

[0071] In an optional embodiment, as shown in Figure 2 , the user inputs a to-be-completed request text from the client, specifically, the input content is "intelligence X1 = customer R and customer P sign……", the client sends the intelligence X1 to the server, and the server converts the intelligence X1 into a query vector C. Further, indexing is performed in the vector database to retrieve a target embedded vector that satisfies a preset condition in similarity with the query vector C. It can be understood that the similarity between the query vector C and the target embedded vector can also be understood as the similarity between two corresponding texts. In an optional embodiment, the similarity can be calculated by cosine similarity, or the similarity can be calculated by a large language model, and the present application does not limit the similarity calculation method.

[0072] S13: Extracting target metadata corresponding to the target embedded vector from the vector database;

[0073] S14: Returning the target metadata as target completion data to the client.

[0074] Further, as shown in Figure 2As shown, after searching in the vector database to obtain the target embedding vector, the target metadata corresponding to the target embedding vector is taken as the target completion data and returned to the client. It should be noted that in specific embodiments, after the target embedding vector is determined, the entire row of complete data where the target embedding vector is located can be returned to the client, or only the target metadata value can be returned to the client, which is not limited by the present application. For example, in Table 1, if the indexed target embedding vector is embedding vector E1, the entire row of data where embedding vector E1 is located can be returned to the client, or only the information of subject party A1, object party B1, time T1 and place D1 corresponding to embedding vector E1 can be returned to the client.

[0075] It should be noted that in an optional embodiment, the target embedding vector can be one or multiple. If the target embedding vector is one, all corresponding target metadata can be directly returned to the client. When the target embedding vector is multiple, all target metadata corresponding to the target embedding vector with the highest similarity can be returned, or all target metadata corresponding to the target embedding vector can be returned, and the user can select the final completion data.

[0076] Of course, in another optional embodiment, when the target embedding vector is multiple, the data with the highest frequency of occurrence can be selected as the target metadata from each data attribute. If the frequency of occurrence is the same, the metadata with the highest similarity can be selected as the target metadata. In order to facilitate understanding, examples will be given below.

[0077] For example, in Table 1, if the index result includes embedding vectors E1 to E5, and embedding vector E5 has the highest similarity. When extracting the target metadata, for the subject party data attribute, since subject party A2 appears twice, subject party A2 is returned as the target metadata. For the object party, the object parties corresponding to embedding vectors E1 to E5 are all different, and at this time, the object party B5 corresponding to the embedding vector E5 with the highest similarity is returned to the client as the target metadata, and other data attributes are the same, which is not described here.

[0078] Therefore, the data completion method provided by the embodiment of the present application can quickly search for a target embedding vector satisfying a region in a pre-constructed vector database based on embedding vector technology, so as to extract high-reliability target metadata as completion data. In addition, even if the data is encrypted or represented by abbreviations, the vector similarity index method provided by the present application can also satisfy efficient and accurate extraction and completion of data, thereby improving data integrity and usability.

[0079] Figure 3This is a schematic diagram illustrating the principle of a data completion method provided in another embodiment of this application. In an optional embodiment, the target embedding vector whose similarity between the index and the query vector meets a preset condition includes:

[0080] Extract the target similarity with a similarity greater than a threshold to obtain the candidate embedding vector corresponding to the target similarity;

[0081] When the target similarity is one, the candidate embedding vector is used as the target embedding vector;

[0082] When there are multiple target similarities, the candidate metadata corresponding to different data attributes of the candidate embedding vector is scored; and the candidate embedding vector corresponding to the candidate metadata with the largest score is used as the target embedding vector.

[0083] Based on the above embodiments, after calculating the similarity between the query vector and each embedded vector in the vector database, further, target similarities with similarities greater than a threshold are extracted. The higher the similarity, the more similar the query vector and the embedded vector are in terms of association, semantics, etc. When the similarity is greater than the threshold, it can be understood that the metadata of the corresponding embedded vector has a high confidence level as target completion data. It is understood that after extracting the target similarity, candidate embedded vectors can be obtained, where there is a one-to-one correspondence between the candidate embedded vectors and the target similarities.

[0084] In one alternative embodiment, such as Figure 3 As shown, when the target similarity is one, that is, there is only one similarity greater than the threshold, there is also only one candidate embedding vector, which can be used as the target embedding vector.

[0085] When there are multiple target similarities, there are also multiple candidate embedding vectors, and the number of both is the same. In this case, to achieve high-precision data completion, in one optional embodiment, the candidate metadata under different data attributes corresponding to each candidate embedding vector is scored. Based on the scored data, the candidate embedding vector corresponding to the highest score is used as the target embedding vector. For ease of understanding, examples from Table 1 and intelligence analysis scenarios will be provided below.

[0086] For example, the threshold of the similarity comparison is 0.5, the user input request is "intelligence X1 = customer R and customer P sign……", the client sends intelligence X1 to the server, and the server converts intelligence X1 into a query vector C. In Table 1, after indexing and similarity calculation of the query vector C, the obtained candidate embedding vectors include embedding vector E1 to embedding vector E5, and the similarity between the query vector C and each candidate embedding vector is greater than 0.5, that is, the target similarity is [0.83, 0.72, 0.65, 0.92, 0.66], wherein the embedding vector E4 corresponds to the highest similarity.

[0087] As can be seen, in this example, the target similarity is multiple, and the candidate embedding vector also includes multiple, at this time, the candidate metadata corresponding to the different data attributes of the embedding vector E1 to the embedding vector E5 needs to be scored. For example, for the subject party, the subject party A1, the subject party A2, the subject party A3 and the subject party A4 need to be scored respectively.

[0088] It should be noted that there may be the same candidate metadata, for example, there are two subject parties A2, which can be scored respectively for the same subject party A2, or only one score, which is not limited by the present application.

[0089] After scoring, if the score value corresponding to the subject party A1 is the highest, the embedding vector E1 corresponding to the subject party A1 is taken as the target embedding vector. The metadata of other data attributes is scored and calculated in the same way, so as to determine the target embedding vector. It can be understood that after the similarity between the threshold and the judgment, and the number of target similarities are determined, the finally obtained target embedding vector may be one or multiple. Figure 3 As shown in the similarity and the threshold, and the number of target similarities are determined, the finally obtained target embedding vector may be one or multiple.

[0090] Therefore, the data completion method provided by the embodiment of the present application determines the target embedding vector when the candidate embedding vector is multiple, scores the candidate metadata of different data attributes respectively, determines the target embedding vector from a smaller granularity, and improves the data completion accuracy.

[0091] On the basis of the above embodiment, as an optional embodiment, the target metadata corresponding to the target embedding vector is extracted from the vector database, including:

[0092] When the target similarity is one, all metadata corresponding to the target embedding vector is taken as the target metadata;

[0093] When the target similarity is multiple, the metadata corresponding to the maximum score value is taken as the target metadata.

[0094] It can be understood that, as Figure 3As shown, when there is only one target similarity, there is only one corresponding target embedding vector. At this time, when the target metadata is extracted, all the target metadata corresponding to the target embedding vector is returned to the client as the target completion data. Of course, the entire row of data where the target embedding vector is located can also be returned to the client.

[0095] In another optional embodiment, as Figure 3 shown, when there are multiple target similarities, since the candidate metadata under different data attributes are scored respectively, and the target embedding vector is determined based on the score value after scoring, the corresponding target embedding vector can be one or multiple.

[0096] When the target embedding vector is one, the score values of the candidate metadata representing different data attributes are all maximum values. At this time, all the metadata corresponding to the target embedding vector can be taken as the target metadata. For example, in Table 1 of the above embodiment, if the candidate embedding vectors include embedding vectors E1 to E5, and the target similarity is [0.83, 0.72, 0.65, 0.92, 0.66], among which the score values corresponding to the subject side A2, the object side B3, the time T3, and the place D3 corresponding to the embedding vector E3 under each data attribute of the subject side to the place are all the highest. At this time, the target embedding vector is only one embedding vector E3. When the target metadata is extracted, all the metadata corresponding to the embedding vector E3 is taken as the target metadata.

[0097] When the target embedding vector is multiple, when the target metadata is extracted, the metadata with the highest score value from different data attributes is extracted as the target metadata. For example, in Table 1 of the above embodiment, if the candidate embedding vectors include embedding vectors E1 to E5, and the target similarity is [0.83, 0.72, 0.65, 0.92, 0.66].

[0098] After scoring the candidate metadata of embedding vectors E1 to E5 under different data attributes, the subject side A4 is the metadata with the highest score value of the subject side data attribute, the object side B5 is the metadata with the highest score value of the object side data attribute, the time T1 is the metadata with the highest score value of the time data attribute, and the place D1 is the metadata with the highest score value of the place data attribute. Correspondingly, the target embedding vector includes embedding vector E1 and embedding vector E5. Therefore, when the target metadata is extracted, the subject side A4, the object side B5, the time T1, and the place D1 are combined to form the target metadata, which is returned to the client as the target completion data.

[0099] That is, in the embodiment of the present application, when the target embedding vector is multiple, the target metadata is extracted, and not all metadata corresponding to the target embedding vector is taken as the target metadata, but a part of metadata in different target embedding vectors is extracted respectively and combined to obtain a complete and non-repeated target metadata.

[0100] Therefore, the data completion method provided by the embodiment of the present application takes the metadata corresponding to the maximum score value as the target metadata, can extract the target completion data with the highest confidence to complete the to-be-completed request text, and thus ensures the data completion accuracy.

[0101] In an optional embodiment, scoring the candidate metadata corresponding to different data attributes of the candidate embedding vector comprises:

[0102] According to the target similarity, determining the Boltzmann score corresponding to each candidate embedding vector;

[0103] Determining whether there is same candidate metadata in the same data attribute;

[0104] If not, taking the Boltzmann score as the score value of the corresponding candidate metadata;

[0105] If yes, performing the following steps:

[0106] Determining that the score value of the target candidate metadata is the corresponding Boltzmann score; the target candidate metadata is other metadata in the same data attribute except the same candidate metadata;

[0107] Aggregating and calculating the Boltzmann scores corresponding to the same candidate metadata; and taking the aggregation calculation result as the score value of the same candidate metadata.

[0108] It can be understood that, on the basis of the above-mentioned embodiments, in order to ensure the data completion accuracy, the reliability of scoring the candidate metadata is crucial. Therefore, in an optional embodiment, based on the screened target similarity, the Boltzmann scores corresponding to each different candidate embedding vector are calculated, and the specific calculation formula can be referred to formula (1):

[0109]

[0110] Wherein, Score m is the Boltzmann score of the mth candidate embedding vector, S m is the target similarity of the mth candidate embedding vector, and T is a temperature parameter. Wherein, the temperature parameter T is used to control the degree of emphasis on similarity, and the smaller the temperature parameter T is, the more the emphasis is on the maximum value. The larger the temperature parameter T is, the more the emphasis is on the uniform value.

[0111] In an alternative embodiment, when the temperature parameter T approaches 0, the output result is more biased towards the maximum value, that is, the larger the value, the higher the confidence. At this time, the higher the Boltzmann score Score m , the higher the confidence of the corresponding candidate metadata. When the temperature parameter T approaches ∞, the output result is more biased towards uniform distribution, at this time, the more the number of occurrences, the higher the confidence. Therefore, the more the same Boltzmann score Score m , the higher the confidence of the corresponding candidate metadata. In a specific embodiment, the temperature parameter T can be selected according to actual needs, and the initial default value can be set to 1.

[0112] In order to further improve the reliability of extracting the target metadata, further, the Boltzmann score Score m is calculated, and then it is determined whether there is the same candidate metadata under the same data attribute. If not, the calculated Boltzmann score Score m is taken as the score value of the corresponding candidate metadata. In order to facilitate understanding, the following will be explained in conjunction with Table 1 of the above embodiment.

[0113] For example, if the candidate embedding vector includes embedding vector E1 to embedding vector E5, and the target similarity is [0.83, 0.72, 0.65, 0.92, 0.66]. Under the object side data attribute, the candidate metadata includes object side B1 to object side B5. Obviously, each candidate metadata is different, therefore, the Boltzmann score Score m is calculated for object side B1 to object side B5 respectively, where m is 5. And the calculated Boltzmann score is taken as the score value of the corresponding object side B1 to object side B5.

[0114] In another alternative embodiment, if there is the same candidate metadata under the same data attribute. At this time, the Boltzmann scores corresponding to the same candidate metadata need to be aggregated and calculated, and the result of the aggregated and calculated is taken as the score value of the same metadata. It should be noted that the aggregated and calculated can include but is not limited to summation, mean value.

[0115] In addition, other metadata without the same candidate metadata, similarly, the Boltzmann score is directly taken as the score value of the corresponding candidate metadata. In order to facilitate understanding, the same Table 1 provided in the above embodiment is taken as an example for illustration.

[0116] For example, if the candidate embedding vectors include embedding vector E1 to embedding vector E5, and the target similarity is [0.83, 0.72, 0.65, 0.92, 0.66]. If the temperature parameter T is 1, the calculated Boltzmann scores of embedding vector E1 to embedding vector E5 are [0.21, 0.19, 0.18, 0.23, 0.18].

[0117] Wherein, for the principal party data attribute, the candidate metadata of embedding vector E2 and embedding vector E3 are the same, which is principal party A2, at this time, the Boltzmann scores corresponding to principal party A2, 0.19 and 0.18, need to be summed up to obtain an aggregated result of 0.37. At this time, the final score value of principal party A2 is 0.37. Obviously, the final score value of principal party A2 is greater than other score values, and principal party A2 will be extracted as the target completion data for the principal party data attribute when extracting the target metadata. Similarly, the candidate metadata of other data attributes are processed in the same way, which is not described here.

[0118] In order to realize efficient data completion, in another optional embodiment, the candidate metadata corresponding to the different data attributes of the candidate embedding vectors are scored, including:

[0119] The relevance of the to-be-completed request text and each candidate metadata is scored by calling a specified large model.

[0120] Specifically, the relevance of the to-be-completed request text and each candidate metadata can be scored by the specified large model, so that the one with the highest score value is taken as the target metadata. Similarly, the specified large model can include but is not limited to ChatGPT, GPT-4, Claude, and Tongyi Qianwen, which are not limited by the present application.

[0121] In an optional embodiment, in the absence of a target similarity, the data completion method provided by the present embodiment includes:

[0122] The target completion data is marked as a NaN value; and the NaN value is returned to the client;

[0123] Determine whether the user inputted artificial completion data within a preset time length;

[0124] If yes, the to-be-completed request text is completed by the artificial completion data;

[0125] If not, every interval of a preset period, the step of indexing and querying the target embedding vector with a similarity between the index and the query vector satisfying a preset condition in the pre-constructed vector database is entered, and the subsequent steps are executed.

[0126] It can be understood that, as Figure 3As shown, in specific embodiments, there may be target embedding vectors that do not meet the preset conditions, i.e., the target similarity is not greater than the threshold. At this time, if the metadata corresponding to the embedding vector with low similarity is returned, the confidence is too low, and the accuracy of the completed data cannot meet the user's demand.

[0127] Therefore, in order to solve the above technical problems, the target completion data can be marked as a NaN (NaN Imputation) value, which can be understood as a null value. That is, if the current recommendation result is too low, the user can choose to complete it by himself, or wait for the subsequent vector database data to be richer and then try to complete it again.

[0128] Specifically, after returning the NaN value to the client, if the user input artificial completion data can be obtained within a preset time period, indicating that the user chooses to complete it by himself, the artificial completion data is used to complete the to-be-requested text.

[0129] If no artificial completion data is received within the preset time period, indicating that the user gives up completing it by himself. At this time, every interval preset period, for example, every week, the step of indexing and querying the target embedding vector with a similarity between the index and the query vector satisfying the preset condition in the pre-constructed vector database is entered again, and the subsequent steps are executed. That is, every interval preset period, try to index again to see if the vector database has stored data with a similarity satisfying the preset condition.

[0130] It can be understood that in specific embodiments, the vector database can be dynamically updated, and if there is a request that is not completed, the query vector and the corresponding to-be-completed request text can be stored first, and an indexing attempt is performed every preset period.

[0131] Of course, it should be noted that in order to save storage resources, if the number of attempts reaches a preset number, the corresponding query vector and the corresponding to-be-completed request text are deleted, and a deletion signal is returned to the client. Of course, a signal can also be sent to the client to decide whether to delete it. In another optional embodiment, when the remaining space for storing the query vector and the corresponding to-be-completed request text is insufficient, the data that has been trying unsuccessfully is removed based on the principle of storing first and deleting first.

[0132] In an optional embodiment, the step of constructing the vector database comprises:

[0133] Obtaining a to-be-processed data set; the to-be-processed data set includes core data content and metadata of different data attributes related to the core data content;

[0134] Batch converting the core data content in the to-be-processed data set into embedding vectors;

[0135] Based on the FAISS vector index, an index number is obtained by constructing an index for the embedding vector.

[0136] A mapping relationship between the embedding vector, the index number and the metadata of different data attributes is established to construct a vector database.

[0137] In a specific embodiment, a to-be-processed data set is obtained, and if it is determined that the to-be-processed data set is complete data, the to-be-processed data set can be stored in the vector database after processing. The to-be-processed data set includes multiple complete data, and each piece of data includes core data content and metadata of different data attributes related to the core data content.

[0138] In order to ensure the integrity of the data, after obtaining the to-be-processed data set, the to-be-processed data set is preprocessed, specifically, the to-be-processed data set is de-duplicated and missing data is removed. After processing, the core data content in the to-be-processed data set is batch converted into an embedding vector by specifying a large model (which can include but is not limited to GPT, BERT, FastText). For example, the core data content in the to-be-processed data set can be batch generated into an embedding vector by using the GPT large model, and the generated dimension is a fixed vector of 1536.

[0139] At the same time, in order to accelerate the subsequent indexing speed, in an optional embodiment, an IVF (Inverted File) index is constructed for the embedding vector based on the FAISS vector index to obtain an index number. It can be understood that the IVF index can reduce memory occupation and use a clustering segmentation algorithm, which is particularly suitable for fast and efficient query scenarios.

[0140] Further, a mapping relationship between the embedding vector, the index number and the metadata of different data attributes is established, so that a vector database as shown in Table 1 in the above embodiment can be obtained.

[0141] Compared with the use of string matching in the past, the matching accuracy of the data completion method provided in the present application can reach more than 83%, which is higher than the actual matching accuracy of the past which needs manual experience. In addition, the calculation of string matching in the past is time-consuming, and if there are 30W pieces of large to-be-completed request texts in the application scenario, the query time is long and the efficiency is low. The present application can realize fast indexing and improve the data completion efficiency based on the pre-constructed vector data in a batch processing manner.

[0142] In the above embodiment, the data completion method is described in detail, and the present application also provides an embodiment of a data completion device.

[0143] Figure 4A structural schematic diagram of a data completion device provided by an embodiment of the present application is shown in Figure 4 The device comprises:

[0144] A request acquisition module 40 is configured to acquire a to-be-completed request text input by a user.

[0145] A query vector acquisition module 41 is configured to convert the to-be-completed request text into a query vector.

[0146] A target embedding vector index module 42 is configured to index, in a pre-constructed vector database, a target embedding vector that has a similarity to the query vector satisfying a preset condition.

[0147] A target metadata extraction module 43 is configured to extract, from the vector database, target metadata corresponding to the target embedding vector.

[0148] A data output module 44 is configured to return the target metadata as target completion data to a client.

[0149] In addition, the data completion device provided by the embodiment of the present application further comprises:

[0150] A candidate embedding vector determination module is configured to extract a target similarity greater than a threshold value to obtain a candidate embedding vector corresponding to the target similarity.

[0151] A target embedding vector determination module is configured to, when the target similarity is one, take the candidate embedding vector as the target embedding vector; when the target similarity is multiple, score candidate metadata of different data attributes corresponding to the candidate embedding vector; and take a candidate embedding vector corresponding to a candidate metadata with a maximum score value as the target embedding vector.

[0152] A target metadata determination module is configured to, when the target similarity is one, take all metadata corresponding to the target embedding vector as the target metadata; and when the target similarity is multiple, take metadata corresponding to a candidate metadata with a maximum score value as the target metadata.

[0153] A Boltzmann score determination module is configured to determine a Boltzmann score corresponding to each candidate embedding vector according to the target similarity.

[0154] A score value determination module is configured to determine whether there is same candidate metadata in a same data attribute; if not, take the Boltzmann score as a score value of the corresponding candidate metadata; and if yes, perform the following steps: determine a score value of a target candidate metadata as a corresponding Boltzmann score; the target candidate metadata is other metadata in the same data attribute except the same candidate metadata; aggregate and calculate Boltzmann scores corresponding to the same candidate metadata; and take an aggregation calculation result as a score value of the same candidate metadata.

[0155] The model calling module is configured to call a specified large model to score the relevance of the to-be-completed request text to each candidate metadata.

[0156] The marking module is configured to mark the target completion data as a NaN value, and return the NaN value to the client.

[0157] The processing module is configured to determine whether the user input artificial completion data is obtained within a preset time length, and if so, complete the to-be-completed request text by using the artificial completion data, and if not, enter the step of indexing and querying the target embedding vector that meets the preset condition between the index and the query vector in the pre-constructed vector database every preset period, and perform the subsequent steps.

[0158] The to-be-processed data set acquisition module is configured to acquire a to-be-processed data set, and the to-be-processed data set includes core data content and metadata of different data attributes related to the core data content.

[0159] The embedding vector generation module is configured to batch convert the core data content in the to-be-processed data set into embedding vectors.

[0160] The index construction module is configured to construct an index for the embedding vectors based on the FAISS vector index to obtain an index number.

[0161] The mapping relationship establishment module is configured to establish a mapping relationship between the embedding vectors, the index number and the metadata of different data attributes to construct a vector database.

[0162] Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1. Figure 5 As shown in FIG. 1, the electronic device includes a memory 50 configured to store a computer program.

[0163] A processor 51 is configured to execute the computer program to implement the steps of the data completion method mentioned in the above embodiments.

[0164] The electronic device provided by the embodiment can include but is not limited to a notebook computer or a desktop computer, etc.

[0165] The processor 51 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 51 can be implemented in at least one of a hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA), etc. The processor 51 can also include a main processor and a coprocessor. The main processor is a processor for processing data in a wake-up state, also referred to as a central processing unit (CPU). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 51 can be integrated with a graphics processor (GPU) for rendering and drawing content to be displayed by the display screen. In some embodiments, the processor 51 can further include an artificial intelligence (AI) processor for processing machine learning-related computing operations.

[0166] The memory 50 can include one or more computer-readable storage media, which can be non-transitory. The memory 50 can further include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In this embodiment, the memory 50 is at least used to store the following computer program 501, wherein the computer program is loaded and executed by the processor 51, and can implement the related steps of the data completion method disclosed in any of the preceding embodiments. In addition, the resources stored in the memory 50 can further include an operating system 502 and data 503, etc., and the storage mode can be temporary storage or permanent storage. The operating system 502 can include Windows, Unix, Linux, etc. The data 503 can include, but is not limited to, related data involved in the data completion method, etc.

[0167] In some embodiments, the electronic device can further include a display screen 52, an input / output interface 53, a communication interface 54, a power supply 55, and a communication bus 56.

[0168] Those skilled in the art can understand that, Figure 5 The structure shown in the figure does not constitute a limitation on the electronic device, and can include more or fewer components than those shown.

[0169] The electronic device provided by the embodiment of the present application comprises a memory and a processor, and the processor can realize the data completion method in the above embodiment when executing the program stored in the memory.

[0170] It should be noted that, although the operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring the operations to be performed in the particular order shown or sequentially, or requiring all of the illustrated operations to be performed to achieve a desired result. In some cases, multi-tasking and parallel processing can be advantageous. In addition, the separation of various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

Claims

1. A method of completing data, characterized by, The method comprises: obtaining a user-inputted request text to be completed; transforming the request text to be completed into a query vector; indexing, in a pre-constructed vector database, a target embedding vector that has a similarity between the query vector satisfying a preset condition; extracting, from the vector database, target metadata corresponding to the target embedding vector; returning the target metadata as target completion data to a client.

2. The data completion method of claim 1, wherein, The target embedding vector that has a similarity between the query vector satisfying a preset condition comprises: extracting a target similarity greater than a threshold to obtain a candidate embedding vector corresponding to the target similarity; when the target similarity is one, taking the candidate embedding vector as the target embedding vector; when the target similarity is multiple, scoring candidate metadata of different data attributes corresponding to the candidate embedding vector; and taking a candidate embedding vector corresponding to a candidate metadata with a maximum score value as the target embedding vector.

3. The data completion method of claim 2, wherein, The extracting, from the vector database, of target metadata corresponding to the target embedding vector comprises: when the target similarity is one, taking all metadata corresponding to the target embedding vector as the target metadata; when the target similarity is multiple, taking metadata corresponding to the maximum score value as the target metadata.

4. The data completion method of claim 2, wherein, The scoring of candidate metadata of different data attributes corresponding to the candidate embedding vector comprises: determining a Boltzmann score corresponding to each candidate embedding vector according to the target similarity; determining whether there is same candidate metadata in a same data attribute; if not, taking the Boltzmann score as a score value of the corresponding candidate metadata; if yes, performing the following steps: determining that a score value of a target candidate metadata is a corresponding Boltzmann score; the target candidate metadata is other metadata in the same data attribute except the same candidate metadata; performing an aggregate calculation on the Boltzmann score corresponding to the same candidate metadata; and taking an aggregate calculation result as a score value of the same candidate metadata.

5. The data completion method of claim 2, wherein, The scoring of candidate metadata of different data attributes corresponding to the candidate embedding vector comprises: calling a specified large model to score an association between the request text to be completed and each candidate metadata.

6. The data completion method of claim 2, wherein, In the absence of the target similarity, the method comprises: labeling the target completion data as a NaN value; and returning the NaN value to the client; determining whether user-inputted artificial completion data is obtained within a preset time length; if yes, completing the request text to be completed through the artificial completion data; if no, entering the step of indexing, in a pre-constructed vector database, a target embedding vector that has a similarity between the query vector satisfying a preset condition, and performing subsequent steps every interval of a preset period.

7. The data completion method of claim 1, wherein, The step of constructing the vector database comprises: obtaining a data set to be processed; the data set to be processed comprises core data content and metadata of different data attributes related to the core data content; bulk converting core data content in the to-be-processed data set into embedding vectors; indexing the embedding vectors based on a FAISS vector index to obtain index numbers; establishing a mapping relationship between the embedding vectors, the index numbers, and metadata of different data attributes to construct the vector database.

8. A data completion device, characterized by comprising: The device comprises: a request acquisition module configured to acquire a to-be-completed request text input by a user; a query vector acquisition module configured to convert the to-be-completed request text into a query vector; a target embedding vector indexing module configured to index, in a pre-constructed vector database, a target embedding vector that has a similarity to the query vector satisfying a preset condition; a target metadata extraction module configured to extract, from the vector database, target metadata corresponding to the target embedding vector; a data output module configured to return the target metadata as target completion data to a client.

9. An electronic device comprising a memory and a processor, said memory having stored thereon a computer program operable to run on said processor, characterized in that, The processor executes the computer program to implement the steps of the data completion method of any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the data completion method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Electric appliance function prediction method and device

    CN110895721A

  • Data asset completion method and device thereof, equipment and medium

    CN113987293A

  • Data directory construction method and device, medium and equipment

    CN115510116A

  • Code completion method and device, electronic equipment and medium

    CN118227106A

  • Data completion method based on position coding adaptive deep belief network

    CN119202557A