Knowledge search method and system based on generative retrieval and readable medium
By constructing a knowledge search method based on generative retrieval, and using coding models and generative models to generate text fragment identifiers with semantic expression capabilities, the problem of neglecting semantic information and insufficient matching in the prior art is solved, and efficient and accurate retrieval effects are achieved and cost is reduced.
Patent Information
- Application Number
- CN202510177214.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, generative search ignores the semantic information of text fragments, the matching degree between the fragment identifier and the text fragment is insufficient, and a large amount of query-fragment identifier data is required, which is extremely cost-consuming.
By constructing a knowledge search method based on generative retrieval, the identifier of the text fragment is first constructed, and a generative model that generates predicted query results based on the query task. The encoding model is used to encode the vectors of text fragments, and combined with dimensionality reduction operations and activation functions, an identifier that can effectively retain the semantic information of text fragments.
It improves the semantic expression ability of text fragment identifiers, enhances the matching degree between identifiers and text fragments, reduces dependence on labeled data, reduces the cost of technical implementation, and improves the overall performance of the retrieval system.
Smart Images

Figure CN120067264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a knowledge search method, system and readable medium based on generative retrieval. Background Art
[0002] Knowledge search refers to retrieving text fragments relevant to the user's query needs according to the query (generally including search conditions such as keywords) input by the user. Currently, the mainstream knowledge search methods include sparse retrieval, dense retrieval, and generative retrieval. Sparse retrieval evaluates the relevance between the query and the text fragment based on the frequency statistics of the query and the keywords in the text fragment, and retrieves the text fragment containing specific keywords. Dense retrieval uses deep learning technologies such as neural networks to map the text fragment and the query into the same vector space, and uses similarity functions such as vector inner product to measure the degree of relevance between the text fragment and the query. Generative retrieval transforms the retrieval task into a sequence-to-sequence generation task. First, fragment identifiers are assigned to the text fragments. The fragment identifiers can be obtained by random generation or by obtaining cluster IDs through hierarchical clustering of the semantic vectors of the text fragments. Then, a generative model is trained to generate fragment identifiers according to the query, and thus the corresponding text fragments are obtained.
[0003] However, the above methods all have problems to varying degrees. Sparse retrieval ignores the semantic connection between the query and the text fragment and cannot handle words with similar semantics but different literal expressions such as synonyms and near-synonyms. Dense retrieval relies on simple similarity function calculations to measure the relevance between the query and the text fragment and cannot effectively utilize the capabilities of deep neural networks.
[0004] Generative retrieval also has the following problems: 1. Whether using random generation or hierarchical clustering to obtain fragment identifiers, the semantic information of the text fragments is ignored to varying degrees; 2. The existing fragment identifier generation methods cannot guarantee the matching degree between the fragment identifier and the text fragment; 3. Currently, generative retrieval requires a large amount of query-fragment identifier data annotation work, resulting in extremely high cost consumption. Summary of the Invention
[0005] In order to overcome the defects of generative retrieval in the above-mentioned prior art, the present invention proposes a knowledge search method, system and readable medium based on generative retrieval.
[0006] A knowledge search method based on generative retrieval proposed by the present invention first constructs identifiers of text segments as known identifiers and constructs a generative model for generating identifiers of predicted query results based on query tasks; then, for a query task to be queried, predicted identifiers are generated through the generative model, and then text segments corresponding to the known identifiers closest to the predicted identifiers are obtained as query results of the query task to be queried.
[0007] The method for constructing identifiers of text segments is as follows: First, map the text segment to a vector, and then encode the vector through a pre-trained encoding model; then process the encoded vector matrix into a semantic feature representation of the text segment; after dimensionality reduction processing of the semantic feature representation and then activation, the identifier of the text segment is obtained.
[0008] Preferably, the training method of the encoding model includes the following steps:
[0009] S11. Obtain a set of text segments and rewrite the text segments using a rewriting template;
[0010] S12. Combine the current encoding model to obtain identifiers of text segments and identifiers of rewritten texts;
[0011] S13. Calculate a first loss function based on the identifiers of text segments and the identifiers of rewritten texts, and update the encoding model according to the first loss function;
[0012] S14. Repeat steps S12 - S13 until the loss function converges, and then fix the encoding model.
[0013] Preferably, the first loss function loss 1 is:
[0014]
[0015] where B represents the number of samples in a batch during training, log(·) represents the calculation of the logarithmic function, exp(·) represents the calculation of the exponential function, id i and idg i respectively represent the identifier of text segment d i and the identifier of its rewritten text; id j represents the identifier of text segment d j ; sim(id i , idg i ) represents the similarity between id i and idg i ; sim(id i , id j ) represents the similarity between id i and id j .
[0016] Preferably, the encoding model adopts Semantic-encoder.
[0017] Preferably, the average value of the vector matrix encoded by the encoding model is calculated to form a semantic feature representation in vector form, and the semantic feature representation is activated after dimensionality reduction processing to obtain an identifier.
[0018] Preferably, the training method of the generative model is as follows:
[0019] St1. Combine the trained encoding model to obtain the identifier of each text segment;
[0020] St2. Generate a pseudo-query task for the text segment with a known identifier, and the pseudo-query task can obtain a query result in the text segment;
[0021] St3. Randomly extract the input of the pseudo-query task and generate a segment identifier corresponding to the pseudo-query task as a query symbol;
[0022] St4. Calculate the second loss function based on the identifier of the text segment and the query symbol, and optimize the generative model through the second loss function;
[0023] St5. Repeat steps St3-St4 until the second loss function converges, then fix the generative model.
[0024] Preferably, the second loss function adopts the cross-entropy loss function.
[0025] Preferably, when constructing the identifier of the text segment, the S-Embedding embedding layer is used to map the text segment into an embedding vector, and then the embedding vector is input into the encoding model for encoding.
[0026] A knowledge search system based on generative retrieval proposed by the present invention includes a memory and a processor. A computer program is stored in the memory, and the processor is connected to the memory. The processor is used to execute the computer program to implement the knowledge search method based on generative retrieval.
[0027] A storage medium proposed by the present invention stores a computer program, and the computer program is used to implement the knowledge search method based on generative retrieval when executed.
[0028] The advantages of the present invention are as follows:
[0029] (1) Enhance the semantic expression ability of the identifiers of text segments: Encode the vectors of text segments through an encoding model, and combine dimensionality reduction operations and activation functions to generate identifiers that can effectively retain the semantic information of text segments. This solves the problem in the prior art that segment identifiers ignore semantic information, making the generated identifiers more representative and distinguishable. It enhances the retrieval system's understanding ability of text segments, thereby improving the accuracy and relevance of retrieval results.
[0030] (2) Enhance the matching degree between identifiers and text segments: Use a rewriting template to rewrite text segments to ensure that the core information and intention of the rewritten segments remain unchanged. At the same time, train the encoding model through a contrastive loss function to make the identifiers highly match the text segments. This solves the problem in the prior art that the matching degree between segment identifiers and text segments is insufficient, reducing the cases of false matching and missed matching. It improves the stability and reliability of the retrieval system, making the retrieval results more accurate.
[0031] (3) Reduce the dependence on labeled data: Use a generative model to generate high-quality pseudo-query tasks, significantly reducing the dependence on a large amount of labeled query-segment identifier data. It reduces the cost and time of data annotation, and at the same time solves the problem of data scarcity, making model training more efficient. The generated pseudo-query tasks are of high quality and can effectively support the training of the generative model, further enhancing the generalization ability and retrieval effect of the model.
[0032] (4) Improve the overall performance of the retrieval system: Through the comprehensive application of technologies such as semantic encoding, dimensionality reduction, large model rewriting, and pseudo-query task generation, the overall performance of the generative retrieval system is significantly improved. The retrieval results are more accurate, relevant, and the system has a stronger understanding ability for complex semantics. It is applicable to large-scale text retrieval scenarios and can meet the requirements of efficient and accurate retrieval.
[0033] (5) Reduce the technical implementation cost: By reducing the dependence on labeled data and optimizing the model training process, the cost and complexity of technical implementation are reduced. The introduction of large models and the application of pseudo-query generation technology make the system more scalable and practical. It is applicable to a variety of application scenarios, such as search engines, intelligent question-and-answer systems, etc., and has broad application prospects.
[0034] (6) Through technologies such as semantic encoding, dimensionality reduction, large model rewriting, and pseudo-query generation, the present invention significantly enhances the semantic understanding ability, matching accuracy, and training efficiency of the generative retrieval system, while reducing the implementation cost. Its beneficial effects are mainly reflected in the improvement of retrieval performance, the reduction of data dependence, and the reduction of technical costs, and it has important practical value and broad application prospects. Brief Description of the Drawings
[0035] Figure 1Flowchart of the method for obtaining the identifier of the text segment;
[0036] Figure 2 Flowchart of the training method for the encoding model;
[0037] Figure 3 Flowchart of the training method for the generative model;
[0038] Figure 4 Flowchart of a knowledge search method based on generative retrieval. Detailed implementation manner
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] A knowledge search method based on generative retrieval proposed in this embodiment first constructs the identifier of the text segment as a known identifier and constructs a generative model for generating the identifier of the predicted query result based on the query task; then, for the query task to be queried, the generative model is used to generate a predicted identifier, and then the text segment corresponding to the known identifier closest to the predicted identifier is obtained as the query result of the query task to be queried.
[0041] Refer to Figure 1 , in this embodiment, the method for obtaining the identifier of the text segment includes the following steps:
[0042] S1. Convert the text segment d i into a vector matrix H i ;
[0043] Specifically, during implementation, the text segment d i can be first mapped to an embedding vector E i by using the S-Embedding embedding layer, and then the embedding vector E i is encoded by using the Semantic-encoder tool to obtain the vector matrix H i ; The formula is expressed as:
[0044] E i = S-Embedding(d i ) #(1)
[0045] H i = Semantic-encoder(E i ) #(2)
[0046]
[0047] Among them, n represents the length of the text segment, respectively representing the 1st, 2nd, and nth characters in the text segment d i ;
[0048] m represents the length of the embedding vector sequence; respectively representing the 1st, 2nd, and mth sub-vectors of the embedding vector E i ; respectively representing the 1st, 2nd, and mth vectors of the vector matrix H i ; the jth vector in H i is the jth sub-vector of E corresponding encoding vector; 1 ≤ j ≤ m; i
[0049] S2. Take the average of each vector in the vector matrix H i to obtain the semantic feature representation s of the text segment d i ; i
[0050]
[0051] S3. After performing a dimensionality reduction operation on the semantic feature representation s i and then activating it, obtain the identifier id of the text segment d i ; i
[0052] The formula is expressed as:
[0053] u i = Umap(s i ) #(3)
[0054] id i = Step(u i ) #(4)
[0055]
[0056] Among them, Umap(·) represents the dimensionality reduction operation, and u i represents the semantic feature representation after dimensionality reduction; Step(·) represents the activation operation.
[0057] Specifically, u i is a vector containing multiple values, and Step(u i ) means that each value in u i needs to be calculated by the Step(·) function to obtain the identifier id i .
[0058] Specifically, Umap adopts an existing dimensionality reduction technique that uses nearest neighbor search to construct a high-dimensional neighborhood graph for optimizing the embedding in the low-dimensional space.
[0059] Referring to Figure 2 , in step S1, the S-Embedding layer is a pre-trained model. The training process of the Semantic-encoder includes the following steps:
[0060] S11. First, obtain a set of text fragments {d i}, and then rewrite the text fragments using a set rewrite template. Let the rewritten text of text fragment d i be denoted as dg i ;
[0061] S12. Obtain the identifier id i of text fragment d i and the identifier idg i of the rewritten text dg i ; For obtaining the identifier, refer to steps S1 - S3;
[0062] S13. Calculate the first loss function loss i based on id i and idg 1 , and update the Semantic-encoder according to the first loss function loss 1 ;
[0063]
[0064] where B represents the number of samples in a batch during the training process, log(·) represents the logarithmic function calculation, exp(·) represents the exponential function calculation, and sim(·) represents the similarity calculation function; sim(id i , idg i ) represents the similarity between id i and idg i ; id j represents the identifier of text fragment d j ; sim(id i , id j ) represents the similarity between id i and id j .
[0065] In this step, the Semantic-encoder is optimized using the backpropagation mechanism. Since Umap(·) and Step(·) are not differentiable, the gradient is not calculated to update the parameters during the backpropagation process.
[0066] S14. Repeat steps S12 - S13 until the loss function converges, then fix the Semantic-encoder.
[0067] Refer to Figure 3 , in this embodiment, the training method of the generative model includes the following steps:
[0068] St1. Combine the trained Semantic-encoder to execute steps S1 - S3 to obtain the final identifier fid of each text segment i ;
[0069] St2. Generate a pseudo-query task for the text segment with the known identifier fid i , and the pseudo-query task can obtain a query result in the text segment;
[0070] St3. Randomly extract the input of the pseudo-query task and generate the segment identifier corresponding to the pseudo-query task by the generative model Generative
[0071]
[0072] where Generative(·) represents generating the segment identifier.
[0073] St4. Optimize the generative model Generative based on fid i and using the cross-entropy loss;
[0074]
[0075] where Cross-Entropy(·) represents the cross-entropy loss function, Loss 2 represents the loss value, and optimize the generative model Generative using the backpropagation mechanism of the neural network.
[0076] St5. Repeat steps St3 - St4 until Loss 2 converges, then fix the generative model Generative.
[0077] Refer to Figure 4 , a knowledge search method based on generative retrieval proposed in this embodiment specifically includes the following steps:
[0078] SA1. For the text segments in the text library, combine the trained Semantic-encoderr to execute steps S1 - S3 to obtain the final identifier fid of each text segment i ;
[0079] SA2. Input the task to be queried into the trained Generative model to obtain the predicted segment identifier
[0080] SA3. Obtain the fid in the text library that is closest to the predicted segment identifier as the target identifier, and obtain the text segment corresponding to the target identifier as the query result of the task to be queried i
[0081] Of course, for those skilled in the art, the present invention is not limited to the details of the above exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claims involved
[0082] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art
[0083] The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies
Claims
1. A knowledge search method based on generative retrieval, characterized in that: First, an identifier of a text segment is constructed as a known identifier, and a generative model is constructed to generate an identifier for predicting query results based on the query task; then, for the task to be queried, a predicted identifier is generated through the generative model, and then a text segment corresponding to the known identifier closest to the predicted identifier is obtained as the query result of the task to be queried; The method for constructing the identifier of the text fragment is as follows: first, the text fragment is mapped into a vector, and then the vector is encoded through a pre-trained encoding model; then the encoded vector matrix is processed into a semantic feature representation of the text fragment; the semantic feature representation is reduced in dimension and then activated to obtain the identifier of the text fragment.
2. The knowledge search method based on generative retrieval as claimed in claim 1, characterized in that: The training method of the encoding model includes the following steps: S11, obtaining a set of text fragments, and rewriting the text fragments using a rewriting template; S12, combining the current encoding model, obtaining an identifier of the text segment and an identifier of the rewritten text; S13, calculating a first loss function based on the identifier of the text segment and the identifier of the rewritten text, and updating the encoding model according to the first loss function; S14. Repeat steps S12-S13 until the loss function converges, and then fix the coding model.
3. The knowledge search method based on generative retrieval as claimed in claim 2, characterized in that: The first loss function loss1 is: Where B represents the number of samples in a batch during training, log(·) represents logarithmic function calculation, exp(·) represents exponential function calculation, and id i and idg i Represents the text fragment d i The identifier of the text it rewrites; id j Represents a text fragment d j Identifier of sim(id i ,idg i ) indicates id i and idg i Similarity; sim(id i ,id j ) indicates id i and id j The similarity.
4. The knowledge search method based on generative retrieval according to claim 1, characterized in that: The encoding model uses Semantic-encoder.
5. The knowledge search method based on generative retrieval according to claim 1, characterized in that: The vector matrix encoded by the encoding model is averaged to form a semantic feature representation in vector form. The semantic feature representation is processed through dimensionality reduction and then activated to obtain the identifier.
6. The knowledge search method based on generative retrieval as claimed in claim 1, characterized in that: The generative model is trained as follows: St1. Combine the trained encoding model to obtain the identifier of each text segment; St2, generate pseudo query tasks for text fragments with known identifiers, and the pseudo query tasks can obtain query results in the text fragments; St3, randomly extract pseudo-query tasks and input them into the generative model to generate fragment identifiers corresponding to the pseudo-query tasks as query symbols; St4, calculating a second loss function based on the identifier of the text segment and the query symbol, and optimizing the generative model through the second loss function; St5. Repeat steps St3-St4 until the second loss function converges, then fix the generative model.
7. The knowledge search method based on generative retrieval according to claim 1, characterized in that: The second loss function uses the cross entropy loss function.
8. The knowledge search method based on generative retrieval as claimed in claim 1, characterized in that: When constructing the identifier of a text fragment, the S-Embedding embedding layer is used to map the text fragment into an embedding vector, and then the embedding vector is input into the encoding model for encoding.
9. A knowledge search system based on generative retrieval, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and the processor is connected to the memory, and the processor is used to execute the computer program to implement the knowledge search method based on generative retrieval as described in any one of claims 1-8.
10. A storage medium, characterized in that: A computer program is stored, and when the computer program is executed, it is used to implement the knowledge search method based on generative retrieval as described in any one of claims 1 to 8.