Multi-dimensional efficient recall method for complex timeliness document
By constructing a document graph and fusing keywords and semantic vectors with the BERT model, the problem of traditional document retrieval methods being unable to identify expired documents and dynamically adjust weights in time-sensitive documents is solved, achieving efficient and accurate document recall and improving the accuracy of search results and user experience.
Patent Information
- Application Number
- CN202511363489.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-23
AI Technical Summary
Traditional document retrieval methods cannot intelligently identify expired documents or dynamically adjust weights when dealing with complex and time-sensitive documents, resulting in insufficient retrieval accuracy. They also cannot identify complex relationships between documents, increasing user screening time and the risk of incorrect decisions.
We construct a document graph, use the BERT model to fuse keywords and semantic vectors, and use the graph to recall and trace alternative documents. We introduce dynamic weight allocation and length-aware gating mechanisms to achieve efficient multi-dimensional recall.
It enables accurate, efficient, and comprehensive retrieval of timely documents, improving the timeliness and accuracy of search results and reducing the risk of user decision-making errors.
Smart Images

Figure CN120849684A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document retrieval technology, and more specifically, to a multi-dimensional and efficient retrieval method for complex and time-sensitive documents. Background Technology
[0002] In fields such as taxation, corporate compliance, and finance, timely documents are characterized by their sheer volume, frequent updates, and high timeliness. These documents are complex, covering multiple levels and areas, forming a vast and sophisticated document system. Specifically, this system includes not only macro-level documents at the national level but also detailed implementation rules at the local level, industry regulatory documents, and internal corporate compliance requirements. Documents at all levels and in all areas are intertwined and work together to achieve their intended effect.
[0003] The relationship between old and new documents is complex and mainly manifests in the following types: 1. Modification Relationship: The new document may partially modify or supplement some clauses of the old document. Such modifications may involve multiple aspects such as document content, scope of implementation, and execution standards. For example, with technological development or market changes, the new energy vehicle subsidy document may be adjusted in terms of subsidy amount, applicable vehicle models, or application conditions to better adapt to new development needs.
[0004] 2. Relationships: Some documents are interconnected; for example, one document may be an interpretation or supplement to another. This relationship helps in a more comprehensive understanding of the document's intent and implementation details, but it also increases the complexity of document retrieval and comprehension.
[0005] 3. Obsolescence: As circumstances change or new documents are introduced, older documents may become obsolete or partially obsolete. For example, some outdated industry support documents may no longer be applicable due to industrial restructuring, but users may still encounter these obsolete documents when searching.
[0006] 4. Extensions: Some documents may be delayed in implementation or have their validity extended due to special circumstances. For example, some local governments may extend the validity of certain tax incentive documents to continue supporting business development in response to downward economic pressure.
[0007] These complex relationships make document management and retrieval extremely difficult. When searching for documents, users not only need to find currently valid content, but also need to understand whether these documents are invalid and their relationships with other documents in order to accurately grasp the evolution and impact of the documents.
[0008] However, traditional document retrieval methods mainly rely on vector retrieval and keyword retrieval, which have many limitations: 1. Insufficient Precision: The common dual-path retrieval system (vector retrieval + keyword retrieval) typically uses a fixed weighting, such as merging the results of vector and keyword searches in a fixed ratio. This fixed weighting cannot be adjusted according to different query scenarios and data characteristics. In some scenarios, vector retrieval may be more suitable, while in others, keyword retrieval may be more advantageous. A fixed fusion method cannot dynamically adapt to these differences, resulting in suboptimal search performance.
[0009] 2. Inability to intelligently identify expired documents and complex relationships: Traditional retrieval methods have significant limitations when handling time-sensitive documents. They cannot intelligently identify expired documents, nor can they provide corresponding alternatives or update suggestions. Users often need to manually search and filter documents when looking for replacements for expired ones, a process that is not only time-consuming and laborious but also prone to missing key information. Furthermore, traditional retrieval methods also fail to effectively identify and display complex relationships between documents, such as modification, relevance, expiration, and extension. Users find it difficult to fully understand the overall picture and evolution of documents through simple keyword searches, often requiring significant time for manual filtering and comparison, which greatly reduces retrieval efficiency and user experience.
[0010] 3. Timeliness issue: Documents are highly time-sensitive, but traditional search methods cannot automatically filter out the latest valid documents. Users need to judge the validity of documents themselves, which not only increases the difficulty of retrieval, but may also lead to wrong decisions due to the use of outdated documents.
[0011] Therefore, traditional document retrieval methods can no longer meet users' needs for accurate, efficient, and comprehensive retrieval. There is an urgent need for more intelligent and efficient retrieval tools and methods to help users quickly and accurately obtain and understand the latest and most effective information. Summary of the Invention
[0012] This invention overcomes the shortcomings of existing technologies and provides a multi-dimensional and efficient retrieval method for complex and time-sensitive documents, enabling accurate, efficient, and comprehensive searching.
[0013] The technical solution of the present invention is as follows: A multi-dimensional and efficient retrieval method for complex and time-sensitive documents includes the following steps: (1) Steps for constructing a document graph: Represent time-sensitive documents as nodes in the graph. Node attributes include document content, effective date and status marker. The status marker includes valid and invalid. Nodes are connected by edges, and the edge connections represent semantic relationship types, including repeal, extension, interpretation and update. (2) Recall phase steps: Based on user queries, the keyword vector and the overall semantic vector of the sentence are fused to generate a new fused vector, and joint recall of keywords and text is performed; (3) Graph Recall Steps: After a user submits a query request, the status of each document node in the recall phase is checked: if the status is invalid, the replacement document node is traced through the semantic relationship edge in the graph; if the status of the replacement document node is deferred, its latest revision is retrieved and the invalid document node is replaced with a valid document node as the final recall result. (4) Output results steps: return valid documents and query basis links.
[0014] Furthermore, the recall model in step (2) is fine-tuned based on the BERT model. The specific process of the recall model is as follows: (2.1) Steps for constructing positive and negative samples: Construct semantic samples and keyword samples; positive samples of semantic samples To select relevant questions from the corpus Relevant document paragraphs, negative samples For random selection and problems Irrelevant document paragraphs; positive samples of keyword samples To start from the problem Core terms are extracted using TF-IDF, and negative samples are used. To start from the problem Randomly selected from non-core words; (2.2) Sample vectorization steps: Before vectorizing positive and negative samples, vectorize the text and special identifier [CLS], and concatenate them separately to represent the overall semantic information of the text; (2.3) Fusion vector step: The keyword vector and semantic vector are fused. The keyword information and semantic information are fused through a length-aware dynamic gating mechanism. Specifically, when the BERT model captures implicit semantics that cannot be represented by length, when the query text is relatively short, it is determined that the query text is more related to the keywords. When the query text is relatively long, it is determined that the query text is more related to the semantics. Thus, a learnable length encoder is set.
[0015] Furthermore, the concatenation representation in the sample vectorization step (2.2) is as follows:
[0016]
[0017] semantic features Encoding the BERT model at position [CLS], keyword vector The average value of the encoded vectors at all positions except [CLS]. This represents a sample text to be vectorized. This represents all position indices in the input sequence except for [CLS].
[0018] Furthermore, in the (2.3) fusion vector step, the query length is mapped to a vector representation with semantic information as follows:
[0019] in It is based on the question The length vector constructed from the text length has dimensions of . Dimensionality is reduced to one dimension through a linear layer, and then... The activation function is normalized to obtain the query length gating weight α mentioned above.
[0020] Based on the query length gating weight α, the semantic vector and keyword vector are fused together, and the final fused vector is as follows: .
[0021] Furthermore, the user query and document text input retrieval model are used to generate semantic vectors and keyword vectors; the semantic vectors and keyword vectors are fused through a dynamic gating mechanism to generate a fused vector; and cosine similarity is calculated based on the fused vector to filter relevant documents.
[0022] Furthermore, the dynamic gating mechanism includes generating a length vector based on the length of the user query text, and using a linear layer and The activation function normalizes the length vector into a weight parameter σ; the contribution of the keyword vector and semantic vector is dynamically adjusted using the weight parameter σ.
[0023] Furthermore, the state markers and semantic relationship edges are automatically updated based on document timeliness rules to ensure the timeliness of the recall results.
[0024] Furthermore, this includes constructing a vector-level contrastive loss function:
[0025] in This indicates the calculation of cosine similarity between two texts. Let be the vector representation of the i-th sample within a training batch. This is a vector representation of the positive samples within a training batch. is the temperature hyperparameter used to adjust the model's attention to difficult negative samples, and B is all the samples included in a training batch.
[0026] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method described above.
[0027] A computer-readable storage medium storing a computer program that, when executed, implements the method described above.
[0028] The advantages of this invention compared to the prior art are: This invention introduces the concept of a "time-sensitive document graph," modeling documents in a structured form. The graph structure uses "valid / invalid" states labeled with semantic edges such as "modified," "relevant," "invalid," and "delayed" to construct evolutionary paths between documents. Based on these semantic edges, multi-hop relationship tracking can be performed, automatically inferring the currently valid document version. This helps users accurately obtain target documents and their historical evolution paths, significantly improving the timeliness, accuracy, and interpretability of document retrieval.
[0029] This invention integrates keyword recall, semantic recall, and graph-based recall to form a multi-strategy fusion mechanism with dynamically adjustable weights. First, keyword recall ensures accurate matching for clear query intents, quickly identifying document nodes containing core terms. Second, semantic recall captures the potential semantic relationships between queries and document text through a deep vectorization model, improving adaptability to complex queries such as fuzzy, synonymous, and rewritten queries. Finally, graph-based recall utilizes multi-hop chain-like relationship reasoning in a timely document graph to achieve cross-node and cross-version document retrieval and automatic replacement.
[0030] During the fusion phase, a dynamic weight allocation and length-aware gating mechanism were introduced, which automatically adjusts the contribution of keyword recall results to the final ranking based on the query text length characteristics. For example, for short text queries with high keyword density, the system tends to increase the weight of keyword recall results; for long text queries or highly descriptive queries, it tends to increase the weight of semantic recall. Through this mechanism, the recall strategy can be adaptively adjusted under different business scenarios and different query characteristics, achieving a balance between high precision and high coverage. At the same time, with the timeliness tracking and status verification functions of graph recall, it is ensured that the final output document results are not only highly relevant but also the latest valid version, thereby effectively improving the user's search experience and decision accuracy. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the overall process of the present invention; Figure 2 This is a diagram of the model training architecture of the present invention. Detailed Implementation
[0032] Embodiments of the present invention are described in detail below, wherein the same or similar reference numerals denote the same or similar elements or elements with similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0033] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.
[0034] The numbering of steps mentioned in the various embodiments is merely for descriptive convenience and does not imply a sequential relationship. Different steps in various specific embodiments can be combined in different orders to achieve the inventive objective of this invention.
[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0036] like Figure 1 , 2 As shown, a multi-dimensional and efficient retrieval method for complex and time-sensitive documents includes the following steps: (1) Steps for constructing a document graph: Represent time-sensitive documents as nodes in the graph. Node attributes include document content, effective date and status marker. The status marker includes valid, invalid and draft. Nodes are connected by edges, and the edge connection represents the semantic relationship type, including repeal, extension, interpretation and update.
[0037] First, a structured graph system is constructed based on expired documents. Each document is treated as a node in the graph, with node attributes including document content, effective date, and current status marker (such as "valid," "expired," "draft," etc.). Different document nodes are interconnected through semantic relationship edges, with relationship types including but not limited to "repealed," "extended," "interpreted," and "updated," used to describe the logical connections and evolutionary paths between documents. This graph structure provides the foundation for subsequent multi-hop relationship reasoning and semantic retrieval.
[0038] (2) Recall Phase Steps: Based on the user query, the keyword vector and the overall semantic vector of the sentence are fused to generate a new fused vector, and a contrastive loss is performed for recall. This process uses vectorization to deeply integrate keyword information and semantic information into the same vector space, thus achieving an effective combination of the two. Specifically, a finely tuned BERT model is used for vectorization, i.e., the recall model. The user query and document text are input into the recall model to generate semantic vectors and keyword vectors; the semantic vectors and keyword vectors are fused using a dynamic gating mechanism to generate a fused vector; and cosine similarity is calculated based on the fused vector to filter relevant documents.
[0039] The recall model efficiently integrates the semantic features of keywords with the overall semantic features of sentences, generating more representative vector representations. In particular, the loss function is carefully designed: it narrows the vector distance between the user's query and relevant documents while simultaneously widening the vector distance between the user's query and irrelevant documents, thus generating an enhanced vector representation that comprehensively considers semantics and keywords. This design not only preserves the explicitness of keywords but also fully captures the flexibility of semantics, significantly improving the accuracy and coverage of search results.
[0040] The recall model is an adjustment based on the BERT model. The specific process of the recall model is as follows: (2.1) Positive and negative sample construction steps: construct semantic samples and keyword samples.
[0041] Positive samples of semantic samples To select relevant questions from the corpus Relevant document paragraphs, negative samples For random selection and problems Irrelevant document segment.
[0042] Positive samples of keyword samples To start from the problem Core terms are extracted using TF-IDF, and negative samples are used. To start from the problem Randomly selected from non-core words.
[0043] The ratio of positive to negative samples is generally set to 1:1, and the corpus typically contains 2,000 positive samples and 2,000 or more negative samples.
[0044] (2.2) Sample Vectorization Steps: Before vectorizing positive and negative samples, the text and special identifier [CLS] are vectorized and concatenated to represent the overall semantic information of the text. The concatenated representation is as follows:
[0045]
[0046] semantic features Encode the BERT model at the [CLS] position. This represents a sample text to be vectorized, specifically a keyword vector. The average value of the encoded vectors at all positions except [CLS]. This represents all position indices in the input sequence except for [CLS].
[0047] (2.3) Vector Fusion Step: The keyword vector and semantic vector are fused using a length-aware dynamic gating mechanism. Specifically, when capturing implicit semantics that cannot be represented by length using a recall model, if the query text is relatively short, it is considered to be more related to the keywords; if the query text is relatively long, it is considered to be more related to the semantics. Thus, a learnable length encoder is set.
[0048] The query length is mapped to a semantically informational vector representation as follows:
[0049] in It is based on the question The length vector constructed from the text length has dimensions of . Dimensionality is reduced to one dimension through a linear layer, and then... The activation function is normalized to obtain the query length gating weight α. During training, the model learns a specific vector representation for each possible text length, generating the query length gating weight α. During inference, the length vector is directly based on the pre-trained text length index.
[0050] Based on the query length gating weight α, the semantic vector and keyword vector are fused together, and the final fused vector is as follows: .
[0051] As a preferred option, a vector-level contrastive loss function is also constructed:
[0052] in This indicates the calculation of cosine similarity between two texts. Let be the vector representation of the i-th sample within a training batch. This is a vector representation of the positive samples within a training batch. is the temperature hyperparameter used to adjust the model's attention to difficult negative samples, and B is all the samples included in a training batch.
[0053] (2.4) Document retrieval steps: For all documents in the knowledge base, their feature representations are obtained through the steps described above: in Let n be the set of feature representations of all documents in the knowledge base, and n be the number of documents in the knowledge base.
[0054] By comparing the similarity between the original text vector and the text vector in the knowledge base, the Top N documents are selected as the recall results.
[0055]
[0056] in, This is the final feature representation of the problem text. A function for selecting the Top N similarity scores.
[0057] (3) Graph Recall Steps: After a user submits a query request, the status of each document node in the recall phase is checked: if the status is invalid, the replacement document node is traced through the semantic relationship edge in the graph; if the status of the replacement document node is deferred, its latest revision is retrieved and the invalid document node is replaced with a valid document node as the final recall result.
[0058] The specific document graph query and retrieval process is as follows: (3.1) Status verification and link tracing steps: After the user submits a document query request, the system performs status verification on the recalled Top N document nodes P1.
[0059] If the status of P1 is "deprecated", the alternative document tracing mechanism is automatically triggered, tracing back to the corresponding alternative document node P2 along the document graph relationship chain.
[0060] (3.2) Alternative document status detection steps: Detect the status of P2. If P2 is in the "delayed" state, it is necessary to find the latest revision and start the document latest revision retrieval sub-process. If the latest revision exists, it is necessary to trace back to the document node p3.
[0061] (3.3) Modified Version Confirmation Steps: Based on the document expiration rules, within the extended period of P2, retrieve and extract its latest modified version P3. Verify the status of P3; if it is in a "valid" state, then use that document text as the final version and replace the initial recall result. The status markers and semantic relationship edges are automatically updated based on the document expiration rules to ensure the timeliness of the recall results.
[0062] (4) Output results steps: return valid documents and query basis links.
[0063] Specific implementation process combined with Figure 1 as follows: Step 1: Start the process, assuming the input is "How do I pay stamp duty?".
[0064] Step 2: The user submits an inquiry, and relevant evidence is retrieved. The evidence retrieval model used is as follows: Figure 2 The model shown has the following recall result: "Article 13 If the taxpayer is an entity, he / she shall declare and pay stamp duty to the competent tax authority at the location of his / her institution; if the taxpayer is an individual, he / she shall declare and pay stamp duty to the competent tax authority at the place where the taxable document is executed or at the place of residence of the taxpayer." Step 3: The system finds the corresponding reference node P1 through the recall results.
[0065] Step 4: Perform a status check based on node P1: If P1 is valid → Step 9.
[0066] If P1 has been deprecated → Step 5.
[0067] Step 5: Trigger alternative document search and find P2.
[0068] Step 6: Perform the P2 status check: If P2 is valid → Step 9.
[0069] If P2 has been postponed → Step 7.
[0070] Step 7: Retrieve the latest revision of P2, P3.
[0071] Step 8: Perform P3 status check: If P3 is valid, replace the initial query result with P3 and proceed to step 9.
[0072] Step 9: Return the valid evidence and the query evidence link.
[0073] Step 10: End the process.
[0074] Examples of the recall model are as follows: Step 2.1 Text Encoding and Feature Extraction: 1. Input the raw text (example: "How do I pay stamp duty?") into the BERT model.
[0075] 2. BERT encodes text into a vector sequence ( to ) and special vectors .in As the semantic representation of the entire sentence (denoted as ).
[0076] 3. Convert the vector output by BERT ( to Keyword representations are generated through average pooling layers. ).
[0077] Step 2.2 Obtaining Length Information Weights: 1. Generate a text length vector based on the text length. During reasoning, it is only necessary to obtain the length vector using the numerical value of the text length. That is, through linear layers and Activation function The dimension is reduced to one dimension, σ, which serves as the weight of the keyword vector.
[0078] Step 2.3 Feature fusion and loss calculation: Keyword representation ( ) and semantic representation ( The weighted fusion is performed, with the weights controlled by σ. = (1 -σ) + σ The final feature representation after output fusion is as follows ,based on The loss is calculated and the model parameters are optimized through backpropagation.
[0079] Step 2.4 Document Recall: For all documents in the knowledge base, their feature representations are obtained through the above steps. The top 5 documents are selected as the recall results by comparing the similarity between the original text vectors and the text in the knowledge base.
[0080] In one embodiment, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method steps provided in the above embodiments.
[0081] In one embodiment, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method steps provided in the above embodiments.
[0082] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0083] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0084] The several embodiments described in this application are quite specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and refinements should also be considered within the scope of protection of this invention. Therefore, the scope of protection of this patent application should be determined by the appended claims.
Claims
1. A multi-dimensional and efficient recall method for complex and time-sensitive documents, characterized in that, Includes the following steps: (1) Steps for constructing a document graph: Represent time-sensitive documents as nodes in the graph. Node attributes include document content, effective date and status marker. The status marker includes valid and invalid. Nodes are connected by edges, and the edge connections represent semantic relationship types, including repeal, extension, interpretation and update. (2) Recall phase steps: Based on user queries, the keyword vector and the overall semantic vector of the sentence are fused to generate a new fusion vector, and the keywords and text are jointly recalled; (3) Graph Recall Steps: After a user submits a query request, the status of each document node in the recall phase is checked: if the status is invalid, the replacement document node is traced through the semantic relationship edge in the graph; if the status of the replacement document node is deferred, its latest revision is retrieved and the invalid document node is replaced with a valid document node as the final recall result. (4) Output results steps: return valid documents and query basis links.
2. The multi-dimensional and efficient recall method for complex and time-sensitive documents according to claim 1, characterized in that, The recall model in step (2) is fine-tuned based on the BERT model. The specific process of the recall model is as follows: (2.1) Steps for constructing positive and negative samples: Construct semantic samples and keyword samples; positive samples of semantic samples To select relevant questions from the corpus Relevant document paragraphs, negative samples For random selection and problems Irrelevant document paragraphs; positive samples of keyword samples To start from the problem Core terms are extracted using TF-IDF, and negative samples are analyzed. To start from the problem Randomly selected from non-core words; (2.2) Sample vectorization steps: Before vectorizing positive and negative samples, vectorize the text and special identifier [CLS], and concatenate them separately to represent the overall semantic information of the text; (2.3) Fusion vector step: The keyword vector and semantic vector are fused. The keyword information and semantic information are fused through a length-aware dynamic gating mechanism. Specifically, when the BERT model captures implicit semantics that cannot be represented by length, when the query text is relatively short, it is determined that the query text is more related to the keywords. When the query text is relatively long, it is determined that the query text is more related to the semantics. Thus, a learnable length encoder is set.
3. The multi-dimensional and efficient recall method for complex time-sensitive documents according to claim 2, characterized in that, The concatenation representation in the sample vectorization step (2.2) is as follows: , , , semantic features For the [CLS] position BERT Model encoding, keyword vectors The average value of the encoded vectors at all positions except [CLS]. This represents a sample text to be vectorized. This represents all position indices in the input sequence except for [CLS].
4. The multi-dimensional and efficient recall method for complex time-sensitive documents according to claim 3, characterized in that, In the (2.3) fusion vector step, the query length is mapped to a vector representation with semantic information as follows: , in It depends on the question The length vector constructed from the text length has dimensions of . Dimensionality is reduced to one dimension through a linear layer, and then... The activation function is normalized to obtain the query length gating weight α mentioned above; Based on the query length gating weight α, the semantic vector and keyword vector are fused together, and the final fused vector is as follows: .
5. The multi-dimensional and efficient recall method for complex time-sensitive documents according to claim 2, characterized in that, The user query and document text are input into the retrieval model to generate semantic vectors and keyword vectors. The semantic vectors and keyword vectors are fused through a dynamic gating mechanism to generate a fused vector. Based on the fused vector, cosine similarity is calculated to filter relevant documents.
6. The multi-dimensional and efficient recall method for complex time-sensitive documents according to claim 5, characterized in that, The dynamic gating mechanism includes generating a length vector based on the length of the user query text, and then using a linear layer and... The activation function normalizes the length vector into weight parameters. σ The contribution of keyword vectors and semantic vectors is dynamically adjusted using the weight parameter σ.
7. The multi-dimensional and efficient recall method for complex and time-sensitive documents according to claim 1, characterized in that, The state markers and semantic relationship edges are automatically updated based on document timeliness rules to ensure the timeliness of the recall results.
8. The multi-dimensional and efficient recall method for complex time-sensitive documents according to claim 4, characterized in that, It also includes constructing a vector-level contrastive loss function: , in This indicates the calculation of cosine similarity between two texts. Let be the vector representation of the i-th sample within a training batch. This is a vector representation of the positive samples within a training batch. is the temperature hyperparameter used to adjust the model's attention to difficult negative samples, and B is all the samples included in a training batch.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the system as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the program is executed, it implements the system as described in claims 1 to 8.
Citation Information
Patent Citations
Retrieval enhancement method based on semantic comprehension and semantic generation model
CN118733715A
Visual intelligent digital archive management method and system
CN120508544A
Standard information automatic acquisition interface calling system
CN120611040A
System and method for exploiting semantic annotations in executing keyword queries over a collection of text documents
US20070088734A1
Extracting, deriving, and using legal matter semantics to generate e-discovery queries in an e-discovery system
US20200134757A1