Government affair official document quotation retrieval method and device based on large model, equipment and medium

Through the large-scale model-based citation search method of government official documents, the Langchain and BERT models are used to slice and identify government official documents, and combined with the Elasticsearch and reranker models for semantic matching, the problem of inefficient noun recognition and retrieval in government official documents is solved, and efficient and accurate citation recommendation is achieved.

CN120296146APending Publication Date: 2025-07-11SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510446020.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the compilation and review of existing government official documents, citing relevant documents or regulations mainly relies on manual search, which is inefficient and insufficient in accuracy. Traditional search systems cannot understand the deep semantics and contextual relationships of political terms, and lack intelligent search tools, resulting in inaccurate search results.

Method used

The large-model-based government official document citation search method is used to divide document slices through the Langchain framework, combine Elasticsearch and BGE models to build a document knowledge base, and use pre-trained BERT models to identify professional terms, combine reranker models to perform semantic correlation scoring and reordering, and output target citation documents.

Benefits of technology

It improves the efficiency and accuracy of the search for citations of government official documents, can quickly locate relevant content, significantly improve the accuracy of noun recognition, and is suitable for official document writing and policy research of government departments at all levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296146A_ABST
    Figure CN120296146A_ABST
Patent Text Reader

Abstract

The invention discloses a large model-based government affair official document quotation retrieval method, device and equipment and a medium, and relates to the technical field of artificial intelligence, the method comprises the following steps: obtaining a government affair official document, determining a document slice corresponding to the government affair official document, and creating a document knowledge base and a vector knowledge base based on the document slice; identifying the terminologies in the government affair document by using a target pre-training natural language processing model to obtain a corresponding identification result; based on the document knowledge base and the vector knowledge base, performing keyword matching on the recognition result by using a target information retrieval algorithm to obtain a plurality of corresponding candidate quotation documents; and performing semantic correlation scoring on each candidate citation document and the recognition result through a Rianker model, reordering each candidate citation document according to a corresponding scoring result, and outputting a target citation document according to a corresponding reordering result. Therefore, the quotation retrieval efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, equipment and medium for retrieving citations of government official documents based on large models. Background Art

[0002] Government official documents are important tools for national governance. Their content covers multiple fields such as politics, economy, and law, and usually contains a large number of political terms and strategic terms of great significance. These terms not only need to be accurately understood, but also may require citation of relevant laws, policies, or strategic documents to enhance the authority and reference value of the official documents. Currently, in the process of writing and reviewing government official documents, citing relevant literature or regulations mainly relies on manual search, and this method has the following deficiencies: 1. Low efficiency: Manual search for cited literature takes a lot of time. Especially when there are many and complex terms in the official document, the efficiency drops significantly. Traditional retrieval systems mainly rely on keyword matching and cannot understand the deep semantics and context relationships of political terms, resulting in inaccurate retrieval results and a large amount of irrelevant content. 2. Insufficient accuracy: Manual retrieval may lead to omissions or improper citations due to understanding deviations or knowledge blind spots. Ordinary retrieval systems lack the ability to understand specific terms and expressions in the government affairs field and cannot accurately identify and extract key content such as political terms and strategic expressions. 3. Low level of intelligence: Although a rich collection of policies, regulations, and strategic documents has been accumulated in existing knowledge bases, there is a lack of intelligent retrieval tools, making it difficult to fully utilize their value. Existing systems are mostly passive retrieval systems, lacking functions such as active recommendation and intelligent analysis, and a large amount of screening and judgment work needs to be done manually.

[0003] In recent years, large pre-trained language models (LPLMs) have shown excellent capabilities in the field of natural language processing. These models possess powerful semantic understanding and generation capabilities, providing new possibilities for the automatic analysis and processing of complex texts. However, directly applying these models to solve the problem of retrieving citations of government official documents still faces the following challenges: 1. Accuracy of noun recognition: The political terms and strategic terms involved in government official documents are highly professional and complex, and the model needs to accurately identify these terms. 2. Efficiency of knowledge base matching: The knowledge base contains a vast amount of documents, and how to efficiently retrieve the citation content related to the terms is a key issue. 3. Usability of result presentation: The citation content returned by the system needs to be clear and definite, and can be directly used by official document writers. 4. Keyword-based full-text retrieval system: Using traditional information retrieval technologies such as inverted indexes, the efficiency is low and the accuracy is insufficient.

[0004] Therefore, how to perform efficient retrieval for political terms and strategic terms in the content of government official documents and provide accurate citation retrieval is an urgent problem to be solved currently. Summary of the Invention

[0005] In view of this, an object of the present invention is to provide a method, device, equipment and medium for retrieving citations of government official documents based on a large model, which can efficiently retrieve political nouns and strategic nouns in the content of government official documents and provide accurate citation retrieval, thereby improving the efficiency and quality of writing and reviewing government official documents. The specific scheme is as follows:

[0006] In the first aspect, the present application discloses a method for retrieving citations of government official documents based on a large model, including:

[0007] Obtain a government official document, determine the document slices corresponding to the government official document, and create a document knowledge base and a vector knowledge base based on the document slices;

[0008] Use a target pre-trained natural language processing model to identify professional terms in the government official document to obtain corresponding identification results;

[0009] Based on the document knowledge base and the vector knowledge base, use a target information retrieval algorithm to perform keyword matching on the identification results to obtain a plurality of candidate citation documents;

[0010] Through the reranker model, perform semantic relevance scoring on each candidate citation document and the identification result, re-rank each candidate citation document according to the corresponding scoring result, and output the target citation document according to the corresponding re-ranking result.

[0011] Optionally, the determining the document slices corresponding to the government official document and creating a document knowledge base and a vector knowledge base based on the document slices includes:

[0012] Use the Langchain framework to divide the government official document text to determine the document slices corresponding to the government official document;

[0013] Build an index based on the document slices through Elasticsearch to create the document knowledge base;

[0014] Use a deep learning model based on bidirectional generation coding to vectorize the document slices to obtain the first processed document slices;

[0015] Build the vector knowledge base based on the first processed document slices.

[0016] Optionally, before using the target pre-trained natural language processing model to identify professional terms in the government official document, it further includes:

[0017] Train an initial pre-trained natural language processing model using a target government official document corpus and a special vocabulary list for the field of government official documents to obtain a pre-trained natural language processing model; the special vocabulary list for the field of government official documents includes political nouns, strategic nouns, and policy terms;

[0018] Supplement the nouns in the special vocabulary list for the field of government official documents using an external database to obtain a supplementation result; the external database includes government documents, white papers, and policy documents;

[0019] Train the pre-trained natural language processing model using the supplementation result to obtain the target pre-trained natural language processing model.

[0020] Optionally, performing keyword matching on the recognition result using the target information retrieval algorithm based on the document knowledge base and the vector knowledge base to obtain a corresponding plurality of candidate citation documents, including:

[0021] Perform word segmentation, stop word removal, and keyword extraction operations on the document slices in the document knowledge base to obtain second-processed document slices;

[0022] Construct a global vocabulary based on each of the second-processed document slices; the global vocabulary is a vocabulary that contains unique words or phrases that appear in all second-processed document slices;

[0023] Perform vectorization processing on the second-processed document slices based on the global vocabulary to obtain corresponding sparse vectors, and store the sparse vectors in the document knowledge base;

[0024] Calculate the correlation score between the recognition result and the keywords corresponding to each document slice in the document knowledge base based on the target information retrieval algorithm;

[0025] Screen initial documents from the document knowledge base according to each of the correlation scores;

[0026] Perform keyword matching on the recognition result based on the vector knowledge base and the initial documents to obtain a corresponding plurality of candidate citation documents.

[0027] Optionally, performing keyword matching on the recognition result based on the vector knowledge base and the initial documents to obtain a corresponding plurality of candidate citation documents, including:

[0028] Perform word segmentation, stop word removal, and keyword extraction operations on the recognition result to obtain a processed recognition result;

[0029] Determine the target vector corresponding to the processed recognition result;

[0030] Retrieve the target document slices in the vector knowledge base that are similar to the target vector;

[0031] Determine the similarity score between the target document slice and the target vector;

[0032] Determine the comprehensive score of each initial document based on the similarity score and the relevance score;

[0033] Re-rank the initial documents according to the comprehensive score to determine multiple candidate citation documents.

[0034] Optionally, the method of semantically correlating each candidate citation document with the recognition result through a reranker model, re-ranking each candidate citation document according to the corresponding scoring result, and outputting the target citation document according to the corresponding re-ranking result includes:

[0035] Semantically correlate each candidate citation document with the recognition result through a reranker model;

[0036] Filter the candidate citation documents corresponding to the target scoring results that do not meet the preset scoring threshold conditions according to the scoring results to determine each initial citation document;

[0037] Re-rank the initial citation documents according to the scoring results, the release time, and the release agency of each initial citation document, and output the target citation document according to the corresponding re-ranking result.

[0038] Optionally, the method further includes:

[0039] When outputting the target citation document, output the context information of the target citation document; the context information includes the context information of the cited paragraph and any combination of the title, release agency, and release time of the target citation document.

[0040] In a second aspect, the present application discloses a government official document citation retrieval device based on a large model, including:

[0041] A knowledge base construction module, configured to obtain government official documents, determine the document slices corresponding to the government official documents, and create a document knowledge base and a vector knowledge base based on the document slices;

[0042] An identification module, configured to identify professional terms in the government official documents by using a target pre-trained natural language processing model to obtain corresponding identification results;

[0043] A keyword matching module, configured to perform keyword matching on the recognition result by using a target information retrieval algorithm based on the document knowledge base and the vector knowledge base, so as to obtain a plurality of corresponding candidate citation documents;

[0044] A citation document output module, configured to perform semantic relevance scoring on each of the candidate citation documents and the recognition result through a reranker model, reorder each of the candidate citation documents according to the corresponding scoring result, and output a target citation document according to the corresponding reordering result.

[0045] In a third aspect, the present application discloses an electronic device, including:

[0046] A memory, configured to store a computer program;

[0047] A processor, configured to execute the computer program to implement the government official document citation retrieval method based on a large model as described above.

[0048] In a fourth aspect, the present application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the government official document citation retrieval method based on a large model as described above is implemented.

[0049] The present application first obtains a government official document, determines a document slice corresponding to the government official document, and creates a document knowledge base and a vector knowledge base based on the document slice; uses a target pre-trained natural language processing model to identify professional terms in the government official document to obtain a corresponding recognition result; then performs keyword matching on the recognition result by using a target information retrieval algorithm based on the document knowledge base and the vector knowledge base to obtain a plurality of corresponding candidate citation documents; finally, performs semantic relevance scoring on each of the candidate citation documents and the recognition result through a reranker model, reorders each of the candidate citation documents according to the corresponding scoring result, and outputs a target citation document according to the corresponding reordering result. It can be seen that the present application uses a natural language processing model to accurately identify proper nouns in the government affairs field, combines semantic matching technology and an efficient indexing mechanism to retrieve documents and paragraphs related to the identified nouns from the knowledge base; finally, filters and sorts the retrieval results to recommend the most relevant documents and paragraphs. In this way, the efficiency and accuracy of citation retrieval are improved, and relevant content can be quickly located. At the same time, by using the semantic understanding ability of the large model, the accuracy of noun recognition is significantly improved. It can be widely applied to work scenarios such as official document writing, policy research, and document verification in government departments at all levels, providing technical support for improving the efficiency and accuracy of government affairs work. Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on the provided drawings.

[0051] Figure 1 Flowchart of a government official document citation retrieval method based on a large model disclosed in the present application;

[0052] Figure 2 Schematic diagram of a specific government official document citation retrieval method based on a large model disclosed in the present application;

[0053] Figure 3 Flowchart of a citation screening method disclosed in the present application;

[0054] Figure 4 Schematic diagram of the structure of a government official document citation retrieval device based on a large model disclosed in the present application;

[0055] Figure 5 Structural diagram of an electronic device disclosed in the present application. Detailed implementation manners

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0057] Currently, in the process of writing and reviewing government official documents, the citation of relevant literature or regulations mainly relies on manual search. This method has the following deficiencies: 1. Low efficiency: Manual search for cited literature takes a lot of time. Especially when there are many complex nouns in the official document, the efficiency drops significantly. Traditional retrieval systems mainly rely on keyword matching and cannot understand the deep semantics and context relationships of political terms, resulting in inaccurate retrieval results and a large amount of irrelevant content. 2. Insufficient accuracy: Manual retrieval may lead to omissions or improper citations due to misunderstandings or knowledge blind spots. Ordinary retrieval systems lack the ability to understand specific terms and expressions in the government affairs field and cannot accurately identify and extract key content such as political nouns and strategic expressions. 3. Low level of intelligence: Although rich policies, regulations, and strategic documents have been accumulated in the existing knowledge bases, there is a lack of intelligent retrieval tools, making it difficult to fully utilize their value. Existing systems are mostly passive retrieval and lack functions such as active recommendation and intelligent analysis, requiring a large amount of manual screening and judgment work. To solve the above technical problems, this application discloses a method for retrieving citations of government official documents based on a large model, which can efficiently retrieve political nouns and strategic nouns in the content of government official documents and provide accurate citation retrieval, thereby improving the efficiency and quality of writing and reviewing government official documents.

[0058] See Figure 1 As shown, an embodiment of the present invention discloses a method for retrieving citations of government official documents based on a large model, including:

[0059] Step S11, obtain a government official document, determine the document slice corresponding to the government official document, and create a document knowledge base and a vector knowledge base based on the document slice.

[0060] In this embodiment, first, accept the draft government official documents uploaded by users, supporting both text format and document format. Use the Langchain framework to divide the text of government official documents, cutting the official documents into multiple small pieces (i.e., document slices). This process ensures that each part of the official document can be processed independently while improving the efficiency of subsequent retrieval and matching. Build an index for the sliced document content through Elasticsearch (ES) to create a document knowledge base. This document knowledge base facilitates the quick retrieval of relevant information in the official document and provides efficient data storage and query support for subsequent citation retrieval. Vectorize the document slices using the BGE (Bidirectional Generative Encoder) model. The BGE model can effectively capture the semantic information of the document content and generate high-dimensional vector representations, which can be better used for semantic matching and retrieval. Build a vector index library for the document slices through Faiss (Facebook AI Similarity Search) to form a vector knowledge base. The Faiss library can support efficient approximate nearest neighbor queries, improving the speed and accuracy of large-scale document retrieval. That is, use the Langchain framework to divide the text of government official documents to determine the document slices corresponding to the government official documents; build an index based on the document slices through Elasticsearch to create the document knowledge base; use a deep learning model based on bidirectional generative encoding to vectorize the document slices to obtain the first processed document slices; build the vector knowledge base based on the first processed document slices.

[0061] Specifically, as Figure 2As shown, the Langchain framework is used to divide government official document texts, cutting the official documents into multiple small pieces (i.e., document slices). This process ensures that each part of the official document can be processed independently while improving the efficiency of subsequent retrieval and matching. The sliced document content is indexed through Elasticsearch (ES) to create a document knowledge base. The role of the document knowledge base is to provide fast retrieval support based on the tokenization dimension for the retrieval module. The full-text retrieval ability of Elasticsearch can quickly locate the document slices relevant to the user's query, providing a basis for subsequent exact matching. The BGE model is used to vectorize the document slices. The BGE model can effectively capture the semantic information of the document content and generate high-dimensional vector representations. These vectors can be better used for semantic matching and retrieval, thus improving the accuracy of retrieval. A vector index library of the document slices is constructed through Faiss to form a vector knowledge base. The role of the vector knowledge base is to provide retrieval support based on semantic vectors for the retrieval module. The Faiss library can support efficient approximate nearest neighbor queries, improving the speed and accuracy of large-scale document retrieval. Among them, the document knowledge base: Based on the tokenization dimension, use the full-text retrieval ability of Elasticsearch to quickly locate the document slices relevant to the user's query. This process can quickly screen out the potentially relevant document fragments and provide a candidate set for subsequent exact matching. The vector knowledge base: Based on semantic vectors, use Faiss for efficient approximate nearest neighbor (ANN, Approximate Nearest Neighbor) queries. This process can evaluate the similarity between the document slices and the user's query at the semantic level, thus providing more accurate matching results. At the same time, based on the LangChain framework, document slices are implemented, a sparse index library is constructed through Elasticsearch (ES), and the BGE model is used to vectorize the documents and construct a FAISS vector index library. Through the dual-index mechanism, the system can dynamically manage and expand the knowledge base to meet retrieval requirements of different scales and complexities, providing a flexible and efficient solution for government document management.

[0062] Generally speaking, this application is based on a deep learning model (such as the BGE model) to perform high-dimensional semantic representations of documents and queries, and can capture deep semantic associations. First, the BGE model is used to generate high-dimensional semantic vectors for the documents and query statements in the knowledge base respectively. These vectors represent the overall semantic features of the documents, not limited to the information at the keyword level. Then, the cosine similarity or Euclidean distance is used to calculate the similarity between the query vector and the document vector to measure the semantic relevance between the two. In addition, the system adopts a vector index library based on FAISS, and accelerates the vector retrieval process through clustering sharding and an efficient approximate nearest neighbor (ANN) algorithm. By dividing the vector space into multiple subspaces and using quantization techniques in each subspace, the vectors are encoded into short codes. During retrieval, the nearest neighbor is approximately calculated by looking up these short codes, which significantly reduces the computational amount, accelerates the retrieval process, and ensures the real-time performance and accuracy of large-scale knowledge base retrieval.

[0063] Step S12: Use the target pre-trained natural language processing model to identify the professional terms in the government official document to obtain the corresponding identification results.

[0064] In this embodiment, before using the target pre-trained natural language processing model to identify the professional terms in the government official document, the initial pre-trained natural language processing model is trained using the target government official document corpus and the special vocabulary list in the field of government official documents to obtain the pre-trained natural language processing model; the special vocabulary list in the field of government official documents includes political nouns, strategic nouns, and policy terms; an external database is used to supplement the nouns in the special vocabulary list in the field of government official documents to obtain a supplement result; the external database includes government documents, white papers, and policy documents; the pre-trained natural language processing model is trained using the supplement result to obtain the target pre-trained natural language processing model. That is to say, this application uses the BERT (Bidirectional Encoder Representations from Transformers) model for semantic parsing, and specifically identifies the political nouns, strategic nouns, and related terms involved in government official documents. It should be noted that this application uses a large-scale pre-trained BERT model as the basic model, and fine-tunes it on the government official document corpus in a specific field to make it better understand professional terms such as political terms and strategic nouns related to official documents. The fine-tuning process adjusts the model parameters based on the labeled data (such as the labeled documents with tags such as political nouns and strategic terms) to adapt to the special context of government official documents. At the same time, a special vocabulary list in the field of government official documents is introduced, and the vocabulary list includes political nouns, strategic nouns, policy terms, etc. This vocabulary list not only covers common terms, but also includes some relatively professional and rare political and economic terms. The introduction of the vocabulary list can effectively improve the accuracy of noun recognition and help the BERT model identify some special terms that are difficult for the standard model to capture.

[0065] To enhance the model's recognition ability for newly emerging nouns or terms, a dynamic expansion mechanism is designed. This mechanism can automatically expand the terms in the vocabulary list by analyzing the newly uploaded official documents or the data feedback by users in real time. The expansion process is based on the matching of the new vocabulary in the data with the context, ensuring that new political or strategic nouns can be identified in a timely manner. Specifically: use an external knowledge base (such as government documents, white papers, policy documents, etc.) to supplement the nouns. Based on the prediction results of the BERT model, the unrecognized nouns are labeled and fed back to the training set for retraining to enhance the model's recognition ability.

[0066] In addition to the recognition of political and strategic nouns, the technology of Named Entity Recognition (NER) is also combined to identify relevant personal names, place names, organization names, etc. from official documents. Through the multi-level feature extraction ability of BERT, this process improves the model's understanding ability of complex official document content. Based on the bidirectional semantic understanding of the BERT model, noun term recognition is not limited to word-level matching, but also includes context understanding. By analyzing semantic information at the sentence level, paragraph level, and even full-text level, the system ensures that the recognition of noun terms is not only accurate but also in line with the overall context of the official document.

[0067] Finally, the target pre-trained natural language processing model can be used to identify the professional terms in the official document to obtain corresponding recognition results. That is to say, in order to further improve the system performance, the comprehensive optimization mechanism has been improved in many aspects. First, through the dynamic query expansion technology, the query content of the user is expanded based on semantic embedding and the domain knowledge graph to improve the retrieval coverage rate. Secondly, an efficient knowledge base incremental update strategy is designed to ensure that newly added or modified documents can be included in the retrieval scope in a timely manner. In addition, to cope with large-scale data, index compression, caching mechanism, and distributed retrieval architecture are adopted, effectively reducing the response time of the system. The application of the comprehensive optimization mechanism enables the system to have a high degree of dynamic adaptability and large-scale retrieval ability.

[0068] Step S13: Based on the document knowledge base and the vector knowledge base, use the target information retrieval algorithm to perform keyword matching on the recognition results to obtain a corresponding number of candidate citation documents.

[0069] In this embodiment, when determining candidate citation documents, word segmentation, stop word removal, and keyword extraction operations are performed on the document slices in the document knowledge base to obtain second-processed document slices; a global vocabulary is constructed based on each of the second-processed document slices; the global vocabulary is a vocabulary containing unique words or phrases that appear in all the second-processed document slices; the second-processed document slices are vectorized based on the global vocabulary to obtain corresponding sparse vectors, and the sparse vectors are stored in the document knowledge base; the correlation scores between the recognition result and the keywords corresponding to each document slice in the document knowledge base are calculated based on the target information retrieval algorithm; initial documents are screened from the document knowledge base according to each of the correlation scores; word segmentation, stop word removal, and keyword extraction operations are performed on the recognition result to obtain a processed recognition result; a target vector corresponding to the processed recognition result is determined; target document slices similar to the target vector are retrieved from the vector knowledge base; the similarity score between the target document slice and the target vector is determined; the comprehensive scores of each of the initial documents are determined based on the similarity score and the correlation score; the initial documents are re-ranked according to the comprehensive scores to determine multiple candidate citation documents. That is to say, first, the documents in the knowledge base are preprocessed in this application, including word segmentation, stop word removal, and keyword extraction, to construct a bag-of-words model for each document. Subsequently, the BM25 (Best Matching 25, a probabilistic ranking model) algorithm evaluates the correlation between the query statement and the document by calculating the matching degree between the keywords in the query statement and the document keywords. This process comprehensively considers the term frequency (TF), inverse document frequency (IDF), and the distribution of words in the document, and can efficiently capture the explicit semantic association between the query and the document. In this way, the candidate documents are re-ranked according to the comprehensive scores to generate the final retrieval result. This process ensures that the retrieval result takes into account both the explicit semantic association and the semantic similarity, thus significantly improving the retrieval accuracy and user experience.

[0070] Step S14: Perform semantic correlation scoring on each of the candidate citation documents and the recognition result through a reranker model, re-rank each of the candidate citation documents according to the corresponding scoring result, and output a target citation document according to the corresponding re-ranking result.

[0071] In this embodiment, a reranker model is used to perform semantic relevance scoring on each of the candidate citation documents and the recognition result; candidate citation documents corresponding to target scoring results that do not meet the preset scoring threshold conditions are filtered according to the scoring result to determine each initial citation document; the initial citation documents are re-ranked according to the scoring result, the release time, and the release institution of each initial citation document, and target citation documents are output according to the corresponding re-ranking result. These operations are mainly the functions of the citation recommendation module. The citation recommendation module is an important part of the system for generating high-quality recommendation results, and its core functions include relevance scoring of retrieval results and final re-ranking to ensure that users obtain the optimal citation content. This module combines the noun recognition result and deep learning technology to achieve accurate citation recommendation through a multi-stage processing flow. The specific process is as Figure 3 shown. First, the system extracts a retrieval result set related to these nouns based on the political noun and strategic noun results output by the noun recognition module. These retrieval results are initially screened from the sparse retrieval and vector retrieval modules and contain documents and paragraphs semantically related to the nouns. To further improve the accuracy of the recommendation results, the system introduces a reranker model as a secondary screening tool. In the re-scoring stage, the system deeply analyzes the candidate documents and paragraphs. Specifically, the reranker model is based on pre-trained language models such as Bert and combines context semantics to perform fine-grained scoring on the relevance of the retrieval results. This model not only focuses on the direct match between the query and the document but also considers the deep semantic association between the document context and the query intention. By inputting the query content and the paragraph text of the candidate document, the reranker model generates a relevance score for each document and filters out low-score documents that do not meet the requirements. Subsequently, the system re-ranks the candidate results according to the relevance score of the reranker model. During the re-ranking process, the system not only considers the semantic matching degree but also comprehensively evaluates the authority and timeliness of the document. For example, for government documents, the system preferentially recommends documents issued by authoritative institutions and at the same time provides more timely citation sources for users based on the release time. To achieve the fairness and efficiency of sorting, the system adopts an algorithm based on weighted sorting and assigns a final sorting position to each document in combination with the scoring index.

[0072] Finally, when outputting the target citation document, output the context information of the target citation document; the context information includes the context information of the cited paragraph and any one or several combinations of the title, publishing institution, and publication time of the target citation document. Specifically, when outputting the recommendation results, the system provides detailed context information for each cited content. The context information includes the context before and after the cited paragraph and the metadata of the document (such as the title, publishing institution, and publication time) so that users can comprehensively understand the citation source. In addition, the system supports users to further screen and customize the recommended content, such as refining and sorting the results according to time, source, or keywords. In this way, the citation recommendation module ensures the high relevance and high authority of the recommended content and provides reliable citation support for users. Classification models such as Bert and reranker technology are used to re-score the retrieval results, and re-rank them in combination with semantic matching degree and document authority. By providing the context information and detailed metadata of the cited content, the system significantly improves the authority and practicability of the recommended content and helps users efficiently complete citation screening and citation.

[0073] This application first obtains government official documents, determines the document slices corresponding to the government official documents, and creates a document knowledge base and a vector knowledge base based on the document slices; uses a target pre-trained natural language processing model to identify professional terms in the government official documents to obtain corresponding identification results; then based on the document knowledge base and the vector knowledge base, uses a target information retrieval algorithm to perform keyword matching on the identification results to obtain a corresponding plurality of candidate citation documents; finally, uses a reranker model to perform semantic relevance scoring on each candidate citation document and the identification result, re-ranks each candidate citation document according to the corresponding scoring result, and outputs the target citation document according to the corresponding re-ranking result. It can be seen that this application uses a natural language processing model to accurately identify proper nouns in the government affairs field, combines semantic matching technology and an efficient indexing mechanism to retrieve documents and paragraphs related to the identified nouns from the knowledge base; finally screens and sorts the retrieval results to recommend the most relevant documents and paragraphs. In this way, the efficiency and accuracy of citation retrieval are improved, and relevant content can be quickly located. At the same time, using the semantic understanding ability of the large model significantly improves the accuracy of noun identification. It can be widely applied to work scenarios such as official document writing, policy research, and document checking in government departments at all levels, providing technical support for improving the efficiency and accuracy of government affairs work.

[0074] Based on the previous embodiment, it can be known that this application can use a target information retrieval algorithm to perform keyword matching on the identification results based on the document knowledge base and the vector knowledge base to obtain a corresponding plurality of candidate citation documents. Next, a detailed description will be given of the specific candidate citation document screening process.

[0075] This application uses the traditional bag-of-words model and BM25 algorithm to achieve preliminary screening of documents. First, the documents in the knowledge base are preprocessed, including word segmentation, stop word removal, and keyword extraction, to construct the bag-of-words model for each document. Subsequently, the BM25 algorithm evaluates the relevance between the query statement and the document by calculating the matching degree between the keywords in the query and the document keywords. This process comprehensively considers the term frequency (TF), inverse document frequency (IDF), and the distribution of words in the document, and can efficiently capture the explicit semantic associations between the query and the document. By screening the documents with higher scores, the sparse retrieval module provides a fast and efficient candidate document set for subsequent processing. Specifically, the sparse retrieval module uses the traditional bag-of-words model and BM25 algorithm to achieve preliminary screening of documents. The detailed steps are as follows:

[0076] 1.1. Document preprocessing:

[0077] Word segmentation: Each document in the knowledge base is segmented into words or phrases. This step is implemented based on the language's word segmentation rules. For example, jieba is used for Chinese word segmentation, and tools such as NLTK are used for English word segmentation.

[0078] Stop word removal: Common stop words (such as "de", "shi", "he", etc.) in the document are removed. These words usually do not carry important information semantically, and removing them can reduce noise and improve retrieval efficiency.

[0079] Keyword extraction: Keywords are extracted from the document, and these keywords will be used for subsequent bag-of-words model construction. Keyword extraction can be achieved through algorithms such as TF-IDF and TextRank to ensure that the extracted keywords have high information content and representativeness.

[0080] 1.2. Construction of the bag-of-words model:

[0081] Vocabulary construction: Based on the preprocessed document set, a global vocabulary is constructed. The vocabulary contains all the unique words or phrases that appear in the documents.

[0082] Document vectorization: Each document is represented as a vector, where each dimension of the vector corresponds to a word or phrase in the vocabulary, and the value of the vector is the term frequency (TF) or TF-IDF value of the word or phrase in the document. In this way, each document is converted into a sparse vector and stored in the document knowledge base.

[0083] 1.3. BM25 algorithm:

[0084] Query preprocessing: The user input query statement is preprocessed in the same way as the document, including word segmentation, stop word removal, etc.

[0085] Keyword Matching: Calculate the matching degree between the keywords in the query statement and the keywords in the document. The BM25 algorithm evaluates the relevance between the query and the document by comprehensively considering the term frequency (TF), inverse document frequency (IDF), and the distribution of words in the document.

[0086] Relevance Scoring: The BM25 algorithm generates a relevance score for each document. The higher the score, the higher the relevance of the document to the query. The specific formula is as follows:

[0087] ;

[0088] where, is the relevance score; n is the total number of terms in the query q; q is the query, d is the document, is the keyword in the query, is the keyword the term frequency of the keyword in document d, is the inverse document frequency of the keyword and b are adjustment parameters, is the length of document d, and avgdl is the average length of the document collection.

[0089] 1.4. Preliminary Screening:

[0090] Candidate Document Set: According to the BM25 scores, select the documents with higher scores to form a preliminary candidate document set. These documents are considered to have a high explicit semantic association with the user's query and provide a basis for further dense vector retrieval.

[0091] In addition, the dense vector retrieval module uses a vector knowledge base and an efficient approximate nearest neighbor algorithm to achieve accurate semantic-based retrieval. The detailed steps are as follows:

[0092] 2.1. Vector Knowledge Base Construction:

[0093] Document Slice Vectorization: Use the BGE (Bidirectional Generative Encoder) model to vectorize each document slice in the document knowledge base. The BGE model can capture the semantic information of the document content and generate a high-dimensional vector representation.

[0094] Vector Index Construction: Build a vector index library for the document slices through the Faiss library to form a vector knowledge base. The Faiss library supports efficient approximate nearest neighbor (ANN) queries and can quickly locate the document slices most similar to the query vector.

[0095] 2.2. Query Vectorization:

[0096] Query preprocessing: Perform the same preprocessing on the query statement entered by the user as on the documents, including word segmentation, stop word removal, etc.

[0097] Query vectorization: Use the BGE model to convert the query statement into a high-dimensional vector representation. This vector will be used for similarity retrieval in the vector knowledge base.

[0098] 2.3. Approximate nearest neighbor retrieval:

[0099] Vector retrieval: Input the query vector into the vector knowledge base and use the approximate nearest neighbor (ANN) algorithm of the Faiss library for retrieval. Faiss quickly locates the document slices most similar to the query vector through clustering and efficient indexing structures.

[0100] Similarity scoring: Generate a similarity score for each retrieved document slice. The higher the score, the higher the semantic similarity between the document slice and the query vector.

[0101] 2.4. Result fusion:

[0102] Comprehensive scoring: Combine the results of the sparse retrieval module and the dense vector retrieval module. For each candidate document, generate a comprehensive score by combining the BM25 score and the similarity score.

[0103] Final ranking: Re-rank the candidate documents according to the comprehensive score to generate the final retrieval results. This process ensures that the retrieval results consider both explicit semantic associations (sparse retrieval) and semantic similarity (dense vector retrieval).

[0104] The hybrid retrieval module realizes efficient and accurate knowledge retrieval by combining the advantages of sparse retrieval and dense vector retrieval. The specific synergy is as follows:

[0105] 3.1. Sparse retrieval module:

[0106] Initial screening: Quickly screen out a set of documents with high explicit semantic association with the query statement through the BM25 algorithm, reducing the number of documents to be processed.

[0107] Candidate document set: Provide an efficient candidate document set for the dense vector retrieval module, improving the retrieval efficiency.

[0108] 3.2. Dense vector retrieval module:

[0109] Semantic matching: Perform precise semantic matching on the candidate document set through the BGE model and the Faiss library, further improving the accuracy of the retrieval results.

[0110] Similarity Scoring: Generate similarity scores for each candidate document to ensure that the retrieval results are not only explicitly semantically relevant but also highly semantically matched.

[0111] 3.3. Result Fusion:

[0112] Comprehensive Scoring: Integrate the BM25 scores of the sparse retrieval module and the similarity scores of the dense vector retrieval module to generate a comprehensive score.

[0113] Final Ranking: Re-rank the candidate documents according to the comprehensive score to generate the final retrieval results. This process ensures that the retrieval results consider both explicit semantic associations and semantic similarity, thus significantly improving the retrieval accuracy and user experience.

[0114] Through this hybrid retrieval method, the system can make full use of the efficiency of sparse retrieval and the accuracy of dense vector retrieval, significantly improving the overall performance of document retrieval and user experience. The multi-way retrieval combination strategy comprehensively utilizes the advantages of sparse retrieval and vector retrieval to achieve a balance between retrieval efficiency and precision. First, the sparse retrieval module quickly screens out a candidate set of documents that are explicitly matched with the query content, providing a basis for subsequent in-depth retrieval. Then, the vector retrieval module performs semantic matching on these candidate documents and can be extended to the global document library to discover potentially semantically related documents. Finally, the system combines the BM25 scores of sparse retrieval with the semantic similarity scores of vector retrieval and uses linear weighting or learning-based ranking methods to fuse and rank the retrieval results to generate the final citation recommendation list. This multi-way strategy not only improves the recall rate of retrieval but also ensures the high precision of the results. Innovatively, a multi-way retrieval strategy combining sparse retrieval and vector retrieval is adopted. The sparse retrieval module uses the BM25 algorithm to quickly locate highly relevant documents, and the vector retrieval module constructs a vector index library through the BGE model to achieve deep matching at the semantic level. The comprehensive optimization mechanism of multi-way retrieval ensures that in a large-scale knowledge base, it is possible to quickly find candidate documents related to the query and capture potential relevant information through deep semantic matching, greatly improving the comprehensiveness and accuracy of retrieval.

[0115] In this way, the efficiency and quality of government official document processing and citation retrieval are significantly improved, providing an innovative solution for government informatization and intelligent office work, with great social value and application prospects.

[0116] See Figure 4 As shown, an official document citation retrieval device based on a large model according to an embodiment of the present invention includes:

[0117] A knowledge base construction module 11, configured to obtain official government documents, determine document slices corresponding to the official government documents, and create a document knowledge base and a vector knowledge base based on the document slices;

[0118] An identification module 12, configured to identify professional terms in the government official document by using a target pre-trained natural language processing model to obtain corresponding identification results;

[0119] A keyword matching module 13, configured to perform keyword matching on the identification results by using a target information retrieval algorithm based on the document knowledge base and the vector knowledge base to obtain a plurality of corresponding candidate citation documents;

[0120] A citation document output module 14, configured to perform semantic relevance scoring on each candidate citation document and the identification results through a reranker model, re-rank each candidate citation document according to the corresponding scoring results, and output a target citation document according to the corresponding re-ranking results.

[0121] This application first obtains a government official document, determines a document slice corresponding to the government official document, and creates a document knowledge base and a vector knowledge base based on the document slice; identifies professional terms in the government official document by using a target pre-trained natural language processing model to obtain corresponding identification results; then performs keyword matching on the identification results by using a target information retrieval algorithm based on the document knowledge base and the vector knowledge base to obtain a plurality of corresponding candidate citation documents; finally, performs semantic relevance scoring on each candidate citation document and the identification results through a reranker model, re-rank each candidate citation document according to the corresponding scoring results, and output a target citation document according to the corresponding re-ranking results. It can be seen that this application uses a natural language processing model to accurately identify proper nouns in the government affairs field, combines semantic matching technology and an efficient indexing mechanism to retrieve documents and paragraphs related to the identified nouns from the knowledge base; finally, filters and sorts the retrieval results to recommend the most relevant documents and paragraphs. In this way, the efficiency and accuracy of citation retrieval are improved, and relevant content can be quickly located. At the same time, by using the semantic understanding ability of the large model, the accuracy of noun identification is significantly improved. It can be widely applied to work scenarios such as official document writing, policy research, and document verification in government departments at all levels, providing technical support for improving the efficiency and accuracy of government affairs work.

[0122] In some specific embodiments, the knowledge base construction module 11 may specifically be configured to divide the government official document text by using the Langchain framework to determine a document slice corresponding to the government official document; construct an index based on the document slice through Elasticsearch to create the document knowledge base; perform vectorization processing on the document slice by using a deep learning model based on bidirectional generation coding to obtain a first processed document slice; and construct the vector knowledge base based on the first processed document slice.

[0123] In some specific embodiments, the device can also be used to train an initial pre-trained natural language processing model by using a target government official document corpus and a special vocabulary list for the field of government official documents to obtain a pre-trained natural language processing model; the special vocabulary list for the field of government official documents includes political nouns, strategic nouns, and policy terms; an external database is used to supplement the nouns in the special vocabulary list for the field of government official documents to obtain a supplement result; the external database includes government documents, white papers, and policy documents; the supplement result is used to train the pre-trained natural language processing model to obtain the target pre-trained natural language processing model.

[0124] In some specific embodiments, the keyword matching module 13 can specifically be used to perform word segmentation, stop word removal, and keyword extraction operations on the document slices in the document knowledge base to obtain second-processed document slices; construct a global vocabulary based on each of the second-processed document slices; the global vocabulary is a vocabulary that contains unique words or phrases that appear in all the second-processed document slices; perform vectorization processing on the second-processed document slices based on the global vocabulary to obtain corresponding sparse vectors, and store the sparse vectors in the document knowledge base; calculate the correlation scores between the recognition result and the keywords corresponding to each document slice in the document knowledge base based on a target information retrieval algorithm; screen initial documents from the document knowledge base according to each of the correlation scores; perform keyword matching on the recognition result based on the vector knowledge base and the initial documents to obtain corresponding multiple candidate citation documents.

[0125] In some specific embodiments, the keyword matching module 13 can specifically be used to perform word segmentation, stop word removal, and keyword extraction operations on the recognition result to obtain a processed recognition result; determine a target vector corresponding to the processed recognition result; retrieve a target document slice similar to the target vector in the vector knowledge base; determine the similarity score between the target document slice and the target vector; determine the comprehensive score of each of the initial documents based on the similarity score and the correlation score; re-rank the initial documents according to the comprehensive score to determine multiple candidate citation documents.

[0126] In some specific embodiments, the citation document output module 14 can specifically be used to perform semantic correlation scoring on each of the candidate citation documents and the recognition result through a reranker model; filter the candidate citation documents corresponding to the target scoring results that do not meet the preset scoring threshold conditions according to the scoring results to determine each initial citation document; re-rank the initial citation documents according to the scoring results, the release time, and the release agency of each initial citation document, and output the target citation document according to the corresponding re-ranking result.

[0127] In some specific embodiments, the apparatus can also be used to output context information of the target citation document when outputting the target citation document; the context information includes context information of the cited paragraph and any one or several combinations of the title, publishing agency, and publishing time of the target citation document.

[0128] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 5 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation on the scope of use of the present application.

[0129] Figure 5 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the method for retrieving government official document citations based on a large model disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0130] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0131] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be short-term storage or permanent storage.

[0132] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the method for retrieving government official document citations based on a large model executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.

[0133] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned government official document citation retrieval method based on a large model. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0134] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for relevant details.

[0135] Those skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0136] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0137] Finally, it should also be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0138] The above has introduced the technical solution provided by the present application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for retrieving citations of government official documents based on large models, characterized in that, Including: Obtain government official documents, determine the document slices corresponding to the government official documents, and create a document knowledge base and a vector knowledge base based on the document slices; Use a target pre-trained natural language processing model to identify professional terms in the government official documents to obtain corresponding identification results; Based on the document knowledge base and the vector knowledge base, use a target information retrieval algorithm to perform keyword matching on the identification results to obtain a corresponding plurality of candidate citation documents; Through a reranker model, perform semantic relevance scoring on each candidate citation document and the identification result, re-rank each candidate citation document according to the corresponding scoring result, and output a target citation document according to the corresponding re-ranking result.

2. The method for retrieving citations of government official documents based on large models according to claim 1, wherein, The determining the document slices corresponding to the government official documents and creating a document knowledge base and a vector knowledge base based on the document slices includes: Use the Langchain framework to divide the text of government official documents to determine the document slices corresponding to the government official documents; Build an index based on the document slices through Elasticsearch to create the document knowledge base; Use a deep learning model based on bidirectional generation encoding to perform vectorization processing on the document slices to obtain a first processed document slice; Build the vector knowledge base based on the first processed document slice.

3. The method for retrieving citations of government official documents based on large models according to claim 1, wherein Before using the target pre-trained natural language processing model to identify professional terms in the government official documents, it further includes: Use a target government official document corpus and a special vocabulary list for the government official document field to train an initial pre-trained natural language processing model to obtain a pre-trained natural language processing model; the special vocabulary list for the government official document field includes political nouns, strategic nouns, and policy terms; Use an external database to supplement the nouns in the special vocabulary list for the government official document field to obtain a supplement result; the external database includes government documents, white papers, and policy documents; Use the supplement result to train the pre-trained natural language processing model to obtain the target pre-trained natural language processing model.

4. The method for retrieving citations of government official documents based on large models according to claim 1, characterized in that, The performing keyword matching on the identification results based on the document knowledge base and the vector knowledge base using a target information retrieval algorithm to obtain a corresponding plurality of candidate citation documents includes: Perform word segmentation, stop word removal, and keyword extraction operations on the document slices in the document knowledge base to obtain a second processed document slice; Build a global vocabulary based on each of the second processed document slices; the global vocabulary is a vocabulary that contains unique words or phrases that appear in all the second processed document slices; Perform vectorization processing on the second processed document slices based on the global vocabulary to obtain corresponding sparse vectors, and store the sparse vectors in the document knowledge base; Calculate the relevance score between the identification result and the keywords corresponding to each document slice in the document knowledge base based on the target information retrieval algorithm; Screen initial documents from the document knowledge base according to each of the relevance scores; Perform keyword matching on the recognition result based on the vector knowledge base and the initial document to obtain a corresponding plurality of candidate citation documents.

5. The method for retrieving citations of government official documents based on a large model according to claim 4, wherein The performing keyword matching on the recognition result based on the vector knowledge base and the initial document to obtain a corresponding plurality of candidate citation documents includes: Perform word segmentation, stop word removal, and keyword extraction operations on the recognition result to obtain a processed recognition result; Determine the target vector corresponding to the processed recognition result; Retrieve the target document slices similar to the target vector in the vector knowledge base; Determine the similarity score between the target document slice and the target vector; Determine the comprehensive score of each initial document based on the similarity score and the relevance score; Re-rank the initial documents according to the comprehensive score to determine a plurality of candidate citation documents.

6. The method for retrieving citations of government official documents based on a large model according to claim 1, wherein, The re-ranking each candidate citation document with the recognition result through a reranker model, re-ranking each candidate citation document according to the corresponding scoring result, and outputting a target citation document according to the corresponding re-ranking result includes: Perform semantic relevance scoring on each candidate citation document and the recognition result through a reranker model; Filter the candidate citation documents corresponding to the target scoring results that do not meet the preset scoring threshold conditions according to the scoring result to determine each initial citation document; Re-rank the initial citation documents according to the scoring result, the publication time, and the publication agency of each initial citation document, and output a target citation document according to the corresponding re-ranking result.

7. The method for retrieving citations of government official documents based on large models according to any one of claims 1 to 6, characterized in that, It further includes: When outputting the target citation document, output the context information of the target citation document; The context information includes the context information of the cited paragraph and any one or several combinations of the title, publication agency, and publication time of the target citation document.

8. An official document citation retrieval device based on a large model, characterized in that, It includes: A knowledge base construction module, configured to obtain government official documents, determine the document slices corresponding to the government official documents, and create a document knowledge base and a vector knowledge base based on the document slices; An identification module, configured to use a target pre-trained natural language processing model to identify professional terms in the government official documents to obtain corresponding identification results; A keyword matching module, configured to perform keyword matching on the recognition result based on the document knowledge base and the vector knowledge base by using a target information retrieval algorithm to obtain a corresponding plurality of candidate citation documents; A citation document output module, configured to perform semantic relevance scoring on each candidate citation document and the recognition result through a reranker model, re-rank each candidate citation document according to the corresponding scoring result, and output a target citation document according to the corresponding re-ranking result.

9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the government official document citation retrieval method based on a large model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on a computer-readable storage medium. When the computer program is executed by a processor, it implements the method for retrieving citations of government official documents based on a large model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-field and multi-disciplinary science and technology policy resource retrieval method and device

    CN115344668A

  • Reordering method and device for improving retrieval performance of AI large language model

    CN117725183A

  • Government affair intelligent response device and method based on intention recognition and large language model

    CN118035419A

  • Interaction method and device, equipment and storage medium

    CN118246540A

  • Information retrieval method and device, electronic equipment and storage medium

    CN118394916A

Cited By

  • Industrial document intelligent retrieval method and system based on large model

    CN120744081A

  • Modeling method and device of power business decision model, equipment and medium

    CN120821832A

  • A modeling method, apparatus, equipment and medium for a power business decision model

    CN120821832B

  • Auxiliary method and auxiliary device for scientific research in medical specialized field

    CN120873154A

  • Intelligent official document generation method, system and equipment based on multi-model fusion and medium

    CN120874801A