A label hierarchy evolution document retrieval method, device, equipment and medium

CN122654292APending Publication Date: 2026-08-28BEIJING DIGITAL YIZHI TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611132509.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

其四,传统检索系统仅依靠预设正向标签完成匹配过滤,新增细分标签及其配套文档入库后,静态正向标签库因未提前收录该标签,相关文档被排除在检索命中范围之外,检索结果完整性不足

Benefits of technology

[0016]The beneficial effects of the tag-hierarchical evolution document retrieval method of this invention are as follows: Based on a hybrid architecture of regular expression matching and a large language model, candidate tags are extracted from the text content and then verified and cleaned against a pre-built tag dictionary to obtain compliant tags. Regular expression matching can quickly capture tags corresponding to explicit features such as document titles and file names, while the large language model can deeply analyze the semantic information implicit in the document body paragraphs, such as the scope of application and applicable ship types. The complementary integration of these two methods ensures that the document tag extraction process balances the high deterministic capture of explicit information with the deep mining of implicit semantics, guaranteeing the complete extraction of key tags such as document application scenarios and constraints. Verification and cleaning against the tag dictionary constrains the extraction results to the range of legal tags, ensuring the accuracy of compliant tags. Based on the hierarchical structure of the tag dictionary, bidirectional hierarchical expansion is performed on the compliant tags to obtain a complete tag set, which is then bound and stored with the corresponding shipbuilding-related documents. Through bidirectional hierarchical expansion, compliant tags trace upwards to complete all ancestor tags and recursively downwards to complete all descendant sub-tags in the tag dictionary tree structure. This ensures that the final tag set bound to each document simultaneously includes fine-grained tags, intermediate-level tags, and coarse-grained root tags, covering the entire path of tags from concrete to abstract. This provides a complete tag index foundation for multi-granularity matching in subsequent retrieval stages. Query tags are extracted from user search requests and bidirectionally expanded to obtain a search tag set. Within the same semantic dimension, a set of semantically mutually exclusive search tags is derived. After hierarchical expansion, query tags form a complete search tag set covering different granularities, ensuring that documents at the corresponding level are hit regardless of whether the user uses coarse-grained or fine-grained keywords for retrieval. Simultaneously, the automatic derivation of the opposite tag set based on semantic dimensions allows for the clear identification of tag ranges that belong to the same dimension as the query tags but are semantically mutually exclusive during the retrieval process, providing precise exclusion criteria for subsequent filtering stages. Based on the search tag set and the search opposite tag set, a dual-mode filtering search is performed on the already bound and stored ship domain documents, outputting the search results for ship domain documents that meet the requirements. On the document library with a complete set of bound tags, the search tag set can be used for forward matching to locate relevant documents, while the search opposite tag set is used to reversely eliminate documents with semantic conflicts. The dual filtering mechanism works synergistically, ensuring that the search results cover all documents related to the query tags while excluding irrelevant documents with mutually exclusive semantics under the same dimension, thus guaranteeing the relevance and purity of the search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654292A_ABST
    Figure CN122654292A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and provides a label level evolution document retrieval method, device, equipment and medium, the method comprises: based on the mixed architecture of regular matching and large language model, extracting candidate labels from text content and verifying and cleaning to obtain compliance labels; according to the hierarchical structure of the label dictionary, performing bidirectional hierarchical expansion on the compliance labels to obtain a complete label set and binding storage; extracting query labels from the retrieval request and performing bidirectional hierarchical expansion to obtain a retrieval label set, and deducing a retrieval opposite label set semantically exclusive to the retrieval label set; based on the retrieval label set and the retrieval opposite label set, performing double-mode filtering retrieval, and outputting the retrieval result. The present application makes the label labeling process of the ship field document get rid of the limitation of manual operation, the label system can be continuously improved, the document retrieval can be based on the forward matching and the reverse exclusion double logic at the same time, and the fullness rate and the precision rate of the ship document retrieval are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a document retrieval method, apparatus, device, and medium based on tag hierarchy evolution. Background Technology

[0002] Core business documents in the shipbuilding sector encompass various types of documents, including international maritime conventions, maritime laws and regulations, and classification society standards. These documents serve as the legal basis and implementation standards for the entire chain of ship design, construction, drawing review, operational inspection, and maritime supervision. With the continuous revision and updating of the shipbuilding industry's regulatory framework and the emergence of new ship types, rapidly completing semantic tagging from massive amounts of unstructured ship documents, building an scalable tagging system, and achieving high-precision document retrieval are directly related to the overall compliance level and operational management efficiency of the shipbuilding industry.

[0003] Existing ship document management systems typically employ manual tagging of documents, relying on a fixed tag library to build a retrieval and filtering mechanism. The tag system is divided according to semantic dimensions such as navigation area, vessel purpose, operating environment, and specification type, with multi-level tree structures within each dimension to express the hierarchical relationships between tags. Document tag extraction mainly relies on keyword matching based on explicit information such as document title and filename, while the retrieval phase performs single-match filtering based on preset positive tags. The tag system is pre-fixed during the initial system construction phase, and subsequent maintenance work such as tag addition, hierarchy adjustment, and relationship reconstruction relies on manual operation. The above approach faces the following technical problems when processing ship documents: First, the sheer number and continuous updates of shipping regulations, conventions, and industry standards place a heavy burden on traditional manual labeling, resulting in long response times and a lagging label update speed compared to the frequency of standard iterations. Second, the labeling system in the shipping industry exhibits a multi-dimensional hierarchical structure, with new ship categories and sub-specifications constantly emerging. The initial system phase did not cover all labels, creating a gap between the fixed labeling system and the evolving business needs. Third, much of the key information in shipping professional documents, such as the scope of application and applicable ship types, is implicit within the main text, relying on indirect descriptions or limiting conditions. Document titles and filenames alone are insufficient to fully extract the semantic tags of a document. Fourth, traditional retrieval systems rely solely on preset positive tags for matching and filtering. When new sub-tags and their associated documents are added to the database, the static positive tag library, which did not include these tags beforehand, excludes relevant documents from the search results, leading to incomplete search results. Summary of the Invention

[0004] The present invention aims to solve at least one of the above-mentioned technical problems.

[0005] To address the aforementioned problems, this invention provides a document retrieval method, apparatus, device, and medium based on tag hierarchy evolution.

[0006] In a first aspect, the present invention provides a document retrieval method based on tag hierarchy evolution, comprising: Based on a hybrid architecture of regular expression matching and large language model, candidate tags are extracted from text content. The candidate tags are then checked and cleaned by comparing them with a pre-built tag dictionary to obtain compliant tags. The text content is obtained by parsing the shipbuilding domain document to be processed. The tag dictionary stores multi-level tags and hierarchical relationships divided according to semantic dimensions. Based on the hierarchical structure of the tag dictionary, the compliance tags are expanded bidirectionally to obtain a complete tag set, and the complete tag set is bound and stored with the corresponding shipbuilding domain documents; In response to a user's search request, query tags are extracted from the search request and bidirectional hierarchical expansion is performed to obtain a search tag set. Within the same semantic dimension, the search tag set is deduced to obtain a search opposite tag set that is semantically mutually exclusive with the search tag set. Based on the search tag set and the search opposite tag set, a dual-mode filtering search is performed on the already bound and stored ship domain documents to output the search results for ship domain documents that meet the requirements.

[0007] Optionally, the hybrid architecture based on regular expression matching and large language model extracts candidate tags from the text content, compares them with a pre-built tag dictionary, and performs verification and cleaning on the candidate tags to obtain compliant tags, including: The explicit features of the text content are matched using preset domain regular expression matching rules to obtain a first candidate tag set; The text content, the first candidate tag set, the dimensional constraints and hierarchical rules of the tag dictionary are input into the large language model to extract the second candidate tag set corresponding to the implicit semantics. The first candidate tag set is forcibly merged into the second candidate tag set, illegal tags not in the tag dictionary are removed and deduplication is performed to obtain the compliant tags.

[0008] Optionally, the step of performing bidirectional hierarchical expansion on the compliant tags according to the hierarchical structure of the tag dictionary to obtain a complete tag set includes: Locate the node position of the compliance label in the label dictionary tree structure; If the compliant label is a non-leaf node, then recursively complete the subdivision labels of all descendant nodes downwards and recursively complete the ancestor labels of all corresponding dimension root nodes upwards to obtain the complete label set; If the compliant label is a leaf node, then only the ancestor labels of the root node in the corresponding dimension are recursively added upwards to obtain the complete label set.

[0009] Optionally, the step of deriving a set of mutually exclusive search tags that are semantically mutually exclusive with the search tag set within the same semantic dimension includes: Determine the target semantic dimension to which the search tag set belongs; In the tag dictionary, all tags within the target semantic dimension, excluding the tags contained in the search tag set and all their descendant node tags, are filtered out and integrated to obtain the search opposing tag set.

[0010] Optionally, the step of performing a dual-mode filtering search on the already bound and stored ship-related documents based on the search tag set and the search opposite tag set, and outputting the ship-related documents that meet the requirements to obtain search results, includes: In response to selecting the lenient mode, all ship domain documents in the already bound and stored ship domain documents that do not match the tags in the search opposition tag set are retained as the search results; In response to selecting strict mode, retain the ship domain documents that are already bound and stored, and that meet the condition of not matching the tags in the search opposite tag set and containing at least one tag in the search tag set, as the search results.

[0011] Optionally, the document retrieval method based on tag hierarchy evolution further includes a tag dynamic maintenance step: Receive tag addition, deletion, and modification operation requests for the tag dictionary, and update the topology of the tag dictionary; Based on the manipulated tag, retrieve the corresponding shipbuilding domain document and synchronously update the associated tags of the shipbuilding domain document.

[0012] Optionally, receiving requests for tag addition, deletion, and modification operations on the tag dictionary, and updating the topology of the tag dictionary, includes: In response to a tag deletion request, after deleting the target tag, update the parent node of all child nodes of the target tag to the parent node of the target tag; Based on the deletion operation, the synchronous update of the associated tags of the shipbuilding document is triggered through an asynchronous message queue.

[0013] Secondly, the present invention provides a document retrieval device with tag hierarchy evolution, comprising: The extraction module is used to extract candidate tags from text content based on a hybrid architecture of regular expression matching and large language model. The candidate tags are verified and cleaned by comparing them with a pre-built tag dictionary to obtain compliant tags. The text content is obtained by parsing the shipbuilding domain document to be processed. The tag dictionary stores multi-level tags and hierarchical relationships divided according to semantic dimensions. The extension module is used to perform bidirectional hierarchical expansion on the compliance tags according to the hierarchical structure of the tag dictionary to obtain a complete tag set, and bind and store the complete tag set with the corresponding shipbuilding domain document; The processing module is used to respond to the user's search request, extract query tags from the search request and perform the bidirectional hierarchical expansion to obtain a search tag set, and deduce a search opposite tag set that is semantically mutually exclusive with the search tag set within the same semantic dimension. The retrieval module is used to perform a dual-mode filtering retrieval on the already bound and stored ship-related documents based on the retrieval tag set and the retrieval opposite tag set, and output the retrieval results of the ship-related documents that meet the requirements.

[0014] Thirdly, the present invention provides an electronic device, including a memory and a processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the document retrieval method with tag hierarchy evolution as described in the first aspect.

[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the tag-hierarchical evolution document retrieval method as described in the first aspect.

[0016] The beneficial effects of the tag-hierarchical evolution document retrieval method of this invention are as follows: Based on a hybrid architecture of regular expression matching and a large language model, candidate tags are extracted from the text content and then verified and cleaned against a pre-built tag dictionary to obtain compliant tags. Regular expression matching can quickly capture tags corresponding to explicit features such as document titles and file names, while the large language model can deeply analyze the semantic information implicit in the document body paragraphs, such as the scope of application and applicable ship types. The complementary integration of these two methods ensures that the document tag extraction process balances the high deterministic capture of explicit information with the deep mining of implicit semantics, guaranteeing the complete extraction of key tags such as document application scenarios and constraints. Verification and cleaning against the tag dictionary constrains the extraction results to the range of legal tags, ensuring the accuracy of compliant tags. Based on the hierarchical structure of the tag dictionary, bidirectional hierarchical expansion is performed on the compliant tags to obtain a complete tag set, which is then bound and stored with the corresponding shipbuilding-related documents. Through bidirectional hierarchical expansion, compliant tags trace upwards to complete all ancestor tags and recursively downwards to complete all descendant sub-tags in the tag dictionary tree structure. This ensures that the final tag set bound to each document simultaneously includes fine-grained tags, intermediate-level tags, and coarse-grained root tags, covering the entire path of tags from concrete to abstract. This provides a complete tag index foundation for multi-granularity matching in subsequent retrieval stages. Query tags are extracted from user search requests and bidirectionally expanded to obtain a search tag set. Within the same semantic dimension, a set of semantically mutually exclusive search tags is derived. After hierarchical expansion, query tags form a complete search tag set covering different granularities, ensuring that documents at the corresponding level are hit regardless of whether the user uses coarse-grained or fine-grained keywords for retrieval. Simultaneously, the automatic derivation of the opposite tag set based on semantic dimensions allows for the clear identification of tag ranges that belong to the same dimension as the query tags but are semantically mutually exclusive during the retrieval process, providing precise exclusion criteria for subsequent filtering stages. Based on the search tag set and the search opposite tag set, a dual-mode filtering search is performed on the already bound and stored ship domain documents, outputting the search results for ship domain documents that meet the requirements. On the document library with a complete set of bound tags, the search tag set can be used for forward matching to locate relevant documents, while the search opposite tag set is used to reversely eliminate documents with semantic conflicts. The dual filtering mechanism works synergistically, ensuring that the search results cover all documents related to the query tags while excluding irrelevant documents with mutually exclusive semantics under the same dimension, thus guaranteeing the relevance and purity of the search results.

[0017] This invention constructs a complete closed loop from automated document tag extraction and dynamic hierarchical system completion to retrieval tag expansion, opposing tag derivation, and dual-mode filtering. This enables the tagging process of documents in the shipbuilding field to break free from the limitations of manual operation. The tag system can be continuously improved as documents are added to the database, and document retrieval can be executed simultaneously based on both forward matching and reverse exclusion logic. Ultimately, this achieves a balance between recall and precision in shipbuilding document retrieval. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the document retrieval method based on tag hierarchy evolution according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the three-layer hierarchical decoupling architecture according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a document retrieval device with tag hierarchy evolution according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the document tag extraction process according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the document retrieval process according to an embodiment of the present invention. Detailed Implementation

[0019] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0020] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0021] As used herein, the term "comprising" and its variations are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; and the term "optionally" means "optional embodiments". Definitions for other terms will be given in the description below.

[0022] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0023] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0024] like Figure 1 , Figure 5 and Figure 6 As shown in the figure, an embodiment of the present invention provides a document retrieval method based on tag hierarchy evolution, comprising: Step S1: Extract candidate tags from the text content based on a hybrid architecture of regular expression matching and large language model. Verify and clean the candidate tags by comparing them with a pre-built tag dictionary to obtain compliant tags. The text content is obtained by parsing the shipbuilding domain document to be processed. The tag dictionary stores multi-level tags and hierarchical relationships divided according to semantic dimensions.

[0025] Semantic tags in shipping documents are scattered across different locations. Explicit locations such as document titles and filenames typically contain direct descriptions of the core scope of application, while the body text implicitly contains a wealth of deeper semantic information, including the scope of application, applicable ship types, and operational environment constraints. Relying solely on regular expression matching can only capture explicit features and cannot parse implicit semantics; relying solely on large language models may miss key explicit information such as document titles, and the output of large models is uncontrollable. Therefore, a fusion architecture is needed that balances the high deterministic capture of explicit information with in-depth mining of implicit semantics, while simultaneously constraining the extracted results within a legal range through tag dictionary verification and cleaning, ensuring the accuracy and controllability of the tags. Specifically, the first step is text content parsing. The system receives shipping documents in various formats, performs format parsing and OCR (Optical Character Recognition) text parsing, and extracts the text content and document metadata (document number, publishing organization, effective date, revision version, etc.) from the documents. A keyword-driven hierarchical content extraction strategy is applied to the text content. Priority is given to locating and extracting titles, chapters, and paragraphs containing limiting keywords such as "ship specifications," "applicable ship types," and "navigation areas" as core semantic analysis fragments. If the target keywords do not appear in the document, the core text content of the first 3 to 5 pages of the document is extracted to control the computational overhead of the subsequent large language model and avoid information loss caused by long text truncation. The core text fragments with tagged semantics, the original text content, and document metadata (such as document number, publishing organization, effective date, etc.) are output to obtain the final text content.

[0026] Based on a hybrid architecture of regular expression matching and a large language model, candidate tags are extracted from the truncated text content. In the regular expression pre-extraction stage, ship-specific regular expression matching rules are constructed to accurately match explicit words in documents such as names and titles, capturing highly deterministic explicit tags from the tag dictionary. In the large model semantic extraction stage, based on a ship-specific customized constraint prompt word project, relevant data such as the truncated text content, candidate tags obtained from regular expression pre-extraction, and dimensional constraints and hierarchical rules of the tag dictionary are input into the large language model. Leveraging the deep semantic understanding capabilities of the large language model, the applicable and opposing tags implicit in the document text are parsed, outputting a preliminary list of positive tags and a list of opposing tags, thus obtaining candidate tags. The candidate tags are then validated and cleaned against the pre-constructed tag dictionary. Validation and cleaning may include forced merging correction, illegal tag filtering, and deduplication, resulting in compliant tags. The tag dictionary is stored using a nested tree key-value pair structure, divided by semantic dimensions (navigation area, vessel purpose, operating environment, standard type, etc.), and includes built-in hierarchical relationships, subordinate rules, and dimensional isolation constraints for tags.

[0027] Step S2: Perform bidirectional hierarchical expansion on the compliance tags according to the hierarchical structure of the tag dictionary to obtain a complete tag set, and bind and store the complete tag set with the corresponding shipbuilding domain documents.

[0028] The tagging system in the shipping industry exhibits strong hierarchical characteristics, with multi-level tree structures existing within the same semantic dimension. The compliance tags initially extracted from a document may only be located at a certain level (e.g., "coastal bulk carrier" is a fine-grained leaf node, or "coastal vessel" is an intermediate-level non-leaf node). Simply binding the extracted tags to the document leads to serious defects in the retrieval stage. Specifically, when a fine-grained tag exists alone, users searching with coarse-grained keywords like "sea vessel" will not be able to find the document, resulting in missed searches; similarly, when a coarse-grained tag exists alone, users searching with fine-grained keywords like "coastal bulk carrier" will also not be able to find the document. Therefore, it is necessary to perform bidirectional hierarchical expansion of compliance tags, ensuring that the tag set bound to the document covers the entire path of tags from fine-grained to coarse-grained. Specifically, this involves locating the node position of the compliance tag in the tag dictionary tree structure. Each tag in the compliance tag list is traversed, and its hierarchical position in the tag dictionary is located, depending on the tag type, such as whether it is a leaf node or a non-leaf node. For compliance tags in non-leaf nodes, a bidirectional expansion process of downward recursive completion and upward recursive tracing is performed. For compliance tags in leaf nodes, only upward recursive tracing and completion are performed; downward expansion is not performed. The completed tag set is then bound to the corresponding shipbuilding-related document, associating metadata such as document name, original document content, publication date, specification type, and effective status. A bidirectional index relationship between tags and documents is established, and the tags are persistently stored in the corresponding storage database, such as the document information storage module.

[0029] Step S3: In response to the user's search request, extract the query tags from the search request and perform the bidirectional hierarchical expansion to obtain a search tag set. Within the same semantic dimension, deduce the search tag set to obtain a search opposing tag set that is semantically mutually exclusive with the search tag set.

[0030] When a user submits a search request, their input search terms may be coarse-grained tags (such as "seagoing vessels"), intermediate-level tags (such as "coastal vessels"), or fine-grained tags (such as "coastal bulk carriers"). Matching only the user's original search terms will result in incomplete search results. Therefore, it is necessary to perform bidirectional hierarchical expansion on the query tags to ensure that the search tag set covers the entire path of tags from fine-grained to coarse-grained, so that any granularity of search terms can hit documents at the corresponding level. At the same time, there may be semantically mutually exclusive relationships between tags under the same semantic dimension (such as "inland waterway vessels" and "coastal vessels" being mutually exclusive within a navigation area). It is necessary to explicitly identify and exclude documents that are semantically mutually exclusive with the search tag set during the search stage to improve search accuracy. Specifically, in response to the user's search request, which includes request information and search text information, query tags are extracted from the search request. Users can input search conditions such as "search for maritime regulations applicable to coastal vessels" through natural language text. Using the methods in steps S1 and S2, query tags are extracted from the search request and bidirectional hierarchical expansion is performed to obtain the search tag set. Within the same semantic dimension, derive and retrieve mutually exclusive tag sets. Determine the target semantic dimension to which the retrieval tag set belongs (if the retrieval tag set involves multiple dimensions, perform the following operations for each dimension separately). Filter all tags within that target semantic dimension except for the tags contained in the retrieval tag set and all their descendant node tags, and integrate these tags into a retrieval tag set. This derivation process uses the semantic dimension as the boundary and does not operate across dimensions, ensuring the logical rigor of the mutual exclusion relationships. For example, "passenger ship" (ship purpose dimension) and "inland waterway ship" (navigation area dimension) are not considered mutually exclusive because they belong to different semantic dimensions and can both be applied to the same document.

[0031] Step S4: Based on the search tag set and the search opposite tag set, perform a dual-mode filtering search on the already bound and stored ship domain documents, and output the ship domain documents that meet the requirements to obtain the search results.

[0032] Specifically, a dual-mode filtering retrieval is performed on the shipbuilding document library that has been bound with stored tags. Each document in the library has been bound with a complete tag set in step S2, establishing a bidirectional index between tags and documents. The dual-mode filtering retrieval supports two retrieval methods: a lenient mode and a strict mode. Users select the appropriate retrieval mode according to their business needs and perform corresponding searches based on the search tag set and the search counterpart tag set. The retrieval results, which meet the requirements, are output as shipbuilding documents and can be sorted by document publication time, relevance, etc., and returned to the user.

[0033] This invention employs a hybrid architecture combining regular expression matching and a large language model to extract candidate tags from text content. These tags are then validated and cleaned against a pre-built tag dictionary to obtain compliant tags. Regular expression matching quickly captures tags corresponding to explicit features such as document titles and filenames, while the large language model deeply analyzes semantic information implicit in the document's body paragraphs, such as the scope of application and applicable ship types. This complementary integration ensures that the document tag extraction process balances highly deterministic capture of explicit information with in-depth mining of implicit semantics, guaranteeing the complete extraction of key tags such as document application scenarios and constraints. Validation and cleaning against the tag dictionary constrains the extracted results to the range of legal tags, ensuring the accuracy of compliant tags. Based on the hierarchical structure of the tag dictionary, a bidirectional hierarchical expansion is performed on the compliant tags to obtain a complete tag set, which is then bound and stored with the corresponding shipbuilding-related documents. Through bidirectional hierarchical expansion, compliant tags trace upwards to complete all ancestor tags and recursively downwards to complete all descendant sub-tags in the tag dictionary tree structure. This ensures that the final tag set bound to each document simultaneously includes fine-grained tags, intermediate-level tags, and coarse-grained root tags, covering the entire path of tags from concrete to abstract. This provides a complete tag index foundation for multi-granularity matching in subsequent retrieval stages. Query tags are extracted from user search requests and bidirectionally expanded to obtain a search tag set. Within the same semantic dimension, a set of semantically mutually exclusive search tags is derived. After hierarchical expansion, query tags form a complete search tag set covering different granularities, ensuring that documents at the corresponding level are hit regardless of whether the user uses coarse-grained or fine-grained keywords for retrieval. Simultaneously, the automatic derivation of the opposite tag set based on semantic dimensions allows for the clear identification of tag ranges that belong to the same dimension as the query tags but are semantically mutually exclusive during the retrieval process, providing precise exclusion criteria for subsequent filtering stages. Based on the search tag set and the search opposite tag set, a dual-mode filtering search is performed on the already bound and stored ship domain documents, outputting the search results for ship domain documents that meet the requirements. On the document library with a complete set of bound tags, the search tag set can be used for forward matching to locate relevant documents, while the search opposite tag set is used to reversely eliminate documents with semantic conflicts. The dual filtering mechanism works synergistically, ensuring that the search results cover all documents related to the query tags while excluding irrelevant documents with mutually exclusive semantics under the same dimension, thus guaranteeing the relevance and purity of the search results.

[0034] This invention constructs a complete closed loop from automated document tag extraction and dynamic hierarchical system completion to retrieval tag expansion, opposing tag derivation, and dual-mode filtering. This frees the tagging process of documents in the shipbuilding field from the limitations of manual operation. The tag system can be continuously improved as documents are added to the database, and document retrieval can be performed simultaneously based on both forward matching and reverse exclusion logic. Ultimately, this achieves a balance between recall and precision in shipbuilding document retrieval.

[0035] Optionally, the hybrid architecture based on regular expression matching and large language model extracts candidate tags from the text content, compares them with a pre-built tag dictionary, and performs verification and cleaning on the candidate tags to obtain compliant tags, including: The explicit features of the text content are matched using preset domain regular expression matching rules to obtain a first candidate tag set.

[0036] Specifically, a dedicated regular expression matching rule base for the shipbuilding field is constructed. This rule base can cover multiple dimensions of explicit feature matching patterns, including navigation area (keywords such as inland waterways, coastal waterways, ocean-going vessels, polar regions, etc.), ship purpose (keywords such as passenger ships, cargo ships, engineering vessels, fishing vessels, etc.), and ship power type (diesel-powered, gas-powered, battery-powered, etc.). A controllable interval fuzzy matching strategy is adopted, allowing a limited number of separator characters between keywords to adapt to non-continuous expressions in documents such as "coastal navigation vessels" and "applicable to inland waterways and coastal areas." Precise matching is prioritized for explicit locations such as document names and chapter titles, capturing highly deterministic explicit tags from a tag dictionary to obtain the first candidate tag set.

[0037] The text content, the first candidate tag set, the dimensional constraints and hierarchical rules of the tag dictionary are input into the large language model to extract the second candidate tag set corresponding to the implicit semantics. Specifically, based on a customized constraint prompt word engineering approach for the shipbuilding domain, input prompt words are constructed. These prompt words may include truncated core text fragments of the document, the content of the first candidate tag set, the complete semantic dimension division of the tag dictionary and the complete tag tree structure under each dimension, the hierarchical relationships between tags, explanations of mutual exclusion relationships within the same level, dimensional isolation constraint rules, and output format requirements. The prompt words explicitly require the large language model to reason within the scope defined by the tag dictionary and not to output tags outside the tag dictionary. Leveraging the deep semantic understanding capabilities of the large language model, the applicable and opposing tags implicit in the document text are parsed, and a second candidate tag set is output.

[0038] The first candidate tag set is forcibly merged into the second candidate tag set, illegal tags not in the tag dictionary are removed and deduplication is performed to obtain the compliant tags.

[0039] Specifically, all tags in the first candidate tag set obtained through regular expression pre-extraction are forcibly added to the second candidate tag set output by the large language model, correcting the potential omission of explicit tags in the large language model. This forced merging operation is not limited by the output of the large language model; that is, tags in the first candidate tag set are retained regardless of whether the large language model outputs corresponding explicit tags. Illegal tags not in the tag dictionary are removed by iterating through all merged candidate tags, matching each tag to the tag dictionary, retaining tags present in the tag dictionary, and deleting illegal tags not defined in the tag dictionary. A deduplication operation is performed by iterating through the remaining tag list after filtering, removing duplicate tags to ensure that each tag appears only once among the compliant tags.

[0040] This invention employs regular expression matching to capture highly deterministic tags corresponding to explicit features such as document titles and filenames. A large language model deeply analyzes semantic tags such as applicable scope and applicable ship types hidden in the document text. The complementary integration of these two approaches enables tag extraction to cover the complete space of explicit information and implicit semantics. The forced merging of the first candidate tag set into the second candidate tag set ensures that key explicit tags are not missed by the large model. Illegal tag filtering and deduplication ensure the legality and uniqueness of the tags, ultimately achieving high accuracy and high completeness in ship document tag extraction.

[0041] Optionally, the step of performing bidirectional hierarchical expansion on the compliant tags according to the hierarchical structure of the tag dictionary to obtain a complete tag set includes: Locate the node position of the compliance label in the label dictionary tree structure.

[0042] Specifically, each compliance tag is traversed, and using the tag dictionary as an index, the node position of each compliance tag in the tree structure is located one by one. This determines whether the compliance tag exists in the tag dictionary, as well as its level depth, parent-child relationship, and dimension in the tag tree. At the same time, it is determined whether the compliance tag is a leaf node or a non-leaf node. If the compliance tag has child nodes in the tag dictionary, it is a non-leaf node; if it has no child nodes, it is a leaf node.

[0043] If the compliant label is a non-leaf node, then recursively complete the subdivision labels of all descendant nodes downwards and recursively complete the ancestor labels of all corresponding dimension root nodes upwards to obtain the complete label set; If the compliant label is a leaf node, then only the ancestor labels of the root node in the corresponding dimension are recursively added upwards to obtain the complete label set.

[0044] Specifically, if the compliance label is a non-leaf node, a bidirectional expansion is performed using downward recursive completion and upward recursive tracing. Downward recursive completion starts from the compliance label node and traverses downwards along the tree structure, recursively obtaining all descendant nodes of the compliance label until a leaf node is reached, and completing the label set with the labels of all descendant nodes. Upward recursive tracing starts from the compliance label node and traverses upwards along the tree structure, recursively tracing back to the root node of the semantic dimension to which the compliance label belongs, and completing the label set with the labels of all ancestor nodes along the path. The combination of these two methods yields the complete label set. If the compliance label is a leaf node, only upward recursive tracing is performed, without downward expansion. Starting from the leaf node, traverse upwards along the tree structure, recursively tracing back to the root node of the semantic dimension to which the compliance label belongs, and completing the label set with the labels of all ancestor nodes along the path, resulting in the complete label set. In this context, the root node of a dimension is the top-level node of a certain semantic dimension in the tag tree. For example, the root node of the navigation area dimension is "navigation area", and the root node of the ship purpose dimension is "ship purpose". Each root node of a dimension carries all the hierarchical tags of that dimension.

[0045] This invention, through its embodiments, locates the node positions of compliant tags in the tag dictionary tree structure and executes differentiated expansion strategies based on the different attributes of leaf nodes and non-leaf nodes. This ensures that the final tag set bound to each document simultaneously includes fine-grained tags, intermediate-level tags, and coarse-grained root tags, covering the entire path of tags from concrete to abstract. This provides a complete tag index foundation for multi-granularity matching in the subsequent retrieval stage. The recursive completion of non-leaf nodes ensures complete coverage of fine-grained scenarios, while the upward tracing completion of leaf nodes and all compliant tags ensures the completeness of coarse-grained subordinate tags. The combination of these two approaches effectively solves the retrieval omission problem caused by the existence of a single granular tag.

[0046] Optionally, the step of deriving a set of mutually exclusive search tags that are semantically mutually exclusive with the search tag set within the same semantic dimension includes: Determine the target semantic dimension to which the search tag set belongs; In the tag dictionary, all tags within the target semantic dimension, excluding the tags contained in the search tag set and all their descendant node tags, are filtered out and integrated to obtain the search opposing tag set.

[0047] Specifically, the search tag set may contain tags from multiple semantic dimensions, requiring identification of the dimension to which each tag belongs. Based on predefined dimension divisions in the tag dictionary (navigation area, ship purpose, hull material, power type, etc.), each tag in the search tag set is traversed to query its dimension affiliation in the tag dictionary. If the search tag set contains tags with a single dimension, that dimension is the target semantic dimension; if the search tag set contains tags with multiple dimensions, each dimension is used as the target semantic dimension for subsequent filtering operations, ultimately merging the opposing tag sets derived from each dimension. In the tag dictionary, using the root node of the target semantic dimension as the boundary, the complete set of all tags under that target semantic dimension is obtained. Two parts of the tags are excluded from this complete set: the first part is the tags themselves belonging to that target semantic dimension in the search tag set; the second part is all descendant node tags (i.e., child tags) of these tags in the tag dictionary tree structure. The remaining tags constitute the search opposing tag set. Here, semantic dimensions are independent classification categories of tags, and different dimensions are orthogonally isolated from each other. The navigation area dimension covers spatial categories such as inland waterways, coastal waters, open ocean, and polar regions; the ship purpose dimension covers functional categories such as passenger ships, cargo ships, engineering vessels, and fishing vessels; the hull material dimension covers material categories such as steel, aluminum alloy, fiberglass, and composite materials; and the power type dimension covers power forms such as diesel engines, electric propulsion, hybrid power, and new energy. There is a hierarchical relationship between labels within the same dimension, and there is no mutual exclusion relationship between different dimensions.

[0048] This invention uses a single semantic dimension as the boundary of the search tag set, automatically filtering and excluding all sibling and related tags after the search tag set and all its descendant nodes in the tag dictionary. This achieves fully automated derivation of mutual exclusion relationships without manual configuration. The processing method of independently calculating different semantic dimensions and merging the results ensures that mutual exclusion judgment is only performed within the same semantic dimension, without misjudgment between different dimensions, thus ensuring the orthogonality of the multi-dimensional tag system and the rigor of the derivation logic of the opposing tags. The derivation of the search opposing tag set provides an accurate set of mutually exclusive tags for subsequent dual-mode filtering retrieval, enabling the retrieval process to automatically exclude semantically mutually exclusive irrelevant documents and ensuring the purity of the retrieval results.

[0049] Optionally, the step of performing a dual-mode filtering search on the already bound and stored ship-related documents based on the search tag set and the search opposite tag set, and outputting the ship-related documents that meet the requirements to obtain search results, includes: In response to the selection of the lenient mode, all ship domain documents that do not match the tags in the search opposition tag set from the already bound and stored ship domain documents are retained as the search results.

[0050] Specifically, before executing the dual-mode filtering retrieval, each document in the pre-bound and stored shipbuilding document library has been bound with a complete set of tags, establishing a bidirectional index relationship between tags and documents. Query tags are extracted from the user's search request and subjected to bidirectional hierarchical expansion to obtain the search tag set. Simultaneously, the search counterpart tag set is derived within the same semantic dimension. The dual-mode filtering retrieval supports both a relaxed mode and a strict mode, allowing users to select the mode through the user interaction layer based on their business needs. Users can simultaneously select the desired search mode when submitting a search request.

[0051] The execution logic of the lenient mode includes, in response to the user's selection of the lenient mode, checking for each document in the bound and stored shipbuilding document library whether its complete set of bound tags matches any tag in the search's opposing tag set. If the document's bound tags do not intersect with the search's opposing tag set (i.e., no opposing tags are matched), the document is retained in the search results; if there is an intersection (i.e., at least one opposing tag is matched), the document is removed. In this mode, it is not required that the document match any tag in the search's tag set. The search results are returned to the user after being sorted by document publication time, relevance score, etc.

[0052] In response to selecting strict mode, retain the ship domain documents that are already bound and stored, and that meet the condition of not matching the tags in the search opposite tag set and containing at least one tag in the search tag set, as the search results.

[0053] Specifically, the execution logic of strict mode includes, in response to the user's selection of strict mode, checking two conditions for each document in the bound and stored shipbuilding document library. Condition 1: The document's bound tags and the search's opposing tag set have no intersection (no opposing tags are matched). Condition 2: The document's bound tags and the search's tag set have an intersection (i.e., it contains at least one tag from the search's tag set). Documents that meet both conditions are retained and output as search results; those that do not meet either condition are discarded. The search results are returned to the user after being sorted by document publication time, relevance score, etc.

[0054] The lenient mode uses "no matching of opposing tags" as the sole retention condition, characterized by high recall and primarily serving to "eliminate incorrect items." It is suitable for exploratory browsing scenarios, such as designers consulting specifications for similar ship types or researchers conducting technical research. The strict mode uses "no matching of opposing tags and at least one matching of search tags" as a dual retention condition, characterized by high accuracy and ensuring that search results are both relevant and conflict-free. It is suitable for business scenarios such as precise compliance review and core document retrieval.

[0055] In this embodiment of the invention, the relaxed mode uses the absence of opposing tags as the sole retention condition, achieving high recall for exploratory retrieval and effectively excluding documents that conflict with the query semantics. It does not require mandatory matching of positive tags, making it suitable for exploratory browsing and technical research scenarios in the shipping industry. The strict mode uses the absence of opposing tags and at least one matching search tag as dual retention conditions, achieving high accuracy for precise retrieval. It ensures that the search results comprehensively cover relevant documents while completely excluding mutually exclusive conflicting documents, making it suitable for precise compliance review and core document retrieval scenarios in the shipping industry. The dual-mode switching capability allows the same retrieval system to flexibly adapt to different business needs, balancing recall and precision.

[0056] Optionally, the document retrieval method based on tag hierarchy evolution further includes a tag dynamic maintenance step: Receive tag addition, deletion, and modification operation requests for the tag dictionary, and update the topology of the tag dictionary.

[0057] Based on the manipulated tag, retrieve the corresponding shipbuilding domain document and synchronously update the associated tags of the shipbuilding domain document.

[0058] Optionally, receiving requests for tag addition, deletion, and modification operations on the tag dictionary, and updating the topology of the tag dictionary, includes: In response to a tag deletion request, after deleting the target tag, update the parent node of all child nodes of the target tag to the parent node of the target tag; Based on the deletion operation, the synchronous update of the associated tags of the shipbuilding document is triggered through an asynchronous message queue.

[0059] The labeling system in the shipping industry is not static. With revisions to international maritime conventions, updates to classification society standards, and the emergence of new ship types (such as polar icebreakers, methanol-powered vessels, and unmanned vessels), the labeling system requires continuous addition, modification, or deletion of label nodes, as well as adjustments to the hierarchical relationships and dimensional affiliations between labels. If the label dictionary changes, the associated labels of already bound and stored shipping documents cannot be updated synchronously, leading to inconsistencies between the label dictionary and document label data. After a label is deleted, some documents may still be bound to the deleted label, which no longer exists in the label dictionary, resulting in orphan labels in the document label data. After a label is modified, some documents may still be bound to the old label name, resulting in two versions of the same semantic label. After a label is added, the new label will not automatically be associated with existing documents, causing inconsistencies between the label dictionary and document label data. This data inconsistency will lead to serious defects such as label matching errors and inaccurate search results in subsequent retrieval stages. Therefore, when the label dictionary changes, it is necessary to automatically and synchronously update all associated labels of documents bound to the manipulated label to ensure data consistency between the label dictionary and documents. Specifically, it receives requests for adding, deleting, and modifying tags in the tag dictionary. It provides a user-friendly interface (such as a pre-defined tag modification module) allowing users to perform these operations via a graphical interface or API. Adding a tag involves the user selecting a target parent node in the tag dictionary's tree structure, creating a new tag node and specifying its position within the tree. The user can select its parent and child nodes and insert the new tag into the corresponding position in the tag dictionary storage module. Modifying a tag involves the user selecting an existing tag and editing its content (tag name, description, etc.) or hierarchical position. Deleting a tag involves the user selecting an existing tag and initiating a deletion request. The topology of the tag dictionary is updated, with different topology update logic executed for different operation types. For adding a tag, a new tag node is created under the user-specified parent node in the relevant storage database, establishing a parent-child relationship between the new tag node and its parent. If the user also specifies a child node, the new tag is inserted between the parent node and the original node, and the parent pointer of the child node is automatically adjusted to point to the new tag node. For tag modification operations, locate the node position of the modified tag in the tag dictionary's related storage database, and update the tag name, description, and other content fields of that node. If the modification involves a hierarchical adjustment (i.e., moving the tag from its original parent node to a new parent node), update the node's parent pointer to point to the new parent node, and recalculate the path information of that node and all its subtrees. For tag deletion operations, locate the node position of the target tag to be deleted in the tag dictionary's related storage database, and remove the node from the tag tree. Before deleting the node, retrieve all child nodes of the target tag.For each child node, update its parent node to the parent node of the deleted target tag (i.e., the child node inherits from its parent's parent, one level up), ensuring that the child node does not become an orphan node. Delete the target tag node and its association with its parent node. The deletion operation can be designed as soft deletion (the tag is logically marked as deleted but physically retained, and the document's associated tags are synchronously marked as expired) or hard deletion (the tag is physically deleted, and the document's associated tags are synchronously deleted), to adapt to the data traceability needs of different business scenarios.

[0060] Specifically, using the operated tag as the query condition, a search is performed in the tag-document bidirectional index of the relevant storage database to obtain a list of unique identifiers for all shipbuilding documents bound to that tag. For add operations, since the tag is not yet bound to any document at the time of addition, this step does not perform document synchronization updates; for modify and delete operations, document synchronization updates are performed. The associated tags of shipbuilding documents are synchronized and updated. For modify operations, all documents bound to the modified tag are traversed, and the old tag name or old node identifier in the associated tags of each document is replaced with the modified new tag name or new node identifier. For delete operations, all documents bound to the deleted tag are traversed, and the deleted tag is removed from the associated tags of each document. For add operations, historical documents are not updated immediately; when a new document is added to the database, the new tag is automatically bound through steps S1 and S2. If it is necessary to batch associate new tags with historical documents, it can be manually or automatically triggered through, for example, the batch binding function of the preset tag modification module. The synchronization update process is triggered and executed through an asynchronous message queue. After the tag dictionary topology structure is updated, the tag management service encapsulates the document synchronization task into a message and sends it to the asynchronous message queue. The message content includes the identifier information of the operated tag, the operation type (modification or deletion), and a list of affected document IDs (if a large number of documents are involved, batch processing can be used, with each batch containing a preset number of document IDs). After receiving the message, the message queue consumer asynchronously executes the batch update operation of the document-associated tags in the background. After execution, the task completion status (success / failure / partial failure, number of documents processed, list of failed document IDs, etc.) is sent back to the tag management service via a callback function. This asynchronous processing mechanism allows the topology update operation of the tag dictionary to respond quickly to the user, while the synchronous update of tags for a large number of documents is performed in the background, avoiding long waiting times for users.

[0061] This invention employs a three-layer, hierarchical, decoupled architecture to realize the document retrieval method described above, which involves the evolution of tag levels. For example... Figure 2As shown, the architecture employs a three-layer, decoupled structure, with independent data interaction and modular collaborative operation between layers. The top layer is the user interaction layer, providing human-computer interaction entry points for document upload, text retrieval, visual tag editing, parameter configuration, and mode selection. It supports manual adjustment of tag content, modification of hierarchical structure, and correction of erroneous tags. The middle layer is the internal processing layer, the core algorithm operation layer, integrating core modules for document preprocessing, hybrid tag extraction, dynamic tag hierarchy completion, automatic derivation of opposing tags, dual-mode retrieval filtering, and full lifecycle tag management, completing intelligent processing throughout the entire process. The bottom layer is the data service layer, containing a document content database, a document tag database, and a dedicated tag dictionary database. These databases store the original ship document text, document-bound multi-dimensional tags, and a standardized hierarchical tag system, providing stable data support for upper-layer algorithm calls, data persistence, and real-time query calls.

[0062] Specifically, the document preprocessing module integrates multi-format file upload, parsing, and encoding adaptation functions, incorporates a lightweight OCR recognition algorithm, and is compatible with commonly used ship documents such as PDF, scanned documents, image documents, Word, and TXT. It batch extracts clean text content and document metadata. The metadata includes key information such as document number, publishing organization, effective date, and revision version, providing a complete basic input for subsequent accurate tag extraction. The content search module features a visual graphical interface that supports users' free text and natural language search input. It also includes a built-in tag tree selection control, allowing users to quickly combine search criteria by selecting dimension tags. Combined with dual-mode input, it lowers the barrier to professional search and improves the efficiency and accuracy of search criterion construction. The tag modification module is responsible for the dynamic operation and maintenance of the tag dictionary database. It supports tag addition, content editing, node deletion, adjustment of hierarchical levels, and modification of dimension affiliation. It enables real-time reconstruction of the tag system topology and ensures the evolvability of tag hierarchy relationships. The document tag extraction module receives preprocessed clean text and document metadata. It adopts a rule engine + large model fusion architecture to complete explicit tag matching, implicit semantic tag mining, output result verification and cleaning, and completes preliminary tag screening by combining tag dictionary hierarchical constraints, and outputs compliant tags. The document tag retrieval module receives user retrieval requests and sequentially completes the following steps: query statement tag extraction, query tag hierarchy expansion, automatic reasoning of opposing tags in the same dimension, and dual-mode filtering to remove irrelevant documents with semantic conflicts or inaccurate scopes, and outputs highly relevant and low-noise retrieval results. The document information storage module persistently stores the full original document content of the ship, basic document metadata, and the relationship between the bound search tags and opposing tags, and establishes a two-way index of tags and documents to ensure fast matching during the retrieval stage; The tag dictionary storage module uses a nested structured format such as JSON / YAML to store a tree-like tag system, solidifying the boundaries of each dimension, hierarchical rules, and basic mutual exclusion constraints; it provides unified standard data constraints for the entire process of tag extraction, hierarchical completion, opposition derivation, and tag editing, ensuring the semantic consistency and logical closed loop of tags throughout the system.

[0063] like Figure 3 As shown, an embodiment of the present invention provides a document retrieval device with tag hierarchy evolution, comprising: The extraction module is used to extract candidate tags from text content based on a hybrid architecture of regular expression matching and large language model. The candidate tags are verified and cleaned by comparing them with a pre-built tag dictionary to obtain compliant tags. The text content is obtained by parsing the shipbuilding domain document to be processed. The tag dictionary stores multi-level tags and hierarchical relationships divided according to semantic dimensions. The extension module is used to perform bidirectional hierarchical expansion on the compliance tags according to the hierarchical structure of the tag dictionary to obtain a complete tag set, and bind and store the complete tag set with the corresponding shipbuilding domain document; The processing module is used to respond to the user's search request, extract query tags from the search request and perform the bidirectional hierarchical expansion to obtain a search tag set, and deduce a search opposite tag set that is semantically mutually exclusive with the search tag set within the same semantic dimension. The retrieval module is used to perform a dual-mode filtering retrieval on the already bound and stored ship-related documents based on the retrieval tag set and the retrieval opposite tag set, and output the retrieval results of the ship-related documents that meet the requirements.

[0064] like Figure 4 As shown, an electronic device 400 provided in this embodiment of the invention includes a memory 410 and a processor 420; the memory 410 is used to store a computer program; the processor 420 is used to implement the document retrieval method with tag hierarchy evolution as described above when the computer program is executed.

[0065] Alternatively, an electronic device 400 includes a memory 410 and a processor 420 coupled to the memory 410; the memory 410 is configured to store a computer program; and the processor 420 is configured to perform the following operations when the computer program is executed: Based on a hybrid architecture of regular expression matching and large language model, candidate tags are extracted from text content. The candidate tags are then checked and cleaned by comparing them with a pre-built tag dictionary to obtain compliant tags. The text content is obtained by parsing the shipbuilding domain document to be processed. The tag dictionary stores multi-level tags and hierarchical relationships divided according to semantic dimensions. Based on the hierarchical structure of the tag dictionary, the compliance tags are expanded bidirectionally to obtain a complete tag set, and the complete tag set is bound and stored with the corresponding shipbuilding domain documents; In response to a user's search request, query tags are extracted from the search request and bidirectional hierarchical expansion is performed to obtain a search tag set. Within the same semantic dimension, the search tag set is deduced to obtain a search opposite tag set that is semantically mutually exclusive with the search tag set. Based on the search tag set and the search opposite tag set, a dual-mode filtering search is performed on the already bound and stored ship domain documents to output the search results for ship domain documents that meet the requirements.

[0066] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the document retrieval method with tag hierarchy evolution as described above.

[0067] Alternatively, a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the following operations: Based on a hybrid architecture of regular expression matching and large language model, candidate tags are extracted from text content. The candidate tags are then checked and cleaned by comparing them with a pre-built tag dictionary to obtain compliant tags. The text content is obtained by parsing the shipbuilding domain document to be processed. The tag dictionary stores multi-level tags and hierarchical relationships divided according to semantic dimensions. Based on the hierarchical structure of the tag dictionary, the compliance tags are expanded bidirectionally to obtain a complete tag set, and the complete tag set is bound and stored with the corresponding shipbuilding domain documents; In response to a user's search request, query tags are extracted from the search request and bidirectional hierarchical expansion is performed to obtain a search tag set. Within the same semantic dimension, the search tag set is deduced to obtain a search opposite tag set that is semantically mutually exclusive with the search tag set. Based on the search tag set and the search opposite tag set, a dual-mode filtering search is performed on the already bound and stored ship domain documents to output the search results for ship domain documents that meet the requirements.

[0068] The present invention will now be described an electronic device 400 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic device 400 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 400 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0069] Electronic device 400 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or a computer program loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0070] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A document retrieval method based on tag hierarchy evolution, characterized in that, include: Based on a hybrid architecture of regular expression matching and large language model, candidate tags are extracted from text content. The candidate tags are then checked and cleaned by comparing them with a pre-built tag dictionary to obtain compliant tags. The text content is obtained by parsing the shipbuilding domain document to be processed. The tag dictionary stores multi-level tags and hierarchical relationships divided according to semantic dimensions. Based on the hierarchical structure of the tag dictionary, the compliance tags are expanded bidirectionally to obtain a complete tag set, and the complete tag set is bound and stored with the corresponding shipbuilding domain documents; In response to a user's search request, query tags are extracted from the search request and bidirectional hierarchical expansion is performed to obtain a search tag set. Within the same semantic dimension, the search tag set is deduced to obtain a search opposite tag set that is semantically mutually exclusive with the search tag set. Based on the search tag set and the search opposite tag set, a dual-mode filtering search is performed on the already bound and stored ship domain documents to output the search results for ship domain documents that meet the requirements.

2. The document retrieval method based on tag hierarchy evolution according to claim 1, characterized in that, The hybrid architecture based on regular expression matching and large language model extracts candidate tags from text content, compares them with a pre-built tag dictionary, and performs verification and cleaning on the candidate tags to obtain compliant tags, including: The explicit features of the text content are matched using preset domain regular expression matching rules to obtain a first candidate tag set; The text content, the first candidate tag set, the dimensional constraints and hierarchical rules of the tag dictionary are input into the large language model to extract the second candidate tag set corresponding to the implicit semantics. The first candidate tag set is forcibly merged into the second candidate tag set, illegal tags not in the tag dictionary are removed and deduplication is performed to obtain the compliant tags.

3. The document retrieval method based on tag hierarchy evolution according to claim 1, characterized in that, The step of performing bidirectional hierarchical expansion on the compliant tags according to the hierarchical structure of the tag dictionary to obtain a complete tag set includes: Locate the node position of the compliance label in the label dictionary tree structure; If the compliant label is a non-leaf node, then recursively complete the subdivision labels of all descendant nodes downwards and recursively complete the ancestor labels of all corresponding dimension root nodes upwards to obtain the complete label set; If the compliant label is a leaf node, then only the ancestor labels of the root node in the corresponding dimension are recursively added upwards to obtain the complete label set.

4. The document retrieval method based on tag hierarchy evolution according to claim 1, characterized in that, The process of deriving a set of mutually exclusive search tags within the same semantic dimension of the search tag set includes: Determine the target semantic dimension to which the search tag set belongs; In the tag dictionary, all tags within the target semantic dimension, excluding the tags contained in the search tag set and all their descendant node tags, are filtered out and integrated to obtain the search opposing tag set.

5. The document retrieval method based on tag hierarchy evolution according to claim 4, characterized in that, The process involves performing a dual-mode filtering search on the already bound and stored ship-related documents based on the search tag set and the search opposite tag set, outputting the search results for ship-related documents that meet the requirements, including: In response to selecting the lenient mode, all ship domain documents in the already bound and stored ship domain documents that do not match the tags in the search opposition tag set are retained as the search results; In response to selecting strict mode, retain the ship domain documents that are already bound and stored, and that meet the condition of not matching the tags in the search opposite tag set and containing at least one tag in the search tag set, as the search results.

6. The document retrieval method based on tag hierarchy evolution according to claim 1, characterized in that, It also includes the dynamic maintenance steps for tags: Receive tag addition, deletion, and modification operation requests for the tag dictionary, and update the topology of the tag dictionary; Based on the manipulated tag, retrieve the corresponding shipbuilding domain document and synchronously update the associated tags of the shipbuilding domain document.

7. The document retrieval method based on tag hierarchy evolution according to claim 6, characterized in that, Receiving requests for tag addition, deletion, and modification operations in the tag dictionary, and updating the topology of the tag dictionary, includes: In response to a tag deletion request, after deleting the target tag, update the parent node of all child nodes of the target tag to the parent node of the target tag; Based on the deletion operation, the synchronous update of the associated tags of the shipbuilding document is triggered through an asynchronous message queue.

8. A document retrieval device with tag hierarchy evolution, characterized in that, include: The extraction module is used to extract candidate tags from text content based on a hybrid architecture of regular expression matching and large language model. The candidate tags are verified and cleaned by comparing them with a pre-built tag dictionary to obtain compliant tags. The text content is obtained by parsing the shipbuilding domain document to be processed. The tag dictionary stores multi-level tags and hierarchical relationships divided according to semantic dimensions. The extension module is used to perform bidirectional hierarchical expansion on the compliance tags according to the hierarchical structure of the tag dictionary to obtain a complete tag set, and bind and store the complete tag set with the corresponding shipbuilding domain document; The processing module is used to respond to the user's search request, extract query tags from the search request and perform the bidirectional hierarchical expansion to obtain a search tag set, and deduce a search opposite tag set that is semantically mutually exclusive with the search tag set within the same semantic dimension. The retrieval module is used to perform a dual-mode filtering retrieval on the already bound and stored ship-related documents based on the retrieval tag set and the retrieval opposite tag set, and output the retrieval results of the ship-related documents that meet the requirements.

9. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the document retrieval method with tag hierarchy evolution as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the document retrieval method based on tag hierarchy evolution as described in any one of claims 1 to 7.