A multi-level matching method and system for construction project documents

CN120429417BActive Publication Date: 2026-09-29NINGXIA JIAOJIAN TRANSPORTATION TECH RES INST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510328362.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-29
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

传统的基于关键词匹配的检索方法,难以充分识别和定位这些层次信息,从而影响匹配精度和效率,尤其在工程项目管理中,当文档包含多个层次(如标题、段落、子段落、图表、施工进度表等)时,传统文档匹配方法,精确度低,适用性差

Benefits of technology

[0028]提高匹配精度:本发明通过自定义分词器、层次化匹配和深度语义分析等多重技术手段,能够在多个层级、多个维度上进行精准的文档匹配。特别是在建筑工程文档中,能够精准处理工程图纸、设计方案、施工进度等多种格式和内容。该系统的匹配精度可达到98%以上。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429417B_ABST
    Figure CN120429417B_ABST
Patent Text Reader

Abstract

The application discloses a multi-level matching method and system for construction engineering project documents, and the method comprises the following steps: setting a word segmenter in the field of construction engineering according to a term dictionary in the field of construction engineering; acquiring a query keyword of a user, and performing word segmentation on the query keyword by using a jieba word segmentation tool and in combination with the word segmenter in the field of construction engineering; performing matching retrieval on construction engineering project documents according to a word segmentation result, if the retrieval result contains corresponding construction engineering project documents, then performing whole-word matching on the corresponding construction engineering project documents, if the matching is successful, then returning a matching result, otherwise, performing matching through a tag process, and if no result is found after the matching through the tag process, then performing semantic matching through a BERT deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of architectural engineering document processing technology, and more specifically, relates to a multi-level matching method and system for architectural engineering project documents. Background Technology

[0002] With the continuous improvement of informatization in the construction industry, project-related documents are widely generated and stored on various terminals, servers, and cloud platforms. In this era of information explosion, how to efficiently and accurately retrieve relevant content from a vast document library has become a key challenge for the application of information technology in the construction industry.

[0003] Existing document matching technologies largely rely on keyword matching or full-text search solutions. While these methods achieve high retrieval efficiency, they fail to fully consider the complex structure and multi-layered information of construction engineering documents. In the construction field, documents not only contain basic text content but also diverse information such as design drawings, construction processes, material standards, and project schedules. These documents have complex hierarchical structures, with each level often containing different keywords and themes. Traditional keyword-based retrieval methods struggle to fully identify and locate these hierarchical information levels, thus affecting matching accuracy and efficiency. This is especially true in project management, where traditional document matching methods suffer from low accuracy and poor applicability when documents contain multiple levels (such as titles, paragraphs, sub-paragraphs, charts, and construction schedules).

[0004] Furthermore, traditional methods often overlook the semantic complexity within documents, leading to inaccurate matching results and prolonged response times when searching complex documents, thus impacting overall system efficiency. Therefore, designing an efficient and accurate document matching technology solution is particularly important, considering the multi-layered and multi-dimensional characteristics of documents in the construction engineering field. Summary of the Invention

[0005] To address the above technical problems, this invention proposes a multi-level matching method for architectural engineering project documents, comprising:

[0006] Based on a terminology dictionary for the field of construction engineering, set up a word segmenter for the field of construction engineering.

[0007] Obtain the user's query keywords, and segment the query keywords using the jieba word segmentation tool in conjunction with a word segmenter for the construction engineering field;

[0008] The document is matched and retrieved based on the word segmentation results. If the search results contain the corresponding document, the document is matched with all words. If the match is successful, the matching result is returned. Otherwise, the matching is performed through the tag process. If no result is found after the tag process, the semantic matching is performed through the BERT deep learning model, and the final matching result is returned.

[0009] Furthermore, matching and retrieving construction project documents based on word segmentation results includes: performing fuzzy matching retrieval of construction project documents using Elasticsearch.

[0010] Furthermore, matching via the tag process also includes matching the document content of the construction project documents according to predefined tag rules in the construction project documents.

[0011] Furthermore, the matching and retrieval of construction project documents based on word segmentation results includes: setting up a multi-level matching retrieval model, calculating the matching score of each level of the construction project documents, and using the construction project documents corresponding to the levels whose matching scores exceed the preset level matching threshold as the retrieval results, thereby performing full word matching on the corresponding construction project documents.

[0012] Furthermore, the multi-level matching retrieval model includes:

[0013]

[0014] in, For the matching score at the i-th level, m i γ is the number of sub-levels at the i-th level. i,k Let be the weight of the document content at the k-th sub-level of the i-th level. The document content of the k-th sub-level of the i-th level. The matching function with the query keyword Q, β i Let be the adjustment factor for the i-th level.

[0015] Furthermore, the document content of the k-th sub-level of the i-th level. Matching function with query keyword Q include:

[0016]

[0017] in, To query the j-th term of keyword Q The document content of the i-th level and the k-th sub-level The lth term The similarity function, vQ To query the semantic embedding vector of the keyword Q, The document content of the k-th sub-level of the i-th level. semantic embedding vector, To query the j-th term of keyword Q TF-IDF value, The document content of the k-th sub-level of the i-th level. The lth term The TF-IDF value.

[0018] Furthermore, query the j-th term of keyword Q. The document content of the i-th level and the k-th sub-level The lth term similarity function include:

[0019]

[0020] Where α′ is the similarity function The first adjustment factor To query the j-th term of keyword Q word frequency, The document content of the k-th sub-level of the i-th level. The lth term The word frequency, β′ is the similarity function. The second adjustment factor, To query the j-th term of keyword Q Location, The document content of the k-th sub-level of the i-th level. The lth term The position, γ′ is the similarity function The third adjustment factor, To query the j-th term of keyword Q depth, The document content of the k-th sub-level of the i-th level. The lth term The depth.

[0021] This invention also proposes a multi-level matching system for construction project documents, comprising:

[0022] The word segmentation module is used to set up a word segmenter for the field of construction engineering based on a terminology dictionary for the construction engineering field.

[0023] The word segmentation module is used to obtain the user's query keywords and segment the query keywords using the jieba word segmentation tool in combination with a word segmenter for the construction engineering field.

[0024] The matching module is used to match and retrieve construction project documents based on the word segmentation results. If the search results contain the corresponding construction project documents, the corresponding construction project documents are matched with full words. If the match is successful, the matching result is returned. Otherwise, the matching is performed through the tag process. If no results are found after the tag process, the semantic matching is performed through the BERT deep learning model, and the final matching result is returned.

[0025] Furthermore, matching and retrieving construction project documents based on word segmentation results includes: performing fuzzy matching retrieval of construction project documents using Elasticsearch.

[0026] Furthermore, matching via the tag process also includes matching the document content of the construction project documents according to predefined tag rules in the construction project documents.

[0027] In summary, the technical solutions conceived by this invention have the following beneficial effects compared with the prior art:

[0028] Improved matching accuracy: This invention utilizes multiple technologies, including a custom word segmenter, hierarchical matching, and deep semantic analysis, to achieve accurate document matching across multiple levels and dimensions. Particularly effective in architectural engineering documents, it can accurately handle various formats and content such as engineering drawings, design schemes, and construction schedules. The system's matching accuracy can reach over 98%.

[0029] Accelerated Response Speed: This invention significantly improves system response speed by reducing redundant calculations through hierarchical processing of the matching process. In building project document retrieval, queries can be completed within 2 seconds, greatly enhancing project management efficiency.

[0030] Adaptable to complex document structures: The matching algorithm of this invention is not only applicable to simple documents, but also capable of handling complex engineering project document structures, such as design drawings, construction flowcharts, schedules, tables, and charts. By dynamically recognizing the hierarchical structure of the document, it ensures accurate matching even in complex construction engineering documents.

[0031] Optimized User Experience: The system described in this invention provides fast and accurate document matching, especially in engineering project management, enabling users to quickly obtain relevant documents, improving work efficiency and decision-making speed. Furthermore, the user interface is user-friendly and conforms to the usage habits of professionals in the construction industry, enhancing the overall user experience.

[0032] Flexible scalability: The system described in this invention adopts a modular design, which can customize the word segmenter and matching algorithm according to the needs of different engineering projects. For example, for different types of construction projects (such as road construction, building construction, municipal projects, etc.), the system can flexibly adjust the matching rules and tags, and has strong scalability. Attached Figure Description

[0033] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention;

[0034] Figure 2 This is a system structure diagram of Embodiment 2 of the present invention. Detailed Implementation

[0035] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0036] The method provided by this invention can be implemented in a terminal environment that may include one or more of the following components: a processor, a storage medium, and a display screen. The storage medium stores at least one instruction, which is loaded and executed by the processor to implement the method described in the following embodiments.

[0037] A processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts of the terminal, and performs various functions and processes data by running or executing instructions, programs, code sets or instruction sets stored in the storage medium, and by calling data stored in the storage medium.

[0038] Storage media can include random access memory (RAM) or read-only memory (ROM). Storage media can be used to store instructions, programs, code, code sets, or instructions.

[0039] The display screen is used to show the user interface of each application.

[0040] In addition, those skilled in the art will understand that the structure of the terminal described above does not constitute a limitation on the terminal. The terminal may include more or fewer components, or combine certain components, or have different component arrangements. For example, the terminal may also include radio frequency circuits, input units, sensors, audio circuits, power supplies, and other components, which will not be described in detail here.

[0041] Definitions:

[0042] FastAPI: It is a modern Python web framework for building APIs, featuring high performance and ease of use, making it particularly suitable for scenarios that require rapid development and efficient request processing.

[0043] Elasticsearch (ES) is a distributed search engine based on Lucene, widely used in scenarios such as log analysis, full-text search, and real-time data analysis. It supports fast text search, complex queries, aggregation analysis, and can scale horizontally to support real-time processing of large-scale data.

[0044] Tag matching process: In Elasticsearch, tag matching usually refers to an operation that matches or filters documents based on tags.

[0045] Example 1

[0046] like Figure 1 This embodiment proposes a multi-level matching method for construction project documents. The technical solution of this embodiment combines the hierarchical matching characteristics of construction project document content. Through multi-level, global, and local matching processing of document content, accurate document retrieval and matching are achieved. The specific steps are as follows:

[0047] Step 1: To improve the accuracy of matching professional terms in construction engineering industry documents, this embodiment designs an industry-specific word segmenter (a word segmenter for the construction engineering field) based on a terminology dictionary for the construction engineering field. The word segmenter assigns scores to construction engineering terms (such as building materials, construction technology, engineering standards, etc.) according to their parts of speech, thereby improving the accuracy of word segmentation and ensuring efficient recognition of terms.

[0048] Step 2: Obtain the user's query keywords, call the FastAPI interface to start the matching environment and initialize the matching process, and prepare to process the user request.

[0049] Step 3: Use the jieba word segmentation tool, combined with a word segmenter for the construction engineering field, to segment the user's query keywords. The word segmenter for the construction engineering field described in this embodiment can optimize for the terminology in the construction engineering field, thereby improving the matching accuracy. In particular, in engineering construction and design-related documents, accurate identification of professional terms is crucial.

[0050] Step 4: Perform a matching search on the word segmentation results. If the matching result is not empty, proceed to Step 5; otherwise, proceed to Step 6.

[0051] Specifically, matching and retrieving construction project documents based on word segmentation results includes: performing fuzzy matching retrieval of construction project documents using Elasticsearch.

[0052] Specifically, to make the matching results more accurate, this embodiment designs a multi-level matching retrieval model based on fuzzy matching, calculates the matching score of each level of the construction project document, and uses the construction project document corresponding to the level whose matching score exceeds the preset level matching threshold as the retrieval result, thereby performing full word matching on the corresponding construction project document.

[0053] Specifically, the multi-level matching retrieval model includes:

[0054]

[0055] in, For the matching score of the i-th level (which can be the title, paragraphs, charts, etc. of the architectural project document), m i γ is the number of sub-levels at the i-th level. i,k This represents the weight of the document content at the k-th sub-level of the i-th level (which can be subheadings, paragraphs, sub-tables, etc. in the architectural project document). The document content of the k-th sub-level of the i-th level. The matching function with the query keyword Q, β i Let be the adjustment factor for the i-th level.

[0056] Specifically, the document content of the k-th sub-level of the i-th level. Matching function with query keyword Q include:

[0057]

[0058] in, To query the j-th term of keyword Q The document content of the i-th level and the k-th sub-level The lth term The similarity function, v Q To query the semantic embedding vector of the keyword Q, The document content of the k-th sub-level of the i-th level. semantic embedding vector, To query the j-th term of keyword Q TF-IDF value, The document content of the k-th sub-level of the i-th level. The lth term The TF-IDF value.

[0059] Specifically, query the j-th term of keyword Q. The document content of the i-th level and the k-th sub-level The lth term similarity function include:

[0060]

[0061] Where α′ is the similarity function The first adjustment factor To query the j-th term of keyword Q word frequency, The document content of the k-th sub-level of the i-th level. The lth term The word frequency, β′ is the similarity function. The second adjustment factor, To query the j-th term of keyword Q Location, The document content of the k-th sub-level of the i-th level. The lth term The position, γ′ is the similarity function The third adjustment factor, To query the j-th term of keyword Q depth, The document content of the k-th sub-level of the i-th level. The lth term The depth.

[0062] Step 5: If the search results from Step 4 contain a document, perform a full-word match. If the match is successful, return the matching result; otherwise, proceed to Step 6.

[0063] Step Six: If Steps Four and Five fail to return valid results, proceed to the Tag matching process. During this process, the document content is matched against predefined tag rules in the architectural engineering document. These tag rules are meticulously designed based on the document's content hierarchy and theme to ensure matching of multi-level information. If the tag matching result is not empty, return; otherwise, proceed to Step Seven.

[0064] The predefined rules in step six are to add extra tags to documents in addition to hierarchical information such as keywords, titles, etc. For example, "cleaning up the scene" and "cleaning up the site" have the same meaning. If there is a document with the title "cleaning up the scene", it will also be tagged with "cleaning up the site". Then, the next time there is a match for "cleaning up the site", it will match the document "cleaning up the scene".

[0065] Step 7: If no valid matching results are returned in the above steps, semantic matching is performed based on the BERT deep learning model. This model determines whether the document content and the semantic similarity of the user query meet the user's needs by deeply analyzing the document content. If the matching similarity exceeds the set threshold (e.g., 0.8), the matching result is returned; otherwise, an empty string is returned.

[0066] Example 2

[0067] like Figure 2 As shown, this embodiment of the invention also provides a multi-level matching system for construction project documents, including:

[0068] The word segmentation module is used to set up a word segmenter for the field of construction engineering based on a terminology dictionary for the construction engineering field.

[0069] The word segmentation module is used to obtain the user's query keywords and segment the query keywords using the jieba word segmentation tool in combination with a word segmenter for the construction engineering field.

[0070] The matching module is used to match and retrieve construction project documents based on the word segmentation results. If the search results contain the corresponding construction project documents, the corresponding construction project documents are matched with full words. If the match is successful, the matching result is returned. Otherwise, the matching is performed through the tag process, and the final matching result is returned.

[0071] Specifically, matching and retrieving construction project documents based on word segmentation results includes: performing fuzzy matching retrieval of construction project documents using Elasticsearch.

[0072] Specifically, to make the matching results more accurate, this embodiment designs a multi-level matching retrieval model based on fuzzy matching, calculates the matching score of each level of the construction project document, and uses the construction project document corresponding to the level whose matching score exceeds the preset level matching threshold as the retrieval result, thereby performing full word matching on the corresponding construction project document.

[0073] Specifically, the multi-level matching retrieval model includes:

[0074]

[0075] in, For the matching score at the i-th level, m i γ is the number of sub-levels at the i-th level. i,k Let be the weight of the document content at the k-th sub-level of the i-th level. The document content of the k-th sub-level of the i-th level. The matching function with the query keyword Q, β iLet be the adjustment factor for the i-th level.

[0076] Specifically, the document content of the k-th sub-level of the i-th level. Matching function with query keyword Q include:

[0077]

[0078] in, To query the j-th term of keyword Q The document content of the i-th level and the k-th sub-level The lth term The similarity function, v Q To query the semantic embedding vector of the keyword Q, The document content of the k-th sub-level of the i-th level. semantic embedding vector, To query the j-th term of keyword Q TF-IDF value, The document content of the k-th sub-level of the i-th level. The lth term The TF-IDF value.

[0079] Specifically, query the j-th term of keyword Q. The document content of the i-th level and the k-th sub-level The lth term similarity function include:

[0080]

[0081] Where α′ is the similarity function The first adjustment factor To query the j-th term of keyword Q word frequency, The document content of the k-th sub-level of the i-th level. The lth term The word frequency, β′ is the similarity function. The second adjustment factor, To query the j-th term of keyword Q Location, The document content of the k-th sub-level of the i-th level. The lth term The position, γ′ is the similarity function The third adjustment factor, To query the j-th term of keyword Q depth, The document content of the k-th sub-level of the i-th level. The lth term The depth.

[0082] Specifically, matching via the tag process also includes matching the document content of the construction project documents according to predefined tag rules in the construction project documents.

[0083] Specifically, if no results are found after matching through the tag process, semantic matching is performed using the BERT deep learning model.

[0084] Example 3

[0085] This invention also proposes a storage medium storing multiple instructions for implementing the multi-level matching method for architectural engineering project documents.

[0086] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0087] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: setting up a word segmenter for the field of construction engineering based on a terminology dictionary for the field of construction engineering;

[0088] Obtain the user's query keywords, and segment the query keywords using the jieba word segmentation tool in conjunction with a word segmenter for the construction engineering field;

[0089] The document is matched and retrieved based on the word segmentation results. If the search results contain the corresponding document, the document is matched with all words. If the match is successful, the matching result is returned. Otherwise, the matching is performed through the tag process and the final matching result is returned.

[0090] Specifically, matching and retrieving construction project documents based on word segmentation results includes: performing fuzzy matching retrieval of construction project documents using Elasticsearch.

[0091] Specifically, to make the matching results more accurate, this embodiment designs a multi-level matching retrieval model based on fuzzy matching, calculates the matching score of each level of the construction project document, and uses the construction project document corresponding to the level whose matching score exceeds the preset level matching threshold as the retrieval result, thereby performing full word matching on the corresponding construction project document.

[0092] Specifically, the multi-level matching retrieval model includes:

[0093]

[0094] in, For the matching score at the i-th level, m i γ is the number of sub-levels at the i-th level. i,k Let be the weight of the document content at the k-th sub-level of the i-th level. The document content of the k-th sub-level of the i-th level. The matching function with the query keyword Q, β i Let be the adjustment factor for the i-th level.

[0095] Specifically, the document content of the k-th sub-level of the i-th level. Matching function with query keyword Q include:

[0096]

[0097] in, To query the j-th term of keyword Q The document content of the i-th level and the k-th sub-level The lth term The similarity function, v Q To query the semantic embedding vector of the keyword Q, The document content of the k-th sub-level of the i-th level. semantic embedding vector, To query the j-th term of keyword Q TF-IDF value, The document content of the k-th sub-level of the i-th level. The lth term The TF-IDF value.

[0098] Specifically, query the j-th term of keyword Q. The document content of the i-th level and the k-th sub-level The lth term similarity function include:

[0099]

[0100] Where α′ is the similarity function The first adjustment factor To query the j-th term of keyword Q word frequency, The document content of the k-th sub-level of the i-th level. The lth term The word frequency, β′ is the similarity function. The second adjustment factor, To query the j-th term of keyword Q Location, The document content of the k-th sub-level of the i-th level. The lth term The position, γ′ is the similarity function The third adjustment factor, To query the j-th term of keyword Q depth, The document content of the k-th sub-level of the i-th level. The lth term The depth.

[0101] Specifically, matching via the tag process also includes matching the document content of the construction project documents according to predefined tag rules in the construction project documents.

[0102] Specifically, if no results are found after matching through the tag process, semantic matching is performed using the BERT deep learning model.

[0103] Example 4

[0104] This invention also proposes an electronic device, including a processor and a storage medium connected to the processor. The storage medium stores multiple instructions, which can be loaded and executed by the processor to enable the processor to execute the multi-level matching method for architectural engineering project documents.

[0105] Specifically, the electronic device in this embodiment can be a computer terminal, which may include one or more processors and a storage medium.

[0106] The storage medium can be used to store software programs and modules, such as the multi-level matching method for construction project documents in this embodiment of the invention. The processor executes the software programs and modules stored in the storage medium to perform various functional applications and data processing, thus realizing the aforementioned multi-level matching method for construction project documents. The storage medium may include high-speed random access storage media, and may also include non-volatile storage media, such as one or more magnetic storage systems, flash memory, or other non-volatile solid-state storage media. In some instances, the storage medium may further include storage media remotely configured relative to the processor, which can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0107] The processor can call the information and application stored in the storage medium through the transmission system to perform the following steps: set up a word segmenter for the field of construction engineering according to a terminology dictionary for the field of construction engineering;

[0108] Obtain the user's query keywords, and segment the query keywords using the jieba word segmentation tool in conjunction with a word segmenter for the construction engineering field;

[0109] The document is matched and retrieved based on the word segmentation results. If the search results contain the corresponding document, the document is matched with all words. If the match is successful, the matching result is returned. Otherwise, the matching is performed through the tag process and the final matching result is returned.

[0110] Specifically, matching and retrieving construction project documents based on word segmentation results includes: performing fuzzy matching retrieval of construction project documents using Elasticsearch.

[0111] Specifically, to make the matching results more accurate, this embodiment designs a multi-level matching retrieval model based on fuzzy matching, calculates the matching score of each level of the construction project document, and uses the construction project document corresponding to the level whose matching score exceeds the preset level matching threshold as the retrieval result, thereby performing full word matching on the corresponding construction project document.

[0112] Specifically, the multi-level matching retrieval model includes:

[0113]

[0114] in, For the matching score at the i-th level, m i γ is the number of sub-levels at the i-th level. i,kLet be the weight of the document content at the k-th sub-level of the i-th level. The document content of the k-th sub-level of the i-th level. The matching function with the query keyword Q, β i Let be the adjustment factor for the i-th level.

[0115] Specifically, the document content of the k-th sub-level of the i-th level. Matching function with query keyword Q include:

[0116]

[0117] in, To query the j-th term of keyword Q The document content of the i-th level and the k-th sub-level The lth term The similarity function, v Q To query the semantic embedding vector of the keyword Q, The document content of the k-th sub-level of the i-th level. semantic embedding vector, To query the j-th term of keyword Q TF-IDF value, The document content of the k-th sub-level of the i-th level. The lth term The TF-IDF value.

[0118] Specifically, query the j-th term of keyword Q. The document content of the i-th level and the k-th sub-level The lth term similarity function include:

[0119]

[0120] Where α′ is the similarity function The first adjustment factor To query the j-th term of keyword Q word frequency, The document content of the k-th sub-level of the i-th level. The lth term The word frequency, β′ is the similarity function. The second adjustment factor, To query the j-th term of keyword Q Location, The document content of the k-th sub-level of the i-th level. The lth term The position, γ′ is the similarity function The third adjustment factor, To query the j-th term of keyword Q depth, The document content of the k-th sub-level of the i-th level. The lth term The depth.

[0121] Specifically, matching via the tag process also includes matching the document content of the construction project documents according to predefined tag rules in the construction project documents.

[0122] Specifically, if no results are found after matching through the tag process, semantic matching is performed using the BERT deep learning model.

[0123] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0124] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0125] In the several embodiments provided by this invention, it should be understood that the disclosed technical content can be implemented in other ways. The system embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0126] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0127] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0128] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, optical disks, and other media capable of storing program code.

[0129] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A multi-level matching method for construction project documents, characterized in that, include: Based on a terminology dictionary for the field of construction engineering, set up a word segmenter for the field of construction engineering. Obtain the user's query keywords, and segment the query keywords using the jieba word segmentation tool in conjunction with a word segmenter for the construction engineering field; The document is matched and retrieved based on the word segmentation results. If the search results contain the corresponding document, the document is matched with the whole word. If the match is successful, the matching result is returned. Otherwise, the matching is performed through the tag process. If no result is found after the tag process, the semantic matching is performed through the BERT deep learning model and the final matching result is returned. The matching and retrieval of construction project documents based on word segmentation results includes: setting up a multi-level matching retrieval model, calculating the matching score of each level of construction project documents, and taking the construction project documents corresponding to the level whose matching score exceeds the preset level matching threshold as the retrieval results, thereby performing full word matching on the corresponding construction project documents; The multi-level matching retrieval model includes: in, For the first Matching scores at each level, For the first The number of sub-levels at each level. For the first The first level The weight of document content at each sub-level. For the first The first level Sub-level document content With search keywords The matching function, For the first Adjustment factors at each level; No. The first level Sub-level document content With search keywords Matching function include: in, To search for keywords The Term With the The first level Sub-level document content The Term The similarity function, To search for keywords semantic embedding vector, For the first The first level Sub-level document content semantic embedding vector, To search for keywords The Term TF-IDF value, For the first The first level Sub-level document content The Term TF-IDF value; Search keywords The Term With the The first level Sub-level document content The Term similarity function include: in, Similarity function The first adjustment factor, To search for keywords The Term word frequency, For the first The first level Sub-level document content The Term word frequency, Similarity function The second adjustment factor, To search for keywords The Term Location, For the first The first level Sub-level document content The Term Location, Similarity function The third adjustment factor, To search for keywords The Term depth, For the first The first level Sub-level document content The Term The depth.

2. The multi-level matching method for construction project documents as described in claim 1, characterized in that, Matching and retrieving construction project documents based on word segmentation results includes: performing fuzzy matching retrieval of construction project documents using Elasticsearch.

3. The multi-level matching method for construction project documents as described in claim 1, characterized in that, Matching via the tag process also includes matching the document content of architectural engineering project documents according to predefined tag rules in the architectural engineering documents.

4. A multi-level matching system for construction project documents, characterized in that, include: The word segmentation module is used to set up a word segmenter for the field of construction engineering based on a terminology dictionary for the construction engineering field. The word segmentation module is used to obtain the user's query keywords and segment the query keywords using the jieba word segmentation tool in combination with a word segmenter for the construction engineering field. The matching module is used to match and retrieve construction project documents based on the word segmentation results. If the search results contain the corresponding construction project documents, the corresponding construction project documents are matched with full words. If the match is successful, the matching result is returned. Otherwise, the matching is performed through the tag process. If no results are found after the tag process, the semantic matching is performed through the BERT deep learning model, and the final matching result is returned. The matching and retrieval of construction project documents based on word segmentation results includes: setting up a multi-level matching retrieval model, calculating the matching score of each level of construction project documents, and taking the construction project documents corresponding to the level whose matching score exceeds the preset level matching threshold as the retrieval results, thereby performing full word matching on the corresponding construction project documents; The multi-level matching retrieval model includes: in, For the first Matching scores at each level, For the first The number of sub-levels at each level. For the first The first level The weight of document content at each sub-level. For the first The first level Sub-level document content With search keywords The matching function, For the first Adjustment factors at each level; No. The first level Sub-level document content With search keywords Matching function include: in, To search for keywords The Term With the The first level Sub-level document content The Term The similarity function, To search for keywords semantic embedding vector, For the first The first level Sub-level document content semantic embedding vector, To search for keywords The Term TF-IDF value, For the first The first level Sub-level document content The Term TF-IDF value; Search keywords The Term With the The first level Sub-level document content The Term similarity function include: in, Similarity function The first adjustment factor, To search for keywords The Term word frequency, For the first The first level Sub-level document content The Term word frequency, Similarity function The second adjustment factor, To search for keywords The Term Location, For the first The first level Sub-level document content The Term Location, Similarity function The third adjustment factor, To search for keywords The Term depth, For the first The first level Sub-level document content The Term The depth.

5. A multi-level matching system for construction project documents as described in claim 4, characterized in that, Matching and retrieving construction project documents based on word segmentation results includes: performing fuzzy matching retrieval of construction project documents using Elasticsearch.

6. The multi-level matching system for construction project documents as described in claim 4, characterized in that, Matching via the tag process also includes matching the document content of architectural engineering project documents according to predefined tag rules in the architectural engineering documents.

Citation Information

Patent Citations

  • Dialogue method and device and electronic equipment

    CN111708872A

  • ES-based associated document retrieval method

    CN113486156A