A retrieval method for engineering consulting reports based on combined semantics and association matching

Through the joint semantic matching and association matching methods, combined with fine-tuned comparative learning model and BERT model, the accuracy and efficiency of title and paragraph search in engineering consulting reports are solved, and better search results are achieved, especially on Chinese data sets.

CN115858813BActive Publication Date: 2025-05-16BEIJING INFORMATION SCI & TECH UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211628660.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-05-16
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate and efficient retrieval of titles and paragraphs in engineering consulting reports, especially on Chinese data sets, which cannot effectively solve the problem of word-level matching and precise matching signals between queries and documents.

Method used

The combined semantic matching and association matching method is adopted to construct a text search corpus for engineering consulting reports, and the fine-tuned simCSE comparison learning model and Vanilla BERT model are used for semantic matching. The word-level granularity meaning vector representation is obtained by combining the SAT model, and the correlation match is performed through the DRMM deep text interaction model, and the final matching score is finally obtained through weighted fusion.

Benefits of technology

It realizes accurate and efficient search of the titles and paragraphs of the engineering consulting report, improves the search effect, especially on the Chinese data set, and improves the processing ability of word-level matching and precise matching signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115858813B_ABST
    Figure CN115858813B_ABST
Patent Text Reader

Abstract

The present invention relates to a text retrieval method for engineering consulting reports, so as to improve the problems of high labor cost and long compilation cycle in the process of writing engineering consulting reports, and comprises the following steps: constructing a text retrieval corpus for engineering consulting reports, using the corpus to fine-tune a simCSE comparative learning model, initializing a Vanilla BERT model with the obtained model parameters, and sending the text information of the corpus into the Vanilla BERT model to obtain a semantic matching score. The text information and keyword information are obtained through a SAT model to obtain a semantic primitive word vector representation of word level granularity and sent to a DRMM deep text interaction model to obtain an associated matching score. The obtained semantic matching score and associated matching score are normalized and weightedly fused to obtain a final matching score, and the text retrieval between the title and the paragraph is completed. The present invention combines a context vector representation and a text interaction matching method to effectively enhance the effect of text retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a text retrieval method in the field of engineering consulting reports, and in particular to retrieving paragraph contents that match the title of an engineering consulting report and are significantly helpful for writing the engineering consulting report. Background Art

[0002] Common text retrieval models can be roughly divided into two categories: classic statistical text retrieval models and deep text matching models. The classic statistical text retrieval model is based on the Bayesian hypothesis, which regards each word in the text as independent of each other, only considers the meaning of the word and loses the context information, so the effect is poor. The deep text matching model mainly includes the deep semantic representation matching model, the deep text interaction matching model and the joint text matching model.

[0003] The matching model based on text semantic representation aims to convert the relevance of query and document into semantic matching. It focuses on the semantic relationship of text and hopes to improve the feature extraction ability of the model by strengthening the language model of the encoding layer, thereby obtaining better text representation. However, this method has obvious disadvantages. Since the query and document are encoded and represented separately, the semantic focus is lost, and it is difficult to measure the importance of word-level granularity between the query and the document. The deep text interaction matching model builds the query and document interaction matrix based on the text embedding vector, focusing on the precise matching signal between the query and the document, the importance of the words in the query, and the diversified matching. It is better than the classic statistical retrieval model, but the retrieval effect of the model is affected by the quality of word representation. The joint text matching model obtains text vector representation through pre-trained models such as BERT, and obtains matching scores through the joint deep text interaction matching model. Combining the advantages of the two deep matching models, it has achieved effective results on English datasets.

[0004] The text retrieval corpus for engineering consulting reports is different from the English dataset. It has word granularity and character granularity. The text interactive matching model needs to consider the word-level matching between the query and the document. The Chinese pre-trained language model cannot effectively represent the word granularity vector, and it cannot solve the problem of accurate matching signals between the query and the document in the retrieval task. In addition, there is no relevant research on the text retrieval task for engineering consulting reports in the academic community. Summary of the invention

[0005] In view of the problems of high labor cost and long compilation cycle in the writing of engineering consulting reports, the present invention proposes a text retrieval method for engineering consulting reports, which realizes accurate and efficient retrieval of titles and paragraphs through joint semantic matching and association matching, and can effectively assist the writing of engineering consulting reports.

[0006] The present invention comprises the following steps:

[0007] 1. Construct a text retrieval corpus for engineering consulting reports;

[0008] 2. Use the corpus to fine-tune the simCSE contrastive learning model and use the obtained model parameters to initialize the Vanilla BERT model;

[0009] 3. Send the text information of the corpus into the Vanilla BERT model to obtain the semantic matching score;

[0010] 4. Use the SAT model to obtain the semantic vector representation of the original word at the word level.

[0011] 5. Send it to the DRMM deep text interaction model to obtain the associated matching score;

[0012] 6. Normalize the obtained semantic matching score and association matching score and perform weighted fusion to obtain the final matching score;

[0013] 7. Train the retrieval network model based on the training data and update the parameters, then extract text features on the test set and test it.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: compared with the existing text retrieval method, based on the engineering consulting report text retrieval corpus, the semantic matching score is obtained based on the character granularity in the semantic representation matching stage, and the association matching score is obtained based on the word granularity in the text interactive matching stage, and the final similarity score is obtained by weighted fusion after normalization, thereby obtaining a better effect; the sentence vector expression is optimized by the comparative learning method, and the trained simSCE model parameters are used to initialize the Vanilla BERT semantic representation matching model, and the effect is better than directly using the Vanilla BERT semantic representation matching model; keywords are annotated for each title text and a title keyword conversion dictionary is constructed, and keyword features are integrated into the text interactive matching stage, thereby improving the retrieval effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] By reading the detailed description of the preferred embodiment below, various other advantages and benefits will become clear to those of ordinary skill in the art. The accompanying drawings are only used for the purpose of illustrating the preferred embodiment and are not considered to be limitations of the present invention. In addition, the same reference symbols are used to represent the same components throughout the accompanying drawings.

[0016] In the attached picture:

[0017] Figure 1 It is a flow chart of an engineering consulting report retrieval method combining semantics and association matching according to the present invention;

[0018] Figure 2 It is a model structure diagram of an engineering consulting report retrieval method combining semantics and association matching according to the present invention;

[0019] Figure 3 It is a schematic diagram of the performance of different models on the engineering consulting report retrieval task. DETAILED DESCRIPTION

[0020] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0021] Figure 1 and Figure 2 They are respectively a flow chart and a model structure diagram of an engineering consulting report retrieval method combining semantics and association matching according to the present invention.

[0022] The steps include:

[0023] 1. Use the pre-trained language model BERT to encode the input sentence, that is ,in are the BERT model parameters, and then the simCSE unsupervised contrastive learning method is used to fine-tune all parameters.

[0024] 2. Fine-tune the BERT model parameters trained by simCSE. First, the title is represented as ,in For the title characters, is the length of the title, and the paragraph is represented by ,in For paragraph characters, is the length of the paragraph. The character sequence form is used as the input of BERT, and its length is , input it into the BERT model to obtain the vectorized representation of the title and paragraph , as shown in formula (1).

[0025]

[0026] 3. The vector representation output by BERT combines the three parts of Token Embedding, Segment Embedding, and Position Embedding, so it contains as much text context information as possible. Represents the input sequence Chinese characters BERT vector representation of . As a special character of BERT, it can obtain the context information representation of the entire text, which plays an important role in classification tasks. Take it out separately and integrate it into the subsequent joint ranking model. A linear combination layer is stacked on the token to obtain the semantic matching score. The optimization target is the pairwise hinge loss, as shown in formula (2):

[0027]

[0028] This formula measures the difference between positive and negative samples in the pairwise scenario, where represents the preset threshold value. Represents the input query, represents the positive sample, represents negative samples, It represents the similarity between two vectors. The meaning of this formula is that the loss will be reduced to 0 only when the input query is similar enough to the positive sample. Otherwise, the loss will become very large if the input query is less similar to the positive sample or more similar to the negative sample.

[0029] 4. Collect keywords corresponding to the title The set of keywords corresponding to the paragraph As semantic supplementary information of the text. In the text interactive matching stage, when the search text is input, the keyword set is spliced ​​after the corresponding text and also passed into the model as an input sequence. The annotation of the keyword set of the title and paragraph text is completed by the company's professionals who write engineering consulting reports all year round. And build a title keyword conversion dictionary. If the keyword of the current title is not in the paragraph keyword set, the title keyword is converted and replaced with its corresponding large tag to improve the matching effect.

[0030] 5. After removing special characters and stop words from the text, the Jieba word segmentation tool is used to segment the text based on the keyword set of the retrieval corpus, which helps the deep text interaction model identify the exact matching signals between the title and the paragraph.

[0031] 6. Use semantic word vectors to achieve distributed text representation of retrieval corpus and keywords in the association matching stage. Use the attention mechanism to consider the semantic information of words while considering the context information, which effectively improves the quality of word representation.

[0032] 7. Use the DRMM model to interact with the semantic vector representations of the title and paragraph after the fusion of keywords to obtain the local interaction matrix, and map the variable-length local interactions into a fixed-length matching histogram. Then use the feedforward matching network to generate the term matching score according to the hierarchical matching mode. Finally, the score of each query item in the acquired title is further aggregated through the term gating unit to generate the final matching score.

[0033] 8. The semantic matching score score1 is obtained through the deep semantic representation part, and the associated matching score score2 is obtained through the deep text interaction part. After normalizing score1 and score2, they are weighted fused to obtain the final matching score score, completing the text retrieval between the title and the paragraph, as shown in formula (3).

[0034]

[0035] Embodiment 1:

[0036] The experimental results in this embodiment are obtained by using a text retrieval corpus for engineering consulting reports as a data set and testing on the data set. The technical effects of the present invention are as follows.

[0037] Figure 3 The following is a diagram of the retrieval effect of different models on the text retrieval corpus for engineering consulting reports. Among them, the joint ranking models Vanilla Bert, DRMM, KNRM and PACRR based on the pre-trained language model BERT in CEDR are used as the baseline models of this experiment. CEDR generates word-level granular context vector representations of titles and paragraphs through the pre-trained model, uses cosine similarity to build a word-level semantic similarity matrix for titles and paragraphs, and combines the deep text interactive matching model (DRMM, PACRR and KNRM) to obtain the matching score.

[0038] from Figure 3It can be seen that CEDR-Vanilla Bert is a semantic representation matching model, which has the worst effect among all the methods; while among the other three joint matching models of CEDR, CEDR-DRMM has the best effect, which is 2.78 percentage points higher than CEDR-VanillaBert. Although this method has achieved good results by combining the semantic matching model and the text interaction matching model, the rationality of using word vectors for correlation matching in Chinese data sets needs to be studied. The present invention obtains the semantic matching score of the text retrieval corpus based on word vector calculation, obtains the association matching score based on word vector calculation, and obtains the final retrieval score by weighted fusion of the semantic matching score and the association matching score. The final retrieval result has an accuracy of 80.14% on the P@20 indicator, which is 4.03 percentage points higher than the CEDR-DRMM model. The effectiveness of the present invention has been verified on the text retrieval corpus for engineering consulting reports.

[0039] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for engineering consulting report retrieval based on combined semantics and association matching, characterized in that: The following steps are involved: Construct a text retrieval corpus for engineering consulting reports, and perform keyword annotation on titles and paragraphs in the retrieval corpus; Use the corpus to fine-tune the simCSE contrastive learning model and use the obtained model parameters to initialize the Vanilla BERT model; Feed the text information of the corpus into the Vanilla BERT model to obtain the semantic matching score; The text information and keyword information are passed through the SAT model to obtain the semantic vector representation of the word level granularity; Send it to the DRMM deep text interaction model to get the associated matching score; The obtained semantic matching score and association matching score are normalized and weighted fused to obtain the final matching score; The retrieval network model is trained and the parameters are updated based on the training data, and then text features are extracted and tested on the test set.

2. The engineering consulting report retrieval method combining semantics and association matching according to claim 1, characterized in that: The titles and paragraphs in the retrieval corpus are encoded using the pre-trained language model BERT, i.e., h = f_θ(x), where f_θ is the BERT model parameter, and then all parameters are fine-tuned using the simCSE unsupervised contrastive learning method.

3. The engineering consulting report retrieval method combining semantics and association matching as claimed in claim 2, characterized in that: Spell the title and paragraph into Tokens = {[CLS], q1, q2, ..., q n , [SEP], d1, d2, ..., d m , [SEP]}, where [CLS] represents the classifier, [SEP] represents the separator, q represents the title, n represents the length of the title, d represents the paragraph, and m represents the length of the paragraph. The BERT model parameters trained by simCSE are fed into the model for fine-tuning to obtain the vectorized representation of the title and paragraph. The E stands for vectorization.

4. The engineering consulting report retrieval method combining semantics and association matching as claimed in claim 3, characterized in that: A linear combination layer is stacked on the classifier [CLS] token to get the semantic matching score, and the model is optimized by pairwise hinge loss, calculated as: loss=max(0,margin-<u,d+> +<u,d-> ) This formula measures the difference between positive and negative samples in the pairwise scenario, where margin represents the preset threshold, u represents the input query, d+ represents the positive sample, d- represents the negative sample, and <> represents the similarity between the two vectors. This formula means that the loss will only drop to 0 when the input query is similar enough to the positive sample. Otherwise, the less similar it is to the positive sample or the more similar it is to the negative sample, the larger the loss will be.

5. The engineering consulting report retrieval method combining semantics and association matching as claimed in claim 1, characterized in that: Set the keywords corresponding to the title The set of keywords corresponding to the paragraph As the semantic supplementary information of the text, qk represents the title keywords, l is the number of title keywords, d k represents the paragraph keywords, and t is the number of paragraph keywords; And build a title keyword conversion dictionary. If the keyword of the current title is not in the paragraph keyword set, the title keyword will be converted and replaced with its corresponding large tag to improve the matching effect.

6. The engineering consulting report retrieval method combining semantics and association matching as claimed in claim 5, characterized in that: The semantic primitive word vector is used to realize the distributed text representation of retrieval corpus and keywords in the association matching stage, which effectively improves the quality of word representation.

7. The engineering consulting report retrieval method combining semantics and association matching as claimed in claim 6, characterized in that: The DRMM model is used to interact with the semantic vector representations of the title and paragraph after the fusion of keywords to obtain the local interaction matrix, and the variable-length local interactions are mapped to a fixed-length matching histogram. Then, the feedforward matching network is used to generate the matching score of the term according to the hierarchical matching mode. Finally, the score of each query item in the acquired title is further aggregated through the term gating unit to generate the final matching score.

8. The engineering consulting report retrieval method combining semantics and association matching as claimed in claim 4 or 7, characterized in that: The semantic matching score is obtained through the deep semantic representation part, and the association matching score is obtained through the deep text interaction part. The two types of scores are normalized and weighted fused to obtain the final matching score, completing the text retrieval between titles and paragraphs.

Citation Information

Patent Citations

  • Long text retrieval model based on comparative learning

    CN114201581A

  • Text semantic retrieval method and device for science and technology resource information of experts and scholars

    CN114840645A