Engineering document classification method and device based on text embedding

Through the text embedding method, vectorized processing and semantic similarity calculation of the title of engineering documents is solved, and the problems of low accuracy and poor universality of document classification in the prior art are achieved, and more efficient and accurate document classification is achieved.

CN120234412APending Publication Date: 2025-07-01MCC CAPITAL ENGINEERING & RESEARCH INC LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311808774.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art relies on word segmentation algorithms in engineering document classification, resulting in loss of semantic expression, low classification accuracy, and inability to apply non-text documents, which are poor universal.

Method used

Using the engineering document classification method based on text embedding, the document title is vectorized by fine-tuning the trained target text embedding model, the title feature vector is generated, and the semantic similarity calculation is performed with the pre-generated document classification matrix to determine the classification name of the document.

Benefits of technology

It improves the accuracy of document classification, has a wide range of applications, can effectively improve the classification efficiency of documents, and is suitable for various document content forms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234412A_ABST
    Figure CN120234412A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an engineering document classification method and device based on text embedding, and the method comprises the steps: carrying out the vectorization processing of a document title of an obtained to-be-classified engineering document through a fine-tuning trained target text embedding model, and generating a title feature vector; performing semantic similarity calculation on the title feature vector and a pre-generated document classification matrix to obtain a semantic similarity matrix; according to the semantic similarity matrix, classification names of the project documents are determined, the summary ability of project document titles on full texts is fully utilized, efficient classification is carried out based on document contents, and the document classification accuracy is improved; the method is not limited by document content forms, the application range is wide, and the document classification efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for classifying engineering documents based on text embedding. Background Art

[0002] Digital delivery has currently been widely applied to the completion delivery in the field of metallurgical engineering. Generally, after the general contractor of a project completes engineering design, procurement, construction, and commissioning, all the materials during the project construction period need to be reclassified according to the regulations of the owner or the digital delivery standard to meet the requirements of quick query and backtracking of materials during the project operation and maintenance period. The current related technologies mainly obtain keywords by performing word segmentation on documents, and then classify the documents according to the keywords. However, this method highly depends on the accuracy of the existing word segmentation algorithms. In terms of semantic expression, the method of representing document features by word segmentation will lose a large amount of document information, and the relationship between word segments is defaulted to be relatively independent, unable to fully reflect the context correlation relationship inside the document, thus resulting in a low classification accuracy. In addition, for non-text documents, the methods in the related technologies are also inapplicable, and the universality is poor. Summary of the Invention

[0003] The first object of the present invention is to provide a method for classifying engineering documents based on text embedding, which makes full use of the summarizing ability of the engineering document title for the full text, performs efficient classification based on the document content, and improves the accuracy of document classification; this method is not limited by the form of the document content, has a wide application range, and effectively improves the classification efficiency of the document. The second object of the present invention is to provide a device for classifying engineering documents based on text embedding. The third object of the present invention is to provide a computer-readable medium. The fourth object of the present invention is to provide a computer device.

[0004] To achieve the above objects, on the one hand, the present invention discloses a method for classifying engineering documents based on text embedding, including:

[0005] Performing vectorization processing on the document title of the obtained engineering document to be classified through a fine-tuned trained target text embedding model to generate a title feature vector;

[0006] Calculating the semantic similarity between the title feature vector and a pre-generated document classification matrix to obtain a semantic similarity matrix;

[0007] Determining the classification name of the engineering document according to the semantic similarity matrix.

[0008] Preferably, before performing vectorization processing on the document title of the obtained engineering document to be classified through a fine-tuned trained target text embedding model to obtain a title feature vector, it further includes:

[0009] Obtain the title information and corresponding type information of the original engineering documents;

[0010] Construct a sentence pair set according to the title information and corresponding type information of multiple original engineering documents;

[0011] Divide the sentence pair set according to a preset ratio to form a fine-tuning training set and a fine-tuning validation set;

[0012] Fine-tune and train the original text embedding model through the fine-tuning training set to generate a pre-trained model;

[0013] Verify and optimize the pre-trained model through the fine-tuning validation set to obtain the target text embedding model.

[0014] Preferably, constructing a sentence pair set according to the title information and corresponding type information of multiple original engineering documents includes:

[0015] Extract the text part of the title information to obtain the title text;

[0016] Format the type information according to the preset multi-level classification directory of engineering documents to obtain a multi-level classification;

[0017] Generate sentence pairs according to the title text and multi-level classification corresponding to the original engineering documents;

[0018] Construct a sentence pair set according to multiple sentence pairs.

[0019] Preferably, the fine-tuning validation set includes multiple validation sentence pairs, and the validation sentence pairs include validation titles and validation classifications;

[0020] Verifying and optimizing the pre-trained model through the fine-tuning validation set to obtain the target text embedding model includes:

[0021] Vectorize the validation title through the pre-trained model to obtain a validation title matrix;

[0022] Vectorize the preset multi-level classification directory of engineering documents through the pre-trained model to obtain an initial classification matrix;

[0023] Calculate the semantic similarity between the validation title matrix and the initial classification matrix to obtain a validation similarity matrix;

[0024] Determine the predicted classification corresponding to each validation title according to the validation similarity matrix;

[0025] Iteratively optimize the pre-trained model according to the predicted classification and validation classification corresponding to each validation title to obtain the target text embedding model.

[0026] Preferably, before calculating the semantic similarity between the title feature vector and the pre-generated document classification matrix to obtain the semantic similarity matrix, it further includes:

[0027] The preset multi-level classification directory of engineering documents is vectorized through the target text embedding model to obtain the document classification matrix.

[0028] Preferably, the title of the engineering document to be classified obtained is vectorized through the target text embedding model after fine-tuning training to obtain the title feature vector, including:

[0029] Extract the text part of the document title to obtain the title to be classified;

[0030] The title to be classified is vectorized through the target text embedding model to obtain the title feature vector.

[0031] Preferably, the semantic similarity matrix includes the similarity between the engineering document to be classified and each document classification;

[0032] According to the semantic similarity matrix, the classification name of the engineering document is determined, including:

[0033] Screen out the matrix element with the maximum similarity from the semantic similarity matrix;

[0034] Determine the document classification corresponding to the matrix element with the maximum similarity as the classification name of the engineering document.

[0035] The present invention also discloses an engineering document classification device based on text embedding, including:

[0036] A title vectorization unit, configured to vectorize the document title of the engineering document to be classified obtained through the target text embedding model after fine-tuning training to generate a title feature vector;

[0037] A similarity calculation unit, configured to calculate the semantic similarity between the title feature vector and the pre-generated document classification matrix to obtain the semantic similarity matrix;

[0038] A document classification determination unit, configured to determine the classification name of the engineering document according to the semantic similarity matrix.

[0039] Preferably, the device further includes:

[0040] An acquisition unit, configured to acquire the title information and the corresponding type information of the original engineering document;

[0041] A sentence pair construction unit, configured to construct a sentence pair set according to the title information and the corresponding type information of multiple original engineering documents;

[0042] A division unit, configured to divide the sentence pair set according to a preset ratio to form a fine-tuning training set and a fine-tuning validation set;

[0043] A fine-tuning training unit, configured to perform fine-tuning training on the original text embedding model through the fine-tuning training set to generate a pre-trained model;

[0044] A verification and optimization unit, configured to perform verification and optimization on the pre-trained model through the fine-tuning validation set to obtain a target text embedding model.

[0045] Preferably, the sentence pair construction unit is specifically configured to extract the text part of the title information to obtain the title text; format the type information according to the preset multi-level classification directory of the engineering documents to obtain multi-level classification; generate sentence pairs according to the title text and the multi-level classification corresponding to the original engineering documents; construct a sentence pair set according to multiple sentence pairs.

[0046] Preferably, the fine-tuning validation set includes multiple validation sentence pairs, and the validation sentence pairs include validation titles and validation classifications;

[0047] The verification and optimization unit is specifically configured to perform vectorization processing on the validation title through the pre-trained model to obtain a validation title matrix; perform vectorization processing on the preset multi-level classification directory of the engineering documents through the pre-trained model to obtain an initial classification matrix; calculate the semantic similarity between the validation title matrix and the initial classification matrix to obtain a validation similarity matrix; determine the predicted classification corresponding to each validation title according to the validation similarity matrix; perform iterative optimization on the pre-trained model according to the predicted classification and the validation classification corresponding to each validation title to obtain a target text embedding model.

[0048] Preferably, the apparatus further includes:

[0049] A classification vectorization unit, configured to perform vectorization processing on the preset multi-level classification directory of the engineering documents through the target text embedding model to obtain a document classification matrix.

[0050] Preferably, the title vectorization unit is specifically configured to extract the text part of the document title to obtain a title to be classified; perform vectorization processing on the title to be classified through the target text embedding model to obtain a title feature vector.

[0051] Preferably, the semantic similarity matrix includes the similarity between the engineering document to be classified and each document classification;

[0052] The document classification determination unit is specifically configured to screen out the matrix element with the maximum similarity from the semantic similarity matrix; determine the document classification corresponding to the matrix element with the maximum similarity as the classification name of the engineering document.

[0053] The present invention also discloses a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method is implemented.

[0054] The present invention also discloses a computer device, including a memory and a processor, where the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the processor executes the program, the above-mentioned method is implemented.

[0055] The present invention also discloses a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the above-mentioned method is implemented.

[0056] The present invention fine-tunes the trained target text embedding model to vectorize the document title of the obtained engineering document to be classified, generating a title feature vector; calculates the semantic similarity between the title feature vector and the pre-generated document classification matrix to obtain a semantic similarity matrix; determines the classification name of the engineering document according to the semantic similarity matrix, makes full use of the summarization ability of the engineering document title for the full text, conducts efficient classification based on the document content, improves the accuracy of document classification; this method is not limited by the form of the document content, has a wide range of applications, and effectively improves the classification efficiency of the document. Description of the Drawings

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0058] Figure 1 It is a flowchart of a method for classifying engineering documents based on text embedding provided by an embodiment of the present invention;

[0059] Figure 2 It is a flowchart of another method for classifying engineering documents based on text embedding provided by an embodiment of the present invention;

[0060] Figure 3 It is a schematic structural diagram of a device for classifying engineering documents based on text embedding provided by an embodiment of the present invention;

[0061] Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed Embodiments

[0062] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0063] To facilitate the understanding of the technical solutions provided in this application, the relevant content of the technical solutions in this application will be described first. With the wide application of natural language processing technology and the booming development of large language model technology, document intelligent classification technology has been popularized in all walks of life. This application uses natural language processing and large language model technology to perform intelligent classification on digital delivery documents, and performs efficient and accurate classification on the corresponding delivery documents during the digital delivery process.

[0064] Next, taking the engineering document classification device based on text embedding as the execution subject as an example, the implementation process of the engineering document classification method based on text embedding provided in the embodiments of the present invention will be described. It can be understood that the execution subject of the engineering document classification method based on text embedding provided in the embodiments of the present invention includes but is not limited to the engineering document classification device based on text embedding.

[0065] Figure 1 The flowchart of an engineering document classification method based on text embedding provided in the embodiments of the present invention is as Figure 1 shown, and the method includes:

[0066] Step 101: Vectorize the document title of the obtained engineering document to be classified through a fine-tuned trained target text embedding model to generate a title feature vector.

[0067] Step 102: Calculate the semantic similarity between the title feature vector and the pre-generated document classification matrix to obtain a semantic similarity matrix.

[0068] Step 103: Determine the classification name of the engineering document according to the semantic similarity matrix.

[0069] In the technical solutions provided in the embodiments of the present invention, the document title of the obtained engineering document to be classified is vectorized through a fine-tuned trained target text embedding model to generate a title feature vector; the semantic similarity between the title feature vector and the pre-generated document classification matrix is calculated to obtain a semantic similarity matrix; the classification name of the engineering document is determined according to the semantic similarity matrix, making full use of the summary ability of the engineering document title for the full text, performing efficient classification based on the document content, and improving the accuracy of document classification; this method is not limited by the form of the document content, has a wide range of applications, and effectively improves the classification efficiency of the document.

[0070] Figure 2 This is a flowchart of another engineering document classification method based on text embedding provided by an embodiment of the present invention. As shown in Figure 2 the following, this method includes:

[0071] Step 201: Obtain the title information and corresponding type information of the original engineering document.

[0072] In the embodiment of the present invention, each step is executed by an engineering document classification device based on text embedding.

[0073] In the embodiment of the present invention, the original engineering document is the completed electronic data. There are many types of engineering documents, including but not limited to design drawings, management documents (text), inspection (test) reports, and operation manuals. As an alternative, the electronic data is all the documents of the metallurgical engineering digital delivery project.

[0074] Further, in order to meet the delivery requirements, the original engineering documents of each type are uniformly formatted. As an alternative, the original engineering documents of each type are uniformly converted into PDF format.

[0075] In the embodiment of the present invention, each original engineering document has title information, and each original engineering document corresponds to type information. The title information includes a text description part and a document number. The text description part can generally describe the content or theme of the document, and the type information can characterize the specific classification of the original engineering document.

[0076] Step 202: Construct a set of sentence pairs according to the title information and corresponding type information of multiple original engineering documents.

[0077] In the embodiment of the present invention, step 202 specifically includes:

[0078] Step 2021: Extract the text part of the title information to obtain the title text.

[0079] In the embodiment of the present invention, the title information includes a text description part and a document number. Among them, the document number does not contain information related to the description of the document content. Filter out the document number and extract the text description part in the title information to obtain the title text, which can generally describe the content or theme of the document.

[0080] Step 2022: Format the type information according to the preset multi-level classification directory of engineering documents to obtain a multi-level classification.

[0081] In the embodiments of the present invention, the multi-level classification directory of engineering documents is pre-set according to the industry's engineering digital delivery standard. The multi-level classification directory of engineering documents includes multi-level directories. In this application, taking a 3-level directory as an example, the multi-level classification directory of engineering documents includes a first-level classification, a second-level classification, and a third-level classification, as shown in Table 1:

[0082] Table 1

[0083]

[0084] Specifically, according to the type information, the corresponding first-level classification name, second-level classification name, and third-level classification name are extracted from the multi-level classification directory of engineering documents, and the first-level classification name, second-level classification name, and third-level classification name are connected by special symbols to obtain a multi-level classification. As an alternative, the special symbol is "-". For example: the multi-level classification is: construction materials - decision-making and project establishment documents - project proposal.

[0085] In the embodiments of the present invention, the multi-level classification is a string combination of the first-level classification name, the second-level classification name, and the third-level classification name, which not only avoids the possibility of duplicate 3-level classification names, but also enriches the text description of document classification, and is more conducive to the accurate mapping and expression of spatial vectors

[0086] Step 2023: Generate sentence pairs according to the title text corresponding to the original engineering document and the multi-level classification.

[0087] In the embodiments of the present invention, according to the title text corresponding to the original engineering document and the multi-level classification, sentence pairs are constructed according to the existing document classification relationship. Taking the title text as the project proposal of the steel project in Area A and the multi-level classification as construction materials - decision-making and project establishment documents - project proposal as an example, the form of constructing sentence pairs (sentence-pair) is as follows:

[0088]

[0089] Among them, the "text" label in the sentence pair corresponds to the title text; "text_pos" corresponds to the multi-level classification to which the original engineering document belongs; "class_code" represents the serial number of this classification, and the value range of the classification serial number depends on the number of classifications in the multi-level classification directory of engineering documents. As an alternative, there are 111 classifications in the multi-level classification directory of engineering documents, then the value range of the classification serial number is 0-110. It should be noted that the classification serial number does not participate in the final training, but only for the accuracy verification of the subsequent classification results.

[0090] Step 2024: Construct a sentence pair set according to multiple sentence pairs.

[0091] In the embodiments of the present invention, title information and corresponding type information are extracted from each original engineering document to obtain multiple sentence pairs, and each sentence pair corresponds to an original engineering document; multiple sentence pairs are generated into a sentence pair set.

[0092] Step 203: Divide the sentence pair set according to a preset ratio to form a fine-tuning training set and a fine-tuning validation set.

[0093] In the embodiments of the present invention, the preset ratio is the ratio between the fine-tuning training set and the fine-tuning validation set, which can be set according to actual needs, and the embodiments of the present invention do not limit this. As an optional solution, the preset ratio is 7:3, that is, the ratio of the number of sentence pairs between the fine-tuning training set and the fine-tuning validation set is 7:3.

[0094] In the embodiments of the present invention, the sentence pairs in the fine-tuning training set are used for model fine-tuning training, and the sentence pairs in the fine-tuning validation set are used for model effect verification.

[0095] Step 204: Fine-tune the original text embedding model through the fine-tuning training set to generate a pre-trained model.

[0096] In the embodiments of the present invention, the original text embedding model is the M3E-base model, which is based on the open-source large model M3E-base and the model is fine-tuned based on the python framework Uniem.

[0097] Specifically, the FineTuner module in the Uniem framework is called to load the M3E model; the fine-tuning training set is loaded, and FineTuner is used for fine-tuning training until the model converges, and the model parameters are adjusted to adapt to this engineering document classification task, and then a pre-trained model customized for a specific scenario is obtained.

[0098] Step 205: Verify and optimize the pre-trained model through the fine-tuning validation set to obtain a target text embedding model.

[0099] In the embodiments of the present invention, the fine-tuning validation set includes multiple verification sentence pairs, and the verification sentence pairs include a verification title and a verification classification.

[0100] In the embodiments of the present invention, step 205 specifically includes:

[0101] Step 2051: Vectorize the verification title through the pre-trained model to obtain a verification title matrix.

[0102] In the embodiments of the present invention, a pre-trained model is called through sentence transformer to perform vectorized encoding on the verification title, and the corresponding verification title vector is obtained; the verification title vectors corresponding to all verification titles in the fine-tuning verification set are selected to construct a verification title matrix X. The dimension of the verification title matrix X is n×d, where n represents the number of documents and d represents the spatial representation dimension of the corresponding document. As an alternative, d is 768.

[0103] Step 2052: Perform vectorization processing on the multi-level classification directory of the engineering document through the pre-trained model to obtain an initial classification matrix.

[0104] In the embodiments of the present invention, a pre-trained model is called through sentence transformer to perform vectorized encoding on each classification in the multi-level classification directory of the engineering document, and the corresponding initial classification vector is obtained; according to the initial classification vectors corresponding to the full-scale classification, an initial classification matrix C’ is constructed. The dimension of the initial classification matrix C’ is m×d, where m represents the number of classifications and d represents the spatial representation dimension of the corresponding document. As an alternative, d is 768. In this application, taking the number of classifications as 111 as an example, the dimension of the initial classification matrix is thus 111×768, denoted as the initial classification matrix C’.

[0105] Step 2053: Calculate the semantic similarity between the verification title matrix and the initial classification matrix to obtain a verification similarity matrix.

[0106] In the embodiments of the present invention, the semantic similarity matrix S’ between all verification titles and all initial classifications is obtained through the cosine similarity between the corresponding verification title vectors in the verification title matrix and all initial classification vectors in the initial classification matrix.

[0107] Specifically, through calculating the verification title vector in the verification title matrix and the initial classification vector in the initial classification matrix, the verification similarity is obtained, where sim' ij is the semantic similarity between the i-th verification title and the j-th multi-level classification, is the verification title vector of the i-th verification title, is the initial classification vector of the j-th multi-level classification.

[0108] In the embodiments of the present invention, a verification similarity matrix S’ is constructed according to multiple verification similarities; the matrix dimension of the verification similarity matrix S’ is n×m, where n is the number of verification titles, that is: the number of documents; m is the number of classifications, that is: 111.

[0109] Step 2054: Determine the predicted classification corresponding to each verification title according to the verification similarity matrix.

[0110] In the embodiments of the present invention, the verification similarity matrix records the semantic similarity degree between each verification title and each multi-level classification.

[0111] Specifically, through Class' = argmax(S', axis = 1), according to the verification similarity matrix, the predicted classification corresponding to each verification title is determined. Wherein, Class' is the set of predicted classifications corresponding to each verification title, S' is the verification similarity matrix, and axis = 1 means comparing horizontally with the horizontal axis as the unit. To know the classification of a certain document, only need to know the serial number of the maximum value (MAX value) in the horizontal dimension of the verification title of the document in the verification similarity matrix, and the specific predicted classification of the document can be confirmed by looking up the corresponding multi-level classification of the corresponding serial number.

[0112] Step 2055: Iteratively optimize the pre-trained model according to the predicted classification and verification classification corresponding to each verification title to obtain the target text embedding model.

[0113] In the embodiments of the present invention, the predicted classification and verification classification corresponding to each verification title are compared, the classification accuracy (Accuracy) is calculated to obtain the classification effect of the pre-trained model; if the classification effect is lower than the expected effect, the pre-trained model is iteratively optimized to further adjust the model parameters until the classification effect reaches the expected effect, and the target text embedding model is output.

[0114] It should be noted that the target text embedding model is a fine-tuned M3E model.

[0115] Step 206: Vectorize the preset multi-level classification directory of engineering documents through the target text embedding model to obtain a document classification matrix.

[0116] In the embodiments of the present invention, the multi-level classification directory of engineering documents is shown in Table 1. Specifically, the target text embedding model is called through sentencetransformer to perform vectorized encoding on each classification in the multi-level classification directory of engineering documents to obtain the corresponding document classification vector; according to the document classification vectors corresponding to all document classifications, a document classification matrix C is constructed. The dimension of the initial classification matrix C is m×d, where m represents the number of classifications and d represents the spatial representation dimension of the corresponding document. As an optional solution, d is 768. In this application, taking the number of classifications as 111 as an example, the dimension of the document classification matrix is 111×768, denoted as matrix C.

[0117] Step 207: Vectorize the document title of the obtained engineering document to be classified through the target text embedding model to obtain a title feature vector.

[0118] In the embodiments of the present invention, step 207 specifically includes:

[0119] Step 2071: extract the text portion of the document title to obtain the title to be classified.

[0120] In an embodiment of the present invention, the document title includes a text description part and a document number, wherein the document number does not contain information related to the description of the document content. The document number is filtered out, and the text description part in the document title is extracted to obtain the title to be classified. The title to be classified can summarize the content or theme of the document.

[0121] Step 2072: vectorize the title to be classified through the target text embedding model to obtain a title feature vector.

[0122] Specifically, the target text embedding model is called through the sentence transformer to vectorize and encode the classified title to obtain the corresponding title feature vector.

[0123] Step 208: Calculate semantic similarity between the title feature vector and the document classification matrix to obtain a semantic similarity matrix.

[0124] Specifically, through The title feature vector and the document classification vector in the document classification matrix are calculated to obtain the semantic similarity, where sim ij is the semantic similarity between the title to be classified i and the jth multi-level classification, is the title feature vector, is the document classification vector of the j-th multi-level classification.

[0125] In the embodiment of the present invention, a semantic similarity matrix Sim is constructed according to a plurality of semantic similarities, and the semantic similarity matrix Sim includes the similarities between the engineering document to be classified and each document classification.

[0126] The present invention uses a complete document naming system to obtain a simplified and effective description of the document content through the title to be classified. This method greatly reduces the input dimension of data and the training workload, and reduces the requirements of the classification algorithm on the computing power hardware configuration.

[0127] Step 209: Filter out the matrix element with the maximum similarity from the semantic similarity matrix.

[0128] Specifically, Class=argmax(Sim,axis=1) is used to determine the matrix element with the maximum similarity corresponding to the engineering document to be classified according to the semantic similarity matrix, where Sim is the semantic similarity matrix and axis=1 indicates horizontal comparison based on the horizontal axis.

[0129] Step 210: Determine the classification name of the engineering document as the document classification corresponding to the matrix element with the maximum similarity.

[0130] In the embodiment of the present invention, the matrix element with the maximum similarity indicates that the engineering document to be classified is most similar to the current corresponding document classification; and the current corresponding document classification is determined as the classification name of the engineering document.

[0131] It should be noted that in the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant channels, and the acquisition, storage, use, processing, etc. of user information have obtained the authorization and consent of the customers.

[0132] In the technical solution of the engineering document classification method based on text embedding provided by the embodiment of the present invention, by fine-tuning the trained target text embedding model, the document title of the obtained engineering document to be classified is vectorized to generate a title feature vector; the semantic similarity between the title feature vector and the pre-generated document classification matrix is calculated to obtain a semantic similarity matrix; according to the semantic similarity matrix, the classification name of the engineering document is determined, making full use of the summary ability of the engineering document title for the full text, performing efficient classification based on the document content, and improving the accuracy of document classification; this method is not limited by the form of document content, has a wide range of applications, and effectively improves the classification efficiency of documents.

[0133] Figure 3 The following is a schematic structural diagram of an engineering document classification device based on text embedding provided by an embodiment of the present invention. This device is used to execute the above-mentioned engineering document classification method based on text embedding, as Figure 3 shown. This device includes: a title vectorization unit 11, a similarity calculation unit 12, and a document classification determination unit 13.

[0134] The title vectorization unit 11 is used to vectorize the document title of the obtained engineering document to be classified by fine-tuning the trained target text embedding model, and generate a title feature vector.

[0135] The similarity calculation unit 12 is used to calculate the semantic similarity between the title feature vector and the pre-generated document classification matrix to obtain a semantic similarity matrix.

[0136] The document classification determination unit 13 is used to determine the classification name of the engineering document according to the semantic similarity matrix.

[0137] In the embodiment of the present invention, the device further includes: an acquisition unit 14, a sentence pair construction unit 15, a division unit 16, a fine-tuning training unit 17, and a verification and optimization unit 18.

[0138] The acquisition unit 14 is used to acquire the title information and corresponding type information of the original engineering documents.

[0139] The sentence pair construction unit 15 is used to construct a sentence pair set according to the title information and corresponding type information of multiple original engineering documents.

[0140] The partitioning unit 16 is used to partition the sentence pair set according to a preset ratio to form a fine-tuning training set and a fine-tuning validation set.

[0141] The fine-tuning training unit 17 is used to fine-tune and train the original text embedding model through the fine-tuning training set to generate a pre-trained model.

[0142] The verification and optimization unit 18 is used to verify and optimize the pre-trained model through the fine-tuning validation set to obtain the target text embedding model.

[0143] In an embodiment of the present invention, the sentence pair construction unit 15 is specifically used to extract the text part of the title information to obtain the title text; format the type information according to a preset multi-level classification directory of engineering documents to obtain a multi-level classification; generate sentence pairs according to the title text and multi-level classification corresponding to the original engineering documents; and construct a sentence pair set according to multiple sentence pairs.

[0144] In an embodiment of the present invention, the fine-tuning validation set includes multiple validation sentence pairs, and the validation sentence pairs include validation titles and validation classifications; the verification and optimization unit 18 is specifically used to perform vectorization processing on the validation titles through the pre-trained model to obtain a validation title matrix; perform vectorization processing on the preset multi-level classification directory of engineering documents through the pre-trained model to obtain an initial classification matrix; calculate the semantic similarity between the validation title matrix and the initial classification matrix to obtain a validation similarity matrix; determine the predicted classification corresponding to each validation title according to the validation similarity matrix; and perform iterative optimization on the pre-trained model according to the predicted classification and validation classification corresponding to each validation title to obtain the target text embedding model.

[0145] In an embodiment of the present invention, the device further includes: a classification vectorization unit 19.

[0146] The classification vectorization unit 19 is used to perform vectorization processing on the preset multi-level classification directory of engineering documents through the target text embedding model to obtain a document classification matrix.

[0147] In an embodiment of the present invention, the title vectorization unit 11 is specifically used to extract the text part of the document title to obtain a title to be classified; and perform vectorization processing on the title to be classified through the target text embedding model to obtain a title feature vector.

[0148] In the embodiment of the present invention, the semantic similarity matrix includes the similarity between the engineering document to be classified and each document classification; the document classification determination unit 13 is specifically configured to screen out the matrix element with the maximum similarity from the semantic similarity matrix; and determine the document classification corresponding to the matrix element with the maximum similarity as the classification name of the engineering document.

[0149] In the solution of the embodiment of the present invention, by fine-tuning the trained target text embedding model, the document title of the obtained engineering document to be classified is vectorized to generate a title feature vector; the semantic similarity between the title feature vector and the pre-generated document classification matrix is calculated to obtain a semantic similarity matrix; according to the semantic similarity matrix, the classification name of the engineering document is determined, making full use of the summary ability of the engineering document title for the full text, performing efficient classification based on the document content, and improving the accuracy of document classification; this method is not limited by the form of the document content, has a wide range of applications, and effectively improves the document classification efficiency.

[0150] The system, device, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with a certain function. A typical implementation device is a computer device. Specifically, the computer device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0151] The embodiment of the present invention provides a computer device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, the steps of the above embodiments of the engineering document classification method based on text embedding are implemented. For specific descriptions, reference can be made to the embodiments of the engineering document classification method based on text embedding above.

[0152] Reference is made below Figure 4 , which shows a schematic structural diagram of a computer device 600 suitable for implementing the embodiments of the present application.

[0153] As Figure 4 shown, the computer device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0154] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required so that a computer program read from it can be installed in the storage section 608 as required.

[0155] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 609, and / or installed from the removable medium 611.

[0156] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0157] For convenience of description, the above device is described by dividing it into various units according to functions. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.

[0158] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means, and the instruction means implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or the functions specified in multiple blocks.

[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or the functions specified in multiple blocks.

[0161] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the said element.

[0162] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.

[0163] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0164] The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0165] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0166] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for classifying engineering documents based on text embedding, characterized in that, The method includes: Vectorize the document title of the acquired engineering document to be classified through fine-tuning the trained target text embedding model, and generate a title feature vector; Calculate the semantic similarity between the title feature vector and the pre-generated document classification matrix to obtain a semantic similarity matrix; Determine the classification name of the engineering document according to the semantic similarity matrix.

2. The engineering document classification method based on text embedding according to claim 1, wherein Before vectorizing the document title of the acquired engineering document to be classified through fine-tuning the trained target text embedding model to obtain a title feature vector, it further includes: Obtain the title information and corresponding type information of the original engineering document; Construct a sentence pair set according to the title information and corresponding type information of multiple original engineering documents; Divide the sentence pair set according to a preset ratio to form a fine-tuning training set and a fine-tuning validation set; Fine-tune and train the original text embedding model through the fine-tuning training set to generate a pre-trained model; Verify and optimize the pre-trained model through the fine-tuning validation set to obtain the target text embedding model.

3. The engineering document classification method based on text embedding according to claim 2, wherein The constructing a sentence pair set according to the title information and corresponding type information of multiple original engineering documents includes: Extract the text part of the title information to obtain the title text; Format the type information according to the preset multi-level classification directory of engineering documents to obtain a multi-level classification; Generate a sentence pair according to the title text and multi-level classification corresponding to the original engineering document; Construct a sentence pair set according to multiple sentence pairs.

4. The engineering document classification method based on text embedding according to claim 2, wherein The fine-tuning validation set includes multiple validation sentence pairs, and the validation sentence pair includes a validation title and a validation classification; The verifying and optimizing the pre-trained model through the fine-tuning validation set to obtain the target text embedding model includes: Vectorize the validation title through the pre-trained model to obtain a validation title matrix; Vectorize the preset multi-level classification directory of engineering documents through the pre-trained model to obtain an initial classification matrix; Calculate the semantic similarity between the validation title matrix and the initial classification matrix to obtain a validation similarity matrix; Determine the predicted classification corresponding to each validation title according to the validation similarity matrix; Iteratively optimize the pre-trained model according to the predicted classification and validation classification corresponding to each validation title to obtain the target text embedding model.

5. The engineering document classification method based on text embedding according to claim 1, characterized in that Before calculating the semantic similarity between the title feature vector and the pre-generated document classification matrix to obtain a semantic similarity matrix, it further includes: Vectorize the preset multi-level classification directory of engineering documents through the target text embedding model to obtain a document classification matrix.

6. The engineering document classification method based on text embedding according to claim 1, characterized in that The vectorizing the document title of the acquired engineering document to be classified through fine-tuning the trained target text embedding model to obtain a title feature vector includes: Extract the text part of the document title to obtain a title to be classified; Vectorize the title to be classified through the target text embedding model to obtain the title feature vector.

7. The engineering document classification method based on text embedding according to claim 1, wherein The semantic similarity matrix includes the similarities between the engineering document to be classified and each document classification; Determining the classification name of the engineering document according to the semantic similarity matrix includes: Filtering out the matrix element with the maximum similarity from the semantic similarity matrix; Determining the document classification corresponding to the matrix element with the maximum similarity as the classification name of the engineering document.

8. An engineering document classification device based on text embedding, characterized in that, The device includes: A title vectorization unit, configured to vectorize the document title of the acquired engineering document to be classified through a fine-tuned pre-trained target text embedding model, and generate a title feature vector; A similarity calculation unit, configured to perform semantic similarity calculation on the title feature vector and a pre-generated document classification matrix to obtain a semantic similarity matrix; A document classification determination unit, configured to determine the classification name of the engineering document according to the semantic similarity matrix.

9. The engineering document classification device based on text embedding according to claim 8, wherein The device further includes: An acquisition unit, configured to acquire the title information and corresponding type information of the original engineering document; A sentence pair construction unit, configured to construct a sentence pair set according to the title information and corresponding type information of multiple original engineering documents; A division unit, configured to divide the sentence pair set according to a preset ratio to form a fine-tuning training set and a fine-tuning validation set; A fine-tuning training unit, configured to perform fine-tuning training on the original text embedding model through the fine-tuning training set to generate a pre-trained model; A verification and optimization unit, configured to perform verification and optimization on the pre-trained model through the fine-tuning validation set to obtain the target text embedding model.

10. The engineering document classification device based on text embedding according to claim 9, characterized in that, The sentence pair construction unit is specifically configured to extract the text part of the title information to obtain the title text; format the type information according to a preset multi-level classification directory of engineering documents to obtain a multi-level classification; generate sentence pairs according to the title text and multi-level classification corresponding to the original engineering document; Construct a sentence pair set according to multiple sentence pairs.

11. The engineering document classification device based on text embedding according to claim 9, wherein, The fine-tuning validation set includes multiple validation sentence pairs, and the validation sentence pairs include validation titles and validation classifications; The verification and optimization unit is specifically configured to vectorize the validation title through the pre-trained model to obtain a validation title matrix; vectorize a preset multi-level classification directory of engineering documents through the pre-trained model to obtain an initial classification matrix; perform semantic similarity calculation on the validation title matrix and the initial classification matrix to obtain a validation similarity matrix; determine the predicted classification corresponding to each validation title according to the validation similarity matrix; and perform iterative optimization on the pre-trained model according to the predicted classification and validation classification corresponding to each validation title to obtain the target text embedding model.

12. The engineering document classification device based on text embedding according to claim 8, wherein The device further includes: A classification vectorization unit, configured to vectorize a preset multi-level classification directory of engineering documents through the target text embedding model to obtain a document classification matrix.

13. The engineering document classification device based on text embedding according to claim 8, characterized in that, The title vectorization unit is specifically configured to extract the text part of the document title to obtain a title to be classified; and vectorize the title to be classified through the target text embedding model to obtain the title feature vector.

14. The engineering document classification device based on text embedding according to claim 8, characterized in that, The semantic similarity matrix includes the similarities between the engineering documents to be classified and each document classification; The document classification determination unit is specifically configured to screen out the matrix element with the maximum similarity from the semantic similarity matrix; and determine the document classification corresponding to the matrix element with the maximum similarity as the classification name of the engineering document.

15. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the engineering document classification method based on text embedding according to any one of claims 1 to 7.

16. A computer device, comprising a memory and a processor, the memory being used for storing information including program instructions, and the processor being used for controlling the execution of the program instructions, characterized in that, When the program instructions are loaded and executed by a processor, it implements the engineering document classification method based on text embedding according to any one of claims 1 to 7.

17. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by a processor, it implements the engineering document classification method based on text embedding according to any one of claims 1 to 7.