Document labeling method and device

By preprocessing and converting key measurement values ​​to the distribution project documents, combined with the document labeling model, the problem of manual labeling is solved, and the problem of poor application based on large models is achieved, efficient and accurate document labeling and classification is achieved.

CN120045937APending Publication Date: 2025-05-27BEIJING CHINA POWER INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063638.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In power distribution project management, manual annotation of documents is time-consuming and labor-intensive, error-prone, and the document processing methods based on large models lack professional knowledge, resulting in poor labeling application performance.

Method used

A document annotation method is proposed. By preprocessing the document to be processed, the key degree value of the word is converted into a key weighted embedding vector, and a pre-constructed document annotation model is input to output the annotation information of the document.

Benefits of technology

It realizes the accurate labeling and classification of distribution project documents, improves the accuracy of labeling and classification results, and improves project management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045937A_ABST
    Figure CN120045937A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a document labeling method and device, and the method comprises the steps: carrying out the preprocessing of a to-be-processed document, and obtaining a preprocessed document; counting the key degree value of each word in the pre-processed document; converting the pre-processed document into a key weighted embedding vector according to the key degree value; and inputting the key weighted embedding vector into a pre-constructed document labeling model, and outputting the labeling information of the document by the document labeling model. According to the invention, labeling and classification of the power distribution engineering project file can be realized, the accuracy of labeling and classification results is improved, and the project management efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a document annotation method and device. Background Art

[0002] In the management of distribution engineering projects, a large number of documents such as project reports, contract documents, technical specifications, etc. carry important project information. To improve project management efficiency, various types of documents need to be annotated and classified. Since such documents contain complex professional terms and industry specifications, manual annotation is time-consuming and laborious, and prone to errors. When using a document processing method based on a large model, due to the lack of professional domain knowledge, the performance in document annotation applications for distribution engineering projects is poor. Summary of the Invention

[0003] In view of this, the purpose of the embodiments of the present application is to propose a document annotation method and device to solve the problem of annotation and classification of distribution engineering project documents.

[0004] Based on the above purpose, embodiments of the present application provide a document annotation method, including:

[0005] Preprocess the document to be processed to obtain the preprocessed document;

[0006] Statistically calculate the key degree values of each word in the preprocessed document;

[0007] According to the key degree values, convert the preprocessed document into a key weighted embedding vector;

[0008] Input the key weighted embedding vector into a pre-constructed document annotation model, and the document annotation model outputs the annotation information of the document.

[0009] Optionally, converting the preprocessed document into a key weighted embedding vector according to the key degree values includes:

[0010] Perform word segmentation on the preprocessed document to obtain word tokens;

[0011] Perform positional encoding on the word tokens to obtain the positional information of the word tokens;

[0012] In combination with the positional information, convert the word tokens into embedding vectors with positional information;

[0013] According to the key degree values and the embedding vectors with positional information, obtain key weighted embedding vectors.

[0014] Optionally, the method for obtaining key weighted embedding vectors according to the key degree values and the embedding vectors with positional information is:

[0015] ETF-IDF(t, d) = TF-IDF(t, d) × E(t) (2)

[0016] Among them, TF-IDF(t, d) is the key degree value of the word t in the document d, E(t) is the embedding vector of the word t with position information, and ETF-IDF(t, d) is the key weighted embedding vector.

[0017] Optionally, the document annotation model is trained based on the Transformer model. The document annotation model includes a multi-head attention module, and the calculation method of the attention score is as follows:

[0018]

[0019] Among them, Q, K, and V are the query matrix, key matrix, and value matrix obtained by linearly transforming the key weighted embedding vector respectively, d k is the dimension of the key matrix, and ⊙ represents element-wise multiplication; is the weight matrix calculated from the TF-IDF value of each word in the query matrix, is the weight matrix calculated from the TF-IDF value of each word in the key matrix.

[0020] Optionally, the method further includes:

[0021] According to the annotation information of the document, an entity relationship extraction method is used to determine the classification result of the document to be processed.

[0022] An embodiment of the present application also provides a document annotation device, including:

[0023] A preprocessing module for preprocessing the document to be processed to obtain a preprocessed document;

[0024] A statistics module for counting the key degree values of each word in the preprocessed document;

[0025] A conversion module for converting the preprocessed document into a key weighted embedding vector according to the key degree value;

[0026] An annotation module for inputting the key weighted embedding vector into a pre-constructed document annotation model, and outputting the annotation information of the document by the document annotation model.

[0027] Optionally, the conversion module is used to perform word segmentation on the preprocessed document to obtain word tokens, perform position encoding on the word tokens to obtain the position information of the word tokens, combine the position information, convert the word tokens into embedding vectors with position information, and obtain key weighted embedding vectors according to the key degree value and the embedding vectors with position information.

[0028] Optionally, the method for the conversion module to obtain the key weighted embedding vector is as follows:

[0029] ETF-IDF(t, d) = TF-IDF(t, d) × E(t) (2)

[0030] Wherein, TF-IDF(t, d) is the key degree value of the word t in the document d, E(t) is the embedding vector of the word t with position information, and ETF-IDF(t, d) is the key weighted embedding vector.

[0031] Optionally, the document annotation model is trained based on the Transformer model. The document annotation model includes a multi-head attention module, and the calculation method of the attention score is as follows:

[0032]

[0033] Wherein, Q, K, and V are respectively the query matrix, key matrix, and value matrix obtained by linearly transforming the key weighted embedding vector, d k is the dimension of the key vector, and ⊙ represents element-wise multiplication; is the weight matrix calculated from the TF-IDF value of each word in the query matrix, is the weight matrix calculated from the TF-IDF value of each word in the key matrix.

[0034] Optionally, the apparatus further includes:

[0035] A classification module, configured to determine the classification result of the to-be-processed document by using an entity relationship extraction method according to the annotation information of the document.

[0036] As can be seen from the above, the document annotation method and apparatus provided by the embodiments of the present application preprocess the to-be-processed document to obtain a preprocessed document, count the key degree values of each word in the document, convert the preprocessed document into a key weighted embedding vector according to the key degree values, input the key weighted embedding vector into a pre-constructed document annotation model, and the document annotation model outputs the annotation information of the document. The present application can implement the annotation and classification of distribution engineering project documents, improve the accuracy of the annotation and classification results, and improve the project management efficiency. Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0038] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present application;

[0039] Figure 2 It is a schematic flowchart of the document processing flow according to another embodiment of the present application;

[0040] Figure 3 It is a block diagram of the device structure according to an embodiment of the present application;

[0041] Figure 4 It is a block diagram of the electronic device structure according to an embodiment of the present application. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0043] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be of the ordinary meaning understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second", and similar terms used in the embodiments of the present application do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0044] As Figure 1 、 2 shown, the embodiments of the present application provide a document annotation method, including:

[0045] S101: Preprocess the document to be processed to obtain the preprocessed document;

[0046] In this embodiment, the document to be processed is a document related to a power distribution engineering project, including various reports, specifications, proposals, etc. in different stages and fields of the project. For example, various documents formed in stages such as project planning, design, construction, and acceptance, including those in fields such as electrical engineering, civil engineering, financial management, and legal norms, all contain relevant professional field terms due to their relevance to the power distribution engineering project.

[0047] Preprocess the document to be processed, including text data cleaning and normalization, removing irrelevant information in the document, such as advertisements, headers, footers, copyright statements, etc. After data cleaning, unify the format of the remaining text data, correct grammar and spelling mistakes, uniformly represent professional terms and industry standard terms, replace synonyms for normalization, and use unified fixed representations, etc.

[0048] S102: Statistically calculate the key degree value of each word in the preprocessed document;

[0049] In this embodiment, for the preprocessed document, use the TF-IDF statistical method to statistically calculate the key degree value of each word in the document, that is, calculate the word frequency and inverse document frequency of each word, and calculate the key degree value of the word according to the word frequency and inverse document frequency. Among them, the number of documents related to the power distribution engineering project to be labeled is known.

[0050] S103: According to the key degree value, convert the preprocessed document into a key weighted embedding vector;

[0051] In this embodiment, for the preprocessed document, after determining the key degree value of each word, combine the key degree values of each word to convert the document into a key weighted embedding vector suitable for model processing, so as to input the model for annotation and classification. Among them, the method for converting the key weighted embedding vector includes:

[0052] Perform word segmentation on the preprocessed document to obtain word tokens;

[0053] Perform positional encoding on the word tokens to obtain the position information of the word tokens;

[0054] Combine the position information to convert the word tokens into embedding vectors with position information;

[0055] According to the key degree value and the embedding vector with position information, obtain the key weighted embedding vector.

[0056] In this embodiment, first use a preset word segmentation method to perform word segmentation on the preprocessed document to obtain multiple word tokens. The word segmentation method can be, for example, BPE (Byte-Pair Encoding) or SentencePiece, etc., which can adapt to professional terms and long-tail words in the power distribution engineering project text and improve accuracy; convert the word tokens into embedding vectors. For each word token, use sine and cosine functions to generate corresponding positional encoding to obtain the position information of the word token, and add the positional encoding to the embedding vector of the word token to obtain the embedding vector with position information of the word token.

[0057] In some ways, the method for generating positional encoding is:

[0058]

[0059] Among them, pos is the position index of the token in the sequence, i is the dimension index of the embedding vector, and d model is the dimension of the embedding vector.

[0060] In some embodiments, to enhance the model's attention to key information, a key degree value is added to the embedding vector of the token with position information. By combining the key degree value of each word and the embedding vector, a key weighted embedding vector is generated, so that the weight of the word embedding can be dynamically adjusted according to the characteristics of the document, improving the expression ability of the model. Among them, according to the key degree value and the embedding vector, the key weighted embedding vector is obtained. The method is as follows:

[0061] ETF-IDF(t, d) = TF-IDF(t, d) × E(t) (2)

[0062] Among them, TF-IDF(t, d) is the key degree value of the word t in the document d, which is used to measure the importance of the word t in the document. E(t) is the embedding vector of the word t with position information. ETF-IDF(t, d) is the key weighted embedding vector of the word t. By multiplying TF-IDF(t, d) and E(t), a key weighted embedding vector that contains both the semantic and position information of the word and reflects its importance in the document is generated. This key weighted embedding vector can enhance the model's attention to key information and improve the expression ability of the model.

[0063] S104: Input the key weighted embedding vector into a pre-constructed document annotation model, and the document annotation model outputs the annotation information of the document.

[0064] In some embodiments, the document annotation model is trained based on the Transformer model. The Transformer model includes an encoder, a decoder, an output layer, etc. The encoder includes a multi-head attention module, a feed-forward neural network, a layer normalization, and a residual connection layer, etc. For the key weighted embedding vector input into the document annotation model, it is processed by the multi-head attention module. First, it is linearly transformed into a query matrix, a key matrix, and a value matrix. The query matrix, the key matrix, and the value matrix are mapped into different subspaces and the corresponding attention scores are calculated. Then, the attention scores in different subspaces are merged and concatenated, and the output of the multi-head attention module is obtained through a linear transformation matrix.

[0065] In some ways, the calculation of the attention score combines the weight of the key information to enhance the attention to the key information. The method for calculating the attention score is as follows:

[0066]

[0067] Among them, Q, K, and V are the query matrix, key matrix, and value matrix respectively obtained by linearly transforming the key weighted embedding vectors, and d k is the dimension of the key vector, which is used as a scaling factor to prevent the dot product result from being too large, resulting in the disappearance of the gradient of the softmax function. ⊙ represents element-wise multiplication.

[0068] is the weight matrix calculated from the TF-IDF values of each word in the query matrix, is the weight matrix calculated from the TF-IDF values of each word in the key matrix. The calculation method is as follows: for each word in the query matrix Q, multiply its TF value by the corresponding IDF value to obtain the weight of this word in W Q , and it is composed of the weights of all words For each word in the key matrix K, multiply its TF value by the corresponding IDF value to obtain the weight of this word in W k , and it is composed of the weights of all words

[0069] The output of the multi-head attention module is expressed as:

[0070] MultiHead(Q, K, V) = Concat(head 1 , …, head h )WO (4)

[0071] Among them, h is the number of attention heads, and WO is the linear transformation matrix.

[0072] The features processed by the multi-head attention module are input to the feed-forward neural network layer for non-linear transformation. The feed-forward neural network layer includes two fully connected layers. The first fully connected layer is used to perform a linear transformation on the input features, and then a non-linear transformation is performed using the ReLU activation function. The output of the first fully connected layer is linearly transformed by the second fully connected layer to obtain the processed features. Among them, the linear transformation method of the second fully connected layer is:

[0073] FFN(x) = max(0, xW 1 + b 1 )W 2 + b 2 (5)

[0074] Among them, x is the feature input to the second fully connected layer, and W 1 , W 2 are weights, and b 1 , b 2 are biases.

[0075] In some embodiments, layer normalization and residual connection layers are used to perform layer normalization on the outputs of the multi-head attention module and the feed-forward neural network layer, and then perform residual connection with the input features to obtain the output result of the encoder. Since the processing of the multi-head self-attention module takes into account the key degree values, that is, the attention that has incorporated key information, the output of the encoder can reflect the influence of key information, and the model can pay more attention to the keywords with higher key degree values in subsequent processing, thereby improving the accuracy of annotation and classification.

[0076] In some embodiments, the output layer includes a classifier. The classifier is used to perform average pooling operation on the features output by the encoder to obtain a vector representation with a fixed length, and then map the vector representation to the category space through a fully connected layer, and use the Softmax function to obtain the category probability distribution, which is expressed as:

[0077] p(y∣x)=softmax(W c x+b c ) (6)

[0078] where p(y∣x) represents the probability that the input feature x belongs to category y, W c is the weight, and b c is the bias.

[0079] In some embodiments, the document annotation method further includes:

[0080] Determining the classification result of the document by using the entity relationship extraction method according to the annotation information of the document.

[0081] In this embodiment, after using the document annotation model to determine the annotation information of the document to be processed, the classification result of the document is determined according to the annotation information. Among them, the annotation information of the document includes keywords in the document, such as project name, location, time, equipment information, etc. The obtained annotation information is stored in a specific storage format for convenient document classification. For example, various annotation information is stored in JSON format, and fields such as text content, annotation type, start position, and end position are added to each annotation information. The stored annotation information is used as an entity, and the multi-head attention module is used to capture the entity relationships in the document to implement the relationship extraction between entities. Based on the relationship extraction results, the document is classified according to the project stage and / or professional field, and a label of the classification result is added to the document to be processed.

[0082] In some ways, the annotation and classification results of a certain document are, for example:

[0083] Project name: XX Distribution Project Proposal

[0084] Project location: XX City XX District State Grid Power Supply Office

[0085] Project Time: Start Date (XX / XX), Estimated Completion Date (XX / XX)

[0086] Project Scale: Distribution Transformer Capacity XX kVA, Total Line Length XX km

[0087] Total Project Investment: XX million yuan

[0088] Equipment Information: Distribution Transformers, Switchgear, Ring Main Units, Cables, etc.

[0089] Classification Tags: Planning Phase, Power Distribution Project

[0090] It can be seen that using the document annotation model of this application can accurately identify the key information in the document, and classify the document accurately and effectively according to the key information of the document, improving the efficiency of project management.

[0091] In some ways, when training the document annotation model, relevant documents of power distribution engineering projects are collected to construct a dataset for annotation and classification. Among them, the collected documents include various types of documents such as design specifications, acceptance specifications, technical guidelines, technical dictionaries, equipment manuals, application guides, reports, product manuals, etc. The improved Transformer model is trained using the constructed dataset, and methods such as cross-validation are used to improve the accuracy of the model. During the training process, the model is optimized by introducing external knowledge bases such as relevant industry standards and specifications, improving the model's understanding ability of industry terms and specific contexts. The hyperparameters such as the number of model layers, the number of heads, and the embedding dimension are adjusted using the validation set, and techniques such as L1 and L2 regularization are used to prevent the model from overfitting. The diversity of training samples is increased through data augmentation techniques, thereby improving the recognition accuracy and classification performance.

[0092] In some ways, during the model training process, the cross-entropy loss function is used to measure the difference between the predicted label result of the model and the true label, expressed as:

[0093]

[0094] where N is the number of samples, C is the number of classes, y ic is the one-hot encoding of the true label, and p ic is the class probability predicted by the model.

[0095] The performance of the trained document annotation model is evaluated on the test set. Compared with general large language models, the document annotation model can effectively identify the importance of professional domain vocabulary. The accuracy of the model on the test set reaches 98%, and compared with the general Transformer model, the accuracy of annotation classification has increased by 9 percentage points, demonstrating the superiority of the document annotation model in the document annotation and classification tasks of power distribution engineering projects.

[0096] In the document annotation method provided by the embodiment of the present application, after preprocessing the document to be processed, the key degree values of each word in the document are statistically calculated. According to the key degree values, the preprocessed document is converted into a key weighted embedding vector. The key weighted embedding vector is input into a document annotation model, and the annotation information of the document is output by the document annotation model. Then, based on the annotation information, through entity relationship extraction, the key information and classification result of the document are obtained. The present application can improve the accuracy of extracting key information from distribution engineering project documents and the classification accuracy. Through automated processing, the efficiency of document processing and information retrieval is improved. By providing structured key information of the document, it supports rapid decision-making and risk management, and overall improves the project management level and decision support ability.

[0097] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario, and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.

[0098] It should be noted that the above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0099] As Figure 3 shown, the embodiment of the present application provides a document annotation device, including:

[0100] A preprocessing module, configured to preprocess the document to be processed to obtain a preprocessed document;

[0101] A statistics module, configured to statistically calculate the key degree values of each word in the preprocessed document;

[0102] A conversion module, configured to convert the preprocessed document into a key weighted embedding vector according to the key degree values;

[0103] A annotation module, configured to input the key weighted embedding vector into a pre-constructed document annotation model, and output the annotation information of the document by the document annotation model.

[0104] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in one or more software and / or hardware.

[0105] The device in the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0106] Figure 4 FIG. shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0107] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0108] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0109] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0110] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0111] The bus 1050 includes a path for transmitting information between various components of the device, such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040.

[0112] It should be noted that although only the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050 are shown in the above device, in the specific implementation process, the device may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0113] The electronic device of the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated herein.

[0114] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device.

[0115] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.

[0116] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present application may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be considered illustrative rather than restrictive.

[0117] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0118] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present disclosure.

Claims

1. A document annotation method, characterized in that: include: Preprocess the document to be processed to obtain a preprocessed document; Count the key values ​​of each word in the preprocessed document; According to the criticality value, converting the preprocessed document into a critical weighted embedding vector; The key weighted embedding vector is input into a pre-built document annotation model, and the document annotation model outputs the annotation information of the document.

2. The method according to claim 1, characterized in that According to the criticality value, converting the preprocessed document into a critical weighted embedding vector comprises: Performing word segmentation processing on the preprocessed document to obtain word units; Performing position encoding on the word unit to obtain position information of the word unit; In combination with the position information, convert the word unit into an embedding vector with the position information; A key weighted embedding vector is obtained according to the criticality value and the embedding vector with position information.

3. The method according to claim 2, characterized in that According to the criticality value and the embedding vector with position information, a critical weighted embedding vector is obtained by: ETF-IDF(t,d)=TF-IDF(t,d)×E(t) (2) Among them, TF-IDF(t,d) is the criticality value of word t in document d, E(t) is the embedding vector of word t with position information, and ETF-IDF(t,d) is the key weighted embedding vector.

4. The method according to claim 1, characterized in that The document annotation model is trained based on the Transformer model. The document annotation model includes a multi-head attention module. The calculation method of the attention score is: Among them, Q, K, and V are the query matrix, key matrix, and value matrix obtained by linear transformation of the key weighted embedding vector, respectively. k is the dimension of the key matrix, ⊙ represents element-wise multiplication; is the weight matrix calculated from the TF-IDF value of each word in the query matrix, is the weight matrix calculated from the TF-IDF value of each word in the key matrix.

5. The method according to claim 1, characterized in that Also includes: According to the annotation information of the document, an entity relationship extraction method is adopted to determine the classification result of the document to be processed.

6. A document annotation device, characterized in that: include: A preprocessing module, used for preprocessing the document to be processed to obtain a preprocessed document; A statistical module, used to count the key value of each word in the preprocessed document; A conversion module, configured to convert the preprocessed document into a key weighted embedding vector according to the key value; The annotation module is used to input the key weighted embedding vector into a pre-built document annotation model, and the document annotation model outputs the annotation information of the document.

7. The device according to claim 6, characterized in that The conversion module is used to perform word segmentation on the preprocessed document to obtain word units; perform position encoding on the word units to obtain position information of the word units; combine the position information to convert the word units into an embedding vector with position information; and obtain a key weighted embedding vector based on the criticality value and the embedding vector with position information.

8. The device according to claim 7, characterized in that The method for the conversion module to obtain the key weighted embedding vector is: ETF-IDF(t,d)=TF-IDF(t,d)×E(t) (2) Among them, TF-IDF(t,d) is the criticality value of word t in document d, E(t) is the embedding vector of word t with position information, and ETF-IDF(t,d) is the key weighted embedding vector.

9. The device according to claim 6, characterized in that The document annotation model is trained based on the Transformer model. The document annotation model includes a multi-head attention module. The calculation method of the attention score is: Among them, Q, K, and V are the query matrix, key matrix, and value matrix obtained by linear transformation of the key weighted embedding vector, respectively. k is the dimension of the key vector, ⊙ represents element-wise multiplication; is the weight matrix calculated from the TF-IDF value of each word in the query matrix, is the weight matrix calculated from the TF-IDF value of each word in the key matrix.

10. The device according to claim 6, characterized in that Also includes: The classification module is used to determine the classification result of the document to be processed by adopting an entity relationship extraction method according to the annotation information of the document.