Document processing method, electronic device and program product
By automatically obtaining and parsing images and content in power documents, identifying named entities, and performing document naming and page adjustments, the problems of inefficient and poor accuracy of document processing in the power industry are solved, and efficient document management and archiving are achieved.
Patent Information
- Application Number
- CN202510085508.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The existing document processing methods in the power industry are inefficient and have poor accuracy, and cannot effectively deal with named entities and structured information in complex power documents.
By responsive to the upload operation of document images in the power system, the document image is acquired and the document content is determined, the content is parsed to identify the named entity, and the document naming and page adjustment is carried out according to the document type and named entity, and automated document processing and storage are achieved.
It realizes automatic acquisition and parsing of power documents, improves the accuracy of document type recognition, and ensures the accuracy and management efficiency of document naming and archiving.
Smart Images

Figure CN120011476A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a document processing method, electronic equipment and program product. Background Art
[0002] With the advancement of digital transformation of power enterprises, the number and complexity of power documents are increasing. In the power industry, document processing relies on manual processing, including scanning, classification, naming and archiving of documents. However, manual processing is inefficient and prone to errors, especially in the process of document classification and archiving. Manual processing can lead to problems such as irregular document management and classification errors, affecting the accuracy and efficiency of information retrieval. At the same time, manual processing cannot effectively deal with complex power documents involving named entities and structured information in professional fields, such as power operating procedures, equipment maintenance records, etc., which increases maintenance costs and the burden of manual intervention. Summary of the invention
[0003] The present invention provides a document processing method, electronic device and program product to solve the problems in the power industry that the existing power document processing methods are low in efficiency, poor in accuracy and cannot effectively deal with complex power documents involving named entities and structured information in professional fields.
[0004] According to one aspect of the present invention, there is provided a document processing method, the method comprising:
[0005] In response to an upload operation of a document image of a first document in the power system, acquiring the document image, and determining the document content of the first document according to the document image;
[0006] Parsing the document content to obtain a named entity of the first document, and determining a document type of the first document according to the named entity and the document content;
[0007] The first document is named according to the document type and the named entity, and the page of the first document is adjusted according to the document type and the named entity to obtain a second document, and the second document is stored according to the document naming.
[0008] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:
[0009] at least one processor; and
[0010] a memory communicatively connected to at least one processor; wherein,
[0011] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by at least one processor so that the at least one processor can execute the document processing method of any embodiment of the present invention.
[0012] According to another aspect of the present invention, an embodiment of the present invention further provides a computer program product, including a computer program, which implements any document processing method in the embodiments of the present invention when executed by a processor.
[0013] The technical solution of the embodiment of the present invention obtains the document image in response to the upload operation of the document image of the first document in the power system, and determines the document content of the first document according to the document image, thereby realizing automatic acquisition of the document image and document content of the first document, reducing manual intervention, and providing data support for subsequent document content parsing and information extraction; performing content parsing on the document content to obtain the named entity of the first document, and determining the document type of the first document according to the named entity and the document content, thereby realizing identification of key information in the document, and determining the named entity and document type of the first document to ensure the accuracy of document type identification; naming the first document according to the document type and the named entity, and adjusting the page of the first document according to the document type and the named entity to obtain the second document, and storing the second document according to the document naming, thereby realizing obtaining the second document after performing document naming and page adjustment on the first document, and storing the second document, thereby improving the archiving accuracy and management efficiency of the documents.
[0014] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 A flowchart of a document processing method provided in Embodiment 1 of the present invention;
[0017] Figure 2 A flowchart of a document processing method provided in Embodiment 2 of the present invention;
[0018] Figure 3 A flowchart of a document processing method provided in Embodiment 3 of the present invention;
[0019] Figure 4 A flowchart of a document processing method provided in Embodiment 4 of the present invention;
[0020] Figure 5 A structural diagram of a document processing device provided in Embodiment 5 of the present invention;
[0021] Figure 6 A structural diagram of an electronic device for implementing a document processing method provided in Embodiment 6 of the present invention. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0024] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0025] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes, and are not used to limit the scope of these messages or information.
[0026] It is understandable that before using the technical solutions disclosed in the embodiments of the present invention, the type, scope of use, usage scenarios, etc. of the personal information involved in the present invention should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0027] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present invention according to the prompt message.
[0028] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0029] It is understandable that the above notification and user authorization process is only illustrative and does not constitute a limitation on the implementation of the present invention. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present invention.
[0030] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0031] Embodiment 1
[0032] Figure 1 A flowchart of a document processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to document processing of power documents in the power industry. The method can be executed by a document processing device, which can be implemented in the form of hardware and / or software, and optionally, by an electronic device, which can be a mobile terminal, a PC, or a server, etc.
[0033] like Figure 1 As shown, the method may specifically include:
[0034] S101 . In response to an upload operation of a document image of a first document in a power system, the document image is acquired, and document content of the first document is determined according to the document image.
[0035] In an embodiment of the present invention, the first document may refer to an electric power document generated and used at various stages in the electric power system. The first document records information such as technical parameters, operating procedures, management regulations, engineering progress, and safety records of the electric power system. Exemplarily, the first document may be a file, a chart, or a report. The document image may refer to image information obtained after image preprocessing of the first document. Image preprocessing includes image denoising, correction, and binarization. The document content may refer to the specific text content recorded in the first document.
[0036] Specifically, in response to the upload operation of the document image of the first document in the power system, the first document can be obtained through a network interface or a local upload method. Furthermore, the document image of the first document can be obtained after image preprocessing of the first document. Finally, the document content of the first document can be obtained after text recognition of the document image of the first document, providing data support for subsequent document content analysis and information extraction.
[0037] Optionally, image preprocessing of the first document may be performed by removing noise in the image using a Gaussian filter or a median filter algorithm, correcting image tilt by edge detection and affine transformation to ensure that the text is in a horizontal state, and converting the image into a black and white binary image using an adaptive threshold method to highlight the text outline.
[0038] Optionally, optical character recognition (OCR) technology can be used to perform text recognition on the document image of the first document, specifically: first, a convolutional neural network (CNN) is used to extract features from the image to identify and locate the text area; then, the identified text area is cut into lines to accurately separate each line of text; finally, a recurrent neural network (RNN) and a long short-term memory network (LSTM) are used to recognize each character.
[0039] For example, if the first document is a maintenance report of an electric equipment, key information such as the name of the electric equipment, the maintenance date, and the fault type can be identified, and the recognition result can be corrected and optimized through the language model and context information to obtain the final document content. If the first document is a multi-page document, the first document is segmented into pages so that the multi-page document can be processed separately.
[0040] S102: parse the document content to obtain a named entity of the first document, and determine the document type of the first document according to the named entity and the document content.
[0041] In the embodiment of the present invention, the named entity may refer to a specific name or identifier related to the power industry identified in the document content of the first document. For example, the named entity may be power equipment, technical terms, fault types, or operating procedures.
[0042] Specifically, by parsing the document content of the first document, the named entity of the first document can be obtained. Further, the document type of the first document can be determined based on the named entity and the document content of the first document. Exemplarily, the document type of the first document can be determined based on a preset document type determination method in combination with the named entity and the document content of the first document.
[0043] S103: Name the first document according to the document type and the named entity, and adjust the page of the first document according to the document type and the named entity to obtain a second document, and store the second document according to the document name.
[0044] Specifically, the first document may be named and page adjusted according to the document type and the named entity to form a second document. Furthermore, the second document may be stored according to the document name to facilitate search and management.
[0045] Exemplarily, the first document can be named based on a preset document naming method in combination with the document type and the named entity; the first document can be page adjusted based on a preset page adjustment method in combination with the document type and the named entity; the second document can be stored based on a preset storage method in combination with the document naming.
[0046] Exemplarily, the second document can be stored through an intelligent filing system to achieve document archiving. The intelligent filing system first uses a predefined document classification method and determines the storage location of the document in the document management system based on the document type, content and metadata information. Then, the index information of the document is automatically generated, including the document name, type, creation time, keywords, etc., and stored in the index database of the intelligent filing system for subsequent retrieval and management. At the same time, version control is performed on the document. If it is an update to an existing document, a new version will be created and the historical version will be retained. The intelligent filing system will set corresponding access rights for the document based on the preset access rights determination method to ensure the security of sensitive information.
[0047] As an optional implementation of an embodiment of the present invention, a first document is page-adjusted according to a document type and a named entity to obtain a second document, including: performing page layout segmentation on the first document through an intelligent segmentation algorithm to obtain multiple page elements of the first document; matching the multiple page elements with template elements in a preset layout template, and adjusting the page elements in the first document according to the matching results to obtain the second document.
[0048] In an embodiment of the present invention, the intelligent segmentation algorithm can use image processing technology, such as edge detection and region segmentation, to identify various page elements in the document, including titles, text, tables, and pictures, etc. The layout template defines the standard layout of different types of documents, including specifications such as page margins, font size, and line spacing.
[0049] Specifically, after the document name of the first document is determined, the first document can be segmented by a smart segmentation algorithm to identify multiple page elements of the first document. Then, the identified multiple page elements are matched with template elements in a preset layout template to analyze and adjust the page layout of the page elements in the first document according to the matching results, thereby obtaining a second document.
[0050] The technical solution of the embodiment of the present invention obtains the document image in response to the upload operation of the document image of the first document in the power system, and determines the document content of the first document according to the document image, thereby realizing automatic acquisition of the document image and document content of the first document, reducing manual intervention, and providing data support for subsequent document content parsing and information extraction; performing content parsing on the document content to obtain the named entity of the first document, and determining the document type of the first document according to the named entity and the document content, thereby realizing identification of key information in the document, and determining the named entity and document type of the first document to ensure the accuracy of document type identification; naming the first document according to the document type and the named entity, and adjusting the page of the first document according to the document type and the named entity to obtain the second document, and storing the second document according to the document naming, thereby realizing obtaining the second document after performing document naming and page adjustment on the first document, and storing the second document, thereby improving the archiving accuracy and management efficiency of the documents.
[0051] Embodiment 2
[0052] Figure 2 A flowchart of a document processing method provided for the second embodiment of the present invention. The technical solution of this embodiment further refines the process of parsing the document content in the above embodiment to obtain the named entity of the first document. For the specific implementation method, please refer to the description of this embodiment. Among them, the technical features that are the same or similar to the above embodiment are not repeated here.
[0053] like Figure 2 As shown, the method may specifically include:
[0054] S201 . In response to an upload operation of a document image of a first document in a power system, acquire the document image, and determine the document content of the first document according to the document image.
[0055] S202: Perform text segmentation processing on the document content, and segment the document content into a text sequence including a plurality of text segments according to a preset text length.
[0056] Specifically, based on the preset text length, the document content can be segmented to form a text sequence. The preset text length may refer to the preset number of characters used when segmenting the document content. The text sequence contains multiple text segments of the preset text length. For example, if the document content contains 10,000 characters and the preset text length is 256 characters, the sliding window technology can be used to segment the document content into a text sequence containing 40 text segments to ensure that there is overlap between the text segments and maintain the coherence of the context.
[0057] S203: Use a predefined electric power field dictionary to perform vocabulary matching on each text segment in the text sequence to obtain vocabulary information corresponding to each character in the text segment.
[0058] Specifically, the predefined electric power domain dictionary can be constructed by electric power professional terms. Exemplarily, electric power professional terms include transformer, circuit breaker, insulation, etc. Then, the predefined electric power domain dictionary is used to perform vocabulary matching on each text segment in the text sequence to obtain vocabulary information corresponding to each character in each text segment.
[0059] Exemplarily, the maximum forward matching algorithm can be used to perform vocabulary matching on each text segment in the text sequence, scanning each text segment from left to right, and giving priority to matching the longest vocabulary. At the same time, for each character, the position of the character in the vocabulary can be marked. For example, the starting position of the character in the vocabulary is marked as B, the middle position is marked as M, the end position is marked as E, and the single word is marked as S. "Main transformer" can be marked as "B-main, M-transformer, E-pressure, E-device".
[0060] S204, using a pre-trained language model and an encoder to encode the text sequence, vocabulary information and predefined entity category description text to obtain a context-related character vector representation and a label knowledge vector.
[0061] In the embodiments of the present invention, a pre-trained language model may refer to a pre-trained language model that can encode text. Exemplarily, the pre-trained language model may be a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model based on the Transformer architecture. The predefined entity category description text may refer to the text that predefines and explains entity categories, such as "Device Name: The name of the power device mentioned in the text". The character vector representation may refer to representing characters in vector form. The label knowledge vector may refer to the vector representing label knowledge.
[0062] Specifically, by inputting the text sequence, lexical information, and predefined entity category description text into the pre-trained language model and the encoder, context-related character vector representations and label knowledge vectors can be obtained after encoding processing.
[0063] Exemplarily, when the pre-trained language model is a BERT model, the BERT model contains 12 layers of Transformer encoders, with 12 attention heads in each layer, and the BERT model can capture long-range dependencies through the self-attention mechanism and incorporate the output of each layer into the lexical information. Furthermore, for the text sequence "CLS, main, transformer, fault, SEP", the BERT model will output a 768-dimensional vector sequence, with each vector corresponding to the context representation of a character. At the same time, for the predefined entity category description text, the BERT model can use independent parameters to distinguish different input types and output each entity category in vector representation.
[0064] As an alternative implementation of the embodiments of the present invention, encoding the text sequence, lexical information, and predefined entity category description text using the pre-trained language model and the encoder to obtain context-related character vector representations and label knowledge vectors includes: performing embedding processing on the text sequence and lexical information through the embedding layer of the pre-trained language model to obtain initial character embeddings and lexical embeddings; performing lexical information fusion processing on each layer in the pre-trained language model according to the initial character embeddings and lexical embeddings through the attention mechanism to obtain intermediate character representations with fused lexical information; performing multi-level feature extraction processing on the intermediate character representations through the multi-head self-attention mechanism and the feed-forward neural network of the pre-trained language model to obtain context-related character vector representations; encoding the predefined entity category description text through the pre-trained language model to obtain initial label knowledge vectors, and performing label knowledge enhancement processing on the initial label knowledge vectors to obtain label knowledge vectors.
[0065] Specifically, the text sequence and vocabulary information can be embedded through the embedding layer of the pre-trained language model to obtain initial character embedding and vocabulary embedding. Furthermore, the initial character embedding and vocabulary embedding can be fused with vocabulary information using a multi-head self-attention mechanism at each layer in the pre-trained language model to obtain an intermediate character representation that fuses vocabulary information. Finally, based on the multi-head self-attention mechanism and feedforward neural network in the pre-trained language model, multi-level feature extraction is performed on the intermediate character representation to obtain a context-related character vector representation. Among them, the multi-head self-attention mechanism allows the model to focus on different representation subspaces at the same time, enhancing the model's performance in capturing complex language features.
[0066] Exemplarily, when the pre-trained language model is a BERT model, the embedding layer of the BERT model includes: character embedding, position embedding, and segment embedding. For a text sequence, each character is mapped to a vector space of fixed dimension, usually 768 dimensions. At the same time, the vocabulary information is also converted into a vector representation of the same dimension. For example, for the input sequence "main transformer failure", each character will get an initial 768-dimensional vector representation, and the corresponding vocabulary information "B-equipment, O-fault" will also be converted into a vector form. Furthermore, each BERT layer in the BERT model performs vocabulary information fusion processing through a multi-head attention mechanism, and obtains an intermediate character representation of the fused vocabulary information. Then, each attention head will generate an attention output, and the final attention layer output can be obtained by concatenating and linearly transforming each attention output, so that the feedforward neural network performs nonlinear feature extraction on the final attention layer output to obtain a context-related character vector representation.
[0067] Specifically, the predefined entity category description text can be encoded through a pre-trained language model to obtain an initial label knowledge vector. Then, the label knowledge enhancement processing of the initial label knowledge vector is implemented through nonlinear transformation to obtain a label knowledge vector to enhance the expression performance of the label knowledge. Exemplarily, the initial label knowledge vector is transformed through a multilayer perceptron (MLP), wherein the MLP includes two hidden layers, each layer uses a linear rectified unit (ReLU) activation function, and the last layer uses a hyperbolic tangent activation function to ensure that the output is in the range of [-1,1], so that the label knowledge vector can better capture the semantic information of the entity category.
[0068] S205 , feature fusion of the character vector representation and the label knowledge vector to obtain a fused feature representation, and use a preset conditional random field to label each character in the text sequence based on the fused feature representation to obtain a named entity of the first document.
[0069] In the embodiment of the present invention, the conditional random field (CRF) may refer to a discriminative probability graph model for sequence labeling tasks. The CRF layer defines a transfer matrix A, where A ij Represents the score of transferring from label i to label j. For a sequence of length n, CRF calculates the score of label sequence y as:
[0070]
[0071] Among them, f i Indicates that the i-th character is marked as y i The CRF uses the Viterbi algorithm to find the one with the highest score among all label sequences.
[0072] Specifically, a fused feature representation can be obtained by fusing the character vector representation and the label knowledge vector. Then, each character in the text sequence of the fused feature representation is labeled using a preset conditional random field to obtain the named entity of the first document. For example, for the sequence "main transformer fault", the label sequence output by CRF can be "B-equipment, I-equipment, I-equipment, O, B-fault".
[0073] For example, the character vector representation and the label knowledge vector can be fused by a gating mechanism. i , the gate value g can be calculated i :
[0074] g i =σ(W g *[w i ;v i ; k]);
[0075] Among them, v i represents the vocabulary information vector, k represents the label knowledge vector, W g Represents the learnable parameter matrix and obtains the fusion feature a i :
[0076] a i =g i ·w i +(1-g i )*(W v *v i +W k *k);
[0077] Among them, W v and W k Both represent learnable parameter matrices.
[0078] S206: Determine the document type of the first document according to the named entity and the document content.
[0079] S207 . Name the first document according to the document type and the named entity, and adjust the page of the first document according to the document type and the named entity to obtain a second document, and store the second document according to the document name.
[0080] The technical solution of the embodiment of the present invention obtains the document image in response to the upload operation of the document image of the first document in the power system, and determines the document content of the first document according to the document image, thereby realizing automatic acquisition of the document image and document content of the first document, reducing manual intervention, and providing data support for subsequent document content parsing and information extraction; performing text segmentation processing on the document content, and dividing the document content into a text sequence containing multiple text fragments according to a preset text length; using a predefined power field dictionary to perform vocabulary matching on each text fragment in the text sequence, and obtaining vocabulary information corresponding to each character in the text fragment, thereby realizing segmentation processing on the document content, thereby accurately identifying key information in the document; using a pre-trained language model and an encoder to encode the text sequence, vocabulary information and predefined entity category description text to obtain context Related character vector representations and label knowledge vectors; feature fusion of the character vector representation and the label knowledge vector to obtain a fused feature representation, and use a preset conditional random field to annotate each character in the text sequence based on the fused feature representation to obtain a named entity of the first document, thereby realizing annotating each character in the text sequence to determine the named entity of the first document, thereby improving the accuracy of subsequent document type recognition; determining the document type of the first document according to the named entity and the document content; naming the first document according to the document type and the named entity, and adjusting the page of the first document according to the document type and the named entity to obtain a second document, and storing the second document according to the document naming, thereby realizing obtaining a second document after naming and adjusting the page of the first document, and storing the second document, thereby improving the archiving accuracy and management efficiency of the documents.
[0081] Embodiment 3
[0082] Figure 3 A flowchart of a document processing method provided for Embodiment 3 of the present invention. The technical solution of this embodiment further refines the process of determining the document type of the first document according to the named entity and the document content in the aforementioned embodiment on the basis of the technical solution of the aforementioned embodiment. For the specific implementation method, please refer to the description of this embodiment. Among them, the technical features that are the same or similar to the aforementioned embodiments are not repeated here.
[0083] like Figure 3 As shown, the method may specifically include:
[0084] S301 . In response to an upload operation of a document image of a first document in a power system, acquire the document image, and determine the document content of the first document according to the document image.
[0085] S302: parse the document content to obtain named entities of the first document.
[0086] S303: inserting the named entity into a preset position in the document content to obtain an enhanced text sequence, and encoding the enhanced text sequence to obtain a text feature representation.
[0087] Specifically, the identified named entities can be inserted as special tags into the preset positions of the document content to form an enhanced text sequence to highlight important entities and provide more information for subsequent text feature extraction. For example, the document content of "110kV main transformer has insulation aging failure" can be converted into "[Equipment] 110kV main transformer[ / Equipment] has [fault type] insulation aging failure[ / fault type]". Furthermore, the enhanced text sequence can be encoded using a pre-trained language model to generate a text feature representation that contains entity information and is relevant to the context.
[0088] S304: Perform feature extraction processing on the image content in the document image through a pre-trained feature extraction model to obtain an image feature representation.
[0089] In an embodiment of the present invention, the pre-trained feature extraction model may refer to a pre-trained model for extracting features from image content. Exemplarily, the pre-trained feature extraction model may be a Visual Geometry Group (VGG-16) consisting of 16 layers. VGG-16 can be used as a deep convolutional neural network model, including 13 convolutional layers and 3 fully connected layers.
[0090] Specifically, the image content in the document image can be input into a pre-trained feature extraction model, and after the feature extraction is processed by each layer in the feature extraction model, the output of the last layer is used as the image feature representation to extract high-level visual features in the image, such as device shape, fault phenomenon, etc. Exemplarily, the image content in the document image can be adjusted to a size of 224×224, and then passed through each layer in the feature extraction model in turn, and a 4096-dimensional vector is output as the image feature representation.
[0091] S305 , performing feature screening on the text feature representation and the image feature representation through a preset multimodal collaborative pooling module to obtain a multimodal representation after dimensionality reduction.
[0092] Specifically, after obtaining the text feature representation and image feature representation of the document content, the text feature representation and the image feature representation can be feature screened by a preset multimodal collaborative pooling module to obtain a multimodal representation after dimensionality reduction. Among them, the multimodal collaborative pooling module may refer to a technology for fusing multimodal features to effectively combine data of different modalities. Exemplarily, the multimodal collaborative pooling module can globally pool the text feature representation and the image feature representation to obtain two global feature representations, so that based on the two global features, a multimodal representation after dimensionality reduction corresponding to the document content can be obtained.
[0093] S306: Perform feature fusion processing on the multimodal representation through a preset feature fusion network to obtain a fused multimodal document representation.
[0094] Specifically, the multimodal representation can be subjected to feature fusion processing through a preset feature fusion network, thereby obtaining a fused multimodal document representation. The feature fusion network may refer to a pre-trained network capable of fusing features of multimodal representations. Exemplarily, the feature fusion network may be a cross-modal multi-granularity interactive fusion network for processing multimodal data.
[0095] As an optional implementation manner of an embodiment of the present invention, a feature fusion network includes a modal mixing offset module, an attention mechanism, and an encoder of a Transformer model; feature fusion processing is performed on the multimodal representation through a preset feature fusion network to obtain a fused multimodal document representation, including: coarse-grained feature offset processing is performed on the text representation and image feature representation in the multimodal representation through the modal mixing offset module in the feature fusion network to obtain initial cross-modal features; text features and image features in the initial cross-modal features are processed through the attention mechanism in the feature fusion network to obtain document-aware visual representation and visually enhanced text representation; multi-level feature extraction processing is performed on the document-aware visual representation and visually enhanced text representation through the encoder of the Transformer model in the feature fusion network to obtain a fused multimodal document representation.
[0096] In the embodiment of the present invention, the feature fusion network includes a modality mixing offset module, an attention mechanism, and an encoder of a Transformer model. The modality mixing offset module is used to perform coarse-grained feature offset processing by representing the text feature V∈R m ×d and image feature representation I∈R m ×l as input, where m represents the number of text paragraphs and d represents the feature dimension, and a gating mechanism is combined to control the flow of information, namely:
[0097] G = sigmoid(W f l[I;V]+bfl );
[0098]
[0099] Among them, W f l∈R 2d×d and b f l∈R d represents a learnable parameter and ⊙ represents element-wise multiplication.
[0100] In the embodiment of the present invention, the attention mechanism is used to perform fine-grained feature interaction by calculating the similarity matrix S∈R between text features and image features. m×l , where l represents the number of image features and the calculation formula is:
[0101] S=H T N 1 W sl V T ;
[0102] Among them, N 1 ∈R l×d represents the image features, W sl ∈R d×d Represents a learnable parameter matrix, and then, through the similarity matrix, document-aware visual representation and visually enhanced text representation can be generated. For example, the attention weight matrix A can be obtained through the softmax operation c =softmax(S T ), and through A document-aware visual representation is computed.
[0103] In an embodiment of the present invention, the Transformer encoder is composed of multiple identical layers stacked together, each of which contains a multi-head self-attention mechanism and a feedforward neural network. Furthermore, for the input sequence, the output of each Transformer layer is added to the input through residual connection and layer normalization. After multiple iterations, the model can capture deeper cross-modal interaction information.
[0104] Specifically, the modal mixing offset module in the feature fusion network can be used to perform coarse-grained feature offset processing, so that the text representation and image feature representation in the multimodal representation are initially fused, thereby obtaining the initial cross-modal feature. For example, for a power document containing a text description of "abnormal oil temperature of the main transformer" and a corresponding equipment picture, the modal mixing offset module will initially fuse the "abnormal oil temperature" information in the text with the visual features in the image. Furthermore, the attention mechanism in the feature fusion network can be used to perform fine-grained interactive processing on the text features and image features in the initial cross-modal features, thereby obtaining a document-aware visual representation and a visually enhanced text representation. Finally, the encoder of the Transformer model in the feature fusion network can be used to perform multi-level feature extraction processing on the document-aware visual representation and the visually enhanced text representation, so that the text feature representation and the image feature representation are effectively fused, thereby obtaining a fused multimodal document representation. Using the feature fusion network, not only the semantic information of the original text and image is retained, but also the complex interactive relationship between text and image is captured, so that a more comprehensive document understanding can be provided for the power document classification task.
[0105] For example, for a transformer fault report, the fused representation not only includes the fault type and parameter information in the text description, but also integrates visual clues in the image, such as abnormal equipment appearance or meter readings, thereby providing rich and accurate feature information for subsequent classification tasks and improving the model's ability to understand complex power documents.
[0106] S307 , classify the multimodal document representation through a preset feature aggregation network and a fully connected layer to obtain a document type of the first document.
[0107] Specifically, the fused multimodal document representation can be classified through a preset feature aggregation network and a fully connected layer to obtain the document type of the first document. The feature aggregation network uses a multi-layer document transformer to capture the relationship between paragraphs, and each layer of transformer contains a multi-head self-attention mechanism and a feedforward neural network. Furthermore, the features after the feature aggregation network aggregates the multimodal document representation can be connected to the fully connected layer and mapped to the predefined document type space to obtain the probability distribution of each type, thereby selecting the type with the highest probability as the document type of the first document.
[0108] S308: Name the first document according to the document type and the named entity, and adjust the page of the first document according to the document type and the named entity to obtain a second document, and store the second document according to the document name.
[0109] The technical solution of the embodiment of the present invention obtains the document image in response to the upload operation of the document image of the first document in the power system, and determines the document content of the first document based on the document image, thereby realizing automatic acquisition of the document image and document content of the first document, reducing manual intervention, and providing data support for subsequent document content parsing and information extraction; performing content parsing on the document content to obtain the named entities of the first document; inserting the named entities into preset positions in the document content to obtain enhanced text sequences, and encoding the enhanced text sequences to obtain text feature representations; performing feature extraction on the image content in the document image through a pre-trained feature extraction model to obtain image feature representations; performing feature screening on the text feature representation and the image feature representation through a preset multimodal collaborative pooling module to obtain a multimodal representation after dimensionality reduction. , realizing the determination of the multimodal representation after dimensionality reduction based on the text feature representation and the image feature representation; performing feature fusion processing on the multimodal representation through a preset feature fusion network to obtain a fused multimodal document representation; performing classification processing on the multimodal document representation through a preset feature aggregation network and a fully connected layer to obtain the document type of the first document, realizing the determination of the document type of the first document based on the multimodal document representation, and further improving the accuracy of document type determination; naming the first document according to the document type and the named entity, and adjusting the page of the first document according to the document type and the named entity to obtain the second document, and storing the second document according to the document naming, realizing the acquisition of the second document after the document naming and page adjustment of the first document, and storing the second document, improving the archiving accuracy and management efficiency of the documents.
[0110] Embodiment 4
[0111] Figure 4 A flowchart of a document processing method provided for Embodiment 4 of the present invention. The technical solution of this embodiment further refines the process of naming the first document according to the document type and the named entity in the aforementioned embodiment based on the technical solution of the aforementioned embodiment. For the specific implementation method, please refer to the description of this embodiment. Among them, the technical features that are the same or similar to the aforementioned embodiments are not repeated here.
[0112] like Figure 4 As shown, the method may specifically include:
[0113] S401 . In response to an upload operation of a document image of a first document in a power system, acquire the document image, and determine the document content of the first document according to the document image.
[0114] S402: parse the document content to obtain a named entity of the first document, and determine the document type of the first document according to the named entity and the document content.
[0115] S403: Obtain a naming template corresponding to the first document from a preset naming rule library according to the document type, filter the information of the named entity, and obtain core elements for document naming.
[0116] In an embodiment of the present invention, the naming rule library may refer to a database containing various power document types and their corresponding naming templates. For example, for the "equipment failure report" type, the naming template may be "[equipment name][fault type][date]".
[0117] Specifically, based on the document type of the first document, a naming template corresponding to the document type can be obtained from a naming rule library. Furthermore, a predefined importance scoring algorithm can be used to screen key information of the identified named entities by considering factors such as entity type, frequency of occurrence, and contextual relevance, and to screen core elements for document naming from multiple named entities.
[0118] S404: Combine the core elements using a naming template and an intelligent filling algorithm to obtain an initial name.
[0119] Specifically, based on the naming template and the intelligent filling algorithm, the core elements are combined and processed to obtain the initial name of the first document. Among them, the intelligent filling algorithm will take into account the special naming habits and specifications of the power industry, such as the unified use of "YYYYMMDD" in the date format and the use of standard abbreviations for equipment names. For example, for a fault report describing the abnormal oil temperature of the main transformer, an initial document name of "Main Transformer_Insulation Aging Fault Report_20240327" can be generated.
[0120] As an optional implementation mode of an embodiment of the present invention, the core elements are combined and processed through a preset naming template and an intelligent filling algorithm to obtain an initial name, including: performing semantic analysis on the core elements through a natural language processing algorithm to obtain the semantic characteristics and importance scores of each element; sorting the core elements according to the semantic characteristics and importance scores to obtain an ordered element sequence; matching the element sequence and the naming template through a dynamic programming algorithm to obtain an element filling scheme; filling the naming template according to the element filling scheme through a context-aware text generation algorithm to obtain an initial name.
[0121] Specifically, the core elements can be semantically analyzed by a natural language processing algorithm. For example, each core element is converted into a dense vector representation using a pre-trained word embedding model, and then the semantic features of the element in the context are captured using an attention mechanism and a bidirectional long short-term memory network. Furthermore, the importance score of each core element is obtained by combining the frequency, position and semantic relevance of each core element in the document with other elements. For example, the importance score of each element is calculated using a hybrid algorithm based on term frequency-inverse document frequency (TF-IDF) and text ranking algorithm (TextRank). Exemplarily, for a transformer fault report, core elements such as "main transformer" and "insulation aging" will receive a higher importance score. Then, based on the semantic features and importance scores, the core elements are prioritized to obtain an ordered sequence of elements. For example, an improved quick sort can be used, with the importance score as the primary sorting key and the semantic similarity as the secondary sorting key to ensure that important core elements are given priority and that elements with similar semantics are clustered together to facilitate the generation of coherent document names. The dynamic programming algorithm can be used to fill the elements in the element sequence into the corresponding positions of the naming template, so as to match the ordered element sequence with the preset naming template and obtain the element filling scheme. The state of the dynamic programming algorithm can be defined as dp[i][j], which represents the score of filling the first j slots of the template with the first i elements; the transfer equation combines the importance score of the element, the semantic matching degree of the slot, and the continuity of the filling; the time complexity of the algorithm is O(mn), where m represents the number of elements and n represents the number of template slots. Finally, the naming template is intelligently filled according to the element filling scheme through the context-aware text generation algorithm to generate the initial name of the first document. For example, using a pre-trained language model based on Transformer, the elements in the filling scheme are used as input, combined with the professional terminology and naming specifications of the power industry, to generate a document name that conforms to the natural language expression habits, so as to ensure that the generated name can accurately convey the document content and meet industry standards.
[0122] S405: Perform a repetitive search on the document library according to the initial name to obtain a search result, and determine the document name of the first document according to the initial name and the search result.
[0123] Specifically, an efficient string matching algorithm, such as a fast pattern matching algorithm (Knuth-Morris-Pratt, KMP), can be used to quickly search whether there is a document name in the document library that is the same as the initial name of the first document, and use it as the search result. If the search result indicates that there is a document name in the document library that is the same as the initial name of the first document, the initial document name is adjusted until it is different from the document names in the document library, and used as the document name of the first document. Exemplarily, an increasing numeric suffix or timestamp is added to the end of the initial name of the first document to ensure the uniqueness of the document name.
[0124] S406: Adjust the page of the first document according to the document type and the named entity to obtain a second document, and store the second document according to the document name.
[0125] The technical solution of the embodiment of the present invention obtains the document image in response to the upload operation of the document image of the first document in the power system, and determines the document content of the first document according to the document image, thereby realizing automatic acquisition of the document image and document content of the first document, reducing manual intervention, and providing data support for subsequent document content parsing and information extraction; performing content parsing on the document content to obtain the named entity of the first document, determining the document type of the first document according to the named entity and the document content, realizing identification of key information in the document, and determining the named entity and document type of the first document to ensure the accuracy of document type identification; obtaining the naming template corresponding to the first document from a preset naming rule library according to the document type, Information of named entities is filtered to obtain core elements for document naming, thereby realizing the determination of core elements for document naming based on document type; the core elements are combined and processed through naming templates and intelligent filling algorithms to obtain initial naming; a document library is repeatedly searched according to the initial naming to obtain search results, and a document name of the first document is determined according to the initial naming and the search results, thereby realizing the determination of the document name of the first document based on the initial naming and the search results, making the document name of the first document unique; the page of the first document is adjusted according to the document type and the named entity to obtain a second document, and the second document is stored according to the document name, thereby realizing the storage of the second document and improving the archiving accuracy and management efficiency of the documents.
[0126] Embodiment 5
[0127] Figure 5 This is a structural diagram of a document processing device provided in Embodiment 5 of the present invention. This embodiment is applicable to document processing of power documents in the power industry. Figure 5As shown, the document processing device includes: a first determination module 501, a second determination module 502 and a document processing module 503. The first determination module 501 is used to respond to the upload operation of the document image of the first document in the power system, obtain the document image, and determine the document content of the first document according to the document image; the second determination module 502 is used to parse the document content to obtain the named entity of the first document, and determine the document type of the first document according to the named entity and the document content; the document processing module 503 is used to name the first document according to the document type and the named entity, and adjust the page of the first document according to the document type and the named entity to obtain the second document, and store the second document according to the document naming.
[0128] The technical solution of the embodiment of the present invention is to obtain the document image in response to the upload operation of the document image of the first document in the power system through the first determination module 501, and determine the document content of the first document according to the document image, thereby realizing the automatic acquisition of the document image and document content of the first document, reducing manual intervention, and providing data support for subsequent document content parsing and information extraction; the second determination module 502 performs content parsing on the document content to obtain the named entity of the first document, and determines the document type of the first document according to the named entity and the document content, thereby realizing the identification of key information in the document, and determining the named entity and document type of the first document to ensure the accuracy of document type identification; the document processing module 503 performs document naming on the first document according to the document type and the named entity, and performs page adjustment on the first document according to the document type and the named entity to obtain the second document, and stores the second document according to the document naming, thereby realizing the second document after the document naming and page adjustment of the first document, and storing the second document, thereby improving the archiving accuracy and management efficiency of the documents.
[0129] On the basis of any of the above optional technical solutions, optionally, the second determination module 502 includes: a text segmentation unit, a first acquisition unit, a second acquisition unit, and a third acquisition unit. Among them, the text segmentation unit is used to perform text segmentation processing on the document content, and divide the document content into a text sequence containing multiple text fragments according to a preset text length; the first acquisition unit is used to use a predefined electric power field dictionary to perform vocabulary matching on each text fragment in the text sequence to obtain vocabulary information corresponding to each character in the text fragment; the second acquisition unit is used to use a pre-trained language model and an encoder to encode the text sequence, vocabulary information, and predefined entity category description text to obtain context-related character vector representation and label knowledge vector; the third acquisition unit is used to perform feature fusion on the character vector representation and the label knowledge vector to obtain a fused feature representation, and use a preset conditional random field to perform labeling processing on each character in the text sequence based on the fused feature representation to obtain the named entity of the first document.
[0130] Based on any of the above optional technical solutions, optionally, the second acquisition unit is specifically used to: embed text sequences and vocabulary information through the embedding layer of the pre-trained language model to obtain initial character embedding and vocabulary embedding; fuse vocabulary information of each layer in the pre-trained language model according to the initial character embedding and vocabulary embedding through the attention mechanism to obtain an intermediate character representation of the fused vocabulary information; perform multi-level feature extraction on the intermediate character representation through the multi-head self-attention mechanism and feedforward neural network of the pre-trained language model to obtain a context-related character vector representation; encode the predefined entity category description text through the pre-trained language model to obtain an initial label knowledge vector, and perform label knowledge enhancement on the initial label knowledge vector to obtain a label knowledge vector.
[0131] On the basis of any of the above optional technical solutions, optionally, the second determination module 502 also includes: a fourth acquisition unit, a fifth acquisition unit, a sixth acquisition unit, a seventh acquisition unit and an eighth acquisition unit. Among them, the fourth acquisition unit is used to insert the named entity into a preset position in the document content to obtain an enhanced text sequence, and encode the enhanced text sequence to obtain a text feature representation; the fifth acquisition unit is used to perform feature extraction processing on the image content in the document image through a pre-trained feature extraction model to obtain an image feature representation; the sixth acquisition unit is used to perform feature screening on the text feature representation and the image feature representation through a preset multimodal collaborative pooling module to obtain a multimodal representation after dimensionality reduction; the seventh acquisition unit is used to perform feature fusion processing on the multimodal representation through a preset feature fusion network to obtain a fused multimodal document representation; the eighth acquisition unit is used to classify the multimodal document representation through a preset feature aggregation network and a fully connected layer to obtain the document type of the first document.
[0132] On the basis of any of the above optional technical solutions, optionally, the feature fusion network includes a modal mixing offset module, an attention mechanism, and an encoder of a Transformer model. Correspondingly, the seventh acquisition unit is specifically used to: perform coarse-grained feature offset processing on the text representation and image feature representation in the multimodal representation through the modal mixing offset module in the feature fusion network to obtain initial cross-modal features; process the text features and image features in the initial cross-modal features through the attention mechanism in the feature fusion network to obtain document-aware visual representation and visually enhanced text representation; perform multi-level feature extraction processing on the document-aware visual representation and visually enhanced text representation through the encoder of the Transformer model in the feature fusion network to obtain a fused multimodal document representation.
[0133] Based on any of the above optional technical solutions, optionally, the document processing module 503 includes: a core element acquisition unit, an initial naming acquisition unit and a document naming determination unit. The core element acquisition unit is used to obtain a naming template corresponding to the first document from a preset naming rule library according to the document type, and to perform information screening on the named entity to obtain a core element for document naming; the initial naming acquisition unit is used to combine and process the core element through the naming template and the intelligent filling algorithm to obtain an initial name; the document naming determination unit is used to perform a repetitive search on the document library according to the initial name to obtain a search result, and determine the document name of the first document according to the initial name and the search result.
[0134] Based on any of the above optional technical solutions, optionally, the initial naming acquisition unit is specifically used to: perform semantic analysis on the core elements through a natural language processing algorithm to obtain the semantic features and importance scores of each element; sort the core elements according to the semantic features and importance scores to obtain an ordered element sequence; match the element sequence and the naming template through a dynamic programming algorithm to obtain an element filling scheme; fill the naming template according to the element filling scheme through a context-aware text generation algorithm to obtain an initial name.
[0135] Based on any of the above optional technical solutions, optionally, the document processing module 503 further includes: a page layout segmentation unit and a page element adjustment unit. The page layout segmentation unit is used to perform page layout segmentation on the first document through an intelligent segmentation algorithm to obtain multiple page elements of the first document; the page element adjustment unit is used to match the multiple page elements with template elements in a preset layout template, and adjust the page elements in the first document according to the matching results to obtain the second document.
[0136] The document processing device provided in the embodiment of the present invention can execute the document processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0137] Embodiment 6
[0138] Figure 6 A schematic diagram of the structure of an electronic device for implementing a document processing method provided for Embodiment 6 of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0139] like Figure 6 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0140] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0141] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a document processing method.
[0142] In some embodiments, the document processing method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the document processing method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the document processing method in any other appropriate manner (e.g., by means of firmware).
[0143] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0144] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0145] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0147] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0148] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0149] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0150] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A document processing method, characterized in that: include: In response to an upload operation of a document image of a first document in the power system, acquiring the document image, and determining the document content of the first document according to the document image; Parsing the document content to obtain a named entity of the first document, and determining a document type of the first document according to the named entity and the document content; The first document is named according to the document type and the named entity, and the first document is page-adjusted according to the document type and the named entity to obtain a second document, and the second document is stored according to the document naming.
2. The method according to claim 1, characterized in that The performing content parsing on the document content to obtain the named entity of the first document includes: Performing text segmentation processing on the document content, and segmenting the document content into a text sequence including a plurality of text segments according to a preset text length; Using a predefined electric power field dictionary to perform vocabulary matching on each of the text segments in the text sequence to obtain vocabulary information corresponding to each character in the text segment; Using a pre-trained language model and an encoder, the text sequence, the vocabulary information, and the predefined entity category description text are encoded to obtain a context-related character vector representation and a label knowledge vector; The character vector representation and the label knowledge vector are feature fused to obtain a fused feature representation, and a preset conditional random field is used to label each character in the text sequence based on the fused feature representation to obtain a named entity of the first document.
3. The method according to claim 2, characterized in that The method of encoding the text sequence, the vocabulary information and the predefined entity category description text using the pretrained language model and the encoder to obtain a context-related character vector representation and a label knowledge vector includes: Embedding the text sequence and the vocabulary information through the embedding layer of the pre-trained language model to obtain initial character embedding and vocabulary embedding; Performing vocabulary information fusion processing on each layer in the pre-trained language model according to the initial character embedding and vocabulary embedding through an attention mechanism to obtain an intermediate character representation of the fused vocabulary information; Performing multi-level feature extraction processing on the intermediate character representation through the multi-head self-attention mechanism and feed-forward neural network of the pre-trained language model to obtain a context-related character vector representation; The predefined entity category description text is encoded by the pre-trained language model to obtain an initial label knowledge vector, and the initial label knowledge vector is subjected to label knowledge enhancement processing to obtain a label knowledge vector.
4. The method according to claim 1, characterized in that The determining the document type of the first document according to the named entity and the document content includes: Inserting the named entity into a preset position in the document content to obtain an enhanced text sequence, and encoding the enhanced text sequence to obtain a text feature representation; Performing feature extraction processing on the image content in the document image by using a pre-trained feature extraction model to obtain an image feature representation; Performing feature screening on the text feature representation and the image feature representation through a preset multimodal collaborative pooling module to obtain a multimodal representation after dimensionality reduction; Performing feature fusion processing on the multimodal representation through a preset feature fusion network to obtain a fused multimodal document representation; The multimodal document representation is classified through a preset feature aggregation network and a fully connected layer to obtain the document type of the first document.
5. The method according to claim 4, characterized in that The feature fusion network includes a modality mixing offset module, an attention mechanism, and an encoder of a Transformer model; the feature fusion processing is performed on the multimodal representation through a preset feature fusion network to obtain a fused multimodal document representation, including: Performing coarse-grained feature shift processing on the text representation and the image feature representation in the multimodal representation through a modality mixing shift module in a feature fusion network to obtain an initial cross-modal feature; Processing the text features and image features in the initial cross-modal features through an attention mechanism in a feature fusion network to obtain a document-aware visual representation and a visually enhanced text representation; The encoder of the Transformer model in the feature fusion network performs multi-level feature extraction processing on the document-aware visual representation and the visually enhanced text representation to obtain a fused multimodal document representation.
6. The method according to claim 1, characterized in that The step of naming the first document according to the document type and the named entity includes: Acquire a naming template corresponding to the first document from a preset naming rule library according to the document type, filter information of the named entity, and obtain core elements for document naming; The core elements are combined and processed by the naming template and the intelligent filling algorithm to obtain an initial name; A repetitive search is performed on the document library according to the initial name to obtain a search result, and a document name of the first document is determined according to the initial name and the search result.
7. The method according to claim 6, characterized in that The core elements are combined and processed by a preset naming template and an intelligent filling algorithm to obtain an initial name, including: Performing semantic analysis on the core elements through a natural language processing algorithm to obtain semantic features and importance scores of each element; Sorting the core elements according to the semantic features and the importance scores to obtain an ordered element sequence; Matching the element sequence and the naming template by a dynamic programming algorithm to obtain an element filling scheme; The naming template is filled in according to the element filling scheme through a context-aware text generation algorithm to obtain an initial name.
8. The method according to claim 1, characterized in that The step of adjusting the page of the first document according to the document type and the named entity to obtain the second document includes: Performing page layout segmentation on the first document by using an intelligent segmentation algorithm to obtain a plurality of page elements of the first document; The plurality of page elements are matched with template elements in a preset layout template, and the page elements in the first document are adjusted according to the matching result to obtain a second document.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the document processing method according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the document processing method according to any one of claims 1 to 8.
Citation Information
Cited By
Clinical test document quality detection and processing method and system, and computer equipment
CN120257941A