Multimodal Document Structuring and Knowledge Extraction Method Based on Large Language Models
By adopting multimodal document structured processing and knowledge extraction methods based on large language models in multimodal document processing, the problem of insufficient multimodal data fusion in the prior art is solved, and efficient automated processing and more accurate semantic information extraction are achieved.
Patent Information
- Application Number
- CN202411366962.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-09-29
AI Technical Summary
The prior art lacks effective multimodal data fusion methods when processing multimodal documents, resulting in limitations in information extraction and knowledge representation, requiring a lot of manual intervention, and unable to achieve efficient automated processing.
The multimodal document structured processing and knowledge extraction method based on large language models are adopted. By pre-processing, feature extraction and fusion of text and non-text data in multimodal documents, the improved BERT model is used for in-depth semantic analysis, and the knowledge graph is constructed to realize automated knowledge extraction and structured processing.
It improves the accuracy of correlation recognition between text and image and graph data, enhances the system's understanding of complex documents, realizes the extraction of more accurate and comprehensive semantic information from multimodal data, reduces manual intervention, and improves processing efficiency.
Smart Images

Figure CN119227794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal documents, and in particular, to a method for structured processing and knowledge extraction of multimodal documents based on a large language model. Background Art
[0002] In the prior art, the processing of unstructured and multimodal documents mainly relies on single-modal data processing methods, such as only processing a certain type of data in text, images, or charts. Although the processing methods in the prior art can effectively extract and process information in some scenarios, when faced with complex documents containing multiple types of data, the prior art seems powerless. The traditional single-modal processing method cannot effectively fuse and analyze multiple types of data, resulting in obvious limitations in information extraction and knowledge representation. Especially when a large number of unstructured and multimodal documents need to be processed automatically, the traditional technology often relies on a large amount of manual intervention and cannot achieve efficient automatic processing.
[0003] In the prior art, text processing technologies such as natural language processing have been relatively well-developed. In particular, the application of the BERT large language model can achieve certain text semantic understanding, word segmentation, and part-of-speech tagging functions. However, most of these are only for pure text data and lack the ability to process non-text data. For feature extraction technologies for images and charts, convolutional neural network visual models are mostly used, but the models cannot be deeply integrated with text data, resulting in insufficient accuracy and comprehensiveness in extracting knowledge from multimodal data.
[0004] In summary, the prior art mainly has the following disadvantages: First, the existing single-modal processing method has poor applicability in multimodal document processing and lacks an effective multimodal data fusion method; second, the existing text processing and visual feature extraction technologies are separated from each other and cannot achieve the joint understanding of multimodal data. Summary of the Invention
[0005] An object of the present invention is to propose a method for structured processing and knowledge extraction of multimodal documents based on a large language model, and the present invention realizes the structured processing and knowledge extraction of unstructured and multimodal documents.
[0006] A method for structured processing and knowledge extraction of multimodal documents based on a large language model according to an embodiment of the present invention includes the following steps:
[0007] S1. Receive an input multimodal document, where the multimodal document includes at least one type of text data and at least one type of non-text data;
[0008] S2. Preprocess the text data in the multimodal document, where the preprocessing includes word segmentation, part-of-speech tagging, syntactic analysis, and entity recognition;
[0009] S3. Extract features from the non-text data in the multimodal document;
[0010] S4. Perform multimodal data fusion on the preprocessed text data and the feature-extracted non-text data;
[0011] S5. Perform in-depth semantic analysis on the fused multimodal data through a pre-trained improved BERT model;
[0012] S6. Automatically construct a knowledge graph based on the results of the in-depth semantic analysis;
[0013] S7. Output the data of the knowledge graph in a format suitable for analysis or application.
[0014] Optionally, step S1 includes:
[0015] S11. Receive a multimodal document input signal provided by multiple data sources, where the data sources include a text data source and a non-text data source. The text data source provides text data containing natural language content, and the non-text data source provides images, charts, or other types of non-text data:
[0016]
[0017] where D m represents the received multimodal document input data set, T di represents the i-th text data unit, N dj represents the j-th non-text data unit, N t and N n represent the quantities of text data and non-text data respectively, represents the parallel reception and combination operation of multimodal data;
[0018] S12. Perform a synchronous input operation on the text data T d and the non-text data N d to simultaneously import the text data and the non-text data into the processing system;
[0019] S13. Perform preliminary data type identification on the input text data T d and the non-text data N d Verify the text data type based on the metadata identifier M t Verify the non-text data type through the image feature matrix I f ;
[0020] S14. According to the identification results, the text data T d and the non-text data N dClassified and stored in the text data cache unit C t and the non-text data cache unit C n .
[0021] Optionally, the step S2 includes:
[0022] S21. Use a vocabulary matching algorithm based on the context attention mechanism to perform word segmentation on the received text data T d and execute the word segmentation operation:
[0023]
[0024] wherein, T d represents the input text data, T di represents the i-th word in the text data, L m (T d ) is the word segmentation result, α i and β i represent the global dictionary matching weight and the context similarity weight respectively, W(T di ) represents the entry matching score in the dictionary, C(T di ) represents the context word set of the word T di , Sim(T di , T j ) represents the similarity between the word T di and the context word T j ;
[0025] S22. Use a part-of-speech tagging model based on conditional random fields to perform part-of-speech tagging on the word-segmented text data. The part-of-speech tagging process is as follows:
[0026]
[0027] wherein, T di represents the i-th word, P di represents the corresponding part-of-speech tag, ψ(T di , P di ) represents the compatibility function between the word T di and its part of speech P di , φ(P di-1 , P di ) represents the part-of-speech transition probability, P s (T d ) represents the part-of-speech tagging set of the word-segmented text data;
[0028] S23. Perform syntactic analysis on the tagged text data and use a dependency tree syntactic analysis model to construct the grammatical structure tree of the text. The specific process is as follows:
[0029]
[0030] Among them, T di and T dj represent two words in the text, δ(T di , T dj ) represents the distance between words, τ(P di , P dj ) represents the dependency relationship weight between part-of-speech tags, Dep(T di , T dj ) represents the dependency relationship strength between words, S t (T d ) represents the syntactic structure tree of the text data;
[0031] S24. Use an entity recognition model based on bidirectional long short-term memory and conditional random field to perform entity recognition on the text data after syntactic analysis, and execute entity recognition:
[0032]
[0033] Among them, T di represents the i-th word in the text, E di represents the corresponding entity label, λ(T di , E di ) represents the compatibility score between the word and the entity label, μ(E di-1 , E di ) represents the transition score between entity labels, γ represents the fusion coefficient between multi-modal data, Att(T di , F j ) represents the attention weight between the text word T di and the non-text data feature F j , and E r (T d ) represents the entity recognition result in the text data.
[0034] Optionally, the step S3 includes:
[0035] S31. Extract visual features from the image data I d in the multi-modal document, and use a convolutional neural network to analyze the spatial features in the image data to extract the image feature vector V i :
[0036]
[0037] Among them, I d represents the input image data, I lk represents the l-th layer and the k-th feature point in the image, w lk represents the weight parameter, f(I lk)The non-linear activation function representing the image feature points, b lk represents the bias value, V i represents the visual feature vector of the image;
[0038] S32. Use the model based on the region proposal network to detect the target object in the image data, generate the candidate region box R i , and perform object classification on each region box:
[0039]
[0040] where, R pi represents the i-th candidate region, O i represents the true target region, P represents the number of candidate regions, φ(R pi ) is the confidence score of the region proposal, IoU(R pi , O i ) represents the intersection over union between the candidate region and the true target region;
[0041] S33. Perform structure recognition on the chart data G d in the multimodal document, and extract the key structure information S g in the chart, including the coordinate axes, data points, labels, and line relationships:
[0042]
[0043] where, G d represents the input chart data, G dj represents the j-th data element in the chart, L j represents the label associated with the data element, γ j represents the weight parameter, Rel(G dj , L j ) represents the association strength between the chart data element and its label, S g represents the structure recognition result of the chart.
[0044] Optionally, the step S4 includes:
[0045] S41. Fuse the feature vector P d (T s ) of the preprocessed text data T d with the non-text data N v after feature extraction to construct the multimodal feature matrix M f :
[0046]
[0047] where, P s (Td ) represents the part-of-speech tagging feature vector of text data, V i represents the visual feature vector of image data, S g represents the structural feature of chart data, λ 1 and λ 2 respectively represent the fusion weights of text and non-text data, M f represents the fused multi-modal feature matrix;
[0048] S42. Calculate the correlation score A between text data and non-text data based on the attention mechanism s (T d , N d ), where the word segmentation result L of text data m (T d ) and the candidate region box R of non-text data i are used to calculate the similarity score:
[0049]
[0050] where, L m (T di ) represents the word segmentation result in text data, R j represents the candidate region in non-text data, α ij represents the attention weight between text data and non-text data, Sim(L m (T di ), R j ) represents the similarity score between the word in text data and the candidate region of non-text data, A s (T d , N d ) represents the correlation score between text data and non-text data;
[0051] S43. Perform information fusion on multi-modal data based on the correlation score A s (T d , N d ) to generate the final multi-modal feature matrix M r :
[0052]
[0053] where, β ij represents the fused weight coefficient, represents the fusion operation of the word in text data with image and chart features, M r represents the finally generated multi-modal representation.
[0054] Optionally, the step S5 includes:
[0055] S51. Input the fused multi-modal feature matrix M r into the pre-trained improved BERT model to perform deep semantic analysis. The improved BERT model jointly encodes the embedded representations of text and non-text data to generate a multi-modal semantic vector S v ;
[0056] S52. Perform key entity recognition in the multi-modal semantic vector M r through the attention mechanism. The improved BERT model calculates the attention weights of each word and non-text feature to identify the key entity E k in the multi-modal data:
[0057]
[0058] where S vi represents the i-th word or feature in the multi-modal semantic vector, T di represents the corresponding text word, and α i represents the attention weight;
[0059] S53. Based on the deep semantic relation extraction module of the improved BERT model, calculate the relation R s (E k ) between entities in the multi-modal semantic vector, and identify the event Ev s in the multi-modal semantic vector through the event detection module.
[0060] Optionally, the step S6 includes:
[0061] S61. Extract the text entity node E v and the non-text entity node E t from the multi-modal semantic vector S n based on the result of the deep semantic analysis, and construct the initial node set N g of the knowledge graph:
[0062] The text entity node E t represents the person, place name, and event extracted from the text data;
[0063] The non-text entity node E n represents the object and data point extracted from the non-text data such as images and charts;
[0064] S62. Generate the edge L g between the nodes according to the semantic relation between the text data and the non-text data. The edge includes the edge representing the relation between the text description and the object in the image and the edge representing the relation between the chart data and the text description;
[0065] S63. Based on the graph structure learning model, the weights ω of the nodes and edgesn and ω l perform optimization to generate the optimized knowledge graph K g .
[0066] Optionally, the step S63 includes:
[0067] S631. Initialize the node weights ω n and edge weights ω 1 in the initial knowledge graph. The node weight ω n is initialized as the degree of the node, and the edge weight ω 1 is initialized as the association strength between nodes;
[0068] S632. Optimize the node weights and edge weights based on the gradient descent algorithm, and iteratively update the weight values of the nodes and edges and
[0069]
[0070]
[0071] where η 2 is the learning rate, and respectively represent the partial derivatives of the loss function L with respect to the node weights and edge weights;
[0072] S633. The loss function L is optimized by minimizing the error between the nodes and edges. The purpose of the optimization is to improve the accuracy and relevance of the knowledge graph by minimizing L:
[0073]
[0074] where Sim(E i , E j ) represents the similarity between node E i and node E j , and L represents the total loss between the nodes and edges;
[0075] S634. During the optimization iteration process, generate the finally optimized knowledge graph K according to the optimal node weights and edge weights g :
[0076]
[0077] where K g represents the finally optimized knowledge graph, which includes the optimized node set N g , edge set L g and node weights and edge weights
[0078] The beneficial effects of the present invention are as follows:
[0079] (1) By fusing text data with non-text data, the present invention proposes a multi-modal fusion method based on an improved BERT model. The feature extraction model is used to extract features from text and non-text data respectively, and then the features are fused through an adaptive weight mechanism to generate a multi-modal feature matrix. This not only improves the recognition accuracy of the correlation between text and image and chart data, but also enhances the system's ability to understand complex documents, enabling the system to extract more accurate and comprehensive semantic information from multi-modal data and overcoming the limitations of single-modal processing in the prior art.
[0080] (2) In the technical solution of the present invention, the node weights and edge weights of the knowledge graph are adaptively optimized based on the graph structure learning model. By introducing the gradient descent algorithm and optimizing the calculation of the loss function, the weight distribution of nodes and edges in the knowledge graph can be dynamically adjusted. The optimized knowledge graph can more accurately express the semantic relationship between text and non-text data. Compared with the existing knowledge graph construction technology, the present invention can not only process pure text data, but also effectively combine image and chart data, making the constructed knowledge graph have higher expression ability and accuracy, and ensuring that the key knowledge in complex documents can be comprehensively represented.
[0081] (3) By introducing an improved BERT model and a self-supervised learning mechanism, and combining multi-modal fusion technology, the present invention significantly improves the automation processing efficiency of complex documents. Different from the traditional technology that requires a large amount of manual intervention, the present invention can realize automated knowledge extraction and structured processing through deep semantic analysis, entity recognition and relationship extraction steps under unsupervised or weakly supervised conditions, reducing the complexity of manual operations, while significantly reducing the time cost of processing multi-modal documents and improving the overall processing efficiency of the system, thus solving the problem of low efficiency in processing large-scale and multi-modal documents in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0083] Figure 1 is a flowchart of a multi-modal document structured processing and knowledge extraction method based on a large language model proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0084] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0085] Reference Figure 1 , a multi-modal document structured processing and knowledge extraction method based on a large language model, comprising the following steps:
[0086] S1. Receive the input multi-modal document, which contains at least one type of text data and at least one type of non-text data;
[0087] S2. Preprocess the text data in the multi-modal document, and the preprocessing includes word segmentation, part-of-speech tagging, syntactic analysis, and entity recognition;
[0088] S3. Extract features from the non-text data in the multi-modal document;
[0089] S4. Perform multi-modal data fusion on the preprocessed text data and the feature-extracted non-text data;
[0090] S5. Perform in-depth semantic analysis on the fused multi-modal data through a pre-trained improved BERT model;
[0091] S6. Automatically construct a knowledge graph based on the results of the in-depth semantic analysis;
[0092] S7. Output the data of the knowledge graph in a format suitable for analysis or application.
[0093] In this embodiment, step S1 includes:
[0094] S11. Receive the multi-modal document input signal provided by multiple data sources. The data sources include a text data source and a non-text data source. The text data source provides text data containing natural language content, and the non-text data source provides images, charts, or other types of non-text data:
[0095]
[0096] Among them, D m represents the received multi-modal document input data set, T di represents the i-th text data unit, N dj represents the j-th non-text data unit, N t and N n respectively represent the quantities of text data and non-text data, represents the parallel reception and combination operation of multi-modal data;
[0097] S12. For the text data T dand non-text data N d Perform a synchronous input operation to import the text data and non-text data into the processing system simultaneously;
[0098] S13. For the input text data T d and non-text data N d Perform preliminary data type identification, and verify the text data type based on the metadata identifier M t Verify the non-text data type through the image feature matrix I f ;
[0099] S14. According to the recognition result, classify and store the text data T d and non-text data N d in the text data cache unit C t and non-text data cache unit C n .
[0100] In this embodiment, step S2 includes:
[0101] S21. Use a vocabulary matching algorithm based on the context attention mechanism to perform word segmentation on the received text data T d and execute the word segmentation operation:
[0102]
[0103] where T d represents the input text data, T di represents the i-th word in the text data, L m (T d ) is the word segmentation result, α i and β i respectively represent the global dictionary matching weight and the context similarity weight, W(T di ) represents the entry matching score in the dictionary, C(T di ) represents the context word set of the word T di , and Sim(T dj , T j ) represents the similarity between the word T di and the context word T j ;
[0104] S22. Use a part-of-speech tagging model based on the conditional random field to perform part-of-speech tagging on the word-segmented text data, and the part-of-speech tagging process is as follows:
[0105]
[0106] where T di represents the i-th word, P di represents the corresponding part-of-speech tag, ψ(Tdi ,P di ) represents the compatibility function between the word T di and its part of speech P di , φ(P di-1 , P di ) represents the part-of-speech transition probability, P s (T d ) represents the set of part-of-speech tags of the text data after word segmentation;
[0107] S23. Perform syntactic analysis on the annotated text data, and use the dependency tree syntactic analysis model to construct the syntactic structure tree of the text. The specific process is as follows:
[0108]
[0109] Among them, T di and T dj represent two words in the text, δ(T di , T dj ) represents the distance between words, τ(P di , P dj ) represents the dependency relationship weight between parts of speech, Dep(T di , T dj ) represents the strength of the dependency relationship between words, S t (T d ) represents the syntactic structure tree of the text data;
[0110] S24. Use the entity recognition model based on bidirectional long short-term memory and conditional random field to perform entity recognition on the text data after syntactic analysis, and perform entity recognition:
[0111]
[0112] Among them, T di represents the i-th word in the text, E di represents the corresponding entity label, λ(T di , E di ) represents the compatibility score between the word and the entity label, μ(E di-1 , E di ) represents the transition score between entity labels, γ represents the fusion coefficient between multi-modal data, Att(T di , F j ) represents the attention weight between the text word T di and the non-text data feature F j , E r (T d ) represents the entity recognition result in the text data.
[0113] In this embodiment, step S3 includes:
[0114] S31. Extract visual features from the image data I in the multimodal document, and use a convolutional neural network to analyze the spatial features in the image data to extract the image feature vector V d : i
[0115]
[0116] where I d represents the input image data, I lk represents the l-th and k-th feature points in the image, w lk represents the weight parameter, f(I lk ) represents the non-linear activation function of the image feature points, b lk represents the bias value, and V i represents the visual feature vector of the image;
[0117] S32. Use a model based on the region proposal network to detect the target objects in the image data, generate the candidate region box R i , and perform object classification on each region box:
[0118]
[0119] where R pi represents the i-th candidate region, O i represents the true target region, P represents the number of candidate regions, φ(R pi ) is the confidence score of the region proposal, and IoU(R pi , O i ) represents the intersection over union between the candidate region and the true target region;
[0120] S33. Perform structure recognition on the chart data G in the multimodal document, and extract the key structure information S d in the chart, including the coordinate axes, data points, labels, and line relationships: g
[0121]
[0122] where G d represents the input chart data, G dj represents the j-th data element in the chart, L j represents the label associated with the data element, γ j represents the weight parameter, Rel(G dj , L j ) represents the strength of the association between the chart data element and its label, and S g represents the result of the structure recognition of the chart.
[0123] In this embodiment, step S4 includes:
[0124] S41. Fuse the feature vector P d of the preprocessed text data T s (T d ) with the non-text data N v after feature extraction to construct a multi-modal feature matrix M f :
[0125]
[0126] where P s (T d ) represents the part-of-speech tagging feature vector of the text data, V i represents the visual feature vector of the image data, S g represents the structural feature of the chart data, λ 1 and λ 2 respectively represent the fusion weights of the text and non-text data, and M f represents the fused multi-modal feature matrix;
[0127] S42. Calculate the relevance score A s (T d , N d ) between the text data and the non-text data based on the attention mechanism, where the word segmentation result L m (T d ) of the text data and the candidate region box R i of the non-text data are used to calculate the similarity score:
[0128]
[0129] where L m (T di ) represents the word segmentation result in the text data, R j represents the candidate region in the non-text data, α ij represents the attention weight between the text data and the non-text data, and Sim(L m (T di ), R j ) represents the similarity score between the text data word and the non-text data candidate region, and A s (T d , N d ) represents the relevance score between the text data and the non-text data;
[0130] S43. Based on the relevance score A s (T d , Nd ) Perform information fusion on multi-modal data to generate the final multi-modal feature matrix M r :
[0131]
[0132] Among them, β ij represents the fused weight coefficient, represents the fusion operation of text data words with image and chart features, and M r represents the finally generated multi-modal representation.
[0133] In this embodiment, step S5 includes:
[0134] S51. Input the fused multi-modal feature matrix M r into a pre-trained improved BERT model to perform in-depth semantic analysis. The improved BERT model jointly encodes the embedded representations of text and non-text data to generate a multi-modal semantic vector S v ;
[0135] S52. Perform key entity recognition in the multi-modal semantic vector M r through the attention mechanism. The improved BERT model calculates the attention weights of each word and non-text feature to identify the key entity E in the multi-modal data k :
[0136]
[0137] Among them, S vi represents the i-th word or feature in the multi-modal semantic vector, T di represents the corresponding text word, and α i represents the attention weight;
[0138] S53. Based on the in-depth semantic relationship extraction module of the improved BERT model, calculate the relationship R s (E k ) between entities in the multi-modal semantic vector, and identify the event Ev in the multi-modal semantic vector through the event detection module s .
[0139] In this embodiment, step S6 includes:
[0140] S61. Extract the text entity node E v and the non-text entity node E t from the multi-modal semantic vector S based on the results of in-depth semantic analysis, and construct the initial node set N n of the knowledge graph: g :
[0141] Text entity node Et represent the persons, place names, and events extracted from the text data;
[0142] Non-text entity node E n represent the objects and data points extracted from the non-text data of images and charts;
[0143] S62. Generate the edge L between nodes according to the semantic relationship between the text data and the non-text data g , and the edges include the edges representing the relationship between the text description and the object in the image and the edges representing the relationship between the chart data and the text description;
[0144] S63. Optimize the weights ω n and ω l of the nodes and edges based on the graph structure learning model to generate the optimized knowledge graph K g .
[0145] In this embodiment, step S63 includes:
[0146] S631. Initialize the node weights ω n and the edge weights ω l in the initial knowledge graph. The node weight ω n is initialized as the degree of the node, and the edge weight ω l is initialized as the association strength between nodes;
[0147] S632. Optimize the node weights and edge weights based on the gradient descent algorithm, and iteratively update the weight values of the nodes and edges and
[0148]
[0149]
[0150] where η 2 is the learning rate, and respectively represent the partial derivatives of the loss function L with respect to the node weights and edge weights;
[0151] S633. The loss function L is optimized by minimizing the error between the nodes and the edges. The purpose of the optimization is to improve the accuracy and relevance of the knowledge graph by minimizing L:
[0152]
[0153] where Sim(E i , E j ) represents the node E i and the node E jThe similarity between them, where L represents the total loss between nodes and edges;
[0154] S634. During the optimization iteration process, according to the optimal node weights and edge weights generate the finally optimized knowledge graph K g :
[0155]
[0156] where K g represents the finally optimized knowledge graph, including the optimized node set N g , edge set L g as well as node weights and edge weights
[0157] Example 1:
[0158] In the daily equipment maintenance of a certain intelligent manufacturing factory, the technical documents used by the factory include various modal data such as text descriptions, equipment photos, parameter tables, and operation diagrams. The maintenance manual plays an important role in equipment maintenance and operation training. However, the document data is huge and the format is complex. Traditional manual processing methods are inefficient and prone to missing key information, affecting the efficient operation and maintenance of equipment.
[0159] In August 2023, when the factory was conducting the annual maintenance of a certain large-scale precision CNC machine tool, the maintenance team needed to refer to the operation manual of the equipment for various parameter inspections. The equipment manual included 25 pages of text descriptions, accompanied by 15 equipment pictures, 10 operation diagrams, and 20 equipment parameter tables. The factory maintenance personnel processed this document in the traditional way, which took about 5 days. During this period, multiple association errors between charts and text information were found, resulting in a decline in maintenance efficiency. To solve this problem, the factory decided to apply the method of the present invention to improve efficiency and accuracy.
[0160] At the beginning of September 2023, the factory input the maintenance manual of this equipment into the system of the present invention. The system first decomposed the multi-modal data in the document and processed the text, image, and chart data respectively. In the text processing stage, the system performed word segmentation, part-of-speech tagging, and syntactic analysis on the equipment operation instructions in the manual, and identified the key equipment names, operation steps, and maintenance requirements. In the image processing stage, the system extracted visual features from 15 equipment images through a convolutional neural network, identified the key components of the machine tool, and through the correlation analysis with the text description, automatically paired the components in the image with the operation steps in the text. In the chart processing stage, the system performed structured recognition on 20 equipment parameter tables, extracted the key parameters (temperature, pressure, and operating speed) and associated them with the operation process in the text.
[0161] To verify the performance of the system, we compared the method of the present invention with the traditional manual processing method. The following are the specific comparison data:
[0162] Table 1 Comparison data between the traditional method and the method of the present invention
[0163]
[0164]
[0165] As can be seen from Table 1 above, it takes about 40 hours to process the equipment manual by the traditional method, including manually analyzing the association between text, images, and charts page by page, while the system of the present invention only needs 3 hours to automatically complete the processing, increasing the processing speed by 13 times. Through in-depth semantic analysis and multi-modal data fusion, the system of the present invention can accurately extract the key knowledge in the maintenance manual. The traditional method is prone to errors in the association annotation between images and text, with an error rate of 8%, while the automated annotation of this system has an error rate of only 1%. In the traditional method, factory technicians need to manually pair the equipment images with the operation steps, which is prone to mispairing. The system of the present invention can automatically complete the pairing through visual feature extraction and text association analysis, and the accuracy rate reaches 99%. The 20 parameter tables included in the equipment manual need to be manually input in the traditional method, while the system of the present invention can automatically extract parameter data and match it with the text description through structured recognition, with an identification accuracy rate of 98.5%. In the traditional method, knowledge extraction is limited to pure text information and cannot effectively process the data in images and charts. The system of the present invention can fuse text, image, and chart information to generate a complete knowledge graph, ensuring the integrity of multi-modal information, with a completeness of 99%. In the traditional method, manual intervention takes a lot of time, especially in the association annotation between images and text and the input of parameter tables, which requires 15 hours of manual time. However, the system of the present invention has a high degree of automation, and only 0.5 hours of manual review is required throughout the process, greatly reducing the manpower input.
[0166] In an actual operation in September 2023, maintenance personnel used the system of the present invention to process the maintenance manual of a precision CNC machine tool. The system first identified and extracted the key information of "equipment model: NC-5000" and "key components: spindle, tool holder" in the text. Subsequently, the system identified the equipment photos related to these components in the maintenance manual through the image processing module and automatically matched them with the operation steps. The system associated and annotated "spindle maintenance" with the corresponding equipment image, and further analyzed the "spindle speed range" and "operating temperature" in the parameter table. Finally, a global information display of the equipment was formed in the knowledge graph.
[0167] During this process, the system also automatically generated a maintenance report, which included the key components of the equipment, operation steps, relevant parameters, and the association results of images and chart information. Subsequently, the report was submitted to the maintenance team as a reference document for equipment maintenance operations.
[0168] Through this application, the maintenance efficiency of the factory has been significantly improved. The maintenance personnel feedback that the information provided by the system has a clear structure and clear operation steps, reducing the workload of manual access and association of images and parameter tables in the past.
[0169] The present invention proposes a multi-modal fusion method based on an improved BERT model by fusing text data and non-text data. The feature extraction model is used to extract features from text and non-text data respectively, and then the features are fused through an adaptive weight mechanism to generate a multi-modal feature matrix. This not only improves the recognition accuracy of the association between text and image and chart data, but also enhances the system's ability to understand complex documents, enabling the system to extract more accurate and comprehensive semantic information from multi-modal data and overcoming the limitations of single-modal processing in the prior art.
[0170] In the technical solution of the present invention, the node weights and edge weights of the knowledge graph are adaptively optimized based on the graph structure learning model. By introducing the gradient descent algorithm and optimizing the calculation of the loss function, the weight distribution of nodes and edges in the knowledge graph can be dynamically adjusted. The optimized knowledge graph can more accurately express the semantic relationship between text and non-text data. Compared with the existing knowledge graph construction technology, the present invention can not only process pure text data, but also effectively combine image and chart data, making the constructed knowledge graph have higher expression ability and accuracy, and ensuring the comprehensive representation of key knowledge in complex documents.
[0171] The present invention significantly improves the automation processing efficiency of complex documents by introducing an improved BERT model and a self-supervised learning mechanism, combined with multi-modal fusion technology. Different from the traditional technology that requires a large amount of manual intervention, the present invention can achieve automated knowledge extraction and structured processing through deep semantic analysis, entity recognition, and relationship extraction steps under unsupervised or weakly supervised conditions, reducing the complexity of manual operations, significantly reducing the time cost of processing multi-modal documents, improving the overall processing efficiency of the system, and solving the problem of low efficiency in processing large-scale and multi-modal documents in the prior art.
[0172] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A multimodal document structured processing and knowledge extraction method based on a large language model, characterized in that: The steps include: S1. receiving an input multimodal document, wherein the multimodal document comprises at least one text data and at least one non-text data; S2, preprocessing the text data in the multimodal document, wherein the preprocessing includes word segmentation, part-of-speech tagging, syntactic analysis and entity recognition; S3, performing feature extraction on the non-text data in the multimodal document; S4, performing multimodal data fusion on the preprocessed text data and the feature-extracted non-text data; S5. performing deep semantic analysis on the fused multimodal data through a pre-trained improved BERT model; S6. Based on the result of the deep semantic analysis, automatically construct a knowledge graph from the extracted information; S7. Output the data of the knowledge graph into a format that can be analyzed or applied; The step S5 comprises: S51, the fused multimodal feature matrix M r The input is fed into the pre-trained improved BERT model to perform deep semantic analysis. The improved BERT model jointly encodes the embedded representations of text and non-text data to generate a multimodal semantic vector S v ; S52, through the attention mechanism in the multimodal semantic vector M r The improved BERT model calculates the attention weights of each word and non-text features to identify key entities E in multimodal data. k : Among them, S vi represents the i-th word or feature in the multimodal semantic vector, T di represents the corresponding text words, α i represents the attention weight; S53, deep semantic relationship extraction module based on the improved BERT model, calculates the relationship R between entities in the multimodal semantic vector s (E k ), and identify the event Ev in the multimodal semantic vector through the event detection module s ; The step S6 comprises: S61. Based on the results of deep semantic analysis, the multimodal semantic vector S v Extract text entity node E t and non-text entity node E n , construct the initial node set N of the knowledge graph g : Text entity node E t Represents people, places, and events extracted from text data; Non-text entity node E n Represent objects and data points extracted from images and charts as non-text data; S62: Generate edges L between nodes based on the semantic relationship between text data and non-text data g , the edges include edges representing the relationship between the text description and the object in the image and edges representing the relationship between the chart data and the text description; S63. Weights ω of nodes and edges based on graph structure learning model n and ω l Optimize and generate the optimized knowledge graph K g .
2. According to claim 1, a multimodal document structured processing and knowledge extraction method based on a large language model is characterized in that: The step S1 comprises: S11, receiving a multimodal document input signal provided by multiple data sources, wherein the data sources include a text data source and a non-text data source, wherein the text data source provides text data containing natural language content, and the non-text data source provides images, charts, or other types of non-text data: Among them, D m represents the received multimodal document input dataset, T di Represents the i-th text data unit, N dj represents the jth non-text data unit, N t and N n Represents the number of text data and non-text data respectively. Represents the parallel reception and combination operations of multimodal data; S12, text data T d and non-text data N d Perform synchronous input operations so that text data and non-text data are simultaneously imported into the processing system; S13, input text data T d and non-text data N d Perform preliminary data type identification based on metadata identifier M t Verify the text data type through the image feature matrix I f Validate non-text data types; S14, according to the recognition result, the text data T d and non-text data N d Classification and storage in text data cache unit C t and non-text data cache unit C n middle.
3. According to claim 1, a multimodal document structured processing and knowledge extraction method based on a large language model is characterized in that: The step S2 comprises: S21, using the vocabulary matching algorithm based on the contextual attention mechanism to receive the text data T d Perform word segmentation processing and execute word segmentation operation: Among them, T d Represents the input text data, T di Represents the i-th word in the text data, L m (T d ) is the result of word segmentation processing, α i and β i Respectively represent the global dictionary matching weight and context similarity weight, W(T di ) represents the matching score of the entries in the dictionary, C(T di ) represents word T di The context word set, Sim(T di ,T j ) represents word T di With context word T j similarity; S22. Use the part-of-speech tagging model based on conditional random fields to perform part-of-speech tagging on the text data after word segmentation. The part-of-speech tagging process is as follows: Among them, T di represents the i-th word, P di represents the corresponding part-of-speech tag, ψ(T di ,P di ) represents word T di Its part of speech P di The compatibility function between di-1 ,P di ) represents the probability of part-of-speech transfer, P s (T d ) represents the part-of-speech tag set of the text data after word segmentation; S23, perform syntactic analysis on the annotated text data, and use the dependency tree syntactic analysis model to construct a grammatical structure tree of the text. The specific process is as follows: Among them, T di and T dj Represents two words in the text, δ(T di ,T dj ) represents the distance between words, τ(P di ,P dj ) represents the dependency weight between parts of speech, Dep(T di ,T dj ) represents the strength of the dependency relationship between words, S t (T d ) represents the syntax structure tree of text data; S24. Use an entity recognition model based on bidirectional long short-term memory and conditional random field to perform entity recognition on the text data after syntactic analysis, and perform entity recognition: Among them, T di represents the i-th word in the text, E di represents the corresponding entity label, λ(T di ,E di ) represents the compatibility score between the word and the entity tag, μ(E di-1 ,E di ) represents the transfer score between entity labels, γ represents the fusion coefficient between multimodal data, Att(T di ,F j ) represents the text word T di and non-text data features F j The attention weight, E r (T d ) represents the entity recognition result in text data.
4. The method for multimodal document structure processing and knowledge extraction based on a large language model according to claim 1, characterized in that: The step S3 comprises: S31, image data I in multimodal document d Perform visual feature extraction and use convolutional neural network to analyze the spatial features in the image data to extract the image feature vector V i : Among them, I d Represents the input image data, I lk represents the lth layer and the kth feature point in the image, w lk represents the weight parameter, f(I lk ) represents the nonlinear activation function of the image feature points, b lk Indicates the bias value, V i A visual feature vector representing an image; S32, use the model based on the region proposal network to detect the target object in the image data and generate the candidate region box R i , and perform target classification on each region box: Among them, R pi represents the i-th candidate region, O i represents the true target area, P represents the number of candidate areas, φ(R pi ) is the confidence score of the region proposal, IoU(R pi ,O i ) represents the intersection-over-union ratio between the candidate region and the true target region; S33, for the chart data G in the multimodal document d Perform structural recognition and extract key structural information S in the graph g , including the relationship between axes, data points, labels and lines: Among them, G d Represents the input chart data, G dj represents the jth data element in the graph, L j Represents the label associated with the data element, γ j Represents the weight parameter, Rel(G dj ,L j ) represents the strength of the association between the chart data element and its label, S g Represents the structure recognition result of the graph.
5. The multimodal document structured processing and knowledge extraction method based on a large language model according to claim 1, characterized in that: The step S4 comprises: S41, the pre-processed text data T d The eigenvector P s (T d ) and the non-text data N after feature extraction v Perform feature vector fusion and construct a multimodal feature matrix M f : Among them, P s (T d ) represents the part-of-speech tagging feature vector of text data, V i Represents the visual feature vector of image data, S g represents the structural features of the graph data, λ1 and λ2 represent the fusion weights of text and non-text data respectively, and M f Represents the fused multimodal feature matrix; S42. Calculate the correlation score A between text data and non-text data based on the attention mechanism s (T d ,N d ), where the word segmentation result of the text data is L m (T d ) and the candidate region box R of non-text data i Used to calculate the similarity score: Among them, L m (T di ) represents the word segmentation result in the text data, R j represents the candidate region in non-text data, α ij represents the attention weight between text data and non-text data, Sim(L m (T di ),R j ) represents the similarity score between the text data word and the non-text data candidate region, A s (T d ,N d ) represents the correlation score between text data and non-text data; S43. Based on the correlation score A s (T d ,N d ) Information fusion is performed on multimodal data to generate the final multimodal feature matrix M r : Among them, β ij represents the weight coefficient after fusion, represents the fusion operation of text data words and image and chart features, M r Represents the final generated multimodal representation.
6. The method for multimodal document structure processing and knowledge extraction based on a large language model according to claim 1, characterized in that: The step S63 comprises: S631, node weights ω in the initial knowledge graph based on graph structure learning model n and edge weight ω l Initialize, node weight ω n Initialized to the node degree and edge weight ω l Initialized to the association strength between nodes; S632: Optimize node weights and edge weights based on the gradient descent algorithm, and iteratively update the weight values of nodes and edges and Among them, η2 is the learning rate, and Respectively represent the partial derivatives of the loss function L with respect to the node weight and edge weight; S633, the loss function L is optimized by minimizing the error between nodes and edges. The purpose of the optimization is to improve the accuracy and relevance of the knowledge graph by minimizing L: Among them, Sim(E i ,E j ) represents node E i With node E j The similarity between them, L represents the total loss between nodes and edges; S634, in the optimization iteration process, according to the optimal node weight and edge weights Generate the final optimized knowledge graph K g : Among them, K g Represents the final optimized knowledge graph, including the optimized node set N g , edge set L g And the node weight and edge weights
Citation Information
Patent Citations
Document-level image-text comment sentiment classification method fused with common knowledge
CN116521818A
Knowledge graph construction method and system
CN117952209A