Document segmentation method based on multi-modal large model
Through the document segmentation method based on multimodal large model, combined with attention mechanism and dynamic window expansion technology, the problem of low segmentation accuracy in traditional methods when dealing with complex documents is solved, and a high-precision and flexible adaptation document segmentation effect is achieved.
Patent Information
- Application Number
- CN202510360928.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
When traditional document segmentation methods process complex documents containing multimodal information, it is difficult to accurately identify the boundaries of each logical part, which affects subsequent information processing and utilization efficiency.
The document segmentation method based on multimodal large model is adopted, and high-precision segmentation of complex documents is achieved through steps such as document preprocessing, multimodal feature coding, modal fusion, document segmentation and post-processing optimization, combined with attention mechanism and dynamic window expansion technology.
It significantly improves the accuracy of segmentation of complex documents, flexibly adapts to different document types, ensures that the segmentation boundary accurately fits the content logic, and continuously improves the segmentation performance through online optimization and user feedback mechanisms.
Smart Images

Figure CN120218018A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document segmentation, and particularly to a document segmentation method based on a multimodal large model. Background Art
[0002] With the explosive growth of digital information, the quantity and variety of documents are increasing day by day. In many fields such as document management, information retrieval, and text analysis, effective segmentation of documents has become a key task. Traditional document segmentation methods mainly rely on single text features, such as based on characters, vocabulary, grammatical structures, etc. However, when facing complex document structures, diverse document contents, and documents containing multiple modal information (such as images, tables, text formats, etc.), these methods often have problems such as low segmentation accuracy and poor adaptability. For example, for a technical report that contains both text descriptions, charts, and different font and color formats, traditional methods are difficult to accurately identify the boundaries of each logical part, thus affecting the efficiency of subsequent information processing and utilization.
[0003] In recent years, large model technology has made remarkable progress in the field of natural language processing, etc. However, its application in multimodal document segmentation research is relatively less, and some existing preliminary attempts still have deficiencies such as insufficient modal fusion and limited improvement in model complexity but with limited effect. Therefore, there is an urgent need for an innovative and efficient document segmentation method based on a multimodal large model to solve the above problems. Summary of the Invention
[0004] The present invention proposes a document segmentation method based on a multimodal large model, which solves the problems in the prior art that traditional document segmentation methods are difficult to accurately identify the boundaries of each logical part for documents containing multimodal information, thus affecting the efficiency of subsequent information processing and utilization, etc.
[0005] The technical solution of the present invention is realized as follows:
[0006] The present invention provides a document segmentation method based on a multimodal large model, including the following steps:
[0007] S1, Document preprocessing, parsing the document to be segmented, and extracting the original features of each modality in the document. The original features of each modality in the document include at least one or more of text content, image information, table structure, and text format;
[0008] S2, Multimodal feature encoding, using an encoder to encode the original features of each modality to generate feature representations that can be recognized and processed by the model;
[0009] S3. Modal fusion: Extract the global features of the document and the comprehensive features of each modality, and input them into the weight generation network to obtain the weights of each modality. Then, fuse each modality according to the weights to obtain the fused document feature representation;
[0010] S4. Document segmentation: Input the fused document feature representation into the hierarchical segmentation model based on the attention mechanism, and output the segmentation boundaries and class labels of each part in the document;
[0011] S5. Post-processing and optimization: Evaluate the accuracy of the segmentation result, and adjust the segmentation result and model parameters according to the evaluation result.
[0012] Specifically, step S2 includes the following steps:
[0013] Use the text encoder to encode the extracted text content to obtain a text feature vector, which is used to reflect the semantic, syntactic, and lexical information of the text;
[0014] Use the image encoder to encode the extracted image information to obtain an image feature map, which is used to reflect the visual features of the image;
[0015] Use the structure encoder to encode the extracted table structure to obtain a table structure feature, which is used to reflect the logical architecture of the table;
[0016] Use the fully connected neural network to encode the extracted text format to obtain a format feature vector, which is used to reflect the characteristics of the text in visual presentation and potential semantic emphasis information.
[0017] Specifically, in step S3, the method for extracting the global features of the document is as follows: Count the total number of words, the total number of images, and the number of tables in the document; Count the distribution of high-frequency words in the text; Analyze the average color complexity of the images and the average complexity index of the tables; Combine the above statistics and indexes to form a vector with a fixed dimension as the global feature of the document;
[0018] The method for extracting the comprehensive features of each modality of the document is as follows: Count the proportion of words with different parts of speech and the average sentence length in the text; Count the average brightness and contrast of the images; Count the proportion of tables with merged cells and the distribution of data types; Combine the above statistical information with the encoded feature representations of each modality to form the comprehensive features of each modality.
[0019] Specifically, in step S3, the construction method of the weight generation network includes:
[0020] Construct a weight generation network with a fully connected neural network as the main body. The number of nodes in the input layer is determined according to the number of modal features to be fused and the global feature dimension of the document. The number of nodes in the output layer is the same as the number of modal features. Set several hidden layers in the middle layer, and the number of neurons in the hidden layers is set in a decreasing manner. Select a non-linear activation function as the activation function for the hidden layers, and according to the value range requirements of the weight coefficients, select the Sigmoid function or the Softmax function as the activation function for the output layer.
[0021] Specifically, in step S3, the method for fusing each modality according to the weights includes:
[0022] Input the global features of the document and the comprehensive features of each modality into the constructed weight generation network. Through the forward propagation calculation of the network, each node in the output layer generates a weight coefficient correspondingly, which respectively corresponds to the weights of multiple modal features.
[0023] Then multiply the encoded features of each modality by the corresponding weight coefficients respectively, and then sum the weighted modal features in the corresponding dimensions to obtain the fused document feature representation.
[0024] Specifically, step S4 includes the following steps:
[0025] Global attention scanning. In the global attention module, apply the multi-head attention mechanism to the feature vector of the entire document. Each attention head respectively focuses on the feature information of different modalities of the document. By weighted summing the outputs of each attention head, a global feature representation is obtained, and then the global feature representation is input into the fully connected neural network classifier to output the segmentation boundaries and class labels of the approximate regions of each part of the document.
[0026] Local attention refinement. For each region, activate the local attention mechanism respectively. In the local attention module, use the multi-head attention mechanism again, and limit the scope of attention within the current approximate region. The output of the local attention passes through a multi-layer perceptron for feature transformation and integration, and finally outputs the more fine-grained segmentation boundaries and class labels within each approximate region.
[0027] Preferably, in step S4, during the document segmentation process, adopt a dynamic window expansion method to adaptively segment the document, including the following steps:
[0028] Window initialization. When starting to segment the unit boundary, select an initial size for the window according to the document type and structure pattern.
[0029] Feature calculation and judgment: For the content covered by the window, calculate the feature strengths in different modalities respectively; use the semantic similarity algorithm to judge the semantic correlation degree between the text in the window and the adjacent text, and at the same time check whether the format features of the text in the window are consistent with those of the adjacent text. If the formats are consistent and the semantic correlation degree is higher than the set threshold, it is judged that the context is coherent.
[0030] Window dynamic adjustment: When the feature strength in the window shows an upward trend and the context is coherent, start the window expansion process, increase the window size according to the preset step length, and then recalculate the feature strength and coherence in the expanded window. Repeat expanding the window size and calculating the feature strength and coherence in the expanded window until the feature strength in the window no longer rises or the context is incoherent; once a feature mutation is detected in the window, immediately terminate the current window expansion process, quickly shrink the window, and roll back the window to the last position where the feature strength in the window no longer rises and the context is coherent, lock this position as the boundary of a segmentation unit, and start a new window to continue scanning.
[0031] Segmentation unit determination: Each time a window contraction operation is completed, a segmentation unit is determined, and its boundary position and category information are recorded until the entire document is traversed.
[0032] Specifically, in step S5, the post - processing and optimization methods include:
[0033] Compare the document segmentation fragments output by the model with the standard segmentation results manually annotated, and use multiple evaluation indicators to preliminarily evaluate the segmentation quality;
[0034] Input each segmentation fragment into a pre - trained language model to score the semantic coherence of the text in each fragment;
[0035] For the fragments with scores lower than the set threshold, further analyze the internal text logic and feature distribution, check the grammar structure, keyword distribution of the text, and the feature differences from adjacent fragments. If the text is semantically incomplete, merge the current segmentation fragment with the adjacent segmentation fragment, and determine the best adjustment plan by comparing the semantic coherence scores before and after merging.
[0036] Preferably, in step S5, the post - processing and optimization methods include:
[0037] Build a user feedback system to evaluate and annotate the segmentation results through the graphical user interface of the user feedback system;
[0038] The system collects the user's feedback data in real time and stores the feedback data in the database; when the amount of feedback data reaches the threshold or after a certain period of time, an online optimization process is started, the wrongly segmented cases and low-scoring cases marked by the user are extracted from the database, and the original documents corresponding to the above cases and their correct segmentation annotation information are sorted into an optimization dataset;
[0039] The optimization dataset and the original training dataset are mixed according to a preset ratio, and the model is continuously trained on the mixed dataset. During the training process, according to the key areas of concern and error types feedback by the user, the loss function of the model is weighted and adjusted.
[0040] Preferably, in step S5, the post-processing and optimization method includes:
[0041] When fine-tuning the parameters of the multi-modal large model, before each iteration, the learning rate is adjusted according to the segmentation performance of the model on the validation set;
[0042] If the performance of the model shows a downward trend in several consecutive iterations, the learning rate is decreased; if the performance of the model continues to improve in several consecutive iterations, the learning rate is increased;
[0043] The model complexity index is calculated according to the parameter quantity and connection method of each layer in the model, and the weight of the regularization term is dynamically adjusted based on the complexity index; for the layers with higher complexity, the weight of the regularization term is increased; for the layers with lower complexity, the weight of the regularization term is decreased; when calculating the total loss function, the regularization term and the loss term based on the segmentation task are added together to jointly guide the update of the model parameters.
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0045] (1) By integrating multi-modal features such as text, images, tables, and formats, and combining an adaptive weighted fusion mechanism to dynamically allocate the weights of each modality, the present invention effectively captures the correlation and importance differences between different modalities, and significantly improves the segmentation accuracy of complex documents;
[0046] (2) The present invention adopts a dynamic window expansion technology, adaptively adjusts the window boundary according to the local feature intensity and context coherence of the document content, flexibly adapts to different document types, and ensures that the segmentation boundary accurately fits the content logic;
[0047] (3) The present invention adopts an innovative weight generation network to combine the global features of the document and the modal statistical information, efficiently generates weight coefficients, avoids feature redundancy or information loss caused by traditional fixed weight allocation, and improves the computational efficiency and effect of the fusion process;
[0048] (4) The present invention introduces a hierarchical prediction architecture based on the attention mechanism and a semantic coherence secondary verification mechanism for post-processing. Through language model scoring and logical analysis, it ensures the integrity and consistency of the segmentation units in terms of semantics, grammar, and format;
[0049] (5) The present invention collects error cases through an online user feedback system, combines incremental training with dynamic loss function adjustment, realizes real-time optimization and personalized adaptation of the model, and gradually improves the segmentation performance and user satisfaction in actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0051] Figure 1 It is a schematic flow chart of a document segmentation method based on a multi-modal large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0053] Referring to Figure 1 , the present invention provides a document segmentation method based on a multi-modal large model, including the following steps:
[0054] S1. Document preprocessing: Parse the document to be segmented (this document can be in various formats, such as PDF, Word, etc.), and use a pre-trained neural network model to automatically extract the original features of each modality in the document. The original features of each modality in the document include text content, image information, table structure, and text format (such as font, font size, color, bold, underline, etc.);
[0055] S2. Multi-modal feature encoding: Use an encoder to encode the original features of each modality to generate feature representations that can be recognized and processed by the model;
[0056] S3. Modality fusion: Extract the global features of the document and the comprehensive features of each modality, and input them into a weight generation network to obtain the weights of each modality. According to the weights, fuse each modality to obtain a fused document feature representation;
[0057] S4, Document segmentation, input the fused document feature representation into a hierarchical segmentation model based on the attention mechanism, and output the segmentation boundaries and class labels of each part in the document;
[0058] S5, Post-processing and optimization, evaluate the accuracy of the segmentation result, and adjust the segmentation result and model parameters according to the evaluation result.
[0059] Multimodal feature encoding refers to the process of converting the original features extracted from different modalities (such as text, images, tables, text formats, etc.) into a unified and more expressive feature representation form suitable for computer processing and model understanding through specific algorithms and models.
[0060] Specifically, step S2 includes the following steps:
[0061] For text content, use the text encoder in the pre-trained multimodal large model to encode the extracted text content to obtain a sequence of text feature vectors. This encoder can capture the semantic, syntactic, and lexical information of the text. For example, use an encoder based on the Transformer architecture to comprehensively model the word relationships in the text through the multi-head attention mechanism;
[0062] For image information, use a dedicated image encoder (such as a variant of the convolutional neural network) to encode the extracted image information, and convert the image into a low-dimensional feature map. These feature maps can reflect the visual features of the image, such as shape, color distribution, texture, etc.;
[0063] For the table structure, use a structure encoder to convert the row-column relationships, header information, etc. of the table into a structured feature representation for the model to understand the logical structure of the table;
[0064] For text formats, use a small fully-connected neural network to encode the extracted text formats to generate corresponding format feature vectors. These vectors can reflect the characteristics of the text in visual presentation and potential semantic emphasis information.
[0065] The fusion module adopts an innovative adaptive weighted fusion mechanism to dynamically assign weights according to the importance of different modality features in a specific document. The following is the specific implementation method of the innovative adaptive weighted fusion mechanism in the fusion module to dynamically assign weights according to the importance of different modality features in a specific document:
[0066] Specifically, in step S3, the method for extracting the global features of the document is as follows: It is achieved by performing a coarse-grained analysis of the entire document. For example, basic statistics such as the total number of words, the total number of images, and the number of tables in the document are counted; the distribution of high-frequency words in the text is calculated to reflect the theme tendency of the document; indicators such as the average color complexity of images and the average complexity of tables (such as the average number of rows and columns) are analyzed. These statistics and indicators are combined to form a vector with a fixed dimension as the global feature representation of the document, which is used to enable the weight generation network to understand the general situation of the entire document.
[0067] The method for extracting the comprehensive features of each modality of the document is as follows: For the features of each modality, in addition to the feature vector encoded by itself, some statistical information needs to be collected. For example, for the text modality, the proportion of words with different parts of speech and the average sentence length of the text are counted; for the image modality, statistical features such as the average brightness and contrast of the image are calculated; for the table modality, the proportion of tables with merged cells and the distribution of data types are counted. This statistical information, together with the encoded modality feature vectors, is provided as input to the weight generation network to help the network more comprehensively grasp the characteristics of each modality feature in the current document.
[0068] Specifically, in step S3, the construction method of the weight generation network includes:
[0069] Network structure selection: A weight generation network with a fully connected neural network (FNN) as the main body is constructed. The number of input layer nodes of this network is determined according to the number of modality features to be fused and the dimension of the global features of the document. For example, if there are 4 modality features including text, image, table, and text format, and each modality feature has a different dimension after encoding (assumed to be d1, d2, d3, d4 respectively), plus the dimension of the global features of the document is d0, then the number of input layer nodes is d0 + d1 + d2 + d3 + d4.
[0070] Several hidden layers (such as 2 - 3 layers) can be set in the middle layer. The number of neurons in the hidden layers is set in a decreasing manner. For example, the number of neurons in the first hidden layer can be half of the number of input layer nodes, and the subsequent hidden layers decrease sequentially. The specific number can be adjusted according to experiments and actual situations. The number of output layer nodes is the same as the number of modality features because weight coefficients corresponding to each modality feature need to be generated. For example, for the above 4 modality features, there are 4 nodes in the output layer.
[0071] Activation function selection: In the hidden layer, common non-linear activation functions can be selected, such as the ReLU (Rectified Linear Unit) function, whose expression is f(x) = max(0, x). It can enhance the non-linear expression ability of the network, avoid the vanishing gradient problem, and enable the weight generation network to better fit the relationship between complex modal features and weights. For the output layer, according to the value range requirements of the weight coefficients, the Sigmoid function or the Softmax function can be selected. If it is desired that the weight coefficients are between 0 and 1 and the sum of the weight coefficients is not necessarily 1, the Sigmoid function (the expression is ) can be chosen; if it is required that the sum of the weight coefficients is 1, presenting a probability distribution form, then the Softmax function is selected (for the output vector z = [z1, z2,..., z n , the result of the i-th element after the Softmax transformation is ).
[0072] Specifically, in step S3, the method of fusing each modality according to the weights includes:
[0073] Weight coefficient calculation: Input the global features of the document and the features of each modality (including their encoded vectors and statistical information) into the constructed weight generation network. Through the forward propagation calculation of the network, each node in the output layer will correspondingly generate a weight coefficient. For example, the output weight coefficient vector may be [ω1, ω2, ω3, ω4], corresponding to the weights of the 4 modality features of text, image, table, and text format respectively.
[0074] Dynamic allocation process: After obtaining the weight coefficients of each modality feature, multiply the encoded feature vectors of each modality (such as the text feature vector sequence, image feature map, table structure feature, text format feature vector) by their corresponding weight coefficients respectively to achieve dynamic weighting according to importance. Taking the text modality as an example, if the text feature vector sequence is represented as V text , and its corresponding weight coefficient is ω1, then the weighted text feature is represented as ω1×V text . Weight the other modality features in the same way, and finally sum the weighted modality features in the corresponding dimensions to obtain the fused document feature representation V fused , and its calculation formula can be expressed as:
[0075] V fused = ω1×V text + ω2×V image + ω3×V table + + ω4×V format
[0076] Among them, ω2×Vimage 、 ω3×V table 、 ω4×V format represent the weighted image feature, table structure feature, and text format feature respectively;
[0077] In this way, through the adaptive weighted fusion mechanism, the importance differences of different modality features in a specific document are fully considered, and the modality features are dynamically fused into a unified document feature representation, providing more comprehensive and targeted feature information for subsequent document segmentation.
[0078] In practical applications, the weight generation network needs to be trained on a training dataset. The training objective can be to minimize the loss function for the document segmentation task based on the fused features (such as cross-entropy loss, which is used to measure the difference between the segmentation prediction result and the true label). By continuously adjusting the parameters of the weight generation network, it can accurately dynamically allocate the weights of modality features according to the characteristics of different documents, thereby improving the performance of the entire document segmentation method.
[0079] Specifically, step S4 includes the following steps:
[0080] Global attention scanning: Input the fused document feature representation into a hierarchical prediction architecture based on the attention mechanism. First, in the global attention module, apply the multi-head attention mechanism to the feature vector sequence of the entire document. Each attention head focuses on the feature information of different aspects of the document. For example, one attention head may focus on the keyword distribution in the text to determine the theme, another attention head may focus on the position and importance of the image features in the overall layout of the document, and some other attention heads will pay attention to the association between the table structure and the text, etc.
[0081] By performing weighted summation on the outputs of each attention head, a global feature representation vector is obtained, which synthesizes the overall theme and macro-structure information of the document. Then, input this global feature representation into a simple fully-connected neural network classifier, which is trained to identify the boundaries and class labels of the main part divisions that may exist in the document, such as the boundaries and category labels of roughly areas like core content, auxiliary explanations, appendices, etc.
[0082] Local attention refinement: After determining the approximate areas, activate the local attention mechanism for each area. In the local attention module, use the multi-head attention mechanism again, but this time the scope of attention is limited to the current approximate area. The attention heads focus on the detailed features within the area, such as the sentence structure in the text, the semantic relationships of words, the detailed elements in the image, and the content and format of the cells in the table.
[0083] The output of local attention undergoes feature transformation and integration through a multi-layer perceptron (MLP), and finally outputs finer-grained segmentation boundaries and class labels within each approximate region, accurately identifying specific segmentation units such as paragraphs, sentences, and chart elements.
[0084] Preferably, in step S4, during the document segmentation process, a dynamic window expansion method is adopted to adaptively segment the document, including the following steps:
[0085] Window initialization, setting the initial window size: When starting to predict the boundaries of segmentation units, based on the document type and common structural patterns, a smaller initial size is selected for the window. For example, for text-based documents, the initial window can be set to contain 3-5 consecutive words; if the document contains more charts, the initial window can be set to cover a small chart and 1-2 adjacent text short sentences, ensuring that the window can focus on local key information to start the analysis.
[0086] Feature calculation and judgment, calculating the feature intensity within the window: For the content covered by the window, calculate the feature intensity in different modalities respectively. In the text modality, with the help of a pre-trained word vector model and semantic analysis tools, statistics such as the frequency of keyword occurrences and the degree of semantic association tightness are used as the text feature intensity; in the image modality, features reflecting visual attractiveness such as color contrast and shape complexity are extracted through an image feature extraction algorithm; in the table modality, the table feature intensity is measured based on the key information density of the table header and the data change gradient of the cells.
[0087] Evaluating context coherence: From the perspective of text coherence, using semantic similarity algorithms in natural language processing (such as calculating based on word vector cosine similarity or BLEU value), judge the semantic association degree between the text within the window and the adjacent text. The higher the value, the better the coherence. At the same time, combined with format features, check whether the font, font size, color, etc. are consistent. If the format is coherent and the semantic coherence score is higher than the set threshold (such as 0.7), it is initially determined that the context is coherent.
[0088] Window dynamic adjustment, window expansion decision: When the feature intensity within the window shows an upward trend and the context coherence is good, start the window expansion process. Increase the window size by a predetermined step length. The text window can be increased by 1-2 words each time, and the window containing charts will appropriately expand the boundary according to the chart complexity, and recalculate the feature intensity and coherence within the expanded window. Repeat this process until the feature intensity tends to be stable or starts to decline, or the coherence score is lower than the threshold.
[0089] Window contraction decision: Once a feature mutation occurs within the window, such as a sudden change in text topic keywords (judged by a keyword frequency change exceeding 50%), format switching (font change, appearance of a new bold style), introduction of new image elements, etc., immediately terminate the current window expansion or maintain the existing size, and quickly contract the window. Retreat the window boundary to the last position where the features are stable and coherent, lock this position as the boundary of a segmentation unit, and then open a new window to continue scanning.
[0090] Segmentation unit determination and marking: Each time a window contraction operation is completed, a segmentation unit is determined, and its boundary position and the content categories it contains (text paragraphs, charts, tables, etc.) are recorded, providing basic information for subsequent document reconstruction and application until the entire document is traversed and all segmentation units are defined.
[0091] Through the above rigorous and refined dynamic window expansion technical steps, it is possible to closely fit the diverse and changing content characteristics of the document, adaptively and accurately segment the document, and inject strong flexibility and high-precision guarantee into the document segmentation method based on the multimodal large model.
[0092] Specifically, in step S5, the post-processing and optimization method is implemented using a secondary verification mechanism based on semantic coherence evaluation, including the following steps:
[0093] In the post-processing stage, first, compare the segmented document fragments with the standard segmentation results manually annotated, and calculate error metrics using common evaluation indicators (such as accuracy, recall rate, F1 value, etc.) to preliminarily evaluate the segmentation quality.
[0094] Next, input each segmented fragment into a pre-trained language model (such as the GPT series models or other language models trained on large-scale corpora), and the language model scores the semantic coherence of the text in each fragment. The scoring can be comprehensively evaluated based on multiple aspects such as the language model's understanding of the logical relationships between sentences in the text, the coherence of vocabulary, and the correctness of grammar.
[0095] For fragments with scores lower than the set threshold (determine a suitable threshold through experiments on the validation set, such as 0.6), further analyze the internal text logic and feature distribution. Use text parsing techniques and feature extraction methods to check the grammar structure of the text, the distribution of keywords, and the feature differences from adjacent fragments. If it is found that the text seems semantically abrupt or incomplete, such as the lack of a connecting word at the beginning of a sentence or the logical incoherence between paragraphs, try to merge it with adjacent fragments or re-divide the boundary. The best adjustment plan can be determined by comparing the semantic coherence scores and other feature metrics (such as format consistency, keyword relevance, etc.) before and after merging to ensure the semantic integrity and coherence of each segmentation unit.
[0096] Preferably, in step S5, the post - processing and optimization method is implemented by an online optimization mechanism based on user feedback, including the following steps:
[0097] Build a user feedback system. After users use the document segmentation results, they can conveniently evaluate and annotate the segmentation results through a graphical interface. The user interface provides simple operation buttons, such as options like "correct segmentation", "wrong segmentation", "partially wrong", etc. for users to choose. For the wrongly segmented parts, users can mark them by dragging the mouse or other interactive methods and can add short notes to explain the problems.
[0098] The system collects the user's feedback data in real - time and stores it in a database. Every certain period (such as 1 hour) or when the amount of feedback data reaches a certain scale (such as 100 pieces), start the online optimization process. Extract the wrongly segmented cases and low - scoring cases marked by users from the database, and organize the original documents corresponding to these cases and their correct segmentation annotation information into an optimization dataset.
[0099] In the optimization dataset, perform special processing on the wrongly segmented cases. For example, mark them as high - weight negative samples and give more attention during the training process. Use this optimization dataset to perform incremental training on the model, mix it with the original training dataset in a certain proportion (such as 1:9), and then continue to train the model on the mixed dataset. During the training process, adjust the weighted loss function of the model according to the key - attention areas and error types feedback by users. For example, if users frequently feedback problems in table segmentation, increase the weight of the table - segmentation part in the loss function, so that the model pays more attention to improving the accuracy of table segmentation during training.
[0100] By continuously collecting user feedback and performing online optimization, the model can gradually adapt to the document - segmentation requirements of different users in various actual application scenarios, continuously improve the segmentation quality and user satisfaction, and achieve a positive interaction and continuous evolution between the model and users.
[0101] Preferably, in step S5, the post - processing and optimization method is implemented by an adaptive learning - rate adjustment strategy and a regularization method based on model complexity, including the following steps:
[0102] When fine - tuning the parameters of the multi - modal large model using the back - propagation algorithm, before each iteration, adjust the learning rate according to the segmentation performance of the model on the validation set. First, record the changes in accuracy and recall rates of the model in the recent several iterations (such as 5 - 10 times).
[0103] If the model's performance improvement is not obvious in several consecutive iterations (for example, the changes in accuracy and recall are less than 0.01) or shows a downward trend, the learning rate is multiplied by a decay factor less than 1 (such as 0.9) to appropriately reduce the learning rate and avoid excessive oscillation near the local optimal solution; conversely, if the model's performance continues to improve rapidly (for example, the changes in accuracy and recall are greater than 0.05), the learning rate is multiplied by a growth factor greater than 1 (such as 1.1) to moderately increase the learning rate, accelerate the pace of parameter update, and improve the optimization efficiency.
[0104] At the same time, calculate the model complexity index according to the number of parameters and connection methods of each layer in the model. For example, weighted sums of the number of parameters (assigning different weights to the parameters of different layers according to their importance) and connection density and other indicators can be used to quantify the model complexity. Based on this complexity index, dynamically adjust the weight of the regularization term. For layers with higher complexity, increase the weight of the regularization term to prevent overfitting during the model optimization process; for layers with lower complexity, appropriately reduce the weight of the regularization term to maintain the fitting ability of the model. When calculating the total loss function, add the regularization term to the loss term based on the segmentation task (such as cross-entropy loss) to jointly guide the update of model parameters, thereby better balancing the fitting ability and generalization performance of the model, especially when dealing with multimodal data, and improving the optimization effect of the model.
[0105] The following uses a specific case to illustrate the working process of the document segmentation method based on the multimodal large model of the present invention:
[0106] Suppose there is a PDF document about medical research, which contains text paragraphs, medical image pictures, experimental data tables, and titles and annotations in different formats. Now it is necessary to segment this document.
[0107] First, in the document preprocessing stage, parse the PDF document to extract the pure text content, image files, table data, and corresponding text format information. After performing conventional processing such as word segmentation and denoising on the text, it is converted into a sequence of word vectors acceptable to the text encoder; the image is input into the image encoder after operations such as scaling and normalization; the table data is organized into a structured matrix form for the structure encoder to process; and the text format information is quantified into format feature vectors.
[0108] Next, in the multimodal feature encoding step, the text encoder encodes the text sequence to generate a sequence of high-dimensional text feature vectors to capture the semantic information of the text; the image encoder extracts visual features from the medical image pictures to obtain image feature maps; the structure encoder converts the logical structure of the table into a feature representation; and the text format features are encoded through a fully connected neural network.
[0109] Then, in the modal fusion stage, the weight generation network calculates the weight coefficients of text, image, table, and format features based on the overall features of the document and the statistical situation of each modal feature. For example, in this medical document, since the imaging pictures are crucial for presenting the research results, the image features may be assigned a higher weight. The modal features are weighted and summed according to the weights to obtain the fused document feature representation.
[0110] After that, the fused features are fed into the segmentation predictor. Through the multi-layer non-linear transformation of the MLP, the segmentation label sequence of the document is output, determining the categories and boundaries of each part in the document. For example, the title part is marked as "title", the body paragraphs are marked as "body", and the medical images and their descriptions are marked as "image and description", etc.
[0111] Finally, the document is actually segmented according to the segmentation prediction results to obtain multiple independent document fragments. These segmentation results are compared with the standard results of manual annotation, the error metrics are calculated, and the parameters of the multi-modal large model are fine-tuned through the backpropagation algorithm. After multiple iterations of training, the model can achieve a high segmentation accuracy and stability when processing similar medical research documents, and can be extended and applied to various types of document segmentation tasks in other fields.
[0112] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A document segmentation method based on a multimodal large model, characterized in that: The following steps are involved: S1, document preprocessing, parsing the document to be segmented, extracting the original features of each modality in the document, the original features of each modality in the document at least including one or more of text content, image information, table structure, and text format; S2, multimodal feature encoding, uses the encoder to encode the original features of each modality to generate feature representations that can be recognized and processed by the model; S3, modality fusion, extracts the global features of the document and the comprehensive features of each modality, and inputs them into the weight generation network to obtain the weight of each modality. The modalities are fused according to the weight to obtain the fused document feature representation; S4, document segmentation, inputs the fused document feature representation into the hierarchical segmentation model based on the attention mechanism, and outputs the segmentation boundaries and category labels of each part of the document; S5, post-processing and optimization, evaluates the accuracy of the segmentation results and adjusts the segmentation results and model parameters based on the evaluation results.
2. A document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: Step S2 includes the following steps: The extracted text content is encoded using a text encoder to obtain a text feature vector, which is used to reflect the semantic, grammatical and lexical information of the text; The extracted image information is encoded using an image encoder to obtain an image feature map, which is used to reflect the visual features of the image; The extracted table structure is encoded using a structure encoder to obtain table structure features, which are used to reflect the logical architecture of the table; A fully connected neural network is used to encode the extracted text format to obtain a format feature vector, which is used to reflect the visual presentation characteristics of the text and the potential semantic emphasis information.
3. A document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: In step S3, the method for extracting the global features of the document is as follows: counting the total number of words, the total number of images and the number of tables in the document; counting the distribution of high-frequency words in the text; analyzing the average color complexity of the image and the average complexity index of the table; combining the above statistics and indicators to form a vector of fixed dimension as the global feature of the document; The method for extracting comprehensive features of each modality of a document is as follows: counting the proportion of words of different parts of speech and the average sentence length in the text; counting the average brightness and contrast of images; counting the proportion of tables containing merged cells and the distribution of data types; and combining the above statistical information with the encoded feature representations of each modality to form comprehensive features of each modality.
4. The document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: In step S3, the method for constructing the weight generation network includes: Construct a weight generation network with a fully connected neural network as the main body. The number of nodes in the input layer is determined by the number of modal features to be fused and the global feature dimension of the document, and the number of nodes in the output layer is the same as the number of modal features. Set up several hidden layers in the middle layer, and the number of neurons in the hidden layer is set in a decreasing manner. Select a nonlinear activation function as the activation function of the hidden layer, and select the Sigmoid function or the Softmax function as the activation function of the output layer according to the value range requirements of the weight coefficient.
5. The document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: In step S3, the method of fusing each modality according to the weight includes: The global features of the document and the comprehensive features of each modality are input into the constructed weight generation network. After the forward propagation calculation of the network, each node in the output layer generates a weight coefficient corresponding to the weights of multiple modal features. The encoded features of each modality are then multiplied by the corresponding weight coefficients, and then the weighted features of each modality are summed up in the corresponding dimensions to obtain the fused document feature representation.
6. A document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: Step S4 includes the following steps: Global attention scanning: In the global attention module, a multi-head attention mechanism is applied to the feature vector of the entire document. Each attention head focuses on the feature information of different modalities of the document. A global feature representation is obtained by weighted summing the outputs of each attention head. The global feature representation is then input into a fully connected neural network classifier to output the segmentation boundaries and category labels of the approximate areas of each part of the document. Local attention refinement: For each area, the local attention mechanism is activated separately. In the local attention module, the multi-head attention mechanism is used again to limit the scope of attention to the current approximate area. The output of the local attention is transformed and integrated through a multi-layer perceptron, and finally a finer-grained segmentation boundary and category label are output in each approximate area.
7. The document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: In step S4, during the document segmentation process, a dynamic window expansion method is used to adaptively segment the document, including the following steps: Window initialization, when starting to split cell boundaries, an initial size is selected for the window based on the document type and structure mode; Feature calculation and judgment: for the content covered by the window, the feature strength under different modes is calculated respectively; the semantic similarity algorithm is used to judge the semantic relevance between the text in the window and the adjacent text, and at the same time, check whether the format features of the text in the window and the adjacent text are consistent. If the format is consistent and the semantic relevance is higher than the set threshold, the context is judged to be coherent; Dynamic adjustment of the window: when the feature strength in the window shows an upward trend and the context is coherent, the window expansion process is started, the window size is increased according to the preset step size, and the feature strength and coherence in the expanded window are recalculated. The window size is repeatedly expanded and the feature strength and coherence in the expanded window are calculated until the feature strength in the window no longer increases or the context is incoherent. Once a sudden change in the feature is detected in the window, the current window expansion process is terminated immediately, and the window is quickly shrunk to the last position in the window where the feature strength no longer increases and the context is coherent. This position is locked as a segmentation unit boundary, and scanning is continued in a new window. The segmentation unit is determined. Each time a window shrinking operation is completed, a segmentation unit is determined and its boundary position and category information are recorded until the entire document is traversed.
8. The document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: In step S5, the post-processing and optimization method includes: Compare the document segmentation fragments output by the model with the manually annotated standard segmentation results, and use multiple evaluation indicators to preliminarily evaluate the segmentation quality; Input each segmented fragment into the pre-trained language model and score the semantic coherence of the text of each fragment; For segments with scores lower than the set threshold, further analyze their internal text logic and feature distribution, check the grammatical structure, keyword distribution and feature differences with adjacent segments of the text, and if the text is semantically incomplete, merge the current segment with the adjacent segment, and determine the best adjustment plan by comparing the semantic coherence scores before and after the merger.
9. The document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: In step S5, the post-processing and optimization method includes: Build a user feedback system to evaluate and annotate the segmentation results through the graphical interactive interface of the user feedback system; The system collects user feedback data in real time and stores it in the database. When the amount of feedback data reaches a threshold or after a certain period of time, the online optimization process is started to extract the incorrect segmentation cases and low-scoring cases marked by users from the database, and organize the original documents corresponding to the above cases and their correct segmentation annotation information into an optimized data set. The optimized dataset is mixed with the original training dataset in a preset ratio, and the model is continued to be trained on the mixed dataset. During the training process, the loss function of the model is weighted and adjusted according to the key focus areas and error types provided by user feedback.
10. The document segmentation method based on a multimodal large model as claimed in claim 1, characterized in that: In step S5, the post-processing and optimization method includes: When fine-tuning the parameters of a large multimodal model, adjust the learning rate before each iteration based on the model's segmentation performance on the validation set. If the performance of the model shows a downward trend in several consecutive iterations, reduce the learning rate; if the performance of the model continues to improve in several consecutive iterations, increase the learning rate; The model complexity index is calculated according to the number of parameters and connection methods of each layer in the model, and the weight of the regularization term is dynamically adjusted based on the complexity index; for layers with higher complexity, the weight of the regularization term is increased; for layers with lower complexity, the weight of the regularization term is reduced; when calculating the total loss function, the regularization term is added to the loss term based on the segmentation task to jointly guide the update of the model parameters.
Citation Information
Cited By
PDF text extraction method and system based on large language model
CN120599643A
Material information processing method and device based on task planning
CN121146539A
Complex table intelligent analysis method based on multi-modal fusion and semantic analysis
CN121303073A
Intelligent collecting and processing method for open source intelligence data crawling
CN121808125A
Document block quality detection method, electronic equipment and storage medium
CN121920321A