A multimodal data processing method, apparatus and electronic device

By performing modality-specific block processing and structural semantic parsing on multimodal data, the problem of lost correlation in multimodal data processing is solved, enabling more accurate and effective data block generation and improving the quality of retrieval enhancement results.

CN120688021BActive Publication Date: 2025-10-28INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511197936.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-28
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing technologies cannot effectively preserve the correlation between different modal data when processing multimodal data, leading to the collapse of cross-modal data structures and the breakage of correlations. In particular, the loss of table row and column relationships and the separation of text and images in documents result in misalignment of descriptive text and images.

Method used

We employ a semantic entropy dynamic segmentation strategy, a table structure extraction strategy, a formula format encapsulation strategy, and an image extraction strategy to segment text, tables, formulas, and images into blocks, respectively. Combined with structural semantic parsing and cross-modal merging processing, we generate structural semantic tags to preserve the relationships between data.

Benefits of technology

It improves the accuracy and effectiveness of multimodal data segmentation, ensures the correlation and integrity between different modal data, and enhances the accuracy of generated answers and the effectiveness of retrieval enhancement results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688021B_ABST
    Figure CN120688021B_ABST
Patent Text Reader

Abstract

This application discloses a multimodal data processing method, apparatus, and electronic device, relating to the field of data processing technology. The method includes segmenting multimodal data into blocks according to a segmentation strategy corresponding to each modality, obtaining modal data blocks for each modality. By employing different segmentation strategies for different modalities, effective segmentation of each modality's data is ensured. Further, structural semantic parsing is performed on the modal data blocks to obtain structural semantic data, and structural semantic tags are generated based on the structural semantic data. Then, cross-modal merging processing is performed on the modal data blocks based on the structural semantic tags to obtain data blocks corresponding to the multimodal data. This cross-modal merging processing from the structural semantic dimension ensures that the obtained data blocks retain the effective correlation between data from different modalities, solving the problem of inaccurate data segmentation in multimodal data; and achieving the technical effect of improving the accuracy and effectiveness of data blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a multimodal data processing method, apparatus and electronic device. Background Technology

[0002] In Retrieval-Augmented Generation (RAG) systems, document chunking is a fundamental preprocessing step. The quality of the chunking results directly affects the relevance of information retrieval and the accuracy of the generated answers. Currently, RAG systems can use fixed-length cutting or fixed sliding windows for document chunking. In this way, efficient segmentation can be achieved by relying on preset rules (fixed length or fixed window size). In addition, semantically coherent paragraphs can be merged through embedding clustering of NLP (Natural Language Processing) models to maintain the integrity of semantic logic in the chunking results.

[0003] However, in the current era of digital information explosion, document form has evolved from traditional plain text to composite carriers containing complex elements such as text, tables, formulas and / or images. The semantic integrity of such multimodal documents is highly dependent on the content structure and logical relationship. If traditional methods are still used for block processing, the multimodal structure will collapse and the relationship between cross-modal data will be broken. For example, the row and column relationship of the table will be lost after block processing, and the separation of text and images will lead to the misalignment of descriptive text and images. Summary of the Invention

[0004] This application provides a multimodal data processing method, apparatus, and electronic device to at least solve the problem of inaccurate data segmentation of multimodal data in related technologies.

[0005] This application provides a multimodal data processing method, including:

[0006] Acquire multimodal data, and divide the corresponding modal data in the multimodal data into blocks according to the block division strategy corresponding to each modality to obtain modal data blocks for each modality;

[0007] Structural semantic parsing is performed on the modal data block to obtain structural semantic data, and structural semantic tags for the modal data block are generated based on the structural semantic data.

[0008] Based on the structural semantic tags, the modal data blocks are merged across modalities to obtain the data blocks corresponding to the multimodal data.

[0009] This application also provides a multimodal data processing apparatus, including:

[0010] The block processing module is used to acquire multimodal data, and to perform block processing on the corresponding modal data in the multimodal data according to the block processing strategy corresponding to each modality to obtain modal data blocks for each modality.

[0011] The structural semantic parsing module is used to perform structural semantic parsing on the modal data block, obtain structural semantic data, and generate structural semantic tags for the modal data block based on the structural semantic data.

[0012] The merging processing module is used to perform cross-modal merging processing on the modal data blocks based on the structural semantic tags to obtain the data blocks corresponding to the multimodal data.

[0013] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described multimodal data processing methods.

[0014] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described multimodal data processing methods.

[0015] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described multimodal data processing methods.

[0016] In processing multimodal data, this application first divides the corresponding modal data in the multimodal data into blocks according to the block division strategy corresponding to each modality, thereby obtaining modal data blocks for each modality. Different block division strategies are then applied to modal data of different modalities to ensure effective block division for each modality. Based on obtaining the modal data blocks for each modality, structural semantic parsing is performed on the modal data blocks to obtain structural semantic data. Based on the structural semantic data, structural semantic tags are generated for the modal data blocks. Then, cross-modal merging processing is performed on the modal data blocks based on the structural semantic tags to obtain the corresponding data blocks for the multimodal data. This cross-modal merging processing of modal data blocks, starting from the structural semantic tags, allows the obtained data blocks to retain the effective correlation between data of different modalities by combining structure and semantics, thereby improving the accuracy and effectiveness of the obtained data blocks. Attached Figure Description

[0017] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 A flowchart illustrating a multimodal data processing method provided in an embodiment of this application;

[0019] Figure 2 A flowchart illustrating a semantic entropy dynamic block processing procedure provided in an embodiment of this application;

[0020] Figure 3 This is a schematic diagram illustrating the process of determining cross-modal association data of an image data block, as provided in an embodiment of this application.

[0021] Figure 4 A flowchart illustrating a multimodal data processing method applied to a multimodal data segmentation scenario, provided in an embodiment of this application;

[0022] Figure 5 A flowchart illustrating a multimodal data processing method applied to a document data processing scenario, provided as an embodiment of this application;

[0023] Figure 6 A flowchart illustrating a multimodal data processing method applied to a user query data processing scenario, provided in an embodiment of this application;

[0024] Figure 7 This application provides a schematic diagram of a multimodal data processing method applied to a user query data processing scenario, as illustrated in an embodiment of the present application.

[0025] Figure 8 This is a schematic diagram of the structure of a multimodal data processing device provided in an embodiment of this application. Detailed Implementation

[0026] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0028] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] The embodiments of this application provide a multimodal data processing method, and the method is described in detail with reference to the execution flow of the multimodal data processing method.

[0030] Reference Figure 1 As shown, this multimodal data processing method includes the following steps S11-S13.

[0031] S11. Acquire multimodal data, and divide the corresponding modal data in the multimodal data into blocks according to the block division strategy corresponding to each modality to obtain modal data blocks for each modality.

[0032] In this embodiment, after obtaining multimodal data, the corresponding modal data in the multimodal data is segmented according to the segmentation strategy corresponding to each modality to obtain modal data blocks for each modality. Thus, different segmentation strategies are used for data of different modalities to ensure effective segmentation of data for each modality. Multimodal data consists of data from at least one modality; for example, multimodal data can consist of text, tables, formulas, and / or images; specifically, multimodal data can be multimodal documents or user query data.

[0033] In practice, to preserve the semantics of the text, the text in the multimodal data can be segmented according to the semantic entropy dynamic segmentation strategy to obtain text data blocks; to divide paginated tables into the same data block and improve the integrity of the table data block, the tables in the multimodal data can be segmented according to the table structure extraction strategy to obtain table data blocks; to prevent formulas from being corrupted by the text segmenter, the formulas in the multimodal data can be segmented according to the formula format encapsulation strategy to obtain formula data blocks; and to achieve image extraction, the images in the multimodal data can be segmented according to the image extraction strategy to obtain image data blocks.

[0034] The following details the process of dividing the multimodal data into blocks according to the block strategy corresponding to each modality to obtain the modal data blocks for each modality.

[0035] (1) Based on the semantic entropy dynamic segmentation strategy, the text in the multimodal data is segmented to obtain text data blocks.

[0036] In the embodiments of this application, reference is made to Figure 2 As shown, the text in multimodal data is segmented according to the semantic entropy dynamic segmentation strategy to obtain text data blocks, which can be achieved through the following steps S21-S24.

[0037] S21. Determine at least one sentence in the text based on regular expression matching and semantic coherence verification.

[0038] In the embodiments of this application, in the process of determining at least one sentence in the text based on regular expression matching and semantic coherence verification, boundary filtering can be performed first based on regular expression matching to obtain candidate sentences; then, semantic coherence verification can be performed on the candidate sentences to obtain at least one sentence that passes the verification.

[0039] For example, by matching the boundaries in the text with boundary symbols from a pre-configured set of boundary symbols, candidate sentences are obtained. Then, semantic coherence is checked on the candidate sentences to obtain at least one sentence that passes the check.

[0040] In the process of performing semantic coherence verification on candidate sentences and obtaining at least one sentence that passes the verification, the semantic coherence verification can be performed on each candidate sentence first to obtain the verification result of each candidate sentence. If adjacent candidate sentences fail the verification, adjacent candidate sentences are merged to obtain a merged candidate sentence, and the semantic coherence verification is performed on the merged candidate sentence. If the verification passes, the merged candidate sentence is regarded as the sentence that passes the verification. If the verification fails, no processing is required.

[0041] S22. Dynamically generate multiple text fragments for each sentence using a sliding window.

[0042] In the embodiments of this application, the number of multiple text segments can be equal to the block granularity, and the size of the sliding window can be an integer less than or equal to the block granularity.

[0043] Specifically, by using a sliding window to generate text fragments of different lengths at a block granularity, the limitations of a fixed window are overcome, local fluctuations are avoided, long-distance dependencies are captured, and robustness is enhanced by generating multiple text fragments of different window sizes.

[0044] For example, the granularity of the segmentation is n, for a sentence The window sizes are 1, 2, 3, ..., n, respectively, for each sentence. Add subsequent sentences to get the result of adding one subsequent sentence. Add two subsequent sentences Add 3 subsequent sentences ..., adding n subsequent sentences .

[0045] In this embodiment, different users have different preferences, and different query data have different levels of complexity. In order to make the granularity of the segmentation more suitable for users, user query data, and / or the system, and to avoid segmenting simple query data into fine-grained blocks and complex query data into coarse-grained blocks, thereby improving the effectiveness of the obtained granularity of the segmentation and thus improving the effectiveness of semantic entropy segmentation based on the granularity of the segmentation, the granularity of the segmentation can be obtained in the following way:

[0046] Obtain user query data and read user historical profiles and / or system load data;

[0047] Input user query data, user historical profiles and / or system load data into the feature extraction model to extract features and obtain feature vectors;

[0048] The feature vector is input into the granularity calculation model to perform granularity calculation and obtain the block granularity.

[0049] Specifically, after acquiring user query data, the system reads user historical profiles and / or system load data. Then, it inputs these data into a feature extraction model for feature extraction, obtaining feature vectors representing dimensions such as user query data, user historical profiles, and / or system load data. Finally, the feature vectors are input into a granularity calculation model for granularity calculation to obtain the block granularity. The user historical profile includes data characterizing user interaction preferences; for example, it includes historical query data and / or user feedback data on historical responses.

[0050] In addition, for documents, the document, user history profile and / or system load data can be input into the feature extraction model to extract features and obtain feature vectors; the feature vectors can be input into the granularity calculation model to calculate granularity and obtain block granularity.

[0051] S23. Calculate the semantic entropy of each sentence based on multiple text fragments.

[0052] Based on the above method of dynamically generating multiple text fragments for each sentence through a sliding window, the semantic entropy of each sentence is calculated based on the multiple text fragments of each sentence. Semantic entropy is used to measure the certainty of a sentence as the end point of a semantic unit in a specific context. By using semantic entropy, the generative capability of the language model is transformed into a quantitative analysis tool for document structure, which solves the fundamental defects of traditional segmentation methods and improves the accuracy of measuring sentences as semantic end units.

[0053] In this embodiment of the application, in order to integrate information from multiple text segments of each sentence to calculate the semantic entropy of each sentence, thereby improving the accuracy of the calculated semantic entropy, the semantic entropy of each sentence based on multiple text segments can be calculated in the following way:

[0054] Each text segment from multiple text segments is input into the termination probability prediction model to predict the termination probability and obtain the termination probability of each text segment.

[0055] Calculate the information entropy of each text segment based on the termination probability, and calculate the sum of the information entropy of each text segment as the target information entropy;

[0056] The negative normalization factor is determined based on the block granularity, and the product of the negative normalization factor and the target information entropy is calculated as the semantic entropy.

[0057] Specifically, sentences The formula for calculating semantic entropy can be found as follows:

[0058]

[0059] in, The sentence indicating the current assessment, Sentence Semantic entropy, used to quantify sentences As the strength of the semantic boundary, n represents the context window size, i.e., the block granularity, which is used to control the scope of semantic judgment. Sentence ;

[0060] This represents the probability of termination, and log2 represents the logarithm to the base 2, which is the standard form for calculating information entropy. It is a negative standardization factor used to convert the result into a positive entropy value.

[0061] in, The end probability can be obtained using an end probability prediction model. For example, the text fragment can be input into the end probability prediction model to predict the end probability and obtain the end probability of the text fragment.

[0062] The termination probability prediction model in this embodiment can be a TinyBERT model, which can be trained using distillation fine-tuning techniques and training data generated by LLaMA-7B. To save model storage space, FP16-INT8 quantization can be used for deployment, reducing the model size by a factor of 4. The trained TinyBERT model can then be used to calculate... When calculating the end probability, the end probability can be predicted by predicting whether the next token is a sentence terminator (such as a period, question mark, etc.). In this way, the computational load of the model is reduced by reasoning only about the last token. The higher the end probability, the more the model believes that the current context has formed a complete semantic unit.

[0063] S24. Determine the block endpoint based on semantic entropy, and perform block processing based on the block endpoint to obtain text data blocks.

[0064] After obtaining the semantic entropy of each sentence through the above calculation, the block endpoint is determined based on the semantic entropy, and the block processing is performed according to the block endpoint to obtain text data blocks.

[0065] Specifically, if the semantic entropy of a sentence is less than the first threshold, the sentence is determined to belong to a strong boundary and is used as the end point of the block. If the semantic entropy of a sentence is greater than the second threshold, the sentence is determined not to belong to the end point of the block. If the semantic entropy of a sentence is greater than or equal to the first threshold and less than or equal to the second threshold, the sentence is determined to belong to the transition zone and is used as a candidate block node.

[0066] For example, when At that time, determine the sentence It belongs to a strong boundary, and the sentence As the end point of the block, when At that time, determine the sentence Not belonging to the end point of the block, when At that time, determine the sentence This is a transitional zone; the sentence needs to be defined. As a candidate block node.

[0067] In the process of segmenting text data into blocks based on the block endpoint, when a block endpoint is detected, the block endpoint and the preceding text can be identified as a single text data block. When a candidate block endpoint is detected, it is checked whether the length of the text formed by the candidate block endpoint and the preceding text is greater than a length threshold. If so, the text formed by the candidate block endpoint and the preceding text can be identified as a single text data block. If not, reading can continue or no processing can be performed.

[0068] (2) Based on the table structure extraction strategy, the tables in the multimodal data are divided into blocks to obtain table data blocks.

[0069] In this embodiment, each page of the document can first be identified to extract the table structure and content within a single page. Then, the header text of the table on the single page can be extracted and converted into a header hash value. This compresses the variable-length header text into a fixed-length string, improving the efficiency and convenience of comparison and matching. After converting the header text into a header hash value, the header hash values ​​of tables on different pages are compared, and the table content with the same header hash value is merged. In this way, the table data blocks of each table are obtained, realizing the merging of cross-page tables into the same table data block.

[0070] (3) According to the formula format encapsulation strategy, the formulas in the multimodal data are divided into blocks to obtain formula data blocks.

[0071] In this embodiment of the application, formulas in multimodal data can be encapsulated in LaTeX to prevent them from being corrupted by the text segmenter.

[0072] (4) Based on the image extraction strategy, the images in the multimodal data are divided into blocks to obtain image data blocks.

[0073] In this embodiment of the application, each image in the multimodal data can be divided into an image data block.

[0074] In this embodiment of the application, in order to further improve the efficiency of block processing, after the data is acquired, it can first be detected whether the acquired data is multimodal data. If so, the corresponding modal data in the multimodal data is block processed according to the block processing strategy corresponding to each modality to obtain modal data blocks of each modality. If not, the acquired data can be block processed according to the block processing strategy corresponding to the modality of the acquired data to obtain data blocks.

[0075] Taking document acquisition as an example, after acquiring the document, in order to obtain the document content, the document can first be input into a multimodal parser for parsing to obtain document data. Then, it is checked whether the document data is multimodal data. If it is, the document data is treated as multimodal data. If not, the document data is divided into blocks according to the semantic entropy dynamic block division strategy to obtain document data blocks.

[0076] Taking the acquisition of user query data as an example, after acquiring the user query data and calculating the granularity of the blocks, it is checked whether the user query data is multimodal data. If so, the corresponding modal data in the user query data is processed into blocks according to the block strategy corresponding to each modality to obtain modal data blocks for each modality. If not, the user query data is processed into blocks according to the block strategy corresponding to the modality of the user query data to obtain data blocks.

[0077] S12. Perform structural semantic parsing on the modal data block to obtain structural semantic data, and generate structural semantic tags for the modal data block based on the structural semantic data.

[0078] In the technical solution of this application, after dividing the modal data in the multimodal data into blocks according to the block division strategy corresponding to each modality to obtain modal data blocks of each modality, structural semantic parsing is performed on the modal data blocks to obtain structural semantic data, and structural semantic labels of the modal data blocks are generated based on the structural semantic data.

[0079] To enable structural semantic data to represent the structural and semantic dimensions of modal data blocks, structural semantic data can include the modality type, semantic content, and / or position of a single modal data block within the multimodal data. Furthermore, to preserve relationships between cross-modal data within the multimodal data and prevent data in a modal data block from being too singular or incomplete, the structural semantic data of a modal data block can also include cross-modal association data, such as the target text fragment associated with an image data block or the title associated with a table data block. Based on this, during the process of structural semantic parsing of multimodal data blocks to obtain structural semantic data, modality type identification can be performed on the modal data blocks to obtain the modality type; semantic extraction can be performed on the modal data to obtain the semantic content; cross-modal matching can be performed on the modal data blocks to obtain the cross-modal association data; and / or, the positional information of the modal data block within the multimodal data can be determined.

[0080] In the specific execution process, semantic extraction is performed on modal data blocks to obtain semantic content. Semantic extraction can be performed on modal data blocks according to the semantic extraction strategy corresponding to the modality type of the modal data blocks to obtain semantic content. Alternatively, the modal data in the modal data blocks can be input into a large language model for semantic extraction to obtain semantic content.

[0081] Taking semantic extraction of text data blocks as an example, during the process of semantic extraction of text data blocks according to the text semantic extraction strategy, text keywords can be extracted as the semantic content of the text data blocks.

[0082] Taking semantic extraction of table data blocks as an example, during the process of semantic extraction of table data blocks according to the table semantic extraction strategy, the table header and / or title can be extracted as the semantic content of the table data block; for example, data that represents the uniqueness of the table, such as the table number, can also be extracted as the semantic content of the table data block.

[0083] Taking semantic extraction of image data blocks as an example, during the process of semantic extraction of image data blocks according to the image semantic extraction strategy, OCR technology can be used to recognize image text and use the image text as the semantic content of the image data block; for example, data that represents the uniqueness of an image, such as the image number, can also be extracted as the semantic content of the image data block.

[0084] Due to their unique format, images and tables often require textual explanations to ensure their completeness and validity. The following sections provide a detailed explanation of the process for determining cross-modal data associations for image and table data blocks.

[0085] The first method involves performing cross-modal matching on image data blocks to obtain cross-modal association data for the image data blocks.

[0086] In the technical solution of this application, the process of performing cross-modal matching on image data blocks to obtain cross-modal association data of image data blocks can be implemented in the following way:

[0087] The images in the image data block are input into the image feature extraction model to extract image features and obtain image features. Similarly, each text fragment is input into the text feature extraction model to extract text features and obtain text features for each text fragment.

[0088] Calculate the similarity between image features and text features of each text segment, and identify target text segments with similarity greater than a similarity threshold;

[0089] The target text fragment is identified as cross-modal associated data of the image data block.

[0090] Specifically, target text fragments with a similarity greater than a first similarity threshold between text features and image features are selected from text segments and used as cross-modal association data for image data blocks corresponding to image features. In this way, target text fragments for image matching are selected by calculating the similarity between image features and text features, improving matching efficiency and convenience. The target text fragments are identified as cross-modal association data for image data blocks, enabling the target text fragments to provide textual descriptions of the images in the image data blocks. This avoids situations where image data blocks only contain images, making it impossible to perceive the meaning of the image representation from the images themselves, thus ensuring the effectiveness of the image data blocks.

[0091] For example, Figure 3 As shown, the image is input into the CLIP (Contrastive Language-Image Pre-training) model to extract image features, and the text fragment is input into the BERT (Bidirectional Encoder Representations from Transformers) model to extract text features, and the similarity between the image features and the text features is calculated. It is then checked whether the similarity is greater than a similarity threshold. If so, the target text fragment with a similarity greater than the similarity threshold is identified as cross-modal association data of the image data block, that is, the association relationship between the target text fragment and the image data block is established. If not, the cross-modal association data of the image data block is determined to be empty.

[0092] Pre-trained language models generally refer to language model training tasks designed based on large-scale corpora (including language training materials such as sentences and paragraphs). A large-scale neural network algorithm structure is trained to learn and implement the model, resulting in a pre-trained language model with the final large-scale neural network algorithm structure and parameters. Subsequent tasks can then use this model for feature extraction or task fine-tuning to achieve specific objectives. The idea behind pre-training is to first train a set of model parameters for one task, then use these parameters to initialize the network model parameters, and finally use the initialized network model to train other tasks, obtaining models adapted for those tasks. By pre-training on large-scale corpora, neural language representation models can learn powerful language representation capabilities, extracting rich syntactic and semantic information from text. Pre-trained language models can provide tokens containing rich semantic information and sentence-level features for downstream tasks. Furthermore, fine-tuning can be directly performed on the pre-trained model for downstream tasks, conveniently and quickly obtaining downstream-specific models. The neural network algorithm structure used to train the pre-trained language model can be CNN, RNN, LSTM, etc., or it can be a model built with attention networks, such as Transformer, BERT, GPT, Clip, etc. This application does not limit it. An attention network is a network model that uses an attention mechanism for training. This model extracts more important feature information from the input sequence by assigning different weights to each part of the input sequence, so that the model can finally obtain a more accurate output.

[0093] The second method involves performing cross-modal matching on the table data blocks to obtain cross-modal associated data for the table data blocks.

[0094] In this embodiment of the application, the process of performing cross-modal matching on table data blocks to obtain cross-modal associated data of the table data blocks can be implemented in the following way:

[0095] Based on the location information of the tables in the tabular data block within the multimodal data, candidate titles are filtered out from the multimodal data;

[0096] Perform semantic matching verification between candidate titles and tables. If the verification passes, the candidate title is identified as cross-modal associated data of the table data block.

[0097] Specifically, candidate titles whose positional difference between their positional information in multimodal data and the positional information of the table in multimodal data is less than a positional threshold can be selected. Semantic matching verification is then performed on the candidate titles and the table. If the verification is successful, the candidate titles are identified as cross-modal associated data of the table data block. In this way, titles are associated with the table data block, thereby improving the completeness and effectiveness of the content in the table data block.

[0098] In the specific process of semantic matching verification between candidate titles and tables, it can be verified whether the table number in the candidate title matches the table number in the table data block. If yes, the verification is considered successful; otherwise, the verification is considered unsuccessful. Alternatively, title keywords and table keywords can be extracted, and the keyword overlap between the title keywords and table keywords can be checked to see if it exceeds the overlap threshold. If yes, the verification is considered successful; otherwise, the verification is considered unsuccessful. In addition, the verification of table numbers can be combined with keyword overlap detection. For example, it can be verified whether the table number in the candidate title matches the table number in the table data block. If no, the verification is considered unsuccessful; if yes, title keywords and table keywords can be extracted, and the keyword overlap between the title keywords and table keywords can be checked to see if it exceeds the overlap threshold. If yes, the verification is considered successful; otherwise, the verification is considered unsuccessful. No action is taken if the verification fails.

[0099] Furthermore, based on the structural semantic data obtained above, in the process of generating structural semantic labels for modal data blocks according to the structural semantic data, modality type, semantic content, location information and / or cross-modal association data can be merged to obtain structural semantic labels.

[0100] The above describes the process of obtaining structural semantic data by parsing the structural semantics of a modal data block using a single modal data block as an example. It should be noted that in this embodiment, structural semantics is parsed for each modal data block to obtain structural semantic data, and structural semantic tags for each modal data block are generated based on the structural semantic data. Specifically, the process of parsing the structural semantics of each modal data block is similar to the above process, and can be referred to the relevant content above. This embodiment will not repeat the details here.

[0101] S13. Based on the structural semantic tags, perform cross-modal merging processing on the modal data blocks to obtain the data blocks corresponding to the multimodal data.

[0102] Based on the structural semantic tags obtained above for the modal data blocks, cross-modal merging processing is performed on the modal data blocks according to the structural semantic tags to obtain the data blocks corresponding to the multimodal data.

[0103] In practice, cross-modal correlation data of modal data blocks can be merged with modal data blocks, and modal data blocks can also be merged with modal data blocks to which cross-modal correlation data belongs, thereby improving the accuracy and effectiveness of the data blocks corresponding to the obtained multimodal data.

[0104] For example, the cross-modal association data of an image data block is "As shown in Figure x, maple leaves are red". This cross-modal association data is added to the image data block to obtain an image data block containing the cross-modal association data. Alternatively, a text data block containing the cross-modal association data is determined, and the text data block is merged with the image data block to obtain a data block containing both text and image.

[0105] Another example is that the cross-modal associated data of a table data block is "xxx sales data". This cross-modal associated data is added to the table data block to obtain a table data block containing the cross-modal associated data. Alternatively, a text data block containing the cross-modal associated data is determined, and the text data block is merged with the table data block to obtain a data block containing both text and a table.

[0106] In another example, if the cross-modal association data of the table data block and the cross-modal association data of the image data block are both "xxx sales data", the table data block and the image data block are merged to obtain a data block containing a table and an image, and the cross-modal association data is added to the data block. Alternatively, the text data block containing the cross-modal association data is determined, and the text data block, the table data block, and the image data block are merged to obtain a data block containing text, an image, and a table.

[0107] In this embodiment, after obtaining the data blocks corresponding to multimodal data, subsequent processing can be performed based on the data blocks. For example, after obtaining the data blocks corresponding to multimodal documents, a knowledge base can be constructed based on the data blocks. Another example is that after obtaining the data blocks corresponding to multimodal user query data, the data blocks are input into a Retrieval-Augmented Generation (RAG) system for retrieval-augmented generation to obtain retrieval-augmented generation results. Furthermore, the retrieval-augmented generation results are sent to the user as response data for the user query data. This improves both the effectiveness and accuracy of the obtained data blocks and the effectiveness and accuracy of the generated retrieval-augmented generation results, enhancing the user's perception of the response data. In addition, documents can also be input into the retrieval-augmented generation system to obtain retrieval-augmented generation results, and a knowledge base can also be constructed based on the data blocks for user query data; this embodiment does not limit the scope of the example.

[0108] Since this embodiment improves the flexibility and accuracy of block granularity by inputting feature vectors into the granularity calculation model for granularity calculation, thereby improving the flexibility and accuracy of semantic entropy dynamic block segmentation, in order to further improve the accuracy of the determined block granularity, this embodiment, after obtaining the retrieval enhancement generation result, can also input the retrieval enhancement generation result into the quality assessment model for quality assessment to obtain an assessment score, input the assessment score into the reward calculation model for reward calculation to obtain a reward signal, and perform reinforcement learning on the granularity calculation model based on the reward signal, thereby improving the performance of the granularity calculation model and thus improving the accuracy of the block granularity obtained by granularity calculation based on the granularity calculation model.

[0109] The following description uses an embodiment of this application to illustrate the application of a multimodal data processing method in a multimodal data segmentation scenario, further explaining the multimodal data processing method provided by the embodiment of this application. (Refer to...) Figure 4 As shown, the multimodal data processing method applied to multimodal data segmentation scenarios includes the following steps.

[0110] S41. Obtain multimodal data.

[0111] S42. Based on the semantic entropy dynamic segmentation strategy, the text in the multimodal data is segmented to obtain text data blocks.

[0112] S43. Based on the table format extraction strategy, the tables in the multimodal data are divided into blocks to obtain table data blocks.

[0113] S44. Based on the formula format encapsulation strategy, the formulas in the multimodal data are divided into blocks to obtain formula data blocks.

[0114] S45. Based on the image extraction strategy, the images in the multimodal data are divided into blocks to obtain image data blocks.

[0115] It should be noted that the execution order of steps S42 to S45 is not limited here.

[0116] S46. Perform structural semantic parsing on text data blocks, table data blocks, formula data blocks, and image data blocks to obtain structural semantic data.

[0117] S47. Generate structural semantic labels for each data block based on the structural semantic data of each data block.

[0118] S48. Perform cross-modal merging processing on the data blocks according to the structural semantic labels of each data block to obtain the data blocks corresponding to the multimodal data.

[0119] It should be noted that any one or more steps from S41 to S48 can be combined with any one or more steps from S11 to S13 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features can be selected in steps S41 to S48 and combined with any one or more technical features provided in steps S11 to S13 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features in steps S41 to S48 can be replaced with any one or more technical features provided in steps S11 to S13 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.

[0120] The following description uses an embodiment of this application to illustrate the application of a multimodal data processing method in a document data processing scenario, further explaining the multimodal data processing method provided by the embodiments of this application. (Refer to...) Figure 5 As shown, the multimodal data processing method applied to document data processing scenarios includes the following steps.

[0121] S51. Obtain the document and input it into the multimodal parser for parsing to obtain document data.

[0122] S52. Detect whether the document data is multimodal document data;

[0123] If so, execute S53 and S58;

[0124] If not, proceed with steps S54 through S58.

[0125] S53. Perform semantic entropy dynamic block processing on the document data to obtain document data blocks.

[0126] S54. Divide the corresponding modal data in the multimodal data into blocks according to the block division strategy corresponding to each modality to obtain modal data blocks for each modality.

[0127] S55. Perform structural semantic parsing on the modal data block to obtain structural semantic data.

[0128] S56. Generate structural semantic labels for modal data blocks based on structural semantic data.

[0129] S57. Based on structural semantic tags, perform cross-modal merging processing on modal data blocks to obtain data blocks corresponding to multimodal document data.

[0130] It should be noted that S54-S57 can be multimodal structure-aware block segmentation; that is, if the document data is multimodal document data, multimodal structure-aware block segmentation processing is performed on the multimodal document data to obtain the data blocks corresponding to the multimodal document data.

[0131] S58, input the document data block or the data block corresponding to the multimodal document data into the retrieval enhancement generation system to perform retrieval enhancement generation and obtain the retrieval enhancement generation result.

[0132] It should be noted that any one or more of steps S51 to S58 can be combined with any one or more of steps S11 to S13 to form a new implementation method according to the needs of implementation and deployment. In addition, according to the actual deployment needs, any one or more technical features in steps S51 to S58 can be selected and combined with any one or more technical features provided in steps S11 to S13 to form a new implementation method. Alternatively, any one or more technical features in steps S51 to S58 can be replaced with any one or more technical features provided in steps S11 to S13 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.

[0133] The following description uses an example of an embodiment of this application to illustrate the application of a multimodal data processing method in a user query data processing scenario, further explaining the multimodal data processing method provided by the embodiments of this application. (Refer to...) Figure 6 and Figure 7 As shown, the multimodal data processing method applied to user query data processing scenarios includes the following steps.

[0134] S61. Obtain user query data and read user historical profiles and system load data.

[0135] S62. Input user query data, user historical profiles and system load data into the feature extraction model to extract features and obtain feature vectors.

[0136] S63. Input the feature vector into the granularity calculation model to perform granularity calculation and obtain the block granularity.

[0137] It should be noted that before inputting the feature vector into the granularity calculation model to perform granularity calculation and obtain the block granularity, the granularity calculation model can also be subjected to reinforcement learning.

[0138] Reference Figure 7As shown, after obtaining user query data, user historical profiles, and system load data, the user query data, user historical profiles, and system load data are input into the feature extraction model for feature extraction to obtain feature vectors. The granularity calculation model is then reinforced by a reinforcement learning agent, and the feature vectors are input into the granularity calculation model for granularity calculation to obtain the block granularity.

[0139] S64. If the user query data is multimodal data, the corresponding modal data in the multimodal data is divided into blocks according to the block division strategy corresponding to each modality to obtain modal data blocks for each modality.

[0140] Among them, the block granularity is used to dynamically divide the text in multimodal data into semantic entropy blocks.

[0141] S65. Perform structural semantic parsing on the modal data block to obtain structural semantic data, and generate structural semantic labels for the modal data block based on the structural semantic data.

[0142] S66. Perform cross-modal merging processing on modal data blocks based on structural semantic tags to obtain data blocks corresponding to multimodal document data.

[0143] S67. Input the data block into the retrieval enhancement generation system to perform retrieval enhancement generation and obtain the retrieval enhancement generation results.

[0144] S68. Input the enhanced search results into the quality assessment model for quality assessment and obtain an assessment score.

[0145] S69. Input the evaluation score into the reward calculation model to calculate the reward and obtain the reward signal, so as to perform reinforcement learning on the granular calculation model based on the reward signal.

[0146] Reference Figure 7 As shown, after obtaining the block granularity, the block granularity and user query data are input into the block segmentation module for block processing according to S64-S66. After obtaining the data blocks, the data blocks are input into the RAG system for retrieval enhancement generation. After obtaining the retrieval enhancement generation results, the retrieval enhancement generation results are input into the quality assessment model for quality assessment to obtain the assessment score. The assessment score is input into the reward calculation model for reward calculation to obtain the reward signal. The reward signal is then fed back to the reinforcement learning agent (intelligent agent or module) to perform reinforcement learning on the granularity calculation model.

[0147] It should be noted that any one or more of steps S61 to S69 can be combined with any one or more of steps S11 to S13 to form a new implementation method according to the needs of implementation and deployment. In addition, according to the actual deployment needs, any one or more technical features in steps S61 to S69 can be selected and combined with any one or more technical features provided in steps S11 to S13 to form a new implementation method. Alternatively, any one or more technical features in steps S61 to S69 can be replaced with any one or more technical features provided in steps S11 to S13 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.

[0148] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0149] Figure 8 This is a schematic diagram of the structure of a multimodal data processing device 800 provided in this disclosure, as shown below. Figure 8 As shown, the device 800 of this embodiment includes:

[0150] The block processing module 810 is used to acquire multimodal data, and to perform block processing on the corresponding modal data in the multimodal data according to the block processing strategy corresponding to each modality to obtain modal data blocks of each modality.

[0151] The structural semantic parsing module 820 is used to perform structural semantic parsing on the modal data block to obtain structural semantic data, and generate structural semantic tags for the modal data block based on the structural semantic data.

[0152] The merging processing module 830 is used to perform cross-modal merging processing on the modal data block based on the structural semantic label to obtain the data block corresponding to the multimodal data.

[0153] As an optional implementation of this application, the block processing module 810 is specifically used to perform the following when dividing the modal data in the multimodal data into blocks according to the block processing strategy corresponding to each modality to obtain modal data blocks for each modality: dividing the text in the multimodal data into blocks according to the semantic entropy dynamic block processing strategy to obtain text data blocks; dividing the tables in the multimodal data into blocks according to the table structure extraction strategy to obtain table data blocks; dividing the formulas in the multimodal data into blocks according to the formula format encapsulation strategy to obtain formula data blocks; and dividing the images in the multimodal data into blocks according to the image extraction strategy to obtain image data blocks.

[0154] As an optional implementation of this application, the block processing module 810 is specifically used to, when performing block processing on the text in the multimodal data according to the semantic entropy dynamic block processing strategy to obtain text data blocks, determine at least one sentence in the text based on regular expression matching and semantic coherence verification; dynamically generate multiple text segments for each sentence through a sliding window; the number of the multiple text segments is equal to the block granularity; the size of the sliding window is successively an integer less than or equal to the block granularity; calculate the semantic entropy of each sentence based on the multiple text segments, and determine the block endpoint based on the semantic entropy, so as to perform block processing based on the block endpoint to obtain text data blocks.

[0155] As an optional implementation of this application, the block processing module 810 is further configured to acquire user query data and read user historical profiles and system load data; input the user query data, the user historical profiles and system load data into a feature extraction model for feature extraction to obtain feature vectors; and input the feature vectors into a granularity calculation model for granularity calculation to obtain the block granularity.

[0156] As an optional implementation of this application, the block processing module 810 is further configured to input the document into a multimodal parser for parsing processing to obtain document data; detect whether the document data is multimodal data; if so, treat the document data as multimodal data; if not, perform block processing on the document data according to the semantic entropy dynamic block strategy to obtain document data blocks.

[0157] As an optional implementation of this application, the structural semantic parsing module 820 is specifically used to, when performing structural semantic parsing on the modal data block to obtain structural semantic data, identify the modal type of the modal data block to obtain the modal type; extract semantics from the modal data block to obtain semantic content; and perform cross-modal matching on the modal data block to obtain cross-modal association data of the modal data block.

[0158] As an optional implementation of this application, the structural semantic parsing module 820 is specifically used to, when performing cross-modal matching on the modal data block to obtain cross-modal association data of the modal data block, input the images in the image data block into an image feature extraction model to extract image features, and input each text segment into a text feature extraction model to extract text features, thereby obtaining text features of each text segment; calculate the similarity between the image features and the text features of each text segment, and determine the target text segment whose similarity is greater than a first similarity threshold; and determine the target text segment as the cross-modal association data of the image data block.

[0159] As an optional implementation of this application, the structural semantic parsing module 820 is specifically used to, when performing cross-modal matching on the modal data block to obtain cross-modal associated data of the modal data block, filter out candidate titles in the multimodal data according to the position information of the table in the table data block in the multimodal data; perform semantic matching verification between the candidate titles and the table; and determine the candidate titles as cross-modal associated data of the table data block if the verification is successful.

[0160] As an optional implementation of this application, the merging processing module 830 is further configured to input the data block into the retrieval enhancement generation system for retrieval enhancement generation, and obtain retrieval enhancement generation results.

[0161] As an optional implementation of this application, the merging processing module 830 is further configured to input the retrieval enhancement generation result into a quality assessment model for quality assessment to obtain an assessment score; input the assessment score into a reward calculation model for reward calculation to obtain a reward signal, and perform reinforcement learning on the granularity calculation model based on the reward signal.

[0162] As an optional implementation of this application, the block processing module 810 is specifically used to, when calculating the semantic entropy of each sentence based on the plurality of text segments, input each text segment in the plurality of text segments into the end probability prediction model to predict the end probability and obtain the end probability of each text segment; calculate the information entropy of each text segment based on the end probability, and calculate the sum of the information entropies as the target information entropy; determine the negative normalization factor based on the block granularity, and calculate the product of the negative normalization factor and the target information entropy as the semantic entropy.

[0163] For a description of the features in the embodiment of a multimodal data processing device, please refer to the relevant description of the embodiment of the multimodal data processing method, which will not be repeated here.

[0164] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the multimodal data processing method.

[0165] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described multimodal data processing method embodiments when running.

[0166] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0167] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the multimodal data processing method.

[0168] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described multimodal data processing method embodiments.

[0169] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0170] The foregoing has provided a detailed description of a multimodal data processing method, apparatus, and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A multimodal data processing method, characterized in that, include: Acquire multimodal data, and divide the corresponding modal data in the multimodal data into blocks according to the block division strategy corresponding to each modality to obtain modal data blocks for each modality; The step of segmenting the corresponding modal data in the multimodal data according to the segmentation strategy corresponding to each modality to obtain modal data blocks for each modality includes: segmenting the text in the multimodal data according to the semantic entropy dynamic segmentation strategy to obtain text data blocks; the step of segmenting the text in the multimodal data according to the semantic entropy dynamic segmentation strategy to obtain text data blocks includes: determining at least one sentence in the text according to regular expression matching and semantic coherence verification; dynamically generating multiple text fragments for each sentence through a sliding window; the number of the multiple text fragments is equal to the segmentation granularity; the size of the sliding window is successively an integer less than or equal to the segmentation granularity; calculating the semantic entropy of each sentence according to the multiple text fragments, and determining the segmentation endpoint according to the semantic entropy, so as to perform segmentation processing according to the segmentation endpoint to obtain the text data blocks; sentence The formula for calculating semantic entropy is: ;in, The sentence indicating the current assessment, Sentence Semantic entropy, used to quantify sentences As the strength of the semantic edge, n represents the context window size, i.e., the block granularity, which is used to control the scope of semantic judgment. Sentence P(EOS) represents the termination probability. The base-2 logarithm represents the standard form of information entropy, and -1 / n is a negative normalization factor used to convert the result into a positive entropy value; P(EOS) is obtained using a termination probability prediction model. Structural semantic parsing is performed on the modal data block to obtain structural semantic data, and structural semantic tags for the modal data block are generated based on the structural semantic data. Based on the structural semantic tags, the modal data blocks are merged across modalities to obtain the data blocks corresponding to the multimodal data.

2. The multimodal data processing method according to claim 1, characterized in that, The step of dividing the multimodal data into blocks according to the block-division strategy corresponding to each modality to obtain the modal data blocks of each modality further includes at least one of the following: Based on the table structure extraction strategy, the tables in the multimodal data are divided into blocks to obtain table data blocks; According to the formula format encapsulation strategy, the formulas in the multimodal data are divided into blocks to obtain formula data blocks; Based on the image extraction strategy, the images in the multimodal data are divided into blocks to obtain image data blocks.

3. The multimodal data processing method according to claim 1, characterized in that, Before determining at least one sentence in the text based on regular expression matching and semantic coherence verification, the method further includes: Obtain user query data, and read user historical profiles and system load data; The user query data, the user historical profile, and the system load data are input into a feature extraction model to extract features and obtain feature vectors. The feature vector is input into the granularity calculation model to perform granularity calculation, thereby obtaining the block granularity.

4. The multimodal data processing method according to claim 1, characterized in that, The method further includes: The document is input into a multimodal parser for parsing and processing to obtain document data; Detect whether the document data is multimodal data; If so, the document data shall be treated as multimodal data; If not, the document data is divided into blocks according to the semantic entropy dynamic block division strategy to obtain document data blocks.

5. The multimodal data processing method according to claim 1, characterized in that, The structural semantic parsing of the modal data block to obtain structural semantic data includes at least one of the following: Modality type identification is performed on the modal data block to obtain the modality type; Semantic extraction is performed on the modal data blocks to obtain semantic content; Cross-modal matching is performed on the modal data block to obtain cross-modal association data of the modal data block.

6. The multimodal data processing method according to claim 5, characterized in that, The step of performing cross-modal matching on the modal data block to obtain cross-modal association data of the modal data block includes: The images in the image data block are input into the image feature extraction model to extract image features and obtain image features. The text fragments are input into the text feature extraction model to extract text features and obtain text features for each text fragment. Calculate the similarity between the image features and the text features of each text segment, and determine the target text segment whose similarity is greater than a similarity threshold; The target text fragment is identified as cross-modal associated data of the image data block.

7. The multimodal data processing method according to claim 5, characterized in that, The step of performing cross-modal matching on the modal data block to obtain cross-modal association data of the modal data block includes: Based on the position information of the tables in the table data block in the multimodal data, candidate titles are filtered out in the multimodal data; The candidate title is semantically matched with the table. If the verification is successful, the candidate title is determined as the cross-modal associated data of the table data block.

8. The multimodal data processing method according to claim 1, characterized in that, The method further includes: The data block is input into the retrieval enhancement generation system to perform retrieval enhancement generation and obtain the retrieval enhancement generation result.

9. The multimodal data processing method according to claim 8, characterized in that, The method further includes: The enhanced search results are input into a quality assessment model for quality assessment to obtain an assessment score. The evaluation score is input into the reward calculation model to calculate the reward and obtain a reward signal, which is then used to perform reinforcement learning on the granularity calculation model.

10. A multimodal data processing device, characterized in that, include: The block processing module is used to acquire multimodal data, and to perform block processing on the corresponding modal data in the multimodal data according to the block processing strategy corresponding to each modality to obtain modal data blocks for each modality. The step of segmenting the corresponding modal data in the multimodal data according to the segmentation strategy corresponding to each modality to obtain modal data blocks for each modality includes: segmenting the text in the multimodal data according to the semantic entropy dynamic segmentation strategy to obtain text data blocks; the step of segmenting the text in the multimodal data according to the semantic entropy dynamic segmentation strategy to obtain text data blocks includes: determining at least one sentence in the text according to regular expression matching and semantic coherence verification; dynamically generating multiple text fragments for each sentence through a sliding window; the number of the multiple text fragments is equal to the segmentation granularity; the size of the sliding window is successively an integer less than or equal to the segmentation granularity; calculating the semantic entropy of each sentence according to the multiple text fragments, and determining the segmentation endpoint according to the semantic entropy, so as to perform segmentation processing according to the segmentation endpoint to obtain the text data blocks; sentence The formula for calculating semantic entropy is: ;in, The sentence indicating the current assessment, Sentence Semantic entropy, used to quantify sentences As the strength of the semantic edge, n represents the context window size, i.e., the block granularity, which is used to control the scope of semantic judgment. Sentence P(EOS) represents the termination probability. The base-2 logarithm represents the standard form of information entropy, and -1 / n is a negative normalization factor used to convert the result into a positive entropy value; P(EOS) is obtained using a termination probability prediction model. The structural semantic parsing module is used to perform structural semantic parsing on the modal data block, obtain structural semantic data, and generate structural semantic tags for the modal data block based on the structural semantic data. The merging processing module is used to perform cross-modal merging processing on the modal data blocks based on the structural semantic tags to obtain the data blocks corresponding to the multimodal data.

11. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the multimodal data processing method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the multimodal data processing method as described in any one of claims 1 to 9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the multimodal data processing method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for realizing efficient semantic understanding of PDF (Portable Document Format) text by using deep learning

    CN119360398A

  • Multi-modal document retrieval enhancement generation method based on large model

    CN119988588A