Multi-modal information generation method, controller, medium and product
Through the method of clustering and graph network construction, data fusion problems in multimodal information processing are solved, and the accuracy and completeness of information generation are improved.
Patent Information
- Application Number
- CN202411789295.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
AI Technical Summary
Existing multimodal information processing methods are difficult to effectively integrate multimodal data such as vision and text, resulting in limited accuracy and completeness of generated information, and ignore the intrinsic connection between tag pairs.
By obtaining business documents, clustering text block content, building a graph network, and updating graph node features using link prediction algorithms to generate the final multimodal information.
It improves the generation effect of document information, enhances the overall correlation of multimodal data, and ensures the accuracy and completeness of information.
Smart Images

Figure CN119942575A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of multimodal information processing, and in particular to a method, controller, medium and product for generating multimodal information. Background Art
[0002] In the related technology, in the field of multimodal information processing, especially in the face of complex documents (such as documents containing mixed documents, images, and tables), traditional information extraction methods face many challenges. These documents are not only complex in layout, but also have various forms of information organization, including but not limited to text, images, and tables, and these elements also contain rich associations. As an important identifier of data, labels not only represent the category of data, but also serve as key constraints in the feature extraction process to improve the accuracy and pertinence of sign extraction.
[0003] However, existing graph network-based methods mostly focus on the combination of label feature extraction and association data, but often ignore the intrinsic connection between label pairs. At the same time, when these methods use OCR recognition results to construct graph networks, they often cannot accurately reflect the complete semantics and positional relationship of text blocks due to the limitations of OCR. The effective fusion of multimodal information is the key to achieving high-quality information generation, but existing methods often find it difficult to achieve overall association when processing multimodal data such as vision and text, resulting in limited accuracy and completeness of generated information. Summary of the invention
[0004] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a method, controller, medium and product for generating multimodal information, aiming to improve the generation effect of document information.
[0005] In a first aspect, an embodiment of the present application provides a method for generating multimodal information, the method comprising: Acquire a business document, and obtain the first text block content of all text blocks according to the business document; Clustering the text blocks according to the first text block contents to obtain second text block contents, wherein the second text block contents include second text block coordinates and second text block labels; Constructing a graph network according to the coordinates of the second text block and the label of the second text block; The graph nodes of the graph network are updated by a link prediction algorithm to obtain indicator information; The final multimodal information is generated according to the indicator information and the third text block content corresponding to the graph node.
[0006] According to some embodiments of the present application, obtaining the first text block content of all text blocks according to the business document includes: Performing layout analysis on the business document to obtain first text block coordinates and first text block labels of all text blocks; Performing OCR recognition on the area of each of the coordinates of the first text block to obtain the corresponding first text block text information; The text information of each of the first text blocks, the coordinates of the first text blocks and the first text block labels are integrated to obtain the corresponding first text block content.
[0007] According to some embodiments of the present application, the first text block label includes a hierarchical title, and clustering the text blocks according to the first text block content to obtain the second text block content includes: Obtaining the angular distance of the text block and the weight of the first text block label; Obtaining the fourth text block content under the hierarchical title through a regularization rule, and transforming the fourth text block content to obtain a word vector; Calculating the similarity of the business document according to the word vector and the weight of the first text block label to obtain a first similarity; The text blocks are clustered according to the first similarity and the angular distance of the text blocks to obtain second text block content.
[0008] According to some embodiments of the present application, constructing a graph network according to the second text block coordinates and the second text block label includes: Combining each of the second text block labels to obtain a text block label pair; Encoding and feature extracting the text block label pair to obtain a first label pair associated feature; Calculating cosine similarity of the text block label pairs to obtain a correlation matrix, and clustering the correlation matrix to obtain cluster labels; Integrating the clustering label and the first label pair association feature to obtain a second label pair association feature; The graph network is constructed according to the second text block coordinates and the second label pair associated features.
[0009] According to some embodiments of the present application, constructing the graph network according to the second text block coordinates and the second label pair association features includes: Abstracting the coordinates of the second text block to obtain a plurality of graph nodes; The graph network is constructed according to the associated features of each of the graph nodes and the second label pairs.
[0010] According to some embodiments of the present application, updating the graph nodes of the graph network by using a link prediction algorithm to obtain indicator information includes: Extracting features of adjacent nodes of the graph node by a link prediction algorithm to obtain adjacent node features; The characteristics of the graph nodes are updated according to the characteristics of the adjacent nodes to obtain indicator information.
[0011] According to some embodiments of the present application, generating final multimodal information according to the indicator information and the third text block content corresponding to the graph node includes: Mapping the indicator information and the third text block content corresponding to the graph node through an encoder; The mapped indicator information and the third text block content are sampled by Gaussian distribution to obtain a first latent vector; Performing specific interpolation and scaling on the first latent vector to obtain a second latent vector; The second latent vector is decoded by a decoder to generate final multimodal information.
[0012] In a second aspect, an embodiment of the present application provides a controller comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method for generating multimodal information of the first aspect when running the computer program.
[0013] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the method for generating multimodal information as described in the first aspect above.
[0014] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising a computer program or computer instructions, characterized in that the computer program or the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer program or the computer instructions from the computer-readable storage medium, and the processor executes the computer program or the computer instructions, so that the computer device executes the method for generating multimodal information as described in the first aspect above.
[0015] According to the technical solution of the embodiment of the present application, at least the following beneficial effects are achieved: the embodiment of the present application proposes a method, controller, medium and product for generating multimodal information, the method comprising: obtaining a business document, obtaining the first text block content of all text blocks according to the business document; clustering the text blocks according to the contents of the multiple first text blocks to obtain the second text block content, wherein the second text block content includes the second text block coordinates and the second text block label; constructing a graph network according to the second text block coordinates and the second text block label; updating the graph nodes of the graph network through a link prediction algorithm to obtain indicator information; finally, generating the final multimodal information according to the indicator information and the third text block content corresponding to the graph node. Therefore, since the present application can obtain the final multimodal information through the indicator information and the third text block content corresponding to the graph node, the generation effect of the document information can be improved.
[0016] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0018] Figure 1 is a flow chart of a method for generating multimodal information provided by an embodiment of the present application; Figure 2 is a flow chart of a method for generating multimodal information provided by another embodiment of the present application; Figure 3 is a flow chart of a method for generating multimodal information provided by another embodiment of the present application; Figure 4 is a flow chart of a method for generating multimodal information provided by another embodiment of the present application; Figure 5 is a flow chart of a method for generating multimodal information provided by another embodiment of the present application; Figure 6 is a flow chart of a method for generating multimodal information provided by another embodiment of the present application; Figure 7 is a flow chart of a method for generating multimodal information provided by another embodiment of the present application; Figure 8 It is an overall flow chart of a method for generating multimodal information provided by an embodiment of the present application; Fig. 9It is a schematic diagram of a controller for executing a method for generating multimodal information provided by an embodiment of the present application. DETAILED DESCRIPTION
[0019] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0020] In the description of the present application, it should be understood that descriptions involving orientation, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0021] In the description of this application, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed", etc. are understood to exclude the number itself, and "above", "below", "within", etc. are understood to include the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0022] In the description of this application, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in this application based on the specific content of the technical solution.
[0023] In some cases, in the field of multimodal information processing, traditional information extraction methods face many challenges, especially when facing complex documents (such as documents containing mixed documents, images, and tables). These documents are not only complex in layout, but also have various forms of information organization, including but not limited to text, images, and tables, and these elements also contain rich associations. As an important identifier of data, labels not only represent the category of data, but also serve as key constraints in the feature extraction process to improve the accuracy and pertinence of feature extraction.
[0024] However, existing graph network-based methods mostly focus on the combination of label feature extraction and association data, but often ignore the intrinsic connection between label pairs. At the same time, when these methods use OCR recognition results to construct graph networks, they often cannot accurately reflect the complete semantics and positional relationship of text blocks due to the limitations of OCR. The effective fusion of multimodal information is the key to achieving high-quality information generation, but existing methods often find it difficult to achieve overall association when processing multimodal data such as vision and text, resulting in limited accuracy and completeness of generated information.
[0025] In addition, although traditional OCR technology can recognize text content in documents, it is difficult to accurately capture the reading order and tag information of text blocks. Especially when the document layout is complex, OCR can often only provide basic line-level information and ignores the recognition of paragraph and text block tags, resulting in one-sided information extraction.
[0026] In addition, although named entity recognition and general information extraction methods are effective in tasks such as table extraction, they are unable to cope with structured hierarchical information and find it difficult to fully capture the deep relationships in the extracted information.
[0027] Based on the above situation, the present application proposes a method, controller, medium and product for generating multimodal information, aiming to improve the generation effect of document information.
[0028] The following further describes various embodiments of a method for generating multimodal information of the present application in conjunction with the accompanying drawings.
[0029] like Figure 1 As shown, Figure 1 It is a flowchart of a method for generating multimodal information provided by an embodiment of the present application; the method for generating multimodal information may include but is not limited to step S110, step S120, step S130, step S140 and step S150.
[0030] Step S110, acquiring a business document, and obtaining the first text block content of all text blocks according to the business document; Step S120: clustering the text blocks according to the contents of the plurality of first text blocks to obtain the contents of the second text block, wherein the contents of the second text block include the coordinates of the second text block and the labels of the second text block; Step S130, constructing a graph network according to the second text block coordinates and the second text block labels; Step S140: updating the graph nodes of the graph network by using a link prediction algorithm to obtain indicator information; Step S150: Generate final multimodal information according to the indicator information and the third text block content corresponding to the graph node.
[0031] In one embodiment, first, a business document is acquired to obtain first text block content of all text blocks according to the business document; secondly, the text blocks are clustered according to the obtained multiple first text block contents to obtain second text block content, wherein the second text block content includes second text block coordinates and second text block labels; thirdly, a graph network is constructed according to the second text block coordinates and the second text block labels; thirdly, the graph nodes of the graph network are updated through a link prediction algorithm to obtain indicator information; finally, the final multimodal information is generated according to the indicator information and the third text block content corresponding to the graph nodes.
[0032] It is worth noting that, since the present application can obtain the final multimodal information through the indicator information and the third text block content corresponding to the graph node, it can improve the generation effect of the document information.
[0033] In addition, if Figure 2 As shown, Figure 2 It is a flowchart of a method for generating multimodal information provided by another embodiment of the present application; regarding the first text block content of all text blocks obtained according to the business document in the above step S110, it may include but is not limited to step S210 and step S220.
[0034] Step S210: parse the layout of the business document to obtain the first text block coordinates and first text block labels of all text blocks; Step S220, performing OCR recognition on the area of each first text block coordinate to obtain corresponding first text block text information; Step S230: Integrate the text information of each first text block, the coordinates of the first text block and the first text block label to obtain the corresponding first text block content.
[0035] It can be understood that the first text block tag includes 1-5 level titles, text, tables, images, lists, etc.
[0036] It can be understood that the present application can introduce a hierarchical title directory tree algorithm to prune and correct the hierarchical title sub-nodes under all nodes, so as to remove redundant, repeated or unnecessary hierarchical titles, further simplify the document structure, and thus better perform OCR recognition on the area of the first text block coordinates to obtain the corresponding first text block text information, improve recognition accuracy and efficiency, and obtain the corresponding first text block content according to each first text block text information, first text block coordinates and first text block label.
[0037] In addition, if Figure 3 As shown, Figure 3It is a flowchart of a method for generating multimodal information provided by another embodiment of the present application; regarding the clustering of text blocks according to multiple first text block contents in the above-mentioned step S120 to obtain second text block content, it may include but is not limited to step S310, step S320, step S330 and step S340.
[0038] Step S310, obtaining the angular distance of the text block and the weight of the first text block label; Step S320: obtaining the fourth text block content under the hierarchical title through the regularization rule, and transforming the fourth text block content to obtain a word vector; Step S330: Calculate the similarity of the business document according to the word vector and the weight of the first text block label to obtain a first similarity; Step S340: cluster the text blocks according to the first similarity and the angular distance of the text blocks to obtain the second text block content.
[0039] It can be understood that the first text block tag includes a hierarchical title.
[0040] In one embodiment, first, the embodiment of the present application obtains the angular distance of the text block and the weight of the first text block label; then, the text block under the hierarchical title is located through the regularization rule to obtain the fourth text block content, and the fourth text block content is transformed to obtain a word vector; then, the similarity of the business document is calculated through the word vector and the weight of the first text block label to obtain the first similarity; finally, the text blocks are clustered through the calculated first similarity and the angular distance of the text blocks to obtain the second text block content.
[0041] It can be understood that the regularization rules can be used to locate the text blocks under the relevant level titles, thereby obtaining the content of the text blocks.
[0042] It can be understood that by calculating the similarity of business documents through the word vectors and the weights of the first text labels, the semantic similarity between different text blocks can be measured more accurately.
[0043] It can be understood that, in the clustering process, by combining the first similarity and the angular distance of the text blocks, the similarities and differences between the text blocks can be considered more comprehensively, thereby obtaining a more accurate clustering result.
[0044] In addition, if Figure 4 As shown, Figure 4It is a flowchart of a method for generating multimodal information provided by another embodiment of the present application; regarding the construction of a graph network according to the second text block coordinates and the second text block label in the above-mentioned step S130, it may include but is not limited to step S410, step S420, step S430, step S440 and step S450.
[0045] Step S410, combining the second text block labels to obtain a text block label pair; Step S420, encoding the text block label pair and extracting features to obtain first label pair associated features; Step S430, performing cosine similarity calculation on the text block label pairs to obtain a correlation matrix, and clustering the correlation matrix to obtain cluster labels; Step S440: Integrate the clustering label and the first label pair associated feature to obtain the second label pair associated feature; Step S450: construct a graph network based on the associated features of the second text block coordinates and the second label pair.
[0046] In one embodiment, first, the embodiment of the present application will combine the clustered second text block labels in pairs to form label pairs such as title-text, title-table, etc., to obtain text block label pairs; secondly, the embodiment of the present application will encode the text block label pairs and perform feature extraction to obtain first label pair associated features; from then on, the embodiment of the present application will perform cosine similarity calculation on the text block label pairs to obtain a correlation matrix, and perform clustering operation on the correlation matrix to obtain cluster labels, thereby adding the obtained cluster labels to the first label pair associated features to obtain second label pair associated features; finally, the embodiment of the present application will construct a graph network based on the second label pair associated features and the second text block coordinates.
[0047] It is understandable that the embodiment of the present application also sets a boundary for extracting component features, and uses a random walk method to capture the skeleton features of the graph network data, and performs unified data encoding on the associated features in the graph network data.
[0048] It can be understood that by combining the clustered second text block labels in pairs, different types of information in the document can be effectively integrated.
[0049] It can be understood that encoding and feature extraction of text block label pairs can obtain the first label pair associated features, thereby being able to more accurately describe the relationship between label pairs and provide strong data support for subsequent analysis.
[0050] It can be understood that by calculating the cosine similarity of text block label pairs, a correlation matrix can be obtained, which can intuitively show the degree of correlation between different label pairs; in addition, by clustering the correlation matrix, cluster labels are obtained, which helps to classify label pairs with high similarity into one category, thereby more clearly revealing the information structure in the document.
[0051] It can be understood that adding clustering labels to the first label pair association features to form the second label pair association features can further improve the accuracy of information extraction.
[0052] In addition, if Figure 5 As shown, Figure 5 It is a flowchart of a method for generating multimodal information provided by another embodiment of the present application; regarding the above-mentioned step S450, it may include but is not limited to step S510 and step S520.
[0053] Step S510, abstracting the coordinates of the second text block to obtain a plurality of graph nodes; Step S520: construct a graph network based on the associated features of each graph node and the second label.
[0054] It can be understood that the embodiment of the present application abstracts the coordinates of the clustered second text block to represent the graph nodes of the graph network, and loads the second label to characterize the associated features, thereby obtaining a constructed graph network.
[0055] In addition, if Figure 6 As shown, Figure 6 It is a flowchart of a method for generating multimodal information provided by another embodiment of the present application; regarding the above-mentioned step S140, it may include but is not limited to step S610 and step S620.
[0056] Step S610: extracting features of adjacent nodes of the graph node through a link prediction algorithm to obtain adjacent node features; Step S620: Update the characteristics of the graph nodes according to the characteristics of the adjacent nodes to obtain indicator information.
[0057] In one embodiment, the embodiment of the present application can extract the features of the adjacent nodes of the graph nodes through the link prediction algorithm, thereby obtaining the adjacent node features, and then update the features of the graph nodes to obtain indicator information.
[0058] It can be understood that the indicator information refers to the multimodal information that needs to be extracted. For example, it can be charity or a donation of 10 yuan.
[0059] In addition, if Figure 7 As shown, Figure 7It is a flowchart of a method for generating multimodal information provided by another embodiment of the present application; regarding the above-mentioned step S150, it may include but is not limited to step S710, step S720, step S730 and step S740.
[0060] Step S710: Map the indicator information and the third text block content corresponding to the graph node through an encoder; Step S720: sampling the mapped indicator information and the third text block content through Gaussian distribution to obtain a first latent vector; Step S730: performing specific interpolation and scaling on the first latent vector to obtain a second latent vector; Step S740: decode the second latent vector through a decoder to generate final multimodal information.
[0061] In one embodiment, first, the embodiment of the present application maps the indicator information and the third text block content corresponding to the graph node to the Gaussian distribution of the latent space through an encoder; then, the embodiment of the present application samples the mapped indicator information and the third text block content through the Gaussian distribution to obtain a first latent vector, thereby performing specific interpolation and scaling on the first latent vector to obtain a second latent vector; finally, the second latent vector is input into the decoder for decoding, thereby obtaining the final multimodal information.
[0062] It can be understood that the embodiment of the present application maps the indicator information and the text corresponding to the graph node to the Gaussian distribution of the latent space through the encoder, and then samples to obtain the latent vector, interpolates and scales it, and finally decodes it by the decoder to generate complete and accurate multimodal information, thereby improving the generation effect of document information.
[0063] Based on the methods for generating multimodal information of the above-mentioned various embodiments, overall embodiments of the methods for generating multimodal information of the present application are respectively proposed below.
[0064] like Figure 8 As shown, Figure 8 It is an overall flow chart of a method for generating multimodal information provided by an embodiment of the present application.
[0065] 1. Obtain information related to business documents (1) Perform layout analysis on the company's business documents to obtain the first text block coordinates and first text block labels of all text blocks. The first text block labels include 1-5 level titles, text, tables, images, lists, etc.; (2) Introduce a hierarchical title directory tree algorithm to prune and correct the hierarchical title subnodes under all nodes; (3) performing OCR recognition on the areas of all the first text block coordinates to obtain the corresponding first text block text information; (4) Integrate the text information of each first text block, the coordinates of the first text block, and the first text block label to obtain the content of the first text block; 2. Cluster the first text block content (1) Locate all fourth text blocks under the relevant level titles through regularization rules and convert word vectors; (2) Calculate the similarity of business documents based on the word vector and the weight of the first text block label; (3) By calculating the similarity of the text block content and the angular distance of the text blocks, the text blocks of the page are clustered to obtain the content of the second text block; 3. Extract label pair features (1) After clustering, the second text block labels are combined in pairs to form label pairs such as title-text and title-table to obtain text block label pairs. The cosine similarity of the text block label pairs is calculated to construct a correlation matrix. (2) Encode the text block label pair information, extract the text block label pair features, and obtain the first label pair associated features; (3) Perform a clustering operation on the correlation matrix to obtain cluster labels, and add the added cluster labels to the associated features of the first label pair to obtain the associated features of the second label pair.
[0066] 4. Build a graph network (1) Abstract the coordinates of the clustered second text block to represent the graph nodes of the graph network, load the second label to encode the associated features and characterize the multimodal associated features; (2) Construct a graph network dataset G = (G1, ..., GN), where a subset of G represents a subgraph constructed from the location information nodes of a certain text block, and G1-GN represents the association features between the nodes of two subgraphs; (3) Set the component feature extraction boundary, use the random walk method to capture the skeleton features of the graph network data, and unify the data encoding of the associated features in the graph network data.
[0067] 5. Use link prediction algorithm to output indicator information (1) Specifically, the algorithm can make predictions by extracting the surrounding subgraph of each target link and encoding the subgraph into an adjacency matrix.
[0068] (2) For any node ∈G: Get graph node { }All the features of adjacent nodes j{ }; Update the node's features { }<-hash{} Repeat the above steps until convergence to obtain the final indicator information (3) This encoding is based on a fast hashing algorithm that labels vertices according to their structural roles in the subgraph while preserving the inherent direction of the subgraph. A neural network is then trained on these adjacency matrices to learn a predictive model. This approach does not assume a specific link formation mechanism (such as common neighbors), but instead learns this mechanism from the graph itself.
[0069] 6. Use the generation algorithm to integrate the indicator information of all output graph nodes to generate the final multimodal information (1) The encoder maps the sample of the third text block corresponding to the indicator information and the graph node to the Gaussian distribution of the latent space; (2) Sampling the first latent vector by Gaussian distribution, performing specific interpolation and scaling operations on the first latent vector to obtain a second latent vector; (3) The second latent vector after sampling and dimension operation is passed through the decoder to generate data with specific attributes and obtain the final multimodal information.
[0070] Based on the methods for generating multimodal information of the above-mentioned embodiments, various embodiments of the controller, computer-readable storage medium and computer program product of the present application are respectively proposed below.
[0071] like Fig. 9 As shown, Fig. 9 700 is a schematic diagram of a controller for executing a method for generating multimodal information provided by an embodiment of the present application. The controller 700 implemented in the present application includes: a processor 710, a memory 720, and a computer program stored in the memory 720 and executable on the processor 710, wherein: Fig. 9 In the figure, a processor 710 and a memory 720 are taken as an example.
[0072] The processor 710 and the memory 720 may be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.
[0073] The memory 720, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory 720 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 720 may optionally include a memory 720 remotely arranged relative to the processor 710, and these remote memories 720 may be connected to the controller 700 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0074] Those skilled in the art will understand that Fig. 9 The device structure shown in the figure does not constitute a limitation on the controller 700, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0075] exist Fig. 9 In the controller 700 shown, the processor 710 can be used to call the control program stored in the memory 720, so as to implement the above-mentioned method for generating multimodal information. Specifically, the non-transient software program and instructions required to implement the method for generating multimodal information of the above-mentioned embodiment are stored in the memory 720, and when executed by the processor 710, the method for generating multimodal information of the above-mentioned embodiment is executed.
[0076] It is worth noting that since the controller 700 of the embodiment of the present application can execute the method for generating multimodal information of any of the above-mentioned embodiments, the specific implementation manner and technical effects of the controller 700 of the embodiment of the present application can refer to the specific implementation manner and technical effects of the method for generating multimodal information of any of the above-mentioned embodiments.
[0077] In addition, an embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to execute the above-described method for generating multimodal information. Figures 1 to 8 The method steps in .
[0078] It is worth noting that since the computer-readable storage medium of the embodiment of the present application can execute the method for generating multimodal information of any of the above-mentioned embodiments, the specific implementation methods and technical effects of the computer-readable storage medium of the embodiment of the present application can refer to the specific implementation methods and technical effects of the method for generating multimodal information of any of the above-mentioned embodiments.
[0079] In addition, an embodiment of the present application further provides a computer program product, including a computer program or computer instructions, the computer program or computer instructions are stored in a computer-readable storage medium, the processor of the computer device reads the computer program or computer instructions from the computer-readable storage medium, and the processor executes the computer program or computer instructions, so that the computer device executes the above-described method for generating multimodal information. Figures 1 to 8 The method steps in .
[0080] It is worth noting that since the computer program product of the embodiments of the present application can execute the method for generating multimodal information of any of the above-mentioned embodiments, the specific implementation methods and technical effects of the computer program product of the embodiments of the present application can refer to the specific implementation methods and technical effects of the method for generating multimodal information of any of the above-mentioned embodiments.
[0081] It will be appreciated by those skilled in the art that all or some of the steps and systems in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically include computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0082] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0083] In several embodiments provided in the present application, it should be understood that the disclosed systems, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of apparatuses or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0084] It should also be understood that the various implementations provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0085] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the above implementation mode. Technical personnel familiar with the field can also make various equivalent modifications or substitutions under the shared conditions without violating the spirit of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A method for generating multimodal information, characterized in that: The method comprises: Acquire a business document, and obtain the first text block content of all text blocks according to the business document; Clustering the text blocks according to the first text block contents to obtain second text block contents, wherein the second text block contents include second text block coordinates and second text block labels; Constructing a graph network according to the coordinates of the second text block and the label of the second text block; The graph nodes of the graph network are updated by a link prediction algorithm to obtain indicator information; The final multimodal information is generated according to the indicator information and the third text block content corresponding to the graph node.
2. A method for generating multimodal information according to claim 1, characterized in that: The obtaining the first text block content of all text blocks according to the business document includes: Performing layout analysis on the business document to obtain first text block coordinates and first text block labels of all text blocks; Performing OCR recognition on the area of each of the coordinates of the first text block to obtain the corresponding first text block text information; The text information of each of the first text blocks, the coordinates of the first text blocks and the first text block labels are integrated to obtain the corresponding first text block content.
3. A method for generating multimodal information according to claim 2, characterized in that: The first text block label includes a hierarchical title, and clustering the text blocks according to the first text block content to obtain the second text block content includes: Obtaining the angular distance of the text block and the weight of the first text block label; Obtaining the fourth text block content under the hierarchical title through a regularization rule, and transforming the fourth text block content to obtain a word vector; Calculating the similarity of the business document according to the word vector and the weight of the first text block label to obtain a first similarity; The text blocks are clustered according to the first similarity and the angular distance of the text blocks to obtain second text block content.
4. The method for generating multimodal information according to claim 1, characterized in that: The constructing a graph network according to the second text block coordinates and the second text block label comprises: Combining each of the second text block labels to obtain a text block label pair; Encoding and feature extracting the text block label pair to obtain a first label pair associated feature; Calculating cosine similarity of the text block label pairs to obtain a correlation matrix, and clustering the correlation matrix to obtain cluster labels; Integrating the clustering label and the first label pair association feature to obtain a second label pair association feature; The graph network is constructed according to the second text block coordinates and the second label pair associated features.
5. A method for generating multimodal information according to claim 4, characterized in that: The step of constructing the graph network according to the second text block coordinates and the second label pair association features includes: Abstracting the coordinates of the second text block to obtain a plurality of graph nodes; The graph network is constructed according to the associated features of each of the graph nodes and the second label pairs.
6. A method for generating multimodal information according to claim 1, characterized in that: The updating of the graph nodes of the graph network by the link prediction algorithm to obtain indicator information includes: Extracting features of adjacent nodes of the graph node by a link prediction algorithm to obtain adjacent node features; The characteristics of the graph nodes are updated according to the characteristics of the adjacent nodes to obtain indicator information.
7. A method for generating multimodal information according to claim 1, characterized in that: Generating final multimodal information according to the indicator information and the third text block content corresponding to the graph node includes: Mapping the indicator information and the third text block content corresponding to the graph node through an encoder; The mapped indicator information and the third text block content are sampled by Gaussian distribution to obtain a first latent vector; Performing specific interpolation and scaling on the first latent vector to obtain a second latent vector; The second latent vector is decoded by a decoder to generate final multimodal information.
8. A controller, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes a method for generating multimodal information as described in any one of claims 1 to 8 when executing the computer program.
9. A computer-readable storage medium, characterized in that: Computer executable instructions are stored, and the computer executable instructions are used to execute a method for generating multimodal information as described in any one of claims 1 to 8.
10. A computer program product comprising a computer program or computer instructions, characterized in that The computer program or the computer instruction is stored in a computer-readable storage medium, and the processor of a computer device reads the computer program or the computer instruction from the computer-readable storage medium. The processor executes the computer program or the computer instruction, so that the computer device executes a method for generating multimodal information as described in any one of claims 1 to 8.