Intelligent OCR dynamic adaptive method and system based on multi-modal fusion
Through the intelligent OCR dynamic adaptive method of multimodal fusion, the problems of low recognition accuracy and poor adaptability in multimodal information OCR processing are solved, and more efficient multimodal information recognition and processing are achieved.
Patent Information
- Application Number
- CN202510501064.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing multimodal information OCR processing has the problem of low recognition accuracy and poor adaptability. It cannot fully utilize the complementarity between multimodal information, and it is difficult to cope with complex and changeable layout structures and information content.
Using an intelligent OCR dynamic adaptive method based on multimodal fusion, by uploading multimodal information to the OCR recognition platform, selecting the first modal information for fuzzy scanning, performing layout analysis to determine the layout structure, introducing hierarchical fusion conditions according to the content complexity, setting a dynamic fusion paradigm, initializing the OCR engine array, performing multi-step focus fusion planning, triggering directional focus and entity alignment focus, dynamically configure the OCR engine and fusion paradigm, and performing scanning and identification management under multimodal fusion.
The recognition accuracy and robustness of multimodal information OCR is improved, and it can better adapt to multimodal information of different complexities and structures, and improve the efficiency and accuracy of information processing.
Smart Images

Figure CN120472481A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field related to multimodal information fusion processing, and in particular to an intelligent OCR dynamic adaptive method and system based on multimodal fusion. Background Art
[0002] In today's digital age, efficient and accurate processing of multimodal information to achieve intelligent optical character recognition (OCR) is crucial for the development of many industries (such as document digitization, automated office, and intelligent information retrieval). It can greatly improve the efficiency and accuracy of information processing and promote the intelligent upgrade of related fields. Currently, the main methods for solving the problem of intelligent OCR of multimodal information are to use traditional single-modal OCR technology to process different modal information separately, or simply fuse multiple modal information directly for recognition, while combining some fixed fusion rules and fixed OCR engine configuration. However, traditional single-modal OCR technology cannot fully utilize the complementarity between multimodal information and has difficulty coping with complex and changing layout structures and information content. The simple fusion of multimodal information lacks in-depth analysis of information content and layout structure, and the fixed fusion rules and OCR engine configuration cannot be dynamically adjusted according to actual conditions. As a result, when faced with multimodal information of different complexity and structure, the recognition accuracy is low and the robustness is poor, which cannot meet the demand for high-precision and highly adaptable OCR technology in practical applications.
[0003] In the current related technologies, multimodal information OCR processing has technical problems such as low recognition accuracy and poor adaptability. Summary of the Invention
[0004] This application provides an intelligent OCR dynamic adaptive method and system based on multimodal fusion, which uploads multimodal information to the OCR recognition platform, selects one modal information for fuzzy scanning, performs layout analysis to determine the layout structure, introduces layered progressive fusion conditions based on content complexity, sets a dynamic fusion paradigm (with complementary and overlapping fusion as the goal), initializes the OCR engine array, performs multi-step focus fusion planning based on the layout structure, triggers directional focus and entity alignment focus, dynamically configures the OCR engine and fusion paradigm, focuses on content characteristics based on the layout structure, and uses modal empowerment and progressive stratification as factors to dynamically allocate engines and adaptively adjust the fusion paradigm of the OCR engine array, thereby achieving the technical effect of improving the recognition accuracy and robustness of multimodal information OCR.
[0005] The present application provides an intelligent OCR dynamic adaptive method based on multimodal fusion, including: uploading multimodal information to an OCR recognition platform, selecting first modal information and performing fuzzy scanning, performing layout analysis to determine the layout structure, wherein the first modal information is any modality in the multimodal information; introducing hierarchical progressive fusion conditions according to content complexity, setting a dynamic fusion paradigm, and initializing the OCR engine array deployed on the OCR recognition platform, wherein complementary fusion and overlapping fusion are used as basic fusion targets; performing multi-step focused fusion planning according to the layout structure, triggering directional focusing based on the first modal information, aligning focusing with entities for additional modal information, dynamically configuring OCR engines and fusion paradigms, and performing scanning and recognition management under multimodal fusion; wherein, based on the focused content characteristics based on the layout structure, the OCR engine array is dynamically allocated engines and the fusion paradigm is adaptively adjusted, with modal empowerment and progressive stratification as adaptive elements.
[0006] In a possible implementation, directional focusing based on the first modal information is triggered, the OCR engine is dynamically configured, and the following processing is performed: a scanning path is planned according to the layout structure of the first modal information; first focusing information is determined according to the scanning path; and the first OCR engine in the OCR engine array is temporarily connected in the display format of the first focusing information.
[0007] In a possible implementation, for entity alignment and focusing of additional modal information, the OCR engine is dynamically configured to perform the following processing: performing entity alignment on the additional modal information based on the first focusing information to determine the additional focusing information, wherein the additional modal information is the remaining modalities in the multimodal information except the first modal information; traversing the display formats of the additional focusing information, and temporarily connecting the additional OCR engines in the OCR engine array, wherein the additional OCR engines correspond to the additional focusing information.
[0008] In a possible implementation, the fusion paradigm is dynamically configured, and multimodal fusion processing is performed, and the following processing is performed: the first focused information and the additional focused information are identified, and information weighting is performed based on the quality of the modal information to determine the information weight; based on the content complexity of the focused information, a fusion layered deployment is performed to determine the fusion level; the focused fusion paradigm is determined based on the information weight and the fusion level; and based on the focused fusion paradigm, the first focused information and the additional focused information are multimodally fused.
[0009] In a possible implementation, according to the focus fusion paradigm, the first focus information and the additional focus information are multimodally fused, and the following processing is performed: the first fusion layer is determined from the bottom up, and the complementary information and the superimposed information are divided, wherein the first fusion layer is the minimum fusion unit; the superimposed information is weightedly fused according to the information weight, and the complementary information is subjected to information prior and enhancement processing to determine a layer of fusion information; based on the second fusion layer, the layer of fusion information is fused, and the hierarchical iterative fusion is performed to determine the fusion result of the first focus information.
[0010] In a possible implementation, after determining the fusion result of the first focus information, the following processing is performed: according to the scanning path, the second focus information based on the first modal information is determined; mutual information analysis is performed on the fusion result of the first focus information and the second focus information to determine the second fusion focus, wherein the second fusion focus is determined based on context association; with the second fusion focus as a constraint, additional modal entity alignment and engine paradigm adaptive fusion based on the second focus information is performed.
[0011] In a possible implementation, the layout structure is determined and the following processing is performed: a pre-check port is deployed and an association is established between the pre-check port and the OCR engine array; based on the pre-check port, a full-area fuzzy scan is performed on the first modal information to determine the content layout characteristics; based on the content layout characteristics, an information framework is determined and the content ranking relationship is identified as the layout structure.
[0012] In a possible implementation, after performing scanning and recognition management under multimodal fusion, the following processing is performed: receiving a display requirement; and deploying the fusion results under N focuses according to the display requirement to determine the multimodal fusion result.
[0013] The present application also provides an intelligent OCR dynamic adaptive system based on multimodal fusion, including: a layout structure determination module, which is used to upload multimodal information to the OCR recognition platform, select the first modal information and perform fuzzy scanning, and perform layout analysis to determine the layout structure, wherein the first modal information is any modality in the multimodal information; an initialization module, which is used to introduce layered progressive fusion conditions according to the complexity of the content, set a dynamic fusion paradigm, and initialize the OCR engine array deployed on the OCR recognition platform, wherein complementary fusion and overlapping fusion are used as the basic fusion goals; a multimodal fusion module, which is used to perform multi-step focused fusion planning according to the layout structure, trigger directional focusing based on the first modal information, align focusing with the entity for additional modal information, dynamically configure the OCR engine and fusion paradigm, and perform scanning and recognition management under multimodal fusion, wherein the OCR engine array is dynamically allocated and the fusion paradigm is adaptively adjusted based on the focused content characteristics based on the layout structure, and modal empowerment and progressive stratification are used as adaptive elements.
[0014] The proposed intelligent OCR dynamic adaptive method and system based on multimodal fusion first uploads multimodal information to the OCR recognition platform, selects the first modal information and performs fuzzy scanning, performs layout analysis to determine the layout structure, wherein the first modal information is any modality in the multimodal information, then introduces hierarchical progressive fusion conditions based on content complexity, sets a dynamic fusion paradigm, and initializes the OCR engine array deployed on the OCR recognition platform, wherein complementary fusion and overlapping fusion are the basic fusion targets. Finally, based on the layout structure, multi-step focused fusion planning is performed, triggering directional focusing based on the first modal information, and entity alignment focusing for the additional modal information. The OCR engine and fusion paradigm are dynamically configured to perform scanning and recognition management under multimodal fusion. The OCR engine array is dynamically allocated and the fusion paradigm is adaptively adjusted based on the focused content characteristics of the layout structure, with modal weighting and progressive stratification as adaptive elements. The technical effect of improving the recognition accuracy and robustness of multimodal information OCR is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0016] Figure 1A flowchart of the intelligent OCR dynamic adaptive method based on multimodal fusion provided in an embodiment of the present application.
[0017] Figure 2 A schematic diagram of the structure of the intelligent OCR dynamic adaptive system based on multimodal fusion provided in an embodiment of the present application.
[0018] Explanation of the accompanying drawings: layout structure determination module 10, initialization module 20, multimodal fusion module 30. DETAILED DESCRIPTION
[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0020] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0021] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict, and the terms “first\second” involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. The terms “including” and “having” and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.
[0022] The embodiment of the present application provides an intelligent OCR dynamic adaptive method based on multimodal fusion, such as Figure 1 As shown, the method includes:
[0023] Step S100 , uploading multimodal information to an OCR recognition platform, selecting first modal information and performing fuzzy scanning, performing layout analysis to determine the layout structure, wherein the first modal information is any modality in the multimodal information.
[0024] Specifically, data containing multiple modalities, such as images, text, audio, etc., are uploaded to the OCR recognition platform through a network interface (such as the HTTP protocol). The OCR recognition platform provides an API interface and supports uploading files in multiple formats. The first modal information refers to the modal information used for fuzzy scanning and layout analysis in multimodal data. It is an image or imaged data with a clear two-dimensional structure and layout features, and can be preliminarily analyzed and parsed through image processing technology. The first modal information (such as an image) is preprocessed, and the image is blurred using a Gaussian blur filter to reduce noise and details and extract low-level features. Use a deep learning model (such as YOLO or Faster R-CNN) to perform layout analysis on the image to identify the position and layout of elements such as text areas, tables, and pictures.
[0025] For example, a user uploads an image containing handwritten text and a diagram through the OCR recognition platform's web interface, along with a text description of the image's contents. After the platform receives the file, it stores it in a designated directory on the server. The system applies a Gaussian blur filter to the uploaded image, smoothing out any small scratches or noise. The system then uses the YOLO model to analyze the blurred image, identifying the title area in the upper left corner, the table area in the middle, and the annotation area in the lower right corner. It then labels the bounding boxes for each area.
[0026] In one possible implementation, the layout structure is determined, and step S100 further includes step S110, deploying a pre-flight port, and establishing an association between the pre-flight port and the OCR engine array. Specifically, the pre-flight port is deployed at the front end of the OCR recognition platform. The pre-flight port is a network interface for receiving and preliminarily processing uploaded multimodal information. The pre-flight port can be a network interface that supports data input in multiple formats. Through the configuration management tool, the pre-flight port is associated with the OCR engine array to ensure that the pre-flight port can pass the processed data directly to the OCR engine array. For example, a user uploads a multimodal file containing images and text through the Web interface, and the pre-flight port receives the file and performs preliminary parsing to extract basic information of the file, such as file type, size, etc. The pre-flight port establishes a connection with the OCR engine array through the configuration file to ensure that the data can be transmitted smoothly. For example, the pre-flight port sends image data to the CNN-based OCR engine and sends text data to the Transformer-based OCR engine.
[0027] In step S120, a global fuzzy scan is performed on the first modal information based on the pre-inspection port to determine the content layout characteristics. Specifically, image processing techniques, such as a Gaussian blur filter, are used to perform a global fuzzy scan on the first modal information (image) to reduce noise and extract low-level texture and shape features. The fuzzy scanned image is analyzed using a deep learning model (such as YOLO or Faster R-CNN) to identify text areas, chart areas, table areas, etc. in the image and determine the content layout characteristics.
[0028] Step S130, determine the information framework according to the content layout characteristics, and identify the content ranking relationship as the layout structure. Specifically, according to the content layout characteristics, construct an information framework, and clarify the hierarchical relationship and logical structure of each area. Analyze the relative position and logical relationship between each area through a deep learning model (such as Transformer), identify the ranking relationship of the content, and form a complete layout structure. For example, the system uses the Transformer model to analyze the relative position and logical relationship between each area, and identifies the ranking relationship of the content as the title is above the text and the table is to the right of the text. This implementation method deploys a pre-inspection port, performs a full-domain fuzzy scan of the first modal information and determines the content layout characteristics, as well as constructs an information framework and identifies the content ranking relationship, which can ensure that the OCR recognition platform can accurately identify and process different areas in the image, so that the system can better understand and process complex multimodal information.
[0029] In step S200 , a hierarchical progressive fusion condition is introduced according to the complexity of the content, a dynamic fusion paradigm is set, and the OCR engine array deployed on the OCR recognition platform is initialized, wherein complementary fusion and overlapping fusion are used as the basic fusion targets.
[0030] Specifically, the content complexity of multimodal information is evaluated by calculating indicators such as the texture complexity of the image (such as edge detection algorithms) and the semantic complexity of the text (such as text length and vocabulary diversity). The fusion strategy is dynamically adjusted according to the content complexity. For complex content, a multi-layer neural network is used to gradually fuse information from different modalities; for simple content, shallow fusion is directly performed. The OCR engine array is initialized, different OCR engines are configured (such as CNN-based OCR engines and Transformer-based OCR engines), and tasks are dynamically assigned according to the fusion strategy.
[0031] For example, the system calculated the texture complexity of the image and found that the image contained complex charts and handwritten text, indicating that the content complexity was high. At the same time, the uploaded text description was analyzed and found to be long and with high vocabulary diversity, further confirming the content complexity. Based on the evaluation results, the system decided to adopt a hierarchical progressive fusion strategy. For the image modality, low-level texture features are first extracted, and then high-level semantic features are gradually integrated; for the text modality, word embedding features are first extracted, and then contextual information is gradually integrated. The system initializes the OCR engine array, configures a CNN-based OCR engine to process the image modality, configures a Transformer-based OCR engine to process the text modality, and assigns a higher weight to the image modality and a lower weight to the text modality based on complexity.
[0032] Step S300: Perform multi-step focus fusion planning based on the layout structure, trigger directional focus based on the first modal information, align focus with the entity for the additional modal information, dynamically configure the OCR engine and fusion paradigm, and perform scanning recognition management under multimodal fusion. Specifically, the OCR engine array is dynamically allocated and the fusion paradigm is adaptively adjusted based on the focus content characteristics of the layout structure, with modal empowerment and progressive stratification as adaptive elements.
[0033] Specifically, the layout is divided into multiple areas according to the layout structure, and different fusion strategies are formulated for each area. For example, for the text area, the focus is on fusing image and text modalities; for the table area, the image, text and structured data modalities are fused. Use an attention mechanism (such as the self-attention mechanism in Transformer) to focus on the first modality information, and perform entity alignment on the additional modality information. For example, optimize multimodal alignment by contrasting learning objective functions. Dynamically adjust the parameters and fusion paradigm of the OCR engine according to the characteristics of the focused content. For example, dynamically adjust the resolution and recognition algorithm of the OCR engine according to the complexity of the text area.
[0034] For example, the system divides the image into text areas and chart areas based on the layout structure. For the text area, the focus is on fusing the image and text modalities; for the chart area, the focus is on fusing the image, text, and structured data modalities. CNN is used to extract image features in the text area, and Transformer is used to extract text features, which are then fused; for the chart area, image features, text features, and table structure features are combined for fusion. The system uses a self-attention mechanism to focus on the image modality and extract key features from the image. At the same time, entity alignment is performed on the text modality to ensure that the text in the image is consistent with the entities in the text description. For example, comparative learning is used to optimize the alignment of the "company name" in the image with the "company name" in the text description. The system dynamically adjusts the resolution of the OCR engine according to the complexity of the text area, using high resolution to process handwritten text areas and low resolution to process printed text areas. At the same time, the fusion paradigm is dynamically adjusted according to the fusion effect, such as increasing or decreasing the fusion level and optimizing the fusion weight. The embodiment of the present application adopts the method of uploading multimodal information to the OCR recognition platform, selecting a modal information for fuzzy scanning, performing layout analysis to determine the layout structure, introducing layered progressive fusion conditions according to the complexity of the content, setting a dynamic fusion paradigm (with complementary and overlapping fusion as the goal), initializing the OCR engine array, and performing multi-step focus fusion planning according to the layout structure, triggering directional focus and entity alignment focus, dynamically configuring the OCR engine and fusion paradigm, focusing on content characteristics based on the layout structure, taking modal empowerment and progressive stratification as factors, and performing dynamic engine allocation and adaptive adjustment of the fusion paradigm on the OCR engine array. These technical means have achieved the technical effect of improving the recognition accuracy and robustness of multimodal information OCR.
[0035] In one possible implementation, directional focusing based on the first modal information is triggered, and the OCR engine is dynamically configured. Step S300 further includes step S310, planning a scanning path according to the layout structure of the first modal information. Specifically, a deep learning model (such as YOLO or Faster R-CNN) is used to analyze the layout structure of the first modal information to identify text areas, chart areas, table areas, etc. Based on the analysis results of the layout structure, a path planning algorithm is used to determine the scanning path, such as the A* algorithm or the Dijkstra algorithm. The goal of path planning is to ensure that the scanning order is logical and improve recognition efficiency. According to the complexity of the content and the importance of the area, the scanning path is dynamically adjusted to give priority to scanning key areas. For example, the system recognizes that the image contains a title area, a text area, and a chart area. The title area is at the top of the image, the text area is in the middle, and the chart area is at the bottom. The system uses the A* algorithm to plan the scanning path, giving priority to scanning the title area, then the text area, and finally the chart area.
[0036] Step S320, determines the first focus information based on the scanning path. Specifically, according to the scanning path, an attention mechanism (such as the self-attention mechanism in Transformer) is used to extract the key information of each area. The extracted key information is fused with the global features to generate the first focus information (the key information extracted according to the scanning path is used for subsequent feature fusion and OCR recognition). According to the content complexity of the area, the extraction depth and range of the focus information are dynamically adjusted. For example, the system uses a self-attention mechanism for the title area to extract key information of the title, such as keywords and phrases. The extracted title key information is fused with the global features of the image to generate the first focus information.
[0037] Step S330, temporarily connect the first OCR engine in the OCR engine array in the display format of the first focused information. Specifically, according to the display format of the first focused information (such as text, geometry, etc.), select a suitable OCR engine. Through the configuration management tool, the first focused information is temporarily connected to the first OCR engine in the OCR engine array (in the OCR engine array, the OCR engine specifically used to process the first modal information) to ensure that the information can be processed correctly. According to the content complexity of the first focused information, the parameters of the OCR engine are dynamically adjusted to improve the recognition efficiency. For example, the system recognizes that the first focused information is in text format, selects the Transformer-based OCR engine for processing, and the system temporarily connects the first focused information to the first OCR engine to ensure that the information can be processed correctly.
[0038] In one possible implementation, the OCR engine is dynamically configured for entity alignment and focusing on the additional modal information. Step S300 further includes step S340, in which entity alignment based on the first focused information is performed on the additional modal information to determine the additional focused information, wherein the additional modal information is the remaining modalities in the multimodal information except the first modal information. Specifically, multimodal fusion technology is used to align the additional modal information (such as text, audio, etc.) with the first modal information (such as an image). The alignment process is based on the first focused information to ensure that entities between different modalities can accurately correspond. Features are extracted from the additional modal information, such as word embeddings of text or acoustic features of audio, and then matched with the features in the first focused information. Based on the alignment results, the processing method of the additional modal information is dynamically adjusted to improve the accuracy and efficiency of the alignment. For example, assume that the first modal information is an image and the additional modal information is a text description. The system uses a pre-trained image-text matching model to calculate the similarity score between the object in the image and the entity in the text description. Word embedding features are extracted from the text, and visual features are extracted from the image, and then the alignment is optimized through comparative learning. If it is found that some text entities are not well aligned with image entities, the system will adjust the alignment strategy, such as adding contextual information or adjusting feature weights.
[0039] Step S350, traverse the display format of the additional focus information, and temporarily connect the additional OCR engine in the OCR engine array, wherein the additional OCR engine corresponds to the additional focus information. Specifically, according to the display format of the additional focus information (such as text format, image format, etc.), select a suitable OCR engine. Through the configuration management tool, the additional focus information is temporarily connected to the additional OCR engine in the OCR engine array to ensure that the information can be processed correctly. According to the content complexity of the additional focus information, the parameters of the OCR engine are dynamically adjusted to improve the recognition efficiency. For example, assuming that the additional focus information is in text format, the system selects a Transformer-based OCR engine for processing. The system temporarily connects the additional focus information to the additional OCR engine to ensure that the information can be processed correctly. If the additional focus information contains a complex table structure, the system will adjust the parameters of the OCR engine to improve the accuracy of table recognition.
[0040] In one possible implementation, the fusion paradigm is dynamically configured to perform multimodal fusion processing, and step S300 further includes step S360, identifying the first focused information and the additional focused information, and weighting the information according to the quality of the modal information to determine the information weight. Specifically, the quality of the first focused information and the additional focused information is evaluated by a pre-trained deep learning model (such as Transformer or CNN). The quality assessment includes the integrity, accuracy, noise level, etc. of the information. Weights are dynamically assigned according to the quality of the modal information. High-quality modal information is assigned a higher weight, and low-quality modal information is assigned a lower weight. The final information weight is determined by an adaptive algorithm (such as an adaptive normalization layer).
[0041] For example, the system evaluates the quality of the first focused information (the text area in the image) and the additional focused information (the text description). The text area in the image is of high quality, while the text description is of medium quality. Based on the quality assessment results, the system assigns a weight of 0.7 to the text area in the image and a weight of 0.3 to the text description. The system adjusts the weights using an adaptive algorithm, ultimately determining a weight of 0.75 for the text area and 0.25 for the text description.
[0042] Step S370: Perform fusion layering deployment based on the content complexity of the focused information and determine the fusion level. Specifically, the content complexity of the focused information is evaluated using a deep learning model (such as CNN or Transformer). Complexity evaluation includes information diversity, structural complexity, etc. Based on content complexity, the focused information is divided into different levels. Information with high complexity is assigned to a high level, and information with low complexity is assigned to a low level. The final fusion level is determined by an adaptive algorithm (such as adaptive pooling).
[0043] For example, the system evaluates the content complexity of the first focused information (the text area in the image) and the additional focused information (the text description). The text area in the image has a higher content complexity, while the text description has a lower content complexity. The system assigns the text area in the image to a high level and the text description to a low level. The system adjusts the levels using an adaptive algorithm, ultimately determining that the text area in the image is at a high level and the text description is at a low level.
[0044] Step S380: Determine a focused fusion paradigm based on the information weight and the fusion level. Specifically, an appropriate fusion paradigm is selected based on the information weight and fusion level. For example, a deep fusion paradigm is used for high-level information, while a shallow fusion paradigm is used for low-level information. The fusion paradigm is optimized using an adaptive algorithm (such as an adaptive normalization layer) to ensure optimal fusion results.
[0045] For example, based on information weight and fusion level, the system selects the deep fusion paradigm for processing text areas in an image and the shallow fusion paradigm for processing text descriptions. The system adjusts the fusion paradigm through an adaptive algorithm, ultimately determining that the deep fusion paradigm is used for text areas in the image and the shallow fusion paradigm is used for text descriptions.
[0046] Step S390: Perform multimodal fusion on the first focused information and the additional focused information according to the focus fusion paradigm. Specifically, the first focused information and the additional focused information are fused according to the focus fusion paradigm. Fusion methods include feature concatenation and weighted feature summation. Post-processing techniques (such as context correction and format restoration) are used to optimize the fusion results to ensure output accuracy and readability.
[0047] For example, the system combines the text regions and text descriptions in an image to generate a fused feature representation. The system then optimizes the fusion results through contextual correction and format restoration techniques to ensure the accuracy and readability of the output.
[0048] In one possible implementation, according to the focus fusion paradigm, the first focus information and the additional focus information are multimodally fused, and step S390 further includes step S391, determining the first fusion layer from bottom to top, dividing the complementary information and the superimposed information, wherein the first fusion layer is the minimum fusion unit. Specifically, the first fusion layer, as the minimum fusion unit, is responsible for processing the most basic feature fusion tasks, such as preliminarily aligning the text area in the image with the corresponding part in the text description. Starting from the most basic feature layer, information fusion is gradually performed upward to ensure that low-level detail information is retained during the fusion process, while gradually integrating high-level semantic information. Feature extraction technology (such as CNN or Transformer) is used to analyze the first focus information and the additional focus information to identify which information is complementary (i.e., unique information provided by different modalities) and which is superimposed (i.e., similar or repeated information provided by different modalities).
[0049] Step S392, weighted fusion is performed on the superimposed information according to the information weight, and information prior and enhancement processing is performed on the complementary information to determine a layer of fused information. Specifically, based on the previously determined information weight, weighted summation or weighted averaging is performed on the superimposed information to ensure that important information occupies a larger proportion in the fusion process. For complementary information, enhancement processing is performed using prior knowledge (such as known modal relationships), for example, highlighting important features through an attention mechanism. The processed superimposed information and complementary information are integrated to form a layer of fused information as the basis for subsequent fusion.
[0050] For example, the system performs a weighted averaging fusion of the text in the image and the text in the text description, based on information weights (e.g., image weight 0.75, text weight 0.25). The system leverages prior knowledge to enhance the correlation between the chart area in the image and the data table in the text, highlighting the key data in the chart through an attention mechanism. The system then integrates the weighted fusion of text information and the enhanced chart information to form a layer of fused information.
[0051] In step S393, the first layer of fusion information is fused based on the second fusion layer, performing hierarchical iterative fusion to determine the fusion result of the first focused information. Specifically, in the second fusion layer, the first layer of fusion information is further processed, such as feature extraction and semantic parsing, to extract deeper semantic information. Through multi-layer iteration, the fusion information at different levels is gradually integrated to ensure effective fusion of information at all levels. The fusion result of the first focused information is ultimately determined as the output of multimodal fusion.
[0052] In one possible implementation, after determining the fusion result of the first focus information, step S390 further includes step S394, determining the second focus information based on the first modal information according to the scanning path. Specifically, according to the previously planned scanning path, the system continues to scan the next focus area in the first modal information (such as an image). The scanning path ensures that each area in the layout structure is processed in a logical order. A deep learning model (such as CNN or Transformer) is used to extract features of the second focus area, which include text content, image features, etc. The extracted features are integrated into the second focus information in preparation for subsequent fusion.
[0053] For example, based on the scan path, the system locates the second text region in the image, which contains an important descriptive text. The system uses a CNN model to extract the image features of this text region and uses OCR technology to identify the text content. The image features and text content are combined into the second focused information, ready for the next step of fusion.
[0054] Step S395, performing mutual information analysis on the fusion result of the first focused information and the second focused information to determine the second fusion focus, wherein the second fusion focus is determined based on context association. Specifically, the correlation between the fusion result of the first focused information and the second focused information is evaluated through mutual information analysis technology. Mutual information analysis can quantify the amount of shared information between the two information sources. Based on the results of the mutual information analysis, the second fusion focus is determined, that is, the direction that needs to be focused on during the fusion process. This is based on context association to ensure the global consistency and logic of the fusion result. According to the results of the mutual information analysis, the fusion strategy is dynamically adjusted to optimize the fusion effect.
[0055] For example, the system performs mutual information analysis on the fusion results of the first focused information (such as the title area in the image) and the second focused information (such as the descriptive text area in the image), and finds that the two have a high correlation in content. The system determines that the second fusion focuses on strengthening the logical connection between the text content and ensuring that the title and descriptive text maintain consistency and coherence in the fusion results. Based on the results of the mutual information analysis, the system adjusts the fusion strategy and increases the weight of the descriptive text to optimize the fusion effect.
[0056] Step S396, with the second fusion focus as a constraint, performs additional modal entity alignment and engine paradigm adaptive fusion based on the second focus information. Specifically, according to the second fusion focus, the entities in the second focus information (such as keywords in the text, objects in the image) are aligned to ensure that the entities in different modalities can accurately correspond. According to the alignment results and the second fusion focus, the additional OCR engine in the OCR engine array is dynamically configured to perform adaptive fusion, including adjusting the parameters and fusion paradigm of the OCR engine to optimize the fusion effect. The fusion results are optimized through post-processing techniques (such as context correction and format restoration) to ensure the accuracy and readability of the output.
[0057] In one possible implementation, after performing multimodal fusion scanning and recognition management, the method further includes step S400 of receiving a display request. Specifically, the user's display request is received via a user interface or API, and the request content is parsed to determine the type and format of information the user desires to be displayed. Based on the parsed results, the display request is categorized as text display, table display, image display, etc. for subsequent processing.
[0058] For example, if a user submits a display request through the web interface, requesting that the recognition results be displayed in a table format and exported as an Excel file, the system will classify this request as a table display request and prepare it for subsequent processing.
[0059] Step S500, deploys the fusion results under N-item focus according to the display requirements and determines the multimodal fusion results. Specifically, according to the display requirements, the results after multimodal fusion are sorted and the information required by the user is extracted. The extracted information is converted into a format specified by the user, such as text, table, image, etc. The converted information is deployed to the user interface or a specified output device, such as a web page, file system, etc. Among them, N-item focus means that during the multimodal fusion process, the system will focus on multiple key areas or feature points, and the number of these areas or feature points is represented by N.
[0060] For example, the system extracts text content and table data from the multimodal fusion results, converts the extracted text content into an Excel spreadsheet format, saves the generated Excel file to the server, and provides a download link to the user.
[0061] In the above, refer to Figure 1 The intelligent OCR dynamic adaptive method based on multimodal fusion according to the embodiment of the present invention is described in detail. Figure 2 An intelligent OCR dynamic adaptive system based on multimodal fusion according to an embodiment of the present invention is described.
[0062] The intelligent dynamic adaptive OCR system based on multimodal fusion according to an embodiment of the present invention is used to address the technical issues of low recognition accuracy and poor adaptability in existing multimodal information OCR processing, thereby achieving the technical effect of improving the recognition accuracy and robustness of multimodal information OCR. The intelligent dynamic adaptive OCR system based on multimodal fusion includes: a layout structure determination module 10, an initialization module 20, and a multimodal fusion module 30.
[0063] The layout structure determination module 10 is used to upload multimodal information to the OCR recognition platform, select the first modal information and perform fuzzy scanning, and perform layout analysis to determine the layout structure, wherein the first modal information is any modality in the multimodal information; the initialization module 20 is used to introduce layered progressive fusion conditions according to the complexity of the content, set a dynamic fusion paradigm, and initialize the OCR engine array deployed by the OCR recognition platform, wherein complementary fusion and overlapping fusion are used as the basic fusion goals; the multimodal fusion module 30 is used to perform multi-step focused fusion planning according to the layout structure, trigger directional focusing based on the first modal information, align focus with the entity for additional modal information, dynamically configure the OCR engine and fusion paradigm, and perform scanning and recognition management under multimodal fusion, wherein the OCR engine array is dynamically allocated and the fusion paradigm is adaptively adjusted based on the focused content characteristics based on the layout structure, and modal empowerment and progressive stratification are used as adaptive elements.
[0064] The specific configuration of the multimodal fusion module 30 will be described in detail below. As described above, to trigger directional focusing based on the first modal information and dynamically configure the OCR engine, the multimodal fusion module 30 may further include: a scanning path planning unit for planning a scanning path based on the layout structure of the first modal information; a first focus information determination unit for determining first focus information based on the scanning path; and a first OCR engine connection unit for temporarily connecting a first OCR engine in the OCR engine array based on the display format of the first focus information.
[0065] Among them, for entity alignment and focusing of additional modal information, the OCR engine is dynamically configured, and the multimodal fusion module 30 may further include: an additional focus information determination unit is used to perform entity alignment on the additional modal information based on the first focus information, and determine the additional focus information, wherein the additional modal information is the remaining modalities in the multimodal information except the first modal information; an additional OCR engine connection unit is used to traverse the display format of the additional focus information, and temporarily connect the additional OCR engine in the OCR engine array, wherein the additional OCR engine corresponds to the additional focus information.
[0066] Among them, the fusion paradigm is dynamically configured to perform multimodal fusion processing. The multimodal fusion module 30 may further include: an information weighting unit for identifying the first focused information and the additional focused information, performing information weighting according to the modal information quality, and determining the information weight; a fusion layered deployment unit for performing fusion layered deployment according to the content complexity of the focused information, and determining the fusion level; a focused fusion paradigm determination unit for determining the focused fusion paradigm based on the information weight and the fusion level; and a multimodal fusion unit for performing multimodal fusion of the first focused information and the additional focused information according to the focused fusion paradigm.
[0067] Among them, according to the focus fusion paradigm, the first focus information and the additional focus information are multimodally fused, and the multimodal fusion unit may further include: a first fusion layer determination subunit is used to determine the first fusion layer from bottom to top, and divide the complementary information and the superimposed information, wherein the first fusion layer is the minimum fusion unit; a layer of fusion information determination subunit is used to perform weighted fusion on the superimposed information according to the information weight, perform information prior and enhancement processing on the complementary information, and determine a layer of fusion information; a hierarchical iterative fusion subunit is used to perform fusion processing on the one layer of fusion information based on the second fusion layer, hierarchical iterative fusion, and determine the fusion result of the first focus information.
[0068] Among them, after determining the fusion result of the first focus information, the multimodal fusion unit may further include: a second focus information determination subunit for determining the second focus information based on the first modal information according to the scanning path; a second fusion focus determination subunit for performing mutual information analysis on the fusion result of the first focus information and the second focus information to determine the second fusion focus, wherein the second fusion focus is determined based on context association; an engine paradigm adaptive fusion subunit for performing additional modal entity alignment and engine paradigm adaptive fusion based on the second focus information with the second fusion focus as a constraint.
[0069] The specific configuration of the layout structure determination module 10 will be described in detail below. As described above, to determine the layout structure, the layout structure determination module 10 may further include: a pre-check port deployment unit for deploying a pre-check port and establishing an association between the pre-check port and the OCR engine array; a global fuzzy scanning unit for performing a global fuzzy scan on the first modal information based on the pre-check port to determine content layout features; and an information framework determination unit for determining an information framework based on the content layout features and identifying the content ranking relationship as the layout structure.
[0070] Among them, after performing scanning and recognition management under multimodal fusion, the system may further include: a display requirement receiving module for receiving display requirements; a multimodal fusion result determination module for deploying the fusion results under N focuses according to the display requirements and determining the multimodal fusion results.
[0071] The intelligent OCR dynamic adaptive system based on multimodal fusion provided by the embodiment of the present invention can execute the intelligent OCR dynamic adaptive method based on multimodal fusion provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0072] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules may be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention.
[0073] The above specific embodiments do not constitute a limitation to the scope of protection of this application. It should be understood by those skilled in the art that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of this application should be included in the scope of protection of this application. In some cases, the actions or steps recorded in this application can be performed in an order different from that in the embodiments and can still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. Intelligent OCR dynamic adaptive method based on multimodal fusion, characterized by: The method comprises: Uploading the multimodal information to the OCR recognition platform, selecting the first modal information and performing fuzzy scanning, and performing layout analysis to determine the layout structure, wherein the first modal information is any modal in the multimodal information; According to the complexity of the content, hierarchical progressive fusion conditions are introduced, a dynamic fusion paradigm is set, and the OCR engine array deployed on the OCR recognition platform is initialized, wherein complementary fusion and overlapping fusion are used as the basic fusion goals; Based on the layout structure, multi-step focus fusion planning is performed, triggering directional focus based on the first modal information, entity alignment focus based on the additional modal information, dynamically configuring the OCR engine and fusion paradigm, and performing scanning recognition management under multimodal fusion; Among them, based on the focused content characteristics of the layout structure, the OCR engine array is dynamically allocated and the fusion paradigm is adaptively adjusted, with modal empowerment and progressive stratification as adaptive elements.
2. The intelligent OCR dynamic adaptive method based on multimodal fusion according to claim 1, characterized in that: Triggering directional focusing based on the first modal information and dynamically configuring the OCR engine, including: Planning a scanning path according to the layout structure of the first modal information; determining first focus information according to the scanning path; A first OCR engine in the OCR engine array is temporarily connected in the display format of the first focused information.
3. The intelligent OCR dynamic adaptive method based on multimodal fusion according to claim 2, characterized in that: Focusing on entities with additional modal information, the OCR engine is dynamically configured, including: Performing entity alignment based on the first focused information on additional modal information to determine additional focused information, wherein the additional modal information is the remaining modalities in the multimodal information except the first modal information; The display formats of the additional focus information are traversed, and an additional OCR engine in the OCR engine array is temporarily connected, wherein the additional OCR engine corresponds to the additional focus information.
4. The intelligent OCR dynamic adaptive method based on multimodal fusion according to claim 3, characterized in that: Dynamically configure fusion paradigms to perform multimodal fusion processing, including: Identifying the first focused information and the additional focused information, weighting the information based on the modal information quality, and determining the information weight; Based on the complexity of the focused information, perform fusion layered deployment and determine the fusion level; determining a focused fusion paradigm based on the information weight and the fusion level; According to the focus fusion paradigm, multimodal fusion is performed on the first focus information and the additional focus information.
5. The intelligent OCR dynamic adaptive method based on multimodal fusion according to claim 4, characterized in that: According to the focus fusion paradigm, performing multimodal fusion on the first focus information and the additional focus information includes: Determine a first fusion layer from bottom to top, and divide the complementary information and the superposition information, wherein the first fusion layer is the minimum fusion unit; Perform weighted fusion on the superimposed information according to the information weight, perform information priori and enhancement processing on the complementary information, and determine a layer of fused information; Based on the second fusion layer, the first layer of fusion information is fused and iteratively fused to determine the fusion result of the first focus information.
6. The intelligent OCR dynamic adaptive method based on multimodal fusion according to claim 5, characterized in that: After determining the fusion result of the first focus information, the method includes: determining, according to the scanning path, second focus information based on the first modal information; Performing mutual information analysis on a fusion result of the first focused information and the second focused information to determine a second fusion focus, wherein the second fusion focus is determined based on context association; Taking the second fusion focus as a constraint, additional modality entity alignment and engine paradigm adaptive fusion based on the second focus information are performed.
7. The intelligent OCR dynamic adaptive method based on multimodal fusion according to claim 1, characterized in that: Determine the layout structure, including: Deploy a pre-check port and establish an association between the pre-check port and the OCR engine array; Performing a full-area fuzzy scan on the first modal information according to the pre-check port to determine content layout features; An information framework is determined according to the content layout characteristics, and the content ranking relationship is identified as the layout structure.
8. The intelligent OCR dynamic adaptive method based on multimodal fusion according to claim 1, characterized in that: After performing scanning recognition management under multimodal fusion, it includes: Receive display requirements; According to the display requirements, the fusion results under N focuses are deployed to determine the multimodal fusion results.
9. Multimodal fusion intelligent OCR dynamic adaptive system, characterized by: The system is used to implement the multimodal fusion intelligent OCR dynamic adaptive method according to any one of claims 1 to 8, and the system includes: A layout structure determination module is used to upload multimodal information to the OCR recognition platform, select the first modal information and perform fuzzy scanning, and perform layout analysis to determine the layout structure, wherein the first modal information is any modal in the multimodal information; The initialization module is used to introduce hierarchical progressive fusion conditions according to the complexity of the content, set the dynamic fusion paradigm, and initialize the OCR engine array deployed by the OCR recognition platform, with complementary fusion and overlapping fusion as the basic fusion goals; The multimodal fusion module is used to perform multi-step focus fusion planning based on the layout structure, trigger directional focus based on the first modal information, align focus with the entity for additional modal information, dynamically configure the OCR engine and fusion paradigm, and perform scanning recognition management under multimodal fusion. Among them, the OCR engine array is dynamically allocated and the fusion paradigm is adaptively adjusted based on the focus content characteristics based on the layout structure, with modal empowerment and progressive stratification as adaptive elements.
Citation Information
Patent Citations
Quick text recognition method
CN101751567A
Multi-source multi-modal data processing system and method for applying system
CN111859451A
Printing text credibility fusion method and device based on three-source OCR result
CN114694152A
OCR error detection method based on multi-modal information fusion
CN117953524A
Wireless table structure identification method and system based on two-stage multi-modal feature fusion and storage medium
CN118279922A
Cited By
File identification processing system based on artificial intelligence model and RAG
CN120894793A
Multi-modal fusion-based OCR (optical character recognition) information dynamic verification method and system and medium
CN121708608A