Intelligent OCR dynamic self-adaption method and system based on multi-modal fusion
By using a multimodal fusion-based intelligent OCR dynamic adaptive method, the problems of low recognition accuracy and poor adaptability in multimodal information OCR processing are solved. This method enables efficient recognition of complex and ever-changing page layouts and information content, improving recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510501064.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-04-21
Smart Images

Figure CN120472481B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of multi-modal information fusion processing, in particular to an intelligent OCR dynamic self-adaptive method and system based on multi-modal fusion. BACKGROUND
[0002] In today's digital age, efficient and accurate processing of multi-modal information to achieve intelligent optical character recognition (OCR) is crucial for the development of many industries (such as document digitization, automated office work, intelligent information retrieval, etc.), which can greatly improve information processing efficiency and accuracy and promote the intelligent upgrading of related fields. The main method to solve the intelligent OCR of multi-modal information at present is to use traditional single modal OCR technology to process different modal information respectively, or to simply fuse multiple modal information and then recognize it, combined with some fixed fusion rules and fixed OCR engine configuration. However, traditional single modal OCR technology cannot fully utilize the complementarity between multi-modal information, and it is difficult to cope with complex and variable layout structures and information content; the simple fusion of multi-modal information lacks in-depth analysis of information content and layout structure, and the fixed fusion rules and OCR engine configuration cannot be dynamically adjusted according to the actual situation, resulting in low recognition accuracy and poor robustness when facing multi-modal information of different complexity and structure, which cannot meet the demand for high-precision and high-adaptability OCR technology in practical applications.
[0003] At present, in the related technology, there are technical problems of low recognition accuracy and poor adaptability in multi-modal information OCR processing. SUMMARY
[0004] The present application provides an intelligent OCR dynamic self-adaptive method and system based on multi-modal fusion, which uploads multi-modal information to an OCR recognition platform, selects one modal information for fuzzy scanning, performs layout analysis to determine the layout structure, introduces hierarchical progressive fusion conditions according to the content complexity, sets a dynamic fusion paradigm (with complementary and overlapping fusion as the target), initializes an OCR engine array, plans multi-step focused fusion according to the layout structure, triggers directional focusing and entity alignment focusing, dynamically configures the OCR engine and the fusion paradigm, and based on the layout structure focusing content characteristics, uses modal weighting and progressive hierarchical as elements to perform engine dynamic distribution and fusion paradigm adaptive adjustment on the OCR engine array, which achieves the technical effects of improving the recognition accuracy and robustness of multi-modal information OCR.
[0005] The application provides an intelligent OCR dynamic self-adaption method based on multi-modal fusion, comprising: uploading multi-modal information to an OCR recognition platform, selecting first modal information and performing fuzzy scanning, and performing layout analysis to determine a layout structure, wherein the first modal information is any modal in the multi-modal information; introducing a hierarchical progressive fusion condition according to content complexity, setting a dynamic fusion paradigm, initializing an OCR engine array deployed by the OCR recognition platform, wherein complementary fusion and overlapping fusion are used as the basis for fusion target; according to the layout structure, multi-step focused fusion planning is performed, directional focusing based on the first modal information is triggered, entity alignment focusing for additional modal information is performed, OCR engines and fusion paradigms are dynamically configured, and scanning and recognition management under multi-modal fusion is performed; wherein, based on the focused content characteristics of the layout structure, the OCR engine array is subjected to engine dynamic distribution and fusion paradigm adaptive adjustment, and modal empowerment and progressive hierarchical are used as adaptive elements.
[0006] In a possible implementation, directional focusing based on the first modal information is triggered, and the OCR engine is dynamically configured to perform the following processing: according to the layout structure of the first modal information, a scanning path is planned; according to the scanning path, first focusing information is determined; and a first OCR engine in the OCR engine array is temporarily coupled in the display format of the first focusing information.
[0007] In a possible implementation, entity alignment focusing for additional modal information is dynamically configured to perform the following processing: entity alignment based on the first focusing information is performed on additional modal information to determine additional focusing information, wherein the additional modal information is the remaining modal in the multi-modal information except the first modal information; the display format of the additional focusing information is traversed, and additional OCR engines in the OCR engine array are temporarily coupled, wherein the additional OCR engines correspond to the additional focusing information.
[0008] In a possible implementation, the fusion paradigm is dynamically configured to perform multi-modal fusion processing, and the following processing is performed: the first focusing information and the additional focusing information are identified, information empowerment is performed according to modal information quality to determine information weight, fusion hierarchical deployment is performed according to the content complexity of the focusing information to determine fusion level, and a focused fusion paradigm is determined according to the information weight and the fusion level; and the first focusing information and the additional focusing information are subjected to multi-modal fusion according to the focused fusion paradigm.
[0009] In a possible implementation, according to the focus fusion paradigm, the first focus information and the additional focus information are fused in a multi-modal manner, and the following processing is performed: a first fusion layer is determined from bottom to top, and complementary information and superimposed information are divided, wherein the first fusion layer is a minimum fusion unit; the superimposed information is fused by weighting according to the information weight, the complementary information is subjected to information priori and enhancement processing, and one layer of fusion information is determined; the one layer of fusion information is subjected to fusion processing based on a second fusion layer, and hierarchical iterative fusion is performed to determine a fusion result of the first focus information.
[0010] In a possible implementation, after the fusion result of the first focus information is determined, the following processing is performed: second focus information based on the first modality information is determined according to the scanning path; mutual information analysis is performed on the fusion result of the first focus information and the second focus information to determine a second fusion emphasis direction, wherein the second fusion emphasis direction is determined based on context association; and additional modality entity alignment based on the second focus information is performed, and engine paradigm adaptive fusion is performed with the second fusion emphasis direction as a constraint.
[0011] In a possible implementation, the layout structure is determined, and the following processing is performed: a pre-check port is deployed, and an association between the pre-check port and the OCR engine array is established; according to the pre-check port, global fuzzy scanning is performed on the first modality information to determine content layout features; an information framework is determined according to the content layout features, and a content position relationship is identified as the layout structure.
[0012] In a possible implementation, after the scanning recognition management under the multi-modal fusion is performed, the following processing is performed: a display requirement is received; according to the display requirement, the fusion results under N focus are deployed to determine a multi-modal fusion result.
[0013] The application also provides an intelligent OCR dynamic adaptive system based on multi-modal fusion, comprising: a layout structure determination module, configured to upload multi-modal information to an OCR recognition platform, select first modal information and perform fuzzy scanning, and perform layout analysis to determine a layout structure, wherein the first modal information is any modal in the multi-modal information; an initialization module, configured to introduce a hierarchical progressive fusion condition according to content complexity, set a dynamic fusion paradigm, and initialize an OCR engine array deployed by the OCR recognition platform, wherein complementary fusion and overlapping fusion are used as the basis fusion target; and a multi-modal fusion module, configured to perform multi-step focus fusion planning according to the layout structure, trigger directional focus based on the first modal information, and entity alignment focus for additional modal information, dynamically configure the OCR engine and the fusion paradigm, and perform scanning and recognition management under multi-modal fusion, wherein focus content characteristics based on the layout structure are used to adaptively adjust the engine dynamic distribution and the fusion paradigm of the OCR engine array, and modal empowerment and progressive layering are used as adaptive elements.
[0014] The intelligent OCR dynamic adaptive method and system based on multi-modal fusion provided in the application first upload multi-modal information to an OCR recognition platform, select first modal information and perform fuzzy scanning, perform layout analysis to determine a layout structure, wherein the first modal information is any modal in the multi-modal information, then introduce a hierarchical progressive fusion condition according to content complexity, set a dynamic fusion paradigm, initialize an OCR engine array deployed by the OCR recognition platform, wherein complementary fusion and overlapping fusion are used as the basis fusion target, and finally perform multi-step focus fusion planning according to the layout structure, trigger directional focus based on the first modal information, and entity alignment focus for additional modal information, dynamically configure the OCR engine and the fusion paradigm, and perform scanning and recognition management under multi-modal fusion, wherein focus content characteristics based on the layout structure are used to adaptively adjust the engine dynamic distribution and the fusion paradigm of the OCR engine array, and modal empowerment and progressive layering are used as adaptive elements. The technical effects of improving the recognition accuracy and robustness of multi-modal information OCR are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings of the embodiments of the application will be briefly introduced as follows. In the present application, flowcharts are used to illustrate the operations performed by the system according to the embodiments of the application. It should be understood that the foregoing or the following operations are not necessarily performed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously according to needs. Meanwhile, other operations can be added to these processes, or one or more steps of operations can be removed from these processes.
[0016] Figure 1A flowchart of an intelligent OCR dynamic self-adaptation method based on multi-modal fusion provided by an embodiment of the present application is shown.
[0017] Figure 2 A structural diagram of an intelligent OCR dynamic self-adaptation system based on multi-modal fusion provided by an embodiment of the present application is shown.
[0018] Legend: layout structure determination module 10, initialization module 20, multi-modal fusion module 30. DETAILED DESCRIPTION
[0019] The above description is only a summary of the technical solutions of the present application. In order to make the technical means of the present application more clear, the following specific embodiments of the present application can be implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described.
[0020] In order to make the purposes, technical solutions and advantages of the present application more clear, the following will combine the drawings to make a further detailed description of the present application. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.
[0021] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict. The term "first\second" is only to distinguish similar objects, and does not represent a specific order of the objects. The terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application.
[0022] The embodiments of the present application provide an intelligent OCR dynamic self-adaptation method based on multi-modal fusion, as shown in Figure 1 The method comprises the following steps:
[0023] In step S100, multi-modal information is uploaded to an OCR recognition platform, first modal information is selected and fuzzy scanning is performed, and layout analysis is performed to determine the layout structure, wherein the first modal information is any modal in the multi-modal information.
[0024] Specifically, the data containing multiple modalities, such as images, texts, audios, etc., are uploaded to the OCR recognition platform through a network interface (such as HTTP protocol). The OCR recognition platform provides an API interface to support file uploads in multiple formats. The first modality information refers to the modality information in the multi-modal data used for fuzzy scanning and layout analysis, which is an image or image data with a clear two-dimensional structure and layout features, and can be analyzed and parsed through image processing techniques. The first modality information (such as image) is preprocessed, and a Gaussian blur filter is used to blur the image to reduce noise and details and extract low-level features. A deep learning model (such as YOLO or Faster R-CNN) is used to analyze the layout of the image, and the positions and layouts of text regions, tables, pictures, and other elements are identified.
[0025] For example, a user uploads an image containing handwritten text and charts through the web interface of the OCR recognition platform, along with a text describing the content of the image. After the platform receives the file, it is stored in a designated directory on the server. The system applies a Gaussian blur filter to the uploaded image, and small scratches or noise points in the image are smoothed. The system uses a YOLO model to analyze the layout of the blurred image, identifies the top-left corner of the image as the title area, the middle part as the table area, and the lower-right corner as the annotation area, and labels the bounding boxes of each area.
[0026] In one possible implementation, determining the layout structure, step S100 further includes step S110, deploying a pre-check port and establishing an association between the pre-check port and the OCR engine array. Specifically, a pre-check port is deployed at the front end of the OCR recognition platform, which is a network interface for receiving and preliminarily processing uploaded multi-modal information. The pre-check port can be a network interface that supports multiple formats of data input. Through a configuration management tool, the pre-check port is associated with the OCR engine array to ensure that the pre-check port can directly deliver processed data to the OCR engine array. For example, a user uploads a multi-modal file containing images and texts through a web interface, and the pre-check port receives the file and performs preliminary analysis to extract basic information such as file type and size. The pre-check port establishes a connection with the OCR engine array through a configuration file to ensure smooth data transmission, for example, the pre-check port sends image data to a CNN-based OCR engine and sends text data to a Transformer-based OCR engine.
[0027] Step S120, according to the pre-check port, the first modal information is globally blurred scanning, determine the content layout features. Specifically, using image processing technology, such as Gaussian blur filter, on the first modal information (image) is globally blurred scanning, reduce noise, extract low-level texture and shape features. Through deep learning model (such as YOLO or Faster R-CNN) analysis blurred scanning image, identify the text area, chart area, table area and so on in the image, determine the content layout features.
[0028] Step S130, according to the content layout features determine information framework, and identify the content bit order relationship, as the layout structure. Specifically, according to the content layout features, construct information framework, clear the hierarchical relationship and logical structure of each area. Through deep learning model (such as Transformer) analysis the relative position and logical relationship between regions, identify the content bit order relationship, form a complete layout structure. For example, the system uses Transformer model analysis the relative position and logical relationship between regions, identify the content bit order relationship as the title is located above the text, table is located on the right side of the text. This implementation way through the deployment of pre-check port, the first modal information is globally blurred scanning and content layout features determination, and construct information framework and identify the content bit order relationship, can ensure that the OCR recognition platform can accurately identify and process different regions in the image, so that the system can better understand and process complex multi-modal information.
[0029] Step S200, according to the content complexity, introduce hierarchical progressive fusion conditions, set dynamic fusion paradigm, initialize the OCR engine array deployed by the OCR recognition platform, wherein, with complementary fusion and overlapping fusion as the basis fusion target.
[0030] Specifically, by calculating the texture complexity (such as edge detection algorithm) and the semantic complexity (such as text length, vocabulary diversity) of the image, evaluate the content complexity of multi-modal information. According to the content complexity, dynamically adjust the fusion strategy. For complex content, use multi-layer neural network to gradually fuse information of different modalities; for simple content, directly perform shallow fusion. Initialize the OCR engine array, configure different OCR engines (such as CNN based OCR engine and Transformer based OCR engine), and dynamically allocate tasks according to the fusion strategy.
[0031] For example, the system calculates the texture complexity of the image and finds that the image contains complex charts and handwritten text, and the content complexity is high. At the same time, by analyzing the uploaded text description, it is found that the text length is long and the vocabulary diversity is high, further confirming the content complexity. According to the evaluation results, the system decides to use a hierarchical progressive fusion strategy. For the image modality, first extract low-level texture features, and then gradually fuse high-level semantic features; for the text modality, first extract word embedding features, and then gradually fuse context information. The system initializes the OCR engine array, configures the CNN-based OCR engine to process the image modality, and configures the Transformer-based OCR engine to process the text modality, and assigns a higher weight to the image modality and a lower weight to the text modality according to the complexity.
[0032] Step S300, according to the layout structure, multi-step focus fusion planning is carried out, directional focus based on the first modality information is triggered, entity alignment focus for additional modality information is carried out, OCR engine and fusion paradigm are dynamically configured, and scanning recognition management under multi-modal fusion is carried out, wherein the focus content characteristics based on the layout structure are used to dynamically allocate the engine to the OCR engine array and adaptively adjust the fusion paradigm, and the modality weighting and progressive layering are adaptive elements.
[0033] Specifically, according to the layout structure, the layout is divided into multiple regions, and different fusion strategies are developed for each region. For example, for the text region, the image and text modalities are focused; for the table region, the image, text and structured data modalities are fused. The attention mechanism (such as the self-attention mechanism in Transformer) is used to focus on the first modality information, and the entity alignment of the additional modality information is carried out. For example, the multi-modal alignment is optimized through the comparison learning objective function. The parameters of the OCR engine and the fusion paradigm are dynamically adjusted according to the focus content characteristics. For example, the resolution and recognition algorithm of the OCR engine are dynamically adjusted according to the complexity of the text region.
[0034] For example, the system divides the image into text regions and chart regions according to the layout structure. For text regions, the focus is on fusing the image and text modalities; for chart regions, the focus is on fusing the image, text, and structured data modalities. For text regions, CNN is used to extract image features, and Transformer is used to extract text features, and then fusion is performed; for chart regions, image features, text features, and table structure features are combined for fusion. The system uses a self-attention mechanism to focus on the image modality in a targeted manner, extracting key features from the image. At the same time, the text modality is aligned with entities to ensure that the text in the image is consistent with the entities in the text description. For example, through contrastive learning, the "company name" in the image is aligned with the "company name" in the text description. The system dynamically adjusts the resolution of the OCR engine according to the complexity of the text region, using high resolution to process handwritten text regions and low resolution to process printed text regions. At the same time, the fusion paradigm is dynamically adjusted according to the fusion effect, such as increasing or decreasing the number of fusion layers and optimizing the fusion weights. The embodiments of the present application upload multi-modal information to the OCR recognition platform, select one modal information for fuzzy scanning, perform layout analysis to determine the layout structure, introduce hierarchical progressive fusion conditions based on content complexity, set dynamic fusion paradigms (aiming at complementary and overlapping fusion), initialize the OCR engine array, plan multi-step focused fusion based on the layout structure, trigger directional focus and entity alignment focus, dynamically configure the OCR engine and the fusion paradigm, focus on content characteristics based on the layout structure, and use modal weighting and progressive hierarchical elements to dynamically allocate the OCR engine array and adaptively adjust the fusion paradigm, thereby improving the recognition accuracy and robustness of multi-modal information OCR.
[0035] In one possible implementation, directional focus based on the first modal information is triggered, and the OCR engine is dynamically configured, step S300 further comprising step S310, planning a scanning path according to the layout structure of the first modal information. Specifically, a deep learning model (such as YOLO or Faster R-CNN) is used to analyze the layout structure of the first modal information, and text regions, chart regions, table regions, etc. are identified. According to the analysis result of the layout structure, a path planning algorithm such as A* algorithm or Dijkstra algorithm is used to determine the scanning path. The goal of path planning is to ensure that the scanning order is logical and to improve recognition efficiency. According to the content complexity and the importance of the region, the scanning path is dynamically adjusted, and the key region is scanned first. For example, the system identifies that the image contains a title region, a text region, and a chart region. The title region is located at the top of the image, the text region is located in the middle, and the chart region is located at the bottom. The system uses the A* algorithm to plan the scanning path, and scans the title region first, then the text region, and finally the chart region.
[0036] At step S320, the first focus information is determined according to the scan path. Specifically, according to the scan path, the key information of each region is extracted using an attention mechanism (such as the self-attention mechanism in Transformer). The extracted key information is fused with the global feature to generate the first focus information (the key information extracted according to the scan path is used for subsequent feature fusion and OCR recognition). According to the content complexity of the region, the extraction depth and range of the focus information are dynamically adjusted. For example, the system uses the self-attention mechanism for the title region to extract the key information of the title, such as keywords and phrases. The extracted title key information is fused with the global feature of the image to generate the first focus information.
[0037] At step S330, the first OCR engine in the OCR engine array is temporarily coupled in the display format of the first focus information. Specifically, according to the display format of the first focus information (such as text, geometry, etc.), a suitable OCR engine is selected. Through the configuration management tool, the first focus information is temporarily coupled to the first OCR engine in the OCR engine array (the OCR engine in the OCR engine array is specially used for processing the first modal information), ensuring that the information can be correctly processed. According to the content complexity of the first focus information, the parameters of the OCR engine are dynamically adjusted to improve the recognition efficiency. For example, the system identifies that the first focus information is in text format, and selects a Transformer-based OCR engine for processing. The system temporarily couples the first focus information to the first OCR engine, ensuring that the information can be correctly processed.
[0038] In one possible implementation, the entity alignment of additional modal information is focused, and the OCR engine is dynamically configured. Step S300 further includes step S340, performing entity alignment based on the first focused information on the additional modal information, and determining additional focused information, wherein the additional modal information is the remaining modal information in the multi-modal information except the first modal information. Specifically, using multi-modal fusion technology, the additional modal information (such as text, audio, etc.) is aligned with the first modal information (such as image). The alignment process is based on the first focused information, ensuring that the entities between different modalities can be accurately matched. Features are extracted from the additional modal information, such as word embeddings of text or acoustic features of audio, and then matched with features in the first focused information. According to the alignment result, the processing mode of the additional modal information is dynamically adjusted to improve the accuracy and efficiency of alignment. For example, assuming that the first modal information is an image and the additional modal information is a text description. The system uses a pre-trained image-text matching model to calculate the similarity score between the objects in the image and the entities in the text description. Word embedding features are extracted from the text and visual features are extracted from the image, and then the alignment is optimized through contrastive learning. If it is found that the alignment effect of some text entities and image entities is not good, the system will adjust the alignment strategy, such as increasing the context information or adjusting the feature weight.
[0039] Step S350, traversing the display format of the additional focused information, temporarily coupling the additional OCR engine in the OCR engine array, wherein the additional OCR engine corresponds to the additional focused information. Specifically, according to the display format of the additional focused information (such as text format, image format, etc.), the appropriate OCR engine is selected. Through the configuration management tool, the additional focused information is temporarily coupled to the additional OCR engine in the OCR engine array, ensuring that the information can be correctly processed. According to the content complexity of the additional focused information, the parameters of the OCR engine are dynamically adjusted to improve the recognition efficiency. For example, assuming that the additional focused information is in text format, the system selects the OCR engine based on Transformer for processing. The system temporarily couples the additional focused information to the additional OCR engine, ensuring that the information can be correctly processed. If the additional focused information contains complex table structures, the system will adjust the parameters of the OCR engine to improve the accuracy of table recognition.
[0040] In one possible implementation, the dynamic configuration fusion paradigm is used for multi-modal fusion processing, and step S300 further includes step S360 of identifying the first focus information and the additional focus information, weighting the information according to the quality of the modal information, and determining the information weight. Specifically, the quality of the first focus information and the additional focus information is evaluated by a pre-trained deep learning model (such as a Transformer or a CNN). The quality evaluation includes the integrity, accuracy, noise level, and the like of the information. According to the quality of the modal information, the weight is dynamically allocated. High-quality modal information is allocated a higher weight, and low-quality modal information is allocated a lower weight. The final information weight is determined by an adaptive algorithm (such as an adaptive normalization layer).
[0041] For example, the system evaluates the quality of the first focus information (the text region in the image) and the additional focus information (the text description). The text region in the image has a high quality, and the text description has a medium quality. The system allocates a weight of 0.7 to the text region in the image and a weight of 0.3 to the text description according to the quality evaluation result. The system adjusts the weights by an adaptive algorithm, and finally determines that the weight of the text region in the image is 0.75 and the weight of the text description is 0.25.
[0042] Step S370, according to the content complexity of the focus information, the fusion is deployed in layers, and the fusion level is determined. Specifically, the content complexity of the focus information is evaluated by a deep learning model (such as a CNN or a Transformer). The complexity evaluation includes the diversity and structural complexity of the information. According to the content complexity, the focus information is divided into different levels. The information with high complexity is allocated to a high level, and the information with low complexity is allocated to a low level. The final fusion level is determined by an adaptive algorithm (such as adaptive pooling).
[0043] For example, the system evaluates the content complexity of the first focus information (the text region in the image) and the additional focus information (the text description). The text region in the image has a high content complexity, and the text description has a low content complexity. The system allocates the text region in the image to a high level and the text description to a low level. The system adjusts the levels by an adaptive algorithm, and finally determines that the text region in the image is in a high level and the text description is in a low level.
[0044] Step S380, the information weight and the fusion level are used to determine the focus fusion paradigm. Specifically, according to the information weight and the fusion level, a suitable fusion paradigm is selected. For example, for high-level information, a deep fusion paradigm is used, and for low-level information, a shallow fusion paradigm is used. The fusion paradigm is optimized by an adaptive algorithm (such as an adaptive normalization layer) to ensure the optimal fusion effect.
[0045] For example, the system selects a deep fusion paradigm to process the text region in the image and a shallow fusion paradigm to process the text description according to the information weight and the fusion level. The system adjusts the fusion paradigm through an adaptive algorithm, and finally determines that the text region in the image adopts the deep fusion paradigm and the text description adopts the shallow fusion paradigm.
[0046] At step S390, the first focus information and the additional focus information are fused according to the focus fusion paradigm. Specifically, the first focus information and the additional focus information are fused according to the focus fusion paradigm. The fusion method includes feature splicing, feature weighted summation, etc. The fusion result is optimized through post-processing techniques (such as context correction and format restoration) to ensure the accuracy and readability of the output.
[0047] For example, the system splices the features of the text region and the text description in the image to generate the fused feature representation. The system optimizes the fusion result through context correction and format restoration techniques to ensure the accuracy and readability of the output.
[0048] In one possible implementation, according to the focus fusion paradigm, the first focus information and the additional focus information are fused, and step S390 further includes step S391 of determining a first fusion layer from bottom to top and dividing complementary information and superimposed information, wherein the first fusion layer is the smallest fusion unit. Specifically, the first fusion layer, as the smallest fusion unit, is responsible for processing the most basic feature fusion task, such as preliminarily aligning the text region in the image with the corresponding part in the text description. Starting from the most basic feature layer, information fusion is gradually performed upwards to ensure that low-level detail information is retained in the fusion process, while high-level semantic information is gradually integrated. Feature extraction techniques (such as CNN or Transformer) are used to analyze the first focus information and the additional focus information to identify which information is complementary (i.e., unique information provided by different modalities) and which information is superimposed (i.e., similar or repeated information provided by different modalities).
[0049] At step S392, the superimposed information is weighted and fused according to the information weight, the complementary information is processed with information priori and enhancement, and one layer of fusion information is determined. Specifically, according to the information weight determined in advance, the superimposed information is subjected to weighted summation or weighted averaging, etc. to ensure that important information occupies a larger proportion in the fusion process. For the complementary information, priori knowledge (such as known modal relationship) is used for enhancement processing, for example, important features are highlighted through attention mechanism. The processed superimposed information and complementary information are integrated to form one layer of fusion information as the basis for subsequent fusion.
[0050] For example, the system performs a weighted average fusion of the text in the image and the text in the text description according to the information weights (e.g., image weight 0.75 and text weight 0.25). The system uses prior knowledge to enhance the relevance of the chart area in the image and the data table in the text, and highlights the key data in the chart through an attention mechanism. The system integrates the weighted and fused text information and the enhanced chart information to form a layer of fused information.
[0051] At step S393, the one layer of fused information is fused based on a second fusion layer, and hierarchical iterative fusion is performed to determine the fusion result of the first focus information. Specifically, at the second fusion layer, the one layer of fused information is further processed, such as feature extraction and semantic analysis, to extract deeper semantic information. Through multi-layer iterative fusion, the fused information at different levels is gradually integrated to ensure effective fusion at different levels. Finally, the fusion result of the first focus information is determined as the output of the multi-modal fusion.
[0052] In one possible implementation, after determining the fusion result of the first focus information, step S390 further includes step S394 of determining second focus information based on the first modal information according to the scanning path. Specifically, according to the previously planned scanning path, the system continues to scan the next focus area in the first modal information (e.g., image). The scanning path ensures that each area in the layout structure is processed in a logical order. Deep learning models (e.g., CNN or Transformer) are used to extract features of the second focus area, including text content, image features, etc. The extracted features are integrated into the second focus information to prepare for subsequent fusion.
[0053] For example, according to the scanning path, the system locates the second text area in the image, which contains an important descriptive text. The system uses a CNN model to extract image features of the text area, and identifies the text content through OCR technology. The image features and text content are integrated into the second focus information to prepare for the next step of fusion.
[0054] At step S395, mutual information analysis is performed on the fusion result of the first focus information and the second focus information to determine a second fusion focus direction, wherein the second fusion focus direction is determined based on context association. Specifically, through mutual information analysis technology, the correlation between the fusion result of the first focus information and the second focus information is evaluated. Mutual information analysis can quantify the amount of shared information between two information sources. Based on the result of mutual information analysis, the second fusion focus direction is determined, i.e., the direction that needs to be focused on during the fusion process. This is based on context association to ensure the global consistency and logicality of the fusion result. According to the result of mutual information analysis, the fusion strategy is dynamically adjusted to optimize the fusion effect.
[0055] For example, the system performs mutual information analysis on the fusion result of the first focus information (such as the title area in the image) and the second focus information (such as the descriptive text area in the image), and finds that the two have high relevance in content. The system determines that the second fusion emphasis is to strengthen the logical association of text content, ensuring that the title and descriptive text remain consistent and coherent in the fusion result. According to the results of mutual information analysis, the system adjusts the fusion strategy and increases the weight of the descriptive text to optimize the fusion effect.
[0056] Step S396, with the second fusion emphasis as a constraint, performs additional modal entity alignment and engine paradigm adaptive fusion based on the second focus information. Specifically, according to the second fusion emphasis, the entities in the second focus information (such as keywords in text, objects in image) are aligned to ensure that the entities in different modalities can be accurately corresponded. According to the alignment result and the second fusion emphasis, the additional OCR engines in the OCR engine array are dynamically configured to perform adaptive fusion, including adjusting the parameters and fusion paradigm of the OCR engines to optimize the fusion effect. The fusion result is optimized through post-processing techniques (such as context correction, format restoration) to ensure the accuracy and readability of the output.
[0057] In one possible implementation, after performing scan recognition management under multi-modal fusion, the method further includes step S400 of receiving display requirements. Specifically, the user's display requirements are received through a user interface or API, and the requirement content is parsed to determine the type and format of information the user wants to display. According to the parsing result, the display requirements are classified into text display, table display, image display, etc. for subsequent processing.
[0058] For example, the user submits a display requirement through a Web interface, requiring the recognition result to be displayed in a table form and exported as an Excel file. The system classifies this requirement as a table display requirement for subsequent processing.
[0059] Step S500, according to the display requirements, deploy the fusion results under N focuses to determine the multi-modal fusion results. Specifically, according to the display requirements, the results after multi-modal fusion are arranged, and the information required by the user is extracted. The extracted information is converted into the format specified by the user, such as text, table, image, etc. The converted information is deployed to the user interface or specified output device, such as Web page, file system, etc. Wherein, N focuses refer to the multiple key areas or feature points that the system pays attention to during multi-modal fusion. The number of these areas or feature points is represented by N.
[0060] For example, the system extracts text content and table data from the multi-modal fusion result. The extracted text content is converted into an Excel table format. The generated Excel file is saved to the server, and a download link is provided to the user.
[0061] In the foregoing, reference is made to Figure 1 The multi-modal fusion-based intelligent OCR dynamic adaptation method according to the embodiment of the application is described in detail. Next, the multi-modal fusion-based intelligent OCR dynamic adaptation system according to the embodiment of the application will be described with reference to Figure 2 The multi-modal fusion-based intelligent OCR dynamic adaptation system according to the embodiment of the application is used to solve the technical problems of low recognition accuracy and poor adaptability existing in the prior art multi-modal information OCR processing, and achieves the technical effect of improving the recognition accuracy and robustness of multi-modal information OCR. The multi-modal fusion-based intelligent OCR dynamic adaptation system comprises a layout structure determination module 10, an initialization module 20, and a multi-modal fusion module 30.
[0062] The layout structure determination module 10 is configured to upload multi-modal information to an OCR recognition platform, select first modal information and perform fuzzy scanning, and execute layout analysis to determine a layout structure, wherein the first modal information is any modal of the multi-modal information. The initialization module 20 is configured to introduce a hierarchical progressive fusion condition according to content complexity, set a dynamic fusion paradigm, and initialize an OCR engine array deployed on the OCR recognition platform, wherein the complementary fusion and overlapping fusion are used as the basis for the fusion target. The multi-modal fusion module 30 is configured to perform multi-step focused fusion planning according to the layout structure, trigger directional focusing based on the first modal information, and perform entity alignment focusing for additional modal information, dynamically configure the OCR engine and the fusion paradigm, and perform scanning and recognition management under multi-modal fusion, wherein the focused content characteristics based on the layout structure are used to dynamically allocate the OCR engine array and adaptively adjust the fusion paradigm, and the modal empowerment and progressive hierarchical are used as adaptive elements.
[0063] The layout structure determination module 10 is configured to upload multi-modal information to an OCR recognition platform, select first modal information and perform fuzzy scanning, and execute layout analysis to determine a layout structure, wherein the first modal information is any modal of the multi-modal information. The initialization module 20 is configured to introduce a hierarchical progressive fusion condition according to content complexity, set a dynamic fusion paradigm, and initialize an OCR engine array deployed on the OCR recognition platform, wherein the complementary fusion and overlapping fusion are used as the basis for the fusion target. The multi-modal fusion module 30 is configured to perform multi-step focused fusion planning according to the layout structure, trigger directional focusing based on the first modal information, and perform entity alignment focusing for additional modal information, dynamically configure the OCR engine and the fusion paradigm, and perform scanning and recognition management under multi-modal fusion, wherein the focused content characteristics based on the layout structure are used to dynamically allocate the OCR engine array and adaptively adjust the fusion paradigm, and the modal empowerment and progressive hierarchical are used as adaptive elements.
[0064] Next, the specific configuration of the multi-modal fusion module 30 will be described in detail. As described above, the directional focusing based on the first modal information is triggered, and the OCR engine is dynamically configured. The multi-modal fusion module 30 can further comprise a scanning path planning unit configured to plan a scanning path according to the layout structure of the first modal information; a first focusing information determination unit configured to determine first focusing information according to the scanning path; and a first OCR engine coupling unit configured to temporarily couple a first OCR engine in the OCR engine array in a display format of the first focusing information.
[0065] Wherein, the entity alignment of the additional modal information is focused, the dynamic configuration of the OCR engine, the multi-modal fusion module 30 can further comprise: an additional focus information determination unit for performing entity alignment based on the first focus information on the additional modal information, determining additional focus information, wherein the additional modal information is the remaining modal information in the multi-modal information except the first modal information; an additional OCR engine coupling unit for traversing the display format of the additional focus information, temporarily coupling the additional OCR engine in the OCR engine array, wherein the additional OCR engine corresponds to the additional focus information.
[0066] Wherein, the dynamic configuration of the fusion paradigm, the multi-modal fusion processing, the multi-modal fusion module 30 can further comprise: an information weighting unit for identifying the first focus information and the additional focus information, performing information weighting according to the quality of the modal information, determining the information weight; a fusion hierarchical deployment unit for performing fusion hierarchical deployment according to the content complexity of the focus information, determining the fusion level; a focus fusion paradigm determination unit for determining the focus fusion paradigm with the information weight and the fusion level; a multi-modal fusion unit for performing multi-modal fusion on the first focus information and the additional focus information according to the focus fusion paradigm.
[0067] Wherein, according to the focus fusion paradigm, the multi-modal fusion unit can further comprise: a first fusion layer determination subunit for determining a first fusion layer from bottom to top, dividing complementary information and superimposed information, wherein the first fusion layer is the smallest fusion unit; a one-layer fusion information determination subunit for weighting fusion on the superimposed information according to the information weight, performing information priori and enhancement processing on the complementary information, determining one-layer fusion information; a hierarchical iterative fusion subunit for performing fusion processing on the one-layer fusion information based on a second fusion layer, hierarchical iterative fusion, determining the fusion result of the first focus information.
[0068] Wherein, after determining the fusion result of the first focus information, the multi-modal fusion unit can further comprise: a second focus information determination subunit for determining a second focus information based on the first modal information according to the scanning path; a second fusion emphasis direction determination subunit for performing mutual information analysis on the fusion result of the first focus information and the second focus information, determining a second fusion emphasis direction, wherein the second fusion emphasis direction is determined based on context association; an engine paradigm adaptive fusion subunit for performing additional modal entity alignment and engine paradigm adaptive fusion based on the second focus information with the second fusion emphasis direction as a constraint.
[0069] Below, the specific configuration of the layout structure determining module 10 will be described in detail. As described above, in order to determine the layout structure, the layout structure determining module 10 can further include: a pre-check port deploying unit for deploying a pre-check port and establishing the association of the pre-check port with the OCR engine array; a global fuzzy scanning unit for performing global fuzzy scanning on the first modality information according to the pre-check port to determine content layout features; and an information framework determining unit for determining an information framework according to the content layout features and identifying a content position relationship as the layout structure.
[0070] After the scanning and recognition management under the multi-modal fusion is performed, the system can further include: a display requirement receiving module for receiving a display requirement; and a multi-modal fusion result determining module for deploying the fusion result under N focuses according to the display requirement to determine a multi-modal fusion result.
[0071] The intelligent OCR dynamic self-adaptive system based on multi-modal fusion provided in the embodiments of the present application can execute the intelligent OCR dynamic self-adaptive method based on multi-modal fusion provided in any of the embodiments of the present application, and has the corresponding functional modules and beneficial effects of the execution method.
[0072] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or server, and the various units and modules are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual differentiation, and do not limit the protection scope of the present application.
[0073] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application. In some cases, the actions or steps described in the present application can be executed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.
Claims
1. An intelligent OCR dynamic self-adaptive method based on multi-modal fusion, characterized in that, The method comprises: uploading multi-modal information to an OCR recognition platform, selecting first modal information and performing fuzzy scanning, and performing layout analysis to determine the layout structure, wherein the first modal information is any modal in the multi-modal information; According to the content complexity, introduce the hierarchical progressive fusion condition, set the dynamic fusion paradigm, initialize the OCR engine array deployed by the OCR recognition platform, wherein the complementary fusion and the overlapping fusion are the basic fusion targets; According to the layout structure, multi-step focused fusion planning is carried out, directional focusing based on the first modal information is triggered, entity alignment focusing for additional modal information is carried out, OCR engine and fusion paradigm are dynamically configured, and scanning and recognition management under multi-modal fusion are carried out; Wherein, based on the focused content characteristics of the layout structure, the OCR engine array is dynamically allocated and the fusion paradigm is adaptively adjusted, and the adaptive elements are modal empowerment and progressive hierarchical; Trigger directional focusing based on the first modal information and dynamically configure the OCR engine, including: According to the layout structure of the first modal information, plan the scanning path; According to the scanning path, determine the first focusing information; Temporarily connect the first OCR engine in the OCR engine array with the display format of the first focusing information; Determine the layout structure, including: Deploy a pre-check port and establish the association between the pre-check port and the OCR engine array; According to the pre-check port, perform global fuzzy scanning on the first modal information to determine the content layout characteristics; According to the content layout characteristics, determine the information framework and identify the content position relationship as the layout structure.
2. The intelligent OCR dynamic self-adaptive method based on multi-modal fusion according to claim 1, characterized in that, Dynamically configure the OCR engine for entity alignment focusing of additional modal information, including: Perform entity alignment based on the first focusing information on additional modal information to determine additional focusing information, wherein the additional modal information is the remaining modal in the multi-modal information except the first modal information; Traverse the display format of the additional focusing information, and temporarily connect the additional OCR engine in the OCR engine array, wherein the additional OCR engine corresponds to the additional focusing information.
3. The intelligent OCR dynamic self-adaptive method based on multi-modal fusion according to claim 2, characterized in that, Dynamically configure the fusion paradigm for multi-modal fusion processing, including: Identify the first focusing information and the additional focusing information, perform information empowerment according to the modal information quality, and determine the information weight; According to the content complexity of the focusing information, perform hierarchical deployment of fusion, and determine the fusion level; Determine the focusing fusion paradigm according to the information weight and the fusion level; According to the focusing fusion paradigm, perform multi-modal fusion on the first focusing information and the additional focusing information.
4. The intelligent OCR dynamic self-adaptive method based on multi-modal fusion of claim 3, wherein, According to the focusing fusion paradigm, perform multi-modal fusion on the first focusing information and the additional focusing information, including: Determine the first fusion layer from bottom to top, divide the complementary information and the superimposed information, wherein the first fusion layer is the smallest fusion unit; According to the information weight, perform empowerment fusion on the superimposed information, perform information priori and enhancement processing on the complementary information, and determine one layer of fusion information; Based on the second fusion layer, the one layer of fusion information is fused and processed, and the fusion result of the first focus information is determined through hierarchical iterative fusion.
5. The intelligent OCR dynamic self-adaptive method based on multi-modal fusion according to claim 4, characterized in that, After determining the fusion result of the first focus information, the method comprises: According to the scanning path, the second focus information based on the first modal information is determined. The fusion result of the first focus information and the second focus information are executed mutual information analysis to determine the second fusion emphasis direction, wherein the second fusion emphasis direction is determined based on context association. With the second fusion emphasis direction as a constraint, additional modal entity alignment based on the second focus information and engine paradigm adaptive fusion are executed.
6. The intelligent OCR dynamic self-adaptive method based on multi-modal fusion of claim 1, wherein, After the scanning recognition management under the multi-modal fusion, the method comprises: Receiving display requirements; According to the display requirements, the fusion results under N focus are deployed to determine the multi-modal fusion result.
7. A multi-modal fused intelligent OCR dynamic adaptive system, characterized in that, The system is used to implement the intelligent OCR dynamic adaptive method of multi-modal fusion according to any one of claims 1-6, and the system comprises: A layout structure determination module is used to upload multi-modal information to an OCR recognition platform, select first modal information and perform fuzzy scanning, and execute layout analysis to determine a layout structure, wherein the first modal information is any modal in the multi-modal information; An initialization module is used to introduce a hierarchical progressive fusion condition according to content complexity, set a dynamic fusion paradigm, and initialize an OCR engine array deployed by the OCR recognition platform, wherein the complementary fusion and overlapping fusion are used as the basic fusion target; a multi-modal fusion module is used to perform multi-step focus fusion planning according to the layout structure, trigger directional focus based on the first modal information, and perform entity alignment focus for additional modal information, dynamically configure the OCR engine and the fusion paradigm, and perform scanning recognition management under multi-modal fusion, wherein the focus content characteristics based on the layout structure are used to dynamically allocate the OCR engine array and adaptively adjust the fusion paradigm, and the modal empowerment and progressive hierarchical are used as adaptive elements.
Citation Information
Patent Citations
Printing text credibility fusion method and device based on three-source OCR result
CN114694152A
End-to-end document image translation method and device
CN118397641A