OCR path selection method and device based on dynamic routing and related equipment
By acquiring multidimensional features of documents and user constraints, and using pre-trained models for OCR path decision-making, the problem of inflexible path selection in existing systems is solved, thereby improving the efficiency and accuracy of OCR processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING TAIXIN TIANCHENG TECHNOLOGY CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-12
AI Technical Summary
Existing OCR systems lack the ability to accurately assess document complexity, cannot adaptively select processing paths based on multi-dimensional features such as text density, geometric deformation, image quality, background complexity, and layout complexity, ignore user constraints, and have unintelligent resource allocation, resulting in low processing efficiency.
By acquiring multidimensional features of the document to be processed, combining them with user constraints to perform vector concatenation, and using a pre-trained path decision model to make OCR processing path decisions, the optimal processing path is selected.
It enables intelligent path selection based on document characteristics and user needs, improving the efficiency and accuracy of OCR processing and adapting to the differentiated needs of different application scenarios.
Smart Images

Figure CN122020060A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to an OCR path selection method, apparatus and related equipment based on dynamic routing. Background Technology
[0002] With the rapid growth in demand for digital office work and document processing, Optical Character Recognition (OCR) technology has become a core technology for document digitization. Traditional OCR systems typically employ fixed processing workflows, using the same recognition algorithms and parameter configurations for all types of documents. This "one-size-fits-all" approach reveals significant limitations when faced with complex and diverse document types.
[0003] In recent years, researchers have begun to explore more intelligent document processing methods. Chinese patent application CN121457462A discloses a content-aware and intelligent routing document parsing method. This method extracts multi-dimensional feature vectors from multimodal documents and uses a pre-defined routing decision model to determine the parsing tool for each page of the document [CN121457462A]. Chinese patent application CN120182989B discloses a multimodal intelligent AI classification system for document organization, which introduces a dual-channel generator and a cross-modal consistency loss function, and uses the PPO algorithm for dynamic routing decisions [CN120182989B]. Furthermore, Chinese patent application CN121009497A proposes a routing method for multimodal problems under an AI platform, which processes various modal data and calculates intent recognition confidence through a multimodal intent fusion model [CN121009497A].
[0004] However, existing OCR systems and document processing methods still have significant technical shortcomings. First, existing systems lack the ability to accurately assess document complexity, failing to adaptively select processing paths based on multi-dimensional features such as text density, geometric deformation, image quality, background complexity, and layout complexity. Second, traditional methods ignore the impact of user constraints on processing strategies, failing to incorporate users' specific needs for accuracy, speed, and power consumption into the path decision-making process. Third, existing systems lack intelligent resource management, especially in environments with limited computing resources, such as mobile and edge devices, making it difficult to find the optimal balance between processing accuracy, latency, and power consumption. Finally, the path selection strategies of existing methods are relatively rigid, lacking dynamic optimization mechanisms based on actual processing results, leading to low overall processing efficiency and failing to meet the differentiated needs of different application scenarios. Summary of the Invention
[0005] The embodiments of the present invention provide a method, apparatus and related equipment for OCR path selection based on dynamic routing, which aims to solve the technical problem that it is difficult to determine OCR processing resources for document processing in traditional technologies.
[0006] In a first aspect, embodiments of the present invention provide an OCR path selection method based on dynamic routing, comprising: Obtain the document to be processed and preprocess it to obtain several pages to be processed; Multidimensional feature extraction is performed on each page to be processed to obtain the multidimensional features of each page to be processed. Obtain the structured text of the user constraints and encode the structured text to obtain the user constraint vector; The multidimensional features and user constraint vectors are concatenated to obtain concatenated features. Based on the pre-trained path decision model, the OCR processing path decision is performed on the processing path of each page to be processed according to the splicing features, so as to obtain the optimal processing path for each page to be processed.
[0007] Secondly, embodiments of the present invention provide an OCR path selection device based on dynamic routing, comprising: The preprocessing module is used to acquire the document to be processed and preprocess the document to obtain several pages to be processed. The feature extraction module is used to perform multi-dimensional feature extraction on each page to be processed, so as to obtain the multi-dimensional features of each page to be processed. The encoding module is used to obtain the structured text of user constraints and encode the structured text to obtain the user constraint vector. The splicing module is used to splice the multidimensional features and the user constraint vector to obtain spliced features; The path selection module is used to make OCR processing path decisions for each page to be processed based on the pre-trained path decision model and the splicing features, so as to obtain the optimal processing path for each page to be processed.
[0008] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the dynamic routing-based OCR path selection method described in the first aspect.
[0009] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the dynamic routing-based OCR path selection method described in the first aspect.
[0010] This invention provides an OCR path selection method, apparatus, and related equipment based on dynamic routing. The method acquires a document to be processed and preprocesses it to obtain several pages to be processed. Multidimensional features are extracted from each page to obtain multidimensional features. Structured text containing user constraints is acquired and encoded to obtain user constraint vectors. The multidimensional features and user constraint vectors are concatenated to obtain concatenated features. Based on a pre-trained path decision model, the OCR processing path for each page is determined according to the concatenated features, resulting in the optimal processing path for each page. This method achieves intelligent selection of OCR processing paths through comprehensive analysis of multidimensional features and personalized consideration of user constraints. It can automatically select the most suitable processing strategy based on the characteristics of different pages and user needs, improving the efficiency and accuracy of OCR processing. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating an embodiment of the OCR path selection method based on dynamic routing provided by this invention; Figure 2 A flowchart illustrating another embodiment of the OCR path selection method based on dynamic routing provided in this invention. Figure 3 This is a schematic block diagram of an OCR path selection device based on dynamic routing provided in an embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0015] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0016] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0017] Please see Figure 1 This is a flowchart illustrating the OCR path selection method based on dynamic routing provided in an embodiment of the present invention. The method includes steps S110 to S150.
[0018] Step S110: Obtain the document to be processed and preprocess the document to be processed to obtain several pages to be processed; In this embodiment, the document to be processed is obtained and preprocessed to obtain several pages to be processed. This step divides the original document into pages, ensuring that each page can be independently processed for subsequent feature extraction and path decision-making.
[0019] Step S120: Extract multidimensional features from each page to be processed to obtain the multidimensional features of each page to be processed; In this embodiment, multi-dimensional feature extraction is performed on each page to be processed to obtain the multi-dimensional features of each page. This step comprehensively analyzes the page characteristics through various technical means: (1) A lightweight text detection network is used to identify the text density of the pagination to be processed, and the text density features are obtained. The lightweight text detection network can quickly locate the text region in the page and count the distribution density of the text in the page, providing text quantification information for subsequent path selection.
[0020] (2) Geometric features of the pagination to be processed are identified by Hough transform to obtain geometric deformation features. Hough transform can detect geometric shapes such as straight lines and circles in the page and identify whether there are geometric deformation problems such as tilting and twisting of the page.
[0021] (3) Extract image quality features from the page to be processed to obtain image quality features. This process analyzes image quality parameters such as page sharpness, contrast, and noise level to evaluate the overall image quality.
[0022] (4) Background identification is performed on the pagination to be processed through color clustering and texture analysis to obtain background complexity features. The color clustering algorithm classifies the colors in the page, and the texture analysis identifies the texture pattern of the background to comprehensively evaluate the complexity of the background.
[0023] (5) Edge detection and connectivity analysis are used to identify the layout of the pagination to be processed and obtain the layout complexity features. The edge detection algorithm identifies the boundaries of each region in the page, and connectivity analysis determines the distribution of different content regions and evaluates the complexity of the page layout.
[0024] Step S130: Obtain the structured text of the user constraints and encode the structured text to obtain the user constraint vector; In this embodiment, the structured text of user constraints is obtained and encoded to obtain a user constraint vector. The user constraints include personalized requirements such as processing accuracy requirements, processing speed requirements, and specific content types, which are converted into numerical vector form through encoding.
[0025] Step S140: Perform vector concatenation on the multidimensional features and user constraint vectors to obtain concatenated features; In this embodiment, multidimensional features and user constraint vectors are concatenated to obtain concatenated features. By combining the objective features of the page with the subjective needs of the user, a comprehensive decision-making basis is formed.
[0026] Step S150: Based on the pre-trained path decision model, perform OCR processing path decision on the processing path of each page to be processed according to the splicing features to obtain the optimal processing path for each page to be processed.
[0027] In this embodiment, based on a pre-trained path decision model, the OCR processing path for each page to be processed is determined according to the splicing features, resulting in the optimal processing path for each page. The path decision model can select the most suitable processing scheme from multiple available OCR processing paths based on page features and user constraints. The training process of the path decision model includes: We acquire sample documents containing features of varying complexity, preprocess them to obtain corresponding sample pagination, and label each sample pagination with a tag indicating the corresponding standard processing path. The sample data covers various types of document pages to ensure that the model can learn the optimal path selection strategy in different scenarios. The concatenated features corresponding to the pagination of the samples are input into the initial path decision model for training, resulting in the predicted processing path output by the model. The initial model learns the mapping relationship between features and the optimal path through training on a large amount of sample data. The decision loss between the predicted processing path and the standard processing path for the corresponding sample pagination is calculated based on a pre-defined loss function. The model parameters of the initial path decision model are then optimized based on this loss to obtain the path decision model. The loss function quantifies the difference between the predicted result and the standard answer. The model parameters are continuously adjusted through backpropagation to improve prediction accuracy.
[0028] In one embodiment, six standard processing paths are preset, each corresponding to different processing combinations and configurations as follows: The standard processing path P0 is suitable for "very simple documents". The processing path is "lightweight detection - fast recognition", and the module is configured as "lowest precision mode". The standard processing path P1 is suitable for "simple documents". The processing path is "standard detection - fast recognition", and the module is configured as "low precision mode". The standard processing path P2 is suitable for "medium-sized documents". The processing path is "enhanced detection - standard recognition", and the module is configured as "standard mode". The standard processing path P3 is suitable for "complex documents". The processing path is "enhanced detection - geometric correction - standard recognition", and the module is configured as "enhanced mode". The standard processing path P4 is suitable for "extremely complex documents". The processing path is "enhanced detection - geometric correction - enhanced recognition - post-processing optimization", and the module is configured as "high-precision mode". The standard processing path P5 is suitable for "professional-level processing". The processing path is "complete pipeline -- multi-model fusion -- iterative optimization", and the module is configured as "highest precision mode".
[0029] In this embodiment, according to the six preset standard processing paths, the splicing features corresponding to the pagination to be processed are input into the pre-trained path decision model. The model outputs the probability distribution of the six processing paths, and the processing path with the highest probability is selected as the output.
[0030] In one embodiment, a lightweight convolutional neural network (LightweightCNN) is used as the feature extraction network. This network employs a depthwise separable convolutional structure, which reduces computational cost and parameter count while maintaining feature extraction capabilities, thus adapting to the deployment requirements of edge devices. After convolution, pooling, activation, and other operations, the image tensor outputs a one-dimensional feature vector with a dimension of 256. This vector contains low-level and high-level features of the document, such as text, geometry, background, and layout.
[0031] In one implementation, as described by Mr. Huang, the path decision model employs a two-layer fully connected neural network structure. The first fully connected layer compresses the 265-dimensional concatenated features to 64 dimensions and enhances feature representation through the ReLU activation function. The second fully connected layer maps the 64-dimensional vector to a 6-dimensional vector, corresponding to the original scores of the six processing paths (P0-P5). The Softmax function converts the 6-dimensional original scores into a probability distribution, where the sum of the probability values of each dimension is 1. A higher probability value indicates that the corresponding path is more suitable for the current input. The path ID corresponding to the dimension with the largest value in the probability distribution is selected as the optimal processing path for the final output. The probability value of this path is also output for confidence verification by the subsequent performance monitoring module.
[0032] In one embodiment, the process of the OCR path selection method based on dynamic routing provided by this invention further includes: Step S210: Perform semantic recognition on each of the recognition units, determine the association relationship between each recognition unit, and label each recognition unit according to the association relationship to obtain the association label; Step S220: Obtain the OCR recognition results corresponding to all recognition units, and fuse and correct the OCR recognition results with related relationships according to the associated tags to obtain the corrected recognition results.
[0033] In this embodiment, considering the possibility of cross-page text content or tables within the context of the document to be processed, semantic recognition is performed on each recognition unit to determine the relationships between them. Based on these relationships, each recognition unit is labeled to obtain association tags. The OCR recognition results corresponding to the related recognition units are then fused and corrected according to the association tags to obtain a complete recognition result. The labeling of association tags can be achieved by constructing a deep learning model for semantic recognition and labeling.
[0034] This method acquires the document to be processed and preprocesses it to obtain several pages to be processed. Multidimensional features are extracted from each page to obtain its multidimensional features. The structured text containing user constraints is obtained and encoded to obtain user constraint vectors. The multidimensional features and user constraint vectors are concatenated to obtain concatenated features. Based on a pre-trained path decision model, the processing path for each page is determined according to the concatenated features, resulting in the optimal processing path for each page. This method achieves intelligent selection of OCR processing paths through comprehensive analysis of multidimensional features and personalized consideration of user constraints. It can automatically select the most suitable processing strategy based on the characteristics of different pages and user needs, improving the efficiency and accuracy of OCR processing.
[0035] This invention also provides a dynamic routing-based OCR path selection device, which is used to execute any of the aforementioned embodiments of the dynamic routing-based OCR path selection method. Specifically, please refer to... Figure 2 , Figure 2 This is a schematic block diagram of an OCR path selection device based on dynamic routing provided in an embodiment of the present invention. The OCR path selection device 100 based on dynamic routing can be configured in a server.
[0036] like Figure 2 As shown, the OCR path selection device 100 based on dynamic routing includes a preprocessing module 110, a feature extraction module 120, an encoding module 130, a splicing module 140, and a path selection module 150.
[0037] The preprocessing module 110 is used to acquire the document to be processed and preprocess the document to be processed to obtain several pages to be processed; Feature extraction module 120 is used to perform multi-dimensional feature extraction on each page to be processed to obtain multi-dimensional features of each page to be processed; The encoding module 130 is used to acquire the structured text of the user constraints and encode the structured text to obtain the user constraint vector. The splicing module 140 is used to perform vector splicing on the multidimensional features and the user constraint vector to obtain spliced features; The path selection module 150 is used to make OCR processing path decisions for each page to be processed based on the pre-trained path decision model and the splicing features, so as to obtain the optimal processing path for each page to be processed.
[0038] In one embodiment, the feature extraction module 120 includes: The first extraction unit is used to perform text density recognition on the pagination to be processed using a lightweight text detection network to obtain text density features; The second extraction unit is used to perform geometric feature recognition on the pagination to be processed by Hough transform to obtain geometric deformation features; The third extraction unit is used to extract image quality features from the pagination to be processed, and obtain image quality features; The fourth extraction unit is used to perform background recognition on the pagination to be processed through color clustering and texture analysis to obtain background complexity features; The fifth extraction unit is used to perform layout recognition on the pagination to be processed using edge detection and connectivity region analysis to obtain layout complexity features.
[0039] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the dynamic routing-based OCR path selection method described above.
[0040] In another embodiment of the invention, a computer-readable storage medium is provided. This computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the dynamic routing-based OCR path selection method as described above.
[0041] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.
[0042] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Units with the same function may be grouped into one unit. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, or it may be an electrical, mechanical, or other form of connection.
[0043] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.
[0044] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0045] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks.
[0046] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for OCR path selection based on dynamic routing, characterized in that, include: Obtain the document to be processed and preprocess it to obtain several pages to be processed; The content of the page to be processed is identified, and different content is divided into regions to obtain multiple identification units; Multidimensional feature extraction is performed on each of the recognition units to obtain the multidimensional features of each recognition unit; Obtain the structured text of the user constraints and encode the structured text to obtain the user constraint vector; The multidimensional features and user constraint vectors are concatenated to obtain concatenated features. Based on the pre-trained path decision model, the OCR processing path decision is performed on the processing path of each page to be processed according to the splicing features, so as to obtain the optimal processing path of each recognition unit.
2. The OCR path selection method based on dynamic routing as described in claim 1, characterized in that, The step of extracting multidimensional features from each page to be processed to obtain multidimensional features for each page to be processed includes: A lightweight text detection network is used to identify the text density of the pagination to be processed, and the text density features are obtained. The geometric features of the pagination to be processed are identified by Hough transform to obtain geometric deformation features; Image quality features are extracted from the pagination to be processed to obtain image quality features; Background complexity features are obtained by performing background recognition on the pagination to be processed through color clustering and texture analysis. The layout of the pagination to be processed is identified by edge detection and connectivity region analysis to obtain layout complexity features.
3. The OCR path selection method based on dynamic routing as described in claim 1, characterized in that, The training process of the path decision model includes: Retrieve the sample pages corresponding to the sample document. Each sample page is marked with a corresponding standard processing path. The splicing features corresponding to the sample pagination are input into the initial path decision model for model training to obtain the predicted processing path output by the model. The decision loss between the predicted processing path and the standard processing path for corresponding sample pagination is calculated based on a preset loss function, and the model parameters of the initial path decision model are optimized based on the decision loss to obtain the path decision model.
4. The OCR path selection method based on dynamic routing as described in claim 1, characterized in that, include: Semantic recognition is performed on each of the recognition units to determine the association between them, and each recognition unit is labeled according to the association to obtain the association label; Obtain the OCR recognition results corresponding to all recognition units, and fuse and correct the OCR recognition results with related relationships according to the associated tags to obtain the complete recognition result.
5. An OCR path selection device based on dynamic routing, characterized in that, include: The preprocessing module is used to acquire the document to be processed and preprocess the document to obtain several pages to be processed. The feature extraction module is used to perform multi-dimensional feature extraction on each page to be processed, so as to obtain the multi-dimensional features of each page to be processed. The encoding module is used to obtain the structured text of user constraints and encode the structured text to obtain the user constraint vector. The splicing module is used to splice the multidimensional features and the user constraint vector to obtain spliced features; The path selection module is used to make OCR processing path decisions for each page to be processed based on the pre-trained path decision model and the splicing features, so as to obtain the optimal processing path for each page to be processed.
6. The OCR path selection device based on dynamic routing as described in claim 5, characterized in that, The feature extraction module includes: The first extraction unit is used to perform text density recognition on the pagination to be processed using a lightweight text detection network to obtain text density features; The second extraction unit is used to perform geometric feature recognition on the pagination to be processed by Hough transform to obtain geometric deformation features; The third extraction unit is used to extract image quality features from the pagination to be processed, and obtain image quality features; The fourth extraction unit is used to perform background recognition on the pagination to be processed through color clustering and texture analysis to obtain background complexity features; The fifth extraction unit is used to perform layout recognition on the pagination to be processed using edge detection and connectivity region analysis to obtain layout complexity features.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the OCR path selection method based on dynamic routing as described in any one of claims 1 to 4.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the OCR path selection method based on dynamic routing as described in any one of claims 1 to 4.