Document identification method and device, equipment and storage medium
By using the preset orientation recognition algorithm to adjust the document orientation and combining the multimodal large model during the document recognition process, the problem of low document recognition accuracy is solved, and higher document recognition accuracy and effective recognition in complex scenarios are achieved.
Patent Information
- Application Number
- CN202510343398.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, document recognition is low accuracy, especially when faced with problems such as font overlap, seal text overlap or font color, OCR technology is prone to error superposition, while multimodal large models have low accuracy in document image recognition in the wrong orientation.
By obtaining the text image and prompt words to be identified, using the preset orientation recognition algorithm to identify the angle to be adjusted and rotate the text image to be adjusted, ensuring that the orientation is correct, the multi-modal model is input for recognition, and special scenes are processed in combination with layout analysis.
It improves the accuracy of document recognition, avoids the identification of OCR error superposition and the error orientation of multimodal large models, and enhances the recognition ability of complex scenes.
Smart Images

Figure CN120260057A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of information processing, and in particular, to a document recognition method, device, equipment and storage medium. Background Art
[0002] In the technical field of information processing, document recognition is a technology that converts a document image into readable, searchable and editable text. In the prior art, the document image is usually converted into editable and searchable text based on OCR (Optical Character Recognition) technology. This technology requires multiple modules to cooperate with each other to convert the document image into editable and searchable text. Taking ID card recognition as an example, it is necessary to use an orientation recognition module, a card positioning module, a text detection module, a text recognition module, and an information extraction module, etc. to perform document image recognition to obtain editable ID card text. However, in OCR technology, the more modules there are, the easier it is to cause error accumulation. Especially when there are problems such as overlapping fonts, overlapping seal characters, or lighter font colors in the document image, when using OCR technology, the text detection module is prone to missed detection, the text recognition module is prone to missed recognition, and the orientation recognition module is prone to incorrect recognition. In this way, the errors of multiple modules are superimposed together, resulting in various errors in the finally recognized document, and the accuracy of OCR for recognizing document images is relatively low in some scenarios. In addition, the multi-modal large model technology can also be used to convert the document image into editable and searchable text. However, this method cannot handle complex scenarios. For example, for a document image with an incorrect orientation, the accuracy of the multi-modal large model technology for recognizing the document image is relatively low. Therefore, it can be seen that in the prior art, the accuracy of document recognition is relatively low. Summary of the Invention
[0003] The purpose of the present invention is to provide at least one document recognition, which can at least solve the technical problem of relatively low accuracy of document recognition, and can at least achieve the technical effect of improving the accuracy of document recognition.
[0004] To solve the above technical problem, at least one embodiment of the present application provides a document recognition method, including: obtaining a text image to be recognized and a prompt word, where the prompt word is used to prompt the information to be recognized in the text image to be recognized; recognizing the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted; rotating the text image to be recognized based on the angle to be adjusted to obtain a target text image; inputting the target text image and the prompt word into a multi-modal large model to obtain a first text, where the first text is editable and searchable text.
[0005] At least one embodiment of the present application further provides a document recognition device, including: an acquisition module, configured to acquire a text image to be recognized and a prompt word, where the prompt word is used to prompt the information to be recognized in the text image to be recognized; a recognition module, configured to recognize the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted; a rotation module, configured to rotate the text image to be recognized based on the angle to be adjusted to obtain a target text image; an input module, configured to input the target text image and the prompt word into a multi-modal large model to obtain a first text, where the first text is an editable and searchable text.
[0006] At least one embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned document recognition method.
[0007] At least one embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned document recognition method is implemented.
[0008] The document recognition method provided by the embodiment of the present application includes: acquiring a text image to be recognized and a prompt word, where the prompt word is used to prompt the information to be recognized in the text image to be recognized; recognizing the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted; rotating the text image to be recognized based on the angle to be adjusted to obtain a target text image; inputting the target text image and the prompt word into a multi-modal large model to obtain a first text, where the first text is an editable and searchable text. Document recognition is performed on the target text image processed by the orientation recognition algorithm using a multi-modal large model. So that for a document image with an incorrect orientation, after being processed by the orientation recognition algorithm to obtain an angle to be adjusted, the document image is rotated based on the angle to be adjusted, thereby obtaining a target text image with a correct orientation, so that the orientation of the target text image input into the multi-modal large model is correct, and further enabling the strong learning ability of the multi-modal large model to be fully utilized for document recognition. It not only avoids the problem of easy error superposition when using OCR to recognize document images, but also avoids the problem of low accuracy in using a multi-modal large model to recognize document images with incorrect orientations. The accuracy of document recognition is improved.
[0009] In some optional embodiments, the recognition of the text image to be recognized based on the preset orientation recognition algorithm to obtain the angle to be adjusted includes: recognizing the text image to be recognized based on the preset orientation recognition algorithm to obtain the current direction of the text in the text image to be recognized; obtaining the angle to be adjusted based on the preset direction and the current direction. For a text image to be recognized that does not conform to the preset direction, the current direction of the text image to be recognized is obtained to obtain the deviation angle between the current direction and the preset direction, that is, the angle to be adjusted is obtained, so that the orientation of the text image to be recognized can be adjusted based on the angle to be adjusted subsequently, facilitating the recognition of the target text image with the correct orientation by the multi-modal large model, and avoiding the problem of low recognition accuracy caused by the recognition of the text image with the wrong orientation by the multi-modal large model. The accuracy of document recognition is improved.
[0010] In some optional embodiments, the method further includes: if the current direction is the same as the preset direction, the angle to be adjusted is zero. When the current direction of the text image to be recognized is the same as the preset direction, it indicates that the orientation of the text image to be recognized is correct and conforms to the orientation of the text image input into the multi-modal large model, so there is no need to perform rotation processing on the text image to be recognized. Rotation resources are saved.
[0011] In some optional embodiments, after rotating the text image to be recognized based on the angle to be adjusted to obtain a target text image, the method further includes: performing layout analysis on the target text image to obtain an intermediate text image; using the intermediate text image as the new target text image; the step of inputting the target text image and the prompt word into the multi-modal large model to obtain a first text includes: inputting the new target text image and the prompt word into the multi-modal large model to obtain a first text. For a target text image with multiple layouts, or for information such as cards, seals, tables, and formulas existing in the target text image
[0012] In some optional embodiments, the layout analysis includes multi-layout analysis, region localization, table detection, and formula detection. Using layout analysis can improve the accuracy of document recognition in special text image recognition scenarios.
[0013] In some optional embodiments, the layout analysis of the target text image to obtain an intermediate text image includes: performing region localization processing on the target region in the target text image to obtain the intermediate text image, where the intermediate text image contains the localization information for locating the target region, the size of the intermediate text image is the same as that of the target text image, and the target region refers to the region to be recognized. When it is necessary to detect small regions such as cards and seals that occupy a small area of the target text image, layout analysis can well locate the small regions. Based on the target text image containing the localization information, the multi-modal large model can better perform text image recognition, avoiding the problem of low recognition accuracy when the multi-modal large model recognizes a target document image in which the target region occupies a small area in the entire target document image, and improving the accuracy of document recognition.
[0014] In some optional embodiments, the inputting of the target text image and the prompt word into the multi-modal large model to obtain a first text includes: the multi-modal large model performing recognition processing on the target text image based on the prompt word to recognize the text information related to the prompt word; and using the text information as the first text. Using the prompt word to determine the content in the target text image to be recognized enables the obtained first text to be the information required by the user, avoiding extracting all information and resulting in unnecessary information in the recognized information, thus saving resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] One or more embodiments are illustrated by way of example in the accompanying drawings, and these exemplary illustrations do not limit the embodiments.
[0016] Figure 1 is a flowchart of a document recognition method provided by an embodiment of the present application;
[0017] Figure 2 is a schematic structural diagram of an apparatus for recognizing a text image provided by an embodiment of the present application;
[0018] Figure 3 is a schematic structural diagram of an apparatus for recognizing a text image provided by another embodiment of the present application;
[0019] Figure 4 is a schematic diagram of a document recognition device provided by another embodiment of the present application;
[0020] Figure 5 is a schematic structural diagram of an electronic device provided by another embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will elaborate on each embodiment of this application in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of this application, many technical details are presented to help readers better understand this application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented. The division of the following embodiments is for convenience of description and should not impose any limitation on the specific implementation of this application. The various embodiments can be combined and cross-referenced with each other on the premise of not being contradictory.
[0022] It should be noted that the acquisition or use of the data in the embodiments of this application requires user consent. Relevant data can only be obtained after the user authorizes and permits it, and the acquisition or use of the data complies with the provisions of relevant laws and regulations.
[0023] In the field of information processing technology, document recognition is a technology that converts document images into readable, searchable, and editable text. In the prior art, the document image is usually converted into editable and searchable text based on OCR (Optical Character Recognition) technology. This technology requires multiple modules to cooperate with each other to convert the document image into editable and searchable text. Taking ID card recognition as an example, it is necessary to use an orientation recognition module, a card positioning module, a text detection module, a text recognition module, and an information extraction module, etc. to perform document image recognition to obtain editable ID card text. However, in OCR technology, the more modules there are, the easier it is for error accumulation to occur. Especially when there are problems such as overlapping fonts, overlapping seal texts, or light font colors in the document image, when using OCR technology, the text detection module is prone to missed detection, the text recognition module is prone to missed recognition, and the orientation recognition module is prone to incorrect recognition. In this way, the errors of multiple modules are superimposed together, resulting in various errors in the finally recognized document, making the accuracy of OCR for recognizing document images relatively low in some scenarios. Moreover, when using OCR technology for document recognition, the more series-connected modules there are, the greater the superimposed error. In addition, multi-modal large model technology can also be used to convert document images into editable and searchable text. However, this method cannot handle complex scenarios. For example, when the multi-modal large model technology recognizes a target document image with an incorrect orientation or the target area occupies a small area in the entire target document image, the recognition accuracy is relatively low. Therefore, it can be seen that in the prior art, the accuracy of document recognition is relatively low.
[0024] Specifically, in the prior art, the relatively effective multi-modal large models are Qwen2-VL and MiniCPM-V. When the document image input given by the user and the prompt input to be parsed are input into the multi-modal large model, the multi-modal large model can give the recognition result. And the multi-modal large model does not require multiple modules to be connected in series to recognize the document image, and there is no problem of cumulative error. However, although the multi-modal large model is trained on a huge amount of data and has good robustness, there are still some drawbacks in the multi-modal large model. The recognition accuracy of document images with abnormal orientations is greatly reduced. In the ID card scenario, if the ID card area is not located in advance and the picture with the background is directly input into the multi-modal large model, the recognition result is also greatly reduced, and the larger the proportion of the background area in the whole picture, the worse the effect. That is, the multi-modal large model has a low accuracy in recognizing documents in some relatively complex scenarios.
[0025] To solve the above technical problem of low accuracy in document recognition, the present invention proposes a document recognition method. The implementation details of the document recognition method in this embodiment will be specifically described below. The following content is only the implementation details provided for easy understanding and is not necessary for implementing this solution.
[0026] Embodiment 1:
[0027] The document recognition method in this embodiment can be applied to an electronic device with communication, computing, and data storage capabilities. Its specific process can be as Figure 1 shown and includes:
[0028] Step 101, obtain the text image to be recognized and the prompt word, where the prompt word is used to prompt the information to be recognized in the text image to be recognized.
[0029] Specifically, the text image to be recognized is an image that needs to be recognized for a document, where the layout included in the document image can be a multi-layout or a single-layout.
[0030] Specifically, the multi-layout refers to multiple independent layouts or pages. For example, in the same page, there are multiple types of content that require multiple independent layouts, or there are multiple pages, where the multiple types of content include table content, graphic content, picture content, formula content, etc.
[0031] Specifically, the single-layout refers to an independent layout or page.
[0032] Specifically, using the prompt word can better determine the information to be recognized, avoid recognizing irrelevant information, and waste recognition resources. It improves the accuracy of obtaining information.
[0033] Step 102: Identify the text image to be recognized based on a preset orientation recognition algorithm to obtain the angle to be adjusted.
[0034] Specifically, the angle to be adjusted refers to the angle information required to adjust the text image to be recognized. Among them, the angle information includes the angle to be adjusted and the adjustment direction. For example, the angle to be adjusted is to rotate 30 degrees counterclockwise or 80 degrees clockwise.
[0035] In some examples, in the aforementioned Step 102, the process of identifying the text image to be recognized based on a preset orientation recognition algorithm to obtain the angle to be adjusted includes: identifying the text image to be recognized based on a preset orientation recognition algorithm to obtain the current direction of the text in the text image to be recognized; obtaining the angle to be adjusted based on the preset direction and the current direction.
[0036] Specifically, the preset orientation recognition algorithm refers to a technology in the fields of image processing and computer vision for identifying and correcting the orientation of documents or other images.
[0037] Specifically, the preset direction refers to the orientation or direction in which the target document image needs to be placed when the multi-modal large model recognizes the target document image with relatively high recognition accuracy.
[0038] In some examples, the method further includes: if the current direction is the same as the preset direction, the angle to be adjusted is zero.
[0039] In some examples, if the angle to be adjusted is not zero, it means that the current orientation of the text image to be recognized is incorrect and cannot be well recognized and processed by the multi-modal large model, and the orientation of the text image to be recognized needs to be adjusted.
[0040] Step 103: Rotate the text image to be recognized based on the angle to be adjusted to obtain a target text image.
[0041] Specifically, the direction in which the target text image is input into the multi-modal large model is the direction in which the recognition effect of the multi-modal large model for recognizing the text image is the best, that is, this direction is the same as the preset direction.
[0042] In some examples, when the text image has multiple layouts, or when the content to be detected and recognized in the text image occupies a relatively small proportion of the entire area of the text image, the recognition accuracy of using a multi-modal large model to recognize the text image is relatively low. For example, when the target text image is a multi-layout text image, or when there are contents such as cards, seals, tables, or formulas in the target text image, the recognition accuracy of the multi-modal large model is relatively low. Therefore, to solve this problem, the present application performs layout analysis on the target text image in some special scenarios. Specifically, refer to the following content: After rotating the text image to be recognized based on the angle to be adjusted to obtain a target text image, the method further includes: performing layout analysis on the target text image to obtain an intermediate text image; using the intermediate text image as the new target text image; and inputting the target text image and the prompt word into the multi-modal large model to obtain a first text, including: inputting the new target text image and the prompt word into the multi-modal large model to obtain a first text.
[0043] In some examples, the layout analysis includes multi-layout analysis, region localization, table detection, and formula detection.
[0044] Specifically, for the content of layout analysis, reference can be made to the prior art, and details will not be elaborated here.
[0045] In some examples, the performing layout analysis on the target text image to obtain an intermediate text image includes: performing region localization processing on the target region in the target text image to obtain an intermediate text image, where the intermediate text image contains the localization information for locating the target region, the size of the intermediate text image is the same as the size of the target text image, and the target region refers to the region to be recognized.
[0046] Step 104: Input the target text image and the prompt word into the multi-modal large model to obtain a first text, where the first text is an editable and searchable text.
[0047] In some examples, the inputting the target text image and the prompt word into the multi-modal large model to obtain a first text in the foregoing step 104 includes: the multi-modal large model performing recognition processing on the target text image based on the prompt word to recognize text information related to the prompt word; and using the text information as the first text.
[0048] To better understand the solution of the present application, when the text image to be recognized is a single-layout image and there is no need to perform card localization, formula detection, or table detection, reference can be made to Figure 2, input the text image to be recognized into the orientation recognition module, use the preset orientation recognition algorithm in the orientation recognition module to recognize the orientation of the text image to be recognized, and rotate the text image to be recognized whose orientation does not conform to the preset direction to obtain the target text image. Input the target text image and the prompt word into the multi-modal large model, and use the multi-modal large model to perform document recognition on the target text image to obtain the first text.
[0049] When the text image to be recognized is a multi-page image, or card positioning is required, or formula detection is required, or table detection is required, reference can be made to Figure 3 , input the text image to be recognized into the orientation recognition module, use the preset orientation recognition algorithm in the orientation recognition module to recognize the orientation of the text image to be recognized, and rotate the text image to be recognized whose orientation does not conform to the preset direction to obtain the target text image. Input the target text image into the layout analysis module for analysis and processing to obtain a new target text image. Input the new target text image and the prompt word into the multi-modal large model, and use the multi-modal large model to perform document recognition on the target text image to obtain the first text.
[0050] In summary, the document recognition method provided by the embodiments of the present application includes: obtaining a text image to be recognized and a prompt word, where the prompt word is used to prompt the information to be recognized in the text image to be recognized; recognizing the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted; rotating the text image to be recognized based on the angle to be adjusted to obtain a target text image; inputting the target text image and the prompt word into a multi-modal large model to obtain a first text, where the first text is an editable and searchable text. Use the multi-modal large model to perform document recognition on the target text image processed by the orientation recognition algorithm. After the document image with an incorrect orientation is processed by the orientation recognition algorithm to obtain the angle to be adjusted, the document image is rotated based on the angle to be adjusted, so as to obtain a target text image with a correct orientation, so that the orientation of the target text image input into the multi-modal large model is correct, and then the strong learning ability of the multi-modal large model can be fully utilized for document recognition. It not only avoids the problem of easy error superposition when using OCR to recognize document images, but also avoids the problem of low accuracy of using the multi-modal large model to recognize document images with incorrect orientations. The accuracy of document recognition is improved.
[0051] In addition, by using layout analysis to solve problems such as multi-layout, card positioning, formula detection, and table detection in the target text image, the recognition accuracy of the multi-modal large model in recognizing the content of a smaller area in the target text image is improved. This solution not only realizes the powerful information extraction ability of the multi-modal large model but also makes up for the limitations of the multi-modal large model in terms of the orientation of the text image input and the limitation in detecting the content of a smaller area in the text image, thereby improving the accuracy of text image recognition.
[0052] Embodiment 2:
[0053] Another embodiment of the present application relates to a document recognition device. The implementation details of the document recognition device in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution. The schematic diagram of the document recognition device in this embodiment can be as Figure 4 shown, including an acquisition module 401, a recognition module 402, a rotation module 403, and an input module 404.
[0054] The acquisition module 401 is used to acquire the text image to be recognized and a prompt word, where the prompt word is used to prompt the information to be recognized in the text image to be recognized;
[0055] The recognition module 402 is used to recognize the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted;
[0056] The rotation module 403 is used to rotate the text image to be recognized based on the angle to be adjusted to obtain a target text image;
[0057] The input module 404 is used to input the target text image and the prompt word into a multi-modal large model to obtain a first text, where the first text is an editable and searchable text.
[0058] In some embodiments, when the device is used to recognize the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted, it is specifically used for: recognizing the text image to be recognized based on a preset orientation recognition algorithm to obtain the current direction of the text in the text image to be recognized; obtaining the angle to be adjusted based on a preset direction and the current direction.
[0059] In some embodiments, the device is further used for: if the current direction is the same as the preset direction, the angle to be adjusted is zero.
[0060] In some embodiments, after the device is used to rotate the text image to be recognized based on the angle to be adjusted to obtain a target text image, the device is further used to: perform layout analysis on the target text image to obtain an intermediate text image; use the intermediate text image as the new target text image; when the device is used to input the target text image and the prompt word into the multi-modal large model to obtain a first text, specifically: input the new target text image and the prompt word into the multi-modal large model to obtain a first text.
[0061] In some embodiments, the layout analysis in the device includes multi-layout analysis, region positioning, table detection, and formula detection.
[0062] In some embodiments, when the device is used to perform layout analysis on the target text image to obtain an intermediate text image, specifically: perform region positioning processing on the target region in the target text image to obtain an intermediate text image, where the intermediate text image contains positioning information for positioning the target region, the size of the intermediate text image is the same as the size of the target text image, and the target region refers to the region that needs to be recognized.
[0063] In some embodiments, when the device is used to input the target text image and the prompt word into the multi-modal large model to obtain a first text, specifically: the multi-modal large model performs recognition processing on the target text image based on the prompt word to recognize text information related to the prompt word; use the text information as the first text.
[0064] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0065] Embodiment Three:
[0066] Another embodiment of the present application relates to an electronic device, as Figure 5 shown, including: at least one processor 901; and a memory 902 communicatively connected to the at least one processor 901; wherein, the memory 902 stores instructions executable by the at least one processor 901, and the instructions are executed by the at least one processor 901 so that the at least one processor 901 can execute the document recognition method in the above embodiments.
[0067] Among them, the memory and the processor are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be one component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0068] The processor is responsible for managing the bus and general processing, and may also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory may be used to store data used by the processor when executing operations.
[0069] Embodiment 4:
[0070] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.
[0071] That is, those skilled in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store program codes.
[0072] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.
Claims
1. A document recognition method, characterized in that, Including: Obtain a text image to be recognized and a prompt word, where the prompt word is used to prompt the information to be recognized in the text image to be recognized; Recognize the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted; Rotate the text image to be recognized based on the angle to be adjusted to obtain a target text image; Input the target text image and the prompt word into a multi-modal large model to obtain a first text, where the first text is an editable and searchable text.
2. The document recognition method according to claim 1, wherein The recognizing the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted includes: Recognize the text image to be recognized based on a preset orientation recognition algorithm to obtain the current direction of the text in the text image to be recognized; Obtain the angle to be adjusted based on a preset direction and the current direction.
3. The document recognition method according to claim 2, characterized in that The method further includes: if the current direction is the same as the preset direction, the angle to be adjusted is zero.
4. The document recognition method according to claim 1, characterized in that After rotating the text image to be recognized based on the angle to be adjusted to obtain a target text image, the method further includes: Perform layout analysis on the target text image to obtain an intermediate text image; Use the intermediate text image as the new target text image; The inputting the target text image and the prompt word into a multi-modal large model to obtain a first text includes: Input the new target text image and the prompt word into a multi-modal large model to obtain a first text.
5. The document recognition method according to claim 4, wherein The layout analysis includes multi-layout analysis, region localization, table detection, and formula detection.
6. The document recognition method according to claim 5, wherein The performing layout analysis on the target text image to obtain an intermediate text image includes: Perform region localization processing on a target region in the target text image to obtain an intermediate text image, where the intermediate text image contains localization information for locating the target region, the size of the intermediate text image is the same as the size of the target text image, and the target region is the region to be recognized.
7. The document recognition method according to claim 1, characterized in that, The inputting the target text image and the prompt word into a multi-modal large model to obtain a first text includes: The multi-modal large model performs recognition processing on the target text image based on the prompt word to recognize text information related to the prompt word; Use the text information as the first text.
8. A document recognition device, characterized in that, Including: An acquisition module, configured to acquire a text image to be recognized and a prompt word, where the prompt word is used to prompt the information to be recognized in the text image to be recognized; A recognition module, configured to recognize the text image to be recognized based on a preset orientation recognition algorithm to obtain an angle to be adjusted; A rotation module, configured to rotate the text image to be recognized based on the angle to be adjusted to obtain a target text image; An input module, configured to input the target text image and the prompt word into a multi-modal large model to obtain a first text, where the first text is an editable and searchable text.
9. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the document recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the document recognition method according to any one of claims 1 to 7.
Citation Information
Cited By
Acquired data processing method and device and medium
CN120612708A
Bill identification method, apparatus and device, and computer program product
CN122200714A