Image processing method, device and system

Through the structure detection model of the two-branch model, the initial model is first trained with unstructured data, and then fine-tuned with structured data, the problem of high training cost of structural detection models in the existing technology is solved, and high-precision image recognition is achieved.

CN114973218BActive Publication Date: 2025-05-16ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110206738.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-24
Publication Date
2025-05-16
Estimated Expiration
2041-02-24

AI Technical Summary

Technical Problem

The training cost of existing structural detection models is high and difficult to effectively reduce.

Method used

The structure detection model of the two-branch model is used to train the initial model by first using a large amount of unstructured annotation data, and then fine-tune the model for training with a small amount of structured annotation data, reducing the training cost.

Benefits of technology

The model training is completed through a small amount of structured processing, which reduces the training cost and improves the recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973218B_ABST
    Figure CN114973218B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, device and system. The method includes: obtaining a text image; using a structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence. The present application solves the technical problem of high training cost of the structure detection model in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular, to an image processing method, device and system. Background Art

[0002] In the information age, data is often not lacking, but what is lacking is structured data. Various manufacturers have a large amount of unstructured data, but these data are often not directly usable. At present, unstructured data can be converted into structured data through annotation, but it requires a lot of manpower and material resources; it is also possible to structure the remaining unstructured data by annotating part of the data for training structured detection algorithms, but training a good structured algorithm model still requires thousands of pictures for each type of data, that is, the training cost of existing structured detection models is high.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present application provide an image processing method, device and system to at least solve the technical problem of high training cost of structure detection models in related technologies.

[0005] According to one aspect of an embodiment of the present application, there is provided an image processing method, comprising: acquiring a text image; recognizing the text image using a structure detection model to obtain a recognition result of the text image, wherein the recognition result comprises: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model comprises: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained sequentially using the first training sample and the second training sample.

[0006] According to another aspect of an embodiment of the present application, there is also provided an image processing method, comprising: displaying a text image; marking a recognition result of the text image on the text image, wherein the recognition result is obtained by recognizing the text image using a structure detection model, and the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0007] According to another aspect of an embodiment of the present application, there is also provided an image processing method, including: obtaining a first training sample and a second training sample; training an initial model using the first training sample to obtain an initial structure detection model; training the initial structure detection model using the second training sample to obtain a structure detection model, wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify a text image and obtain the position of text contained in the text image in the text image, and the second branch model is used to identify a text image and obtain attributes of the text.

[0008] According to another aspect of an embodiment of the present application, there is also provided an image processing method, including: acquiring an ID image; identifying the ID image using a structure detection model to obtain a recognition result of the ID image, wherein the recognition result includes: the attributes of the text contained in the ID image, and the position of the text in the ID image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the ID image and obtain the position of the text in the ID image, the second branch model is used to identify the ID image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0009] According to another aspect of an embodiment of the present application, there is also provided an image processing method, comprising: receiving a text image uploaded by a client; using a structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; outputting the recognition result to the client; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0010] According to another aspect of an embodiment of the present application, an image processing method is also provided, including: receiving a text image by calling a first interface, wherein the first interface includes: a first parameter and a second parameter, the parameter value of the first parameter is the text image, and the parameter value of the second parameter is a target type corresponding to the text image; calling a structure detection model based on the target type, and using the structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; outputting the recognition result by calling a second interface, wherein the second interface includes: a third parameter, and the parameter value of the third parameter is the recognition result; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, and the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0011] According to another aspect of an embodiment of the present application, an image processing device is also provided, including: an acquisition module for acquiring a text image; a recognition module for using a structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0012] According to another aspect of an embodiment of the present application, there is also provided an image processing device, including: a display module, for displaying a text image; a marking module, for marking a recognition result of the text image on the text image, wherein the recognition result is obtained by recognizing the text image using a structure detection model, and the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0013] According to another aspect of an embodiment of the present application, an image processing device is also provided, including: an acquisition module, used to acquire a first training sample and a second training sample; a first training module, used to train an initial model using the first training sample to obtain an initial structure detection model; a second training module, used to train the initial structure detection model using the second training sample to obtain a structure detection model, wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify a text image and obtain the position of the text in the text image, and the second branch model is used to identify a text image and obtain the attributes of the text.

[0014] According to another aspect of an embodiment of the present application, there is also provided an image processing device, including: an acquisition module, used to acquire an ID image; a recognition module, used to recognize the ID image using a structure detection model to obtain a recognition result of the ID image, wherein the recognition result includes: the attributes of the text contained in the ID image, and the position of the text in the ID image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the ID image and obtain the position of the text in the ID image, the second branch model is used to recognize the ID image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0015] According to another aspect of an embodiment of the present application, an image processing device is also provided, including: a receiving module, used to receive a text image uploaded by a client; a recognition module, used to use a structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; an output module, used to output the recognition result to the client; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0016] According to another aspect of an embodiment of the present application, an image processing device is also provided, including: a first calling module, used to receive a text image by calling a first interface, wherein the first interface includes: a first parameter and a second parameter, the parameter value of the first parameter is the text image, and the parameter value of the second parameter is the target type corresponding to the text image; a second calling module, used to call a structure detection model based on the target type, and use the structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; a third calling module, used to output the recognition result by calling the second interface, wherein the second interface includes: a third parameter, and the parameter value of the third parameter is the recognition result; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0017] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned image processing method.

[0018] According to another aspect of an embodiment of the present application, a computer terminal is further provided, including: a memory and a processor, the processor being used to run a program stored in the memory, wherein the above-mentioned image processing method is executed when the program is run.

[0019] According to another aspect of an embodiment of the present application, there is also provided an image processing system, comprising: a processor; and a memory, connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining a text image; recognizing the text image using a structure detection model to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model comprises: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence.

[0020] In an embodiment of the present application, after acquiring a text image, the text image can be identified using a structural detection model to obtain a recognition result of the text image, that is, to obtain the attributes of the text contained in the text image and the position of the text in the text image, thereby achieving the purpose of image recognition. It is easy to notice that the structural detection model includes two branch models, which are used to identify the position of the text in the text image and the attributes of the text respectively. In addition, the structural detection model is obtained by training with the first training sample and the second training sample in sequence, so that the purpose of training the structural detection model can be completed through a small amount of structural processing. Therefore, after acquiring a text image, the text image can be recognized with high precision through the structural detection model, so that the recognition result of the obtained text image is more accurate, achieving the technical effect of reducing the annotation cost of the training sample and improving the recognition accuracy of the structural detection model, thereby solving the technical problem of high training cost of the structural detection model in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0022] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method according to an embodiment of the present application;

[0023] Figure 2 is a flowchart of a first image processing method according to an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of an optional interactive interface according to an embodiment of the present application;

[0025] Figure 4 is a flowchart of an optional image processing method according to an embodiment of the present application;

[0026] Figure 5a is a structural diagram of an optional structural detection model according to an embodiment of the present application;

[0027] Figure 5b is a schematic diagram of an optional training classification branch according to an embodiment of the present application;

[0028] Figure 6 is a flowchart of a second image processing method according to an embodiment of the present application;

[0029] Figure 7 is a flowchart of a third image processing method according to an embodiment of the present application;

[0030] Figure 8 is a flowchart of a fourth image processing method according to an embodiment of the present application;

[0031] Fig. 9 is a flowchart of a fifth image processing method according to an embodiment of the present application;

[0032] Fig.10 is a flowchart of a sixth image processing method according to an embodiment of the present application;

[0033] Fig.11 is a schematic diagram of a first image processing device according to an embodiment of the present application;

[0034] Fig.12 is a schematic diagram of a second image processing device according to an embodiment of the present application;

[0035] Fig.13 is a schematic diagram of a third image processing device according to an embodiment of the present application;

[0036] Fig.14 is a schematic diagram of a fourth image processing device according to an embodiment of the present application;

[0037] Fig.15 is a schematic diagram of a fifth image processing device according to an embodiment of the present application;

[0038] Fig.16 is a schematic diagram of a sixth image processing device according to an embodiment of the present application;

[0039] Fig.17 It is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0042] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:

[0043] Structuring: Organize scattered, isolated information into related, hierarchical information.

[0044] Structured detection: Using detection algorithms, the structured attributes of important objects can be identified while detecting their locations, achieving end-to-end detection and structuring of targets.

[0045] The current scheme is mainly based on matching or data enhancement methods to achieve few samples to save the training cost of the structure detection model. Specifically, the matching-based algorithm first trains a general text detection algorithm to obtain the position of each field, then uses a general text recognition model to identify the text content, and then uses the text position and content to match the template to obtain the attributes of each field to achieve structuring. However, the matching algorithm actually adds rules for matching, which makes matching errors prone to occur, and is very dependent on the accuracy of the detection box and the accuracy of text recognition. Using data enhancement methods such as synthesis or adding noise, a lot of data can be generated for training, but due to the certain differences between synthetic data and real data, some noise synthesis methods cannot be simulated, so the trained structure detection model is not robust, resulting in low image recognition accuracy.

[0046] In order to solve the above problems, the present application provides the following implementation scheme, so that a structure detection model with higher accuracy can be trained at a lower cost, thereby improving the accuracy of image recognition.

[0047] Example 1

[0048] According to an embodiment of the present application, an embodiment of an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0049] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0050] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0051] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the image processing method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned image processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0052] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0053] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0054] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. It should be noted that Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the above-described computer device (or mobile device).

[0055] Under the above operating environment, this application provides Figure 2 The image processing method shown. Figure 2 is a flow chart of a first image processing method according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0056] Step S202, obtaining a text image.

[0057] The text image in the above step may be one or more text images that require text content recognition.

[0058] In an optional embodiment, the text can be directly photographed by a camera, mobile phone, tablet computer, laptop computer or other shooting device to obtain a text image with the text; the screen where the text to be detected is located can also be directly captured by a mobile phone, tablet computer, laptop computer or other terminal device to obtain a text image with the text, wherein only the part of the screen where the text is located can be captured to reduce irrelevant factors in the text image, thereby improving the accuracy of image recognition; one or more text images that need to be recognized can also be directly obtained from the terminal device.

[0059] It should be noted that the text image in the above steps may contain a large amount of text.

[0060] In another optional embodiment, the text image may be an image of an ID card, bank card, business license or other card taken by a camera, or an image of a ticket such as a train ticket, invoice, itinerary or other bill taken by a camera, or an image of a form such as a physical examination form or a logistics form taken by a camera, but is not limited thereto. For example, in an educational scenario, the above-mentioned text image may be an image of a test paper, an image of a student's homework, an image of a teacher's blackboard writing, etc.; in an e-commerce scenario, the above-mentioned text image may be an image of a product poster, a product live broadcast, a video, etc.; in a medical scenario, the above-mentioned text image may be an image of a patient's medical record, an image of a diagnosis, etc.

[0061] In another optional embodiment, a facial image containing multiple faces can be obtained, and the multiple faces in the facial image can be identified to obtain a recognition result of the facial image, wherein the recognition result includes: identity information corresponding to the face, and the position of the face in the facial image, thereby realizing the recognition of multiple faces contained in the facial image and quickly determining the identity information corresponding to the faces at various positions in the facial image. It should be noted that the above-mentioned facial image can contain a large number of faces.

[0062] Step S204: using the structure detection model to recognize the text image to obtain a recognition result of the text image.

[0063] The recognition result includes: the attributes of the characters contained in the text image, and the positions of the characters in the text image.

[0064] Optionally, the structure detection model may include: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, and the second branch model is used to recognize the text image and obtain the attributes of the text.

[0065] The structural detection model in the above steps can use the detection algorithm to detect the location of important objects while identifying the structured attributes of the object, and achieve end-to-end detection and structuring of the target. In order to achieve the purpose of detecting the attributes and the location of the text, the structural detection model can include a feature extraction model, and two branch models connected to the feature extraction model, namely a regression branch model and a classification branch model. The feature extraction model and the regression branch model constitute the above-mentioned first branch model, and the feature extraction model and the classification branch model constitute the above-mentioned second branch model. Among them, the feature extraction model can be used to extract features of the input text image, and the regression branch model can be used to regress the features output by the feature extraction model to obtain the location of the text in the text image. At the same time, the classification branch model can be used to regress the features output by the feature extraction model to obtain the attributes of the text.

[0066] The attributes of the text in the above steps can be the characteristics of the text itself, or the type of text pre-set for different types of text images. For example, the characteristics of the text itself can be text color, text font, text size, text spacing, text structure, text category (such as English, Chinese characters, numbers, symbols, etc.); for example, for an identity card, the type of text can be name, date of birth, address, document number, etc.

[0067] The position of the text in the above steps can be the specific position of the text in the entire text. In the text image, the position of the text can be represented by position coordinates. The lower left corner of the text image can be used as the origin, and a two-dimensional plane coordinate system can be established on the text image based on the origin, with the bottom of the text image as the X-axis and the left side of the text image as the Y-axis, so as to determine the coordinate position of the text in the text image based on the two-dimensional plane coordinate system.

[0068] In an optional embodiment, when the text image is an ID card image, the ID card image is recognized using a structure detection model to obtain the properties of the text in the ID card image, such as name, ID number, validity period, issuing unit and other text information, as well as the position of the text in the ID card image, so that when the user needs to use the ID card information, the user can directly paste the recognition result in the area to be filled in when the ID card information needs to be filled in, without manually entering the information against the ID card; when the text image is an image of a bank card, the bank card image is recognized using a structure detection model to extract the bank card number, bank name and other text information, as well as the position of the text in the bank card, so that when the user needs to fill in the bank card information, the user can directly paste the recognition result in the area to be filled in, without manually entering the information against the bank card, thereby improving the user experience; when the text image is an examination paper image, the examination paper image is recognized using a structure detection model to obtain the properties of the text in the examination paper image, such as name, student number, question, answer, etc., as well as the position of the text in the examination paper image, so that based on the recognition result of the examination paper image, a generated Electronic test papers can be electronicized. Furthermore, for objective questions, the papers can be graded directly based on the recognition results to obtain corresponding scores, which simplifies the marking pressure of teachers and improves the teacher's experience. For subjective questions, teachers can manually grade the papers and obtain corresponding scores. When the text image is a product poster image, the structural detection model is used to recognize the poster image to obtain the attributes of the text in the poster image, such as: product name, merchant, spokesperson, etc., as well as the position of the text in the poster image, so that based on the recognition result of the poster image, a poster template can be generated. board, which makes it convenient for users to use the model to generate posters for their own products without the need for manual generation by users, thereby improving the user experience; when the text image is a patient medical record image, the medical record image can be recognized using the structure detection model to obtain the attributes of the text in the medical record image, such as: patient name, medical card number, chief complaint, current medical history, past history, etc., as well as the position of the text in the medical record image, so that electronic medical records can be generated based on the recognition results of the medical record image, realizing the electronicization of medical records, which is convenient for doctors to have a more comprehensive and detailed understanding of the patient's entire condition and ensure the accuracy of diagnosis.

[0069] The structure detection model is trained in sequence using a first training sample and a second training sample, the first training sample contains unstructured labeled data, the second training sample contains structured labeled data, and the number of the second training samples is less than a preset number.

[0070] The structured annotated data in the above steps refers to organizing scattered and isolated data into related and hierarchical data, such as the attributes of text, but not limited to this.

[0071] The unstructured annotation data in the above steps are essentially all annotation data except structured annotation data. The unstructured annotation data has an internal structure, but is not structured through a predefined digital model or pattern, and is not convenient to represent using a two-dimensional logical table in a database, for example, the position of text in a text image, but not limited to this.

[0072] The first training sample in the above steps can be a text image that has been position-annotated, and the unstructured annotation data can be the position annotation of the text in the text image, for example, the position coordinates of the text in the text image; the second training sample can be a text image that has been position-annotated and attribute-annotated, and the structured annotation data can be the attribute annotation of the text in the text image, for example: the structure, font and size of the text in the text image.

[0073] In an optional embodiment, the text image used in the second training sample can be the same as the text image used in the first training sample, but the first training sample contains unstructured annotation data of the text image, and the second training sample contains unstructured annotation data and structured annotation data of the text image. The first training sample and the second training sample use the same text image to train the structure detection model, which can improve the accuracy of the structure detection model. Specifically, in the first training sample, only the position of the text in the text image can be annotated, while the attributes of the text in the text image are not annotated. In the second training sample, only the attributes of the text in the text image can be annotated, while the position of the text in the text image is not annotated.

[0074] The preset number in the above steps can be set by the user, and can also be the number that the structure detection model determined through multiple experiments can achieve a better recognition effect. That is, the number can be as small as possible on the basis of ensuring the accuracy of the structure detection model, thereby reducing the labeling cost.

[0075] In an embodiment of the present application, the number of first training samples can be much larger than the number of second training samples; since text images with unstructured annotation data are relatively easy to obtain, that is, the acquisition cost is relatively low, therefore, the structure detection model can be first trained with a large number of text images with unstructured annotation data, thereby obtaining a general text detection algorithm to achieve accurate positioning of text in text images; and the cost of acquiring text images with structured annotation data is relatively high, therefore, a small number of text images with structured annotation data can be used to fine-tune the structure detection model so that the structure detection model can classify the text in the text images.

[0076] Furthermore, since the structure detection model has been previously trained with a large number of first training samples containing unstructured annotated data, the structure detection model is already robust enough and will not suffer from low recognition accuracy due to too few second training samples containing structured annotated data. Therefore, re-classifying the structure detection model with a small number of second training samples can reduce the cost of training the model while ensuring the accuracy of the structure detection model, and can also improve the simplicity of the training model, thereby making it convenient for users to train the structure detection model using a small number of second training samples with structured annotated data.

[0077] In another optional embodiment, the text image in the second training sample may be completely different from the text image in the first training sample. Specifically, the user may train the structure detection model trained using the first training sample, and then use the relevant text images that the user actually wants to recognize for training according to the user's needs. This can reduce the training cost of the structured training model while specifically improving the accuracy of the text in the text images that the user needs to recognize.

[0078] For example, when the user needs to use the structure detection model mainly to recognize text in bills, the user can first use a large number of first training samples to train the structure detection model, and then use the bills containing structured annotation data as the second training samples, and use a small number of second training samples to fine-tune the structure detection model, so that the structure detection model can achieve the highest accuracy when classifying the text in the bills. At the same time, since only a small number of bills with structured annotation data are used, the cost of training the structure detection model can be greatly reduced.

[0079] It should be noted that in order to reduce the local computing pressure of the client, the structure detection model can be deployed in the cloud server, and the cloud server provides services to the outside. The cloud server can receive the model training request sent by the client, and obtain the corresponding first training sample and second training sample for training. After the structure detection model training is completed, the client can send the text image that needs to be recognized to the cloud server for processing. The cloud server calls the structure detection model to process the received text image, and returns the recognition result to the client, which is displayed to the user by the client. Figure 3As shown, the client can provide the user with an interactive interface. The user can select the text image to be recognized by clicking the "Select Image" button. The selected text image can be displayed in the "Image Display" area. After the user confirms that the text image is correct, the selected text image can be uploaded to the cloud server for recognition by clicking the "Upload" button. After the cloud server returns the recognition result to the client, the recognition result can be displayed in the "Image Display" area for the user to view.

[0080] In addition, if the user is not satisfied with the recognition result, or the user believes that the recognition result is wrong, the user can modify or edit the recognition result directly on the client, and feedback the corrected recognition result to the cloud server through the client, so that the cloud server can adjust and update the structure detection model based on the user's feedback result. Specifically, the text image and the feedback result uploaded by the user can be used to construct a second training sample, and the structure detection model can be trained using the newly constructed second training sample to further improve the recognition accuracy of the structure detection model.

[0081] Through the scheme provided by the above-mentioned embodiment of the present application, after acquiring the text image, the text image can be identified by using the structure detection model to obtain the recognition result of the text image, that is, the attributes of the text contained in the text image and the position of the text in the text image are obtained, thereby achieving the purpose of image recognition. It is easy to notice that the structure detection model includes two branch models, which are respectively used to identify the position of the text in the text image and the attributes of the text. In addition, the structure detection model is obtained by training with the first training sample and the second training sample in sequence, so that the purpose of training the structure detection model can be completed by a small amount of structured processing. Therefore, after acquiring the text image, the text image can be recognized with high precision by the structure detection model, so that the recognition result of the obtained text image is more accurate, achieving the technical effect of reducing the annotation cost of the training sample and improving the recognition accuracy of the structure detection model, thereby solving the technical problem of the relatively high training cost of the structure detection model in the related technology.

[0082] In the above embodiment of the present application, a text image is recognized using a structure detection model, and the recognition result of the text image is obtained, including: inputting the text image into a feature extraction model to obtain feature information of the text image; inputting the feature information of the text image into a regression branch model to determine the position of the text in the text image; inputting the feature information of the text image into a classification branch model to determine the attributes of the text.

[0083] The characteristic information in the above steps may be text features that can be distinguished from other patterns, such as the shape, size, color, etc. of the text.

[0084] The feature extraction model in the above steps may be a network that can extract features related to text in a text image; the regression branch model in the above steps may be a network that can locate text in a text image; the classification branch model in the above steps may be a network that can classify text in a text image. It should be noted that the specific types and network structures of the feature extraction model, regression branch model, and classification branch model can be implemented using existing solutions, and this application does not make specific limitations on this. For example, the feature extraction model may be VGG (Visual Geometry Group), Shuffle Net (lightweight neural network), etc.

[0085] In an optional embodiment, the text image is input into a feature extraction model, and any data (text or image) in the text image can be converted into numerical features that can be used for machine learning, that is, feature information, and the feature information related to the text is extracted. The feature information related to the text can be a feature vector or feature sequence of the text.

[0086] In another optional embodiment, feature information of the text image may be input into a regression branch model, and the regression branch model may be used to locate the position of the text in the text image.

[0087] Furthermore, after determining the position of the text in the text image, the located text can be marked using a preset text box; wherein the size, shape and tilt angle of the preset text box can be adaptively adjusted according to the size, shape, tilt angle and arrangement of the text.

[0088] In yet another optional embodiment, the feature information of the text image may be classified by a classification branch model to determine the attributes of the text corresponding to the feature information.

[0089] In the above embodiment of the present application, the method also includes: obtaining a first training sample and a second training sample, wherein the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image; using the first training sample to train the initial model to obtain an initial structure detection model; using the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0090] In the above steps, the plurality of second text images may be extracted from the plurality of first text images, or may be images provided by the user according to detection requirements.

[0091] In an optional embodiment, the marking position of the first text in the first text image can be marked by a preset text box, wherein the size, color, shape, thickness of the preset text box can be adaptively adjusted according to the size, color, shape, thickness, etc. of the text. The text blocks composed of adjacent first texts in the first text image can be uniformly marked using the preset text box, or the first texts in the same row or column in the first text image can be uniformly marked using the preset text box.

[0092] Since the first text image with the first text annotation position is relatively easy to obtain, a large number of first text images can be used to train the initial model first, so that the initial structure detection model can be exposed to a large number of first text images, so that the trained initial structure detection model has the ability to accurately detect text, thereby improving the initial structure detection model's recognition accuracy of text positions in text images.

[0093] It should be noted that, in the process of training with the first training sample, the classification results output by the classification branch model can be assumed to meet the training requirements, or the attributes of all characters in the first training sample can be assumed to be the same. Therefore, the training process of the first training sample does not affect the parameters of the classification branch model.

[0094] In another optional embodiment, the annotation attributes of the second text in the second text image can be marked on the text image by annotation. For example, a second text image can record the names of multiple fruits: such as bananas, apples, oranges, etc., then the font color, structure, size, etc. of the fruit name can be marked on each fruit name by text annotation.

[0095] Since the acquisition cost of the second text image with the second text annotation attribute is relatively high, therefore, under the premise that the initial structure detection model has achieved a very high recognition accuracy for the text position in the text image, a small amount of second text images with the second text annotation attribute can be used to train the initial structure detection model, which can also be called fine-tuning of the initial structure detection model, so as to further improve the recognition accuracy of the structure detection model when recognizing text images, and at the same time, it can also reduce the training cost of training the structure detection model.

[0096] The network parameters in the above steps may be weight parameters in each network in the initial structure detection model.

[0097] In yet another optional embodiment, in the process of training the initial structure detection model using the second training sample, the weights of the previously trained feature extraction model and regression branch model remain fixed.

[0098] In the above embodiment of the present application, obtaining the second training sample includes: obtaining a plurality of second text images; and processing the plurality of second text images using a data enhancement algorithm to generate a second training sample.

[0099] The data enhancement algorithm in the above steps can change the multiple second text images in the second training sample so as to make the generalization ability of the structure detection model stronger.

[0100] In an optional embodiment, the second text image is processed by a data enhancement algorithm, which may be rotation, flipping, scaling, translation, scale, contrast, noise disturbance, color change, etc. The second text image is processed by a data enhancement algorithm, so that the number of second text images can be greatly increased, so that the generated second training samples can train the recognition ability of the structure detection model more accurately. At the same time, using fewer second text images to generate second training samples can reduce the cost required for training the structure detection model.

[0101] In the above embodiment of the present application, the initial model is trained using the first training sample to obtain the initial structure detection model, including: inputting each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; inputting the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; inputting the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; based on the predicted position and annotation position of the first text, as well as the classification result, the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model, and the classification branch model of the initial model are updated to obtain the initial structure detection model.

[0102] In an optional embodiment, each first text image can be input into a feature extraction model, and feature information about the first text in the first text image can be extracted using the feature extraction model. The feature information about the first text in each first text image can be input into a regression branch model, and the position of the first text in the first text image can be predicted based on the feature information of the first text, that is, the predicted position of the first text by the regression branch model. At this time, a position error can be determined based on the predicted position of the first text and the marked position of the first text in the first text image, and the network parameters in the feature extraction model, the regression branch model and the classification branch model can be updated based on the position error.

[0103] Furthermore, the classification branch model can be used to determine whether the first character is present at the current position where the first character extracted by the feature extraction model is located. If the classification branch model outputs a classification result for the current position, it is determined that the first character is present at the current position; if the classification branch model does not output a classification result for the current position, it is determined that the first character is not present at the current position. At this time, the network parameters in the structure detection model can be adjusted according to the classification result of the classification branch model, thereby improving the recognition accuracy of the structure detection model.

[0104] In the above embodiment of the present application, the initial structure detection model is trained using the second training sample to obtain the structure detection model, including: inputting each second text image into the feature extraction model of the initial structure detection model to obtain feature information of each second text image; inputting the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; based on the annotation attribute and prediction attribute of the second text, updating the network parameters of the classification branch model of the initial structure detection model to obtain the structure detection model.

[0105] In an optional embodiment, each second text image can be input into a feature extraction model, and feature information about the second text in the second text image can be extracted using the feature extraction model. The feature information about the second text in each second text image can be input into a regression branch model. The attribute of the second text in the second text image can be predicted based on the feature information of the second text, that is, the predicted attribute of the second text by the classification branch model. At this time, an attribute error can be determined based on the predicted attribute of the second text and the annotation attribute of the second text in the second text image, and the network parameters in the classification branch model can be updated based on the attribute error.

[0106] In the above embodiment of the present application, after the text image is recognized using the structure detection model to obtain the recognition result of the text image, the method also includes: determining the confidence level corresponding to the recognition result based on the recognition result of the text image; determining the target labeling method of the recognition result based on the confidence level; and outputting the recognition result according to the target labeling method.

[0107] In an optional embodiment, the structure detection model can give a recognition probability of the recognition result during the process of recognizing the text image, and then the recognition probability can be used as the above-mentioned confidence. For example, for the structure detection model, the position of the text in the text image and the attributes of the text can be identified. Therefore, the recognition probabilities of the two recognition results can be weighted as the confidence corresponding to the final recognition result.

[0108] Since the higher the confidence level, the more accurate the recognition result, in order to facilitate users to determine the recognition accuracy of text images, multiple confidence intervals can be set in advance, and different marking methods can be set for different confidence intervals. For example, for confidence intervals with lower confidence levels, they can be marked with highlight colors, flashing, and other marking methods, and for confidence intervals with higher confidence levels, they can be marked with regular colors, so that users can more easily notice recognition results with lower confidence levels and confirm the recognition results.

[0109] In another optional embodiment, after determining the confidence level corresponding to the recognition result, the confidence level interval to which the confidence level belongs can be determined, and the annotation method corresponding to the confidence level interval is used as the target annotation method of the recognition result, and then the recognition result can be annotated in the text image according to the target annotation method to achieve the purpose of outputting the recognition result. For example, in Figure 3 The text image is displayed in the "Image Display" area shown, and the recognition result is annotated according to the target annotation method.

[0110] In the above embodiment of the present application, after outputting the recognition result, the method further includes: receiving response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and updating the structure detection model based on the response data.

[0111] In an optional embodiment, after viewing the recognition result, the user can confirm the recognition result. If the user confirms that there is an error in the recognition result, the recognition result can be directly modified to obtain the above-mentioned response data, and the response data is returned to the cloud server. The cloud server adjusts the structure detection model based on the response data to further improve the recognition accuracy of the structure detection model.

[0112] It should be noted that since the recognition result includes the position and attributes of the text, when the response data includes the modified position, a new training sample can be constructed based on the text image and the modified position, and used as the first training sample to update the structure detection model; when the response data includes the modified attribute, a new training sample can be constructed based on the text image and the modified attribute, and used as the second training sample to update the structure detection model.

[0113] In the above embodiment of the present application, updating the structure detection model based on the response data includes: generating a new second training sample based on the response data; and training the structure detection model using the new second training sample to obtain an updated structure detection model.

[0114] In an optional embodiment, in order to achieve the purpose of updating the structure detection model, a new training sample can be generated as a second training sample based on the text image and the response data, and the structure detection model can be trained according to the training process of the second training sample to achieve the purpose of updating the structure detection model.

[0115] Combine the following Figure 4 , Figure 5a and Figure 5b A preferred embodiment of the present application is described in detail. Figure 4 As shown, the method can be executed by a front-end client or a back-end server. In the embodiment of the present application, the cloud server is used as an example for explanation. The method includes the following steps:

[0116] Step S41, using a large amount of text detection data sets to train a robust and universal text detection algorithm, wherein the text detection algorithm includes: a backbone (feature extraction model), a regression branch model and a classification branch model;

[0117] like Figure 5a As shown in the model structure diagram, after the image is extracted through backbone features, one branch is used to regress the text position, and another branch is used to determine whether the current position is text.

[0118] Step S42, a small number of structured samples with labeled text positions and attributes are enhanced by noise, blurring, resizing (image size conversion), etc., to obtain a batch of data that can be used to fine-tune the model;

[0119] Step S43, fixing the network parameters of the backbone and regression branch models trained in step S41, and using the network parameters of the classification branch model of the structured sample fintune model produced in step S42;

[0120] like Figure 5b As shown, in the process of training with structured samples, the network parameters in the trained backbone and regression branch models can be fixed and remain unchanged.

[0121] The above steps can enable the model to judge the attributes of each text block.

[0122] In step S44, a relatively small learning rate is used to fintune the backbone, regression branch model, and classification branch model as a whole to obtain an accurate structure detection model.

[0123] Through the above steps, the present application can use a general text detection data set to train the model so that the model has the ability to accurately detect text, and because the data set is large enough, the model sees enough data sets, so the model is robust enough and will not be overly sensitive to noise due to the small size of the structured data used for fine-tuning, thereby causing false detection and missed detection. In addition, compared with the existing training process that uses a matching algorithm with added rules for matching, which makes matching errors prone to occur and is very dependent on the accuracy of the detection box and the accuracy of text recognition, this application does not use any rule parsing, only uses algorithms for training, and uses an end-to-end model to avoid losses caused by errors in text recognition, rule matching, etc.

[0124] Example 2

[0125] According to an embodiment of the present application, an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0126] Figure 6 FIG. 1 is a flow chart of a second image processing method according to an embodiment of the present application. Figure 6 As shown, the method may include the following steps:

[0127] Step S602: display the text image.

[0128] In an optional embodiment, the text image can be displayed in the operation interface on the display screen of the mobile terminal, or the text image can be displayed in the operation interface on the display screen of the computer terminal. Figure 3 The text image is displayed in the "Image Display" area of ​​the interactive interface shown.

[0129] Step S604: marking the recognition result of the text image on the text image.

[0130] Among them, the recognition result is obtained by using a structure detection model to recognize a text image, and the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, the structure detection model is trained using a first training sample and a second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than a preset number.

[0131] In an optional embodiment, the recognition result of the text image may be marked by annotation, text box, etc.

[0132] Exemplarily, the position of the text in the text image may be marked by a preset text box; the properties of the text contained in the text image may be marked by annotation.

[0133] In the above embodiment of the present application, a text image is recognized using a structure detection model, and the recognition result of the text image is obtained, including: inputting the text image into a feature extraction model to obtain feature information of the text image; inputting the feature information of the text image into a regression branch model to determine the position of the text in the text image; inputting the feature information of the text image into a classification branch model to determine the attributes of the text.

[0134] In the above embodiment of the present application, the method also includes: obtaining a first training sample and a second training sample, wherein the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image; using the first training sample to train the initial model to obtain an initial structure detection model; using the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0135] In the above embodiment of the present application, obtaining the second training sample includes: obtaining a plurality of second text images; and processing the plurality of second text images using a data enhancement algorithm to generate a second training sample.

[0136] In the above embodiment of the present application, the initial model is trained using the first training sample to obtain the initial structure detection model, including: inputting each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; inputting the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; inputting the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; based on the predicted position and annotation position of the first text, as well as the classification result, the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model, and the classification branch model of the initial model are updated to obtain the initial structure detection model.

[0137] In the above embodiment of the present application, the initial structure detection model is trained using the second training sample to obtain the structure detection model, including: inputting each second text image into the feature extraction model of the initial structure detection model to obtain feature information of each second text image; inputting the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; based on the annotation attribute and prediction attribute of the second text, updating the network parameters of the classification branch model of the initial structure detection model to obtain the structure detection model.

[0138] In the above embodiment of the present application, marking the recognition result of the text image on the text image includes: determining the confidence level corresponding to the recognition result based on the recognition result of the text image; determining the target labeling method of the recognition result based on the confidence level; and marking the recognition result of the text image on the text image according to the target labeling method.

[0139] In the above embodiment of the present application, after marking the recognition result of the text image on the text image, the method also includes: receiving response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and updating the structure detection model based on the response data.

[0140] In the above embodiment of the present application, updating the structure detection model based on the response data includes: generating a new second training sample based on the response data; and training the structure detection model using the new second training sample to obtain an updated structure detection model.

[0141] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0142] Example 3

[0143] According to an embodiment of the present application, an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0144] Figure 7 FIG. 1 is a flowchart of a third image processing method according to an embodiment of the present application. Figure 7 As shown, the method may include the following steps:

[0145] Step S702: Acquire a first training sample and a second training sample.

[0146] The first training sample includes unstructured labeled data, the second training sample includes structured labeled data, and the number of the second training samples is less than a preset number.

[0147] Step S704: train the initial model using the first training sample to obtain an initial structure detection model.

[0148] Step S706, use the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text contained in the text image in the text image, and the second branch model is used to recognize the text image and obtain the attributes of the text.

[0149] In the above embodiment of the present application, the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image.

[0150] In the above embodiment of the present application, obtaining the second training sample includes: obtaining a plurality of second text images; and processing the plurality of second text images using a data enhancement algorithm to generate a second training sample.

[0151] In the above embodiment of the present application, the initial model is trained using the first training sample to obtain the initial structure detection model, including: inputting each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; inputting the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; inputting the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; based on the predicted position and annotation position of the first text, as well as the classification result, the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model, and the classification branch model of the initial model are updated to obtain the initial structure detection model.

[0152] In the above embodiment of the present application, during the training process using the second training sample, the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged.

[0153] In the above embodiment of the present application, the initial structure detection model is trained using the second training sample to obtain the structure detection model, including: inputting each second text image into the feature extraction model of the initial structure detection model to obtain feature information of each second text image; inputting the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; based on the annotation attribute and prediction attribute of the second text, updating the network parameters of the classification branch model of the initial structure detection model to obtain the structure detection model.

[0154] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0155] Example 4

[0156] According to an embodiment of the present application, an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0157] Figure 8 FIG. 4 is a flowchart of a fourth image processing method according to an embodiment of the present application. Figure 8 As shown, the method may include the following steps:

[0158] Step S802, obtaining a document image.

[0159] The document image in the above step can be an image of various cards or tickets, for example, it can be an image of an ID card, bank card, business license or other card, or it can be an image of a ticket such as a train plate, invoice, itinerary, etc., but is not limited thereto.

[0160] Step S804: Use the structure detection model to identify the document image to obtain a recognition result of the document image.

[0161] The recognition result includes: the attributes of the text contained in the document image, and the position of the text in the document image.

[0162] Optionally, the structure detection model may include: a first branch model and a second branch model, the first branch model is used to identify the document image and obtain the position of the text in the document image, and the second branch model is used to identify the document image and obtain the attributes of the text.

[0163] The structural detection model in the above steps can use the detection algorithm to detect the location of important objects and identify the structured attributes of the object at the same time, and realize the detection and structuring of the target end-to-end. In order to achieve the purpose of detecting the attributes and the location of the text, the structural detection model can include a feature extraction model, and two branch models connected to the feature extraction model, which are a regression branch model and a classification branch model. The feature extraction model and the regression branch model constitute the above-mentioned first branch model, and the feature extraction model and the classification branch model constitute the above-mentioned second branch model. Among them, the feature extraction model can be used to extract features of the input document image, and the regression branch model can be used to regress the features output by the feature extraction model to obtain the location of the text in the document image. At the same time, the classification branch model can be used to regress the features output by the feature extraction model to obtain the attributes of the text.

[0164] The structure detection model is trained in sequence using a first training sample and a second training sample, the first training sample contains unstructured labeled data, the second training sample contains structured labeled data, and the number of the second training samples is less than a preset number.

[0165] In the above embodiment of the present application, after the ID image is recognized using the structure detection model to obtain the recognition result of the ID image, the method also includes: determining the target format of the ID image; and generating text data corresponding to the ID image based on the target format and the recognition result.

[0166] In an optional embodiment, different types of cards or bills are often typeset in different formats or standards, so the target format of the document image can be determined based on the type of the document image. After the target format is determined, the corresponding text can be typeset based on the target format to obtain the final electronic text, that is, the above-mentioned text data, thereby realizing the electronicization of the card or bill.

[0167] In the above embodiment of the present application, the structure detection model is used to identify the ID image, and the identification result of the ID image is obtained, which includes: inputting the ID image into the feature extraction model to obtain the feature information of the ID image; inputting the feature information of the ID image into the regression branch model to determine the position of the text in the ID image; inputting the feature information of the ID image into the classification branch model to determine the attributes of the text.

[0168] In the above embodiment of the present application, the method also includes: obtaining a first training sample and a second training sample, wherein the first training sample includes: a plurality of first ID images, and the annotation position of the first text contained in each first ID image, and the second training sample includes: a plurality of second ID images, and the annotation attributes of the second text contained in each second ID image; using the first training sample to train the initial model to obtain an initial structure detection model; using the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0169] In the above embodiment of the present application, obtaining the second training sample includes: obtaining a plurality of second document images; and processing the plurality of second document images using a data enhancement algorithm to generate a second training sample.

[0170] In the above embodiment of the present application, the initial model is trained using the first training sample to obtain the initial structure detection model, including: inputting each first ID document image into the feature extraction model of the initial model to obtain feature information of each first ID document image; inputting the feature information of each first ID document image into the regression branch model of the initial model to determine the predicted position of the first text in each first ID document image; inputting the feature information of each first ID document image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; based on the predicted position and annotation position of the first text, as well as the classification result, the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model, and the classification branch model of the initial model are updated to obtain the initial structure detection model.

[0171] In the above embodiment of the present application, the initial structure detection model is trained using the second training sample to obtain the structure detection model, which includes: inputting each second ID document image into the feature extraction model of the initial structure detection model to obtain feature information of each second ID document image; inputting the feature information of each second ID document image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; based on the annotation attribute and prediction attribute of the second text, updating the network parameters of the classification branch model of the initial structure detection model to obtain the structure detection model.

[0172] In the above embodiment of the present application, after the ID image is identified using the structure detection model to obtain the identification result of the ID image, the method also includes: determining the confidence level corresponding to the identification result based on the identification result of the ID image; determining the target labeling method of the identification result based on the confidence level; and outputting the identification result according to the target labeling method.

[0173] In the above embodiment of the present application, after outputting the recognition result, the method further includes: receiving response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and updating the structure detection model based on the response data.

[0174] In the above embodiment of the present application, updating the structure detection model based on the response data includes: generating a new second training sample based on the response data; and training the structure detection model using the new second training sample to obtain an updated structure detection model.

[0175] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0176] Example 5

[0177] According to an embodiment of the present application, an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0178] Fig. 9 is a flowchart of a fifth image processing method according to an embodiment of the present application. Fig. 9 As shown, the method may include the following steps:

[0179] Step S902: receiving a text image uploaded by the client.

[0180] The client in the above steps can be a mobile terminal such as a smart phone (such as an Android phone, an iOS phone), a tablet computer, a PDA, or a computer terminal such as a laptop computer, a personal computer, but is not limited thereto.

[0181] Step S904: using the structure detection model to recognize the text image and obtain a recognition result of the text image.

[0182] The recognition result includes: the attributes of the characters contained in the text image, and the positions of the characters in the text image.

[0183] The structural detection model in the above steps can use the detection algorithm to detect the location of important objects while identifying the structured attributes of the object, and achieve end-to-end detection and structuring of the target. In order to achieve the purpose of detecting the attributes and the location of the text, the structural detection model can include a feature extraction model, and two branch models connected to the feature extraction model, namely a regression branch model and a classification branch model. The feature extraction model and the regression branch model constitute the above-mentioned first branch model, and the feature extraction model and the classification branch model constitute the above-mentioned second branch model. Among them, the feature extraction model can be used to extract features of the input text image, and the regression branch model can be used to regress the features output by the feature extraction model to obtain the location of the text in the text image. At the same time, the classification branch model can be used to regress the features output by the feature extraction model to obtain the attributes of the text.

[0184] The structure detection model is trained in sequence using a first training sample and a second training sample, the first training sample contains unstructured labeled data, the second training sample contains structured labeled data, and the number of the second training samples is less than a preset number.

[0185] Step S906: output the recognition result to the client.

[0186] In the above embodiment of the present application, a text image is recognized using a structure detection model, and the recognition result of the text image is obtained, including: inputting the text image into a feature extraction model to obtain feature information of the text image; inputting the feature information of the text image into a regression branch model to determine the position of the text in the text image; inputting the feature information of the text image into a classification branch model to determine the attributes of the text.

[0187] In the above embodiment of the present application, the method also includes: obtaining a first training sample and a second training sample, wherein the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image; using the first training sample to train the initial model to obtain an initial structure detection model; using the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0188] In the above embodiment of the present application, obtaining the second training sample includes: obtaining a plurality of second text images; and processing the plurality of second text images using a data enhancement algorithm to generate a second training sample.

[0189] In the above embodiment of the present application, the initial model is trained using the first training sample to obtain the initial structure detection model, including: inputting each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; inputting the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; inputting the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; based on the predicted position and annotation position of the first text, as well as the classification result, the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model, and the classification branch model of the initial model are updated to obtain the initial structure detection model.

[0190] In the above embodiment of the present application, the initial structure detection model is trained using the second training sample to obtain the structure detection model, including: inputting each second text image into the feature extraction model of the initial structure detection model to obtain feature information of each second text image; inputting the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; based on the annotation attribute and prediction attribute of the second text, updating the network parameters of the classification branch model of the initial structure detection model to obtain the structure detection model.

[0191] In the above embodiment of the present application, after the text image is recognized using the structure detection model to obtain the recognition result of the text image, the method also includes: determining the confidence level corresponding to the recognition result based on the recognition result of the text image; determining the target labeling method of the recognition result based on the confidence level; and outputting the recognition result according to the target labeling method.

[0192] In the above embodiment of the present application, after outputting the recognition result, the method further includes: receiving response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and updating the structure detection model based on the response data.

[0193] In the above embodiment of the present application, updating the structure detection model based on the response data includes: generating a new second training sample based on the response data; and training the structure detection model using the new second training sample to obtain an updated structure detection model.

[0194] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0195] Example 6

[0196] According to an embodiment of the present application, an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0197] Fig.10 is a flowchart of a sixth image processing method according to an embodiment of the present application. Fig.10 As shown, the method may include the following steps:

[0198] Step S1002: receiving a text image by calling a first interface, wherein the first interface includes: a first parameter and a second parameter, the parameter value of the first parameter is the text image, and the parameter value of the second parameter is the target type corresponding to the text image.

[0199] The first interface in the above steps may be an interface for data interaction between the cloud server and the client. The client may pass the text image and target type into the interface function as a parameter of the interface function to achieve the purpose of uploading the text image to the cloud server.

[0200] The target type in the above steps can be the type of text content corresponding to the text image, for example, it can be a card such as an ID card, bank card, business license, or a ticket such as a train ticket, invoice, itinerary, or a form such as a physical examination form or logistics form, but is not limited to this. For example, in an educational scenario, the target type can be a test paper, student homework, teacher's blackboard writing, etc.; in an e-commerce scenario, the target type can be a product poster, product live broadcast, video, etc.; in a medical scenario, the target type can be a patient's medical record, diagnosis, etc.

[0201] In an optional embodiment, the user can directly upload the text image and specify the target type of the text image, so that the cloud server can directly obtain the text image and the target type by calling the first interface. In another optional embodiment, the user can upload the storage path and target type of the text image, so that the cloud server can obtain the storage path and target type by calling the first interface, and then obtain the text image from the storage path. In another optional embodiment, the user can directly upload the text image, and after the cloud server obtains the text image through the first interface, it can identify the text image and determine the target type.

[0202] Step S1004: calling a structure detection model based on the target type, and using the structure detection model to recognize the text image to obtain a recognition result of the text image.

[0203] The recognition result includes: the attributes of the characters contained in the text image, and the positions of the characters in the text image.

[0204] For different target types, different structure detection models can be pre-trained to make the recognition of the structure detection model more targeted and more accurate. In an optional embodiment, different structure detection models can be pre-deployed in the cloud server to recognize different types of text images. Therefore, after determining the target type of the text image, the structure detection model corresponding to the target type can be called, and the text image can be recognized using the structure detection model to obtain a recognition result.

[0205] The structural detection model in the above steps can use the detection algorithm to detect the location of important objects while identifying the structured attributes of the object, and achieve end-to-end detection and structuring of the target. In order to achieve the purpose of detecting the attributes and the location of the text, the structural detection model can include a feature extraction model, and two branch models connected to the feature extraction model, namely a regression branch model and a classification branch model. The feature extraction model and the regression branch model constitute the above-mentioned first branch model, and the feature extraction model and the classification branch model constitute the above-mentioned second branch model. Among them, the feature extraction model can be used to extract features of the input text image, and the regression branch model can be used to regress the features output by the feature extraction model to obtain the location of the text in the text image. At the same time, the classification branch model can be used to regress the features output by the feature extraction model to obtain the attributes of the text.

[0206] The structure detection model is trained in sequence using a first training sample and a second training sample, the first training sample contains unstructured labeled data, the second training sample contains structured labeled data, and the number of the second training samples is less than a preset number.

[0207] Step S1006: output the recognition result by calling the second interface, wherein the second interface includes: a third parameter, and the parameter value of the third parameter is the recognition result.

[0208] The second interface in the above steps may be an interface for data interaction between the cloud server and the client. The cloud server may pass the recognition result into the interface function as a parameter of the interface function to achieve the purpose of sending the recognition result to the client.

[0209] In the above embodiment of the present application, a text image is recognized using a structure detection model, and the recognition result of the text image is obtained, including: inputting the text image into a feature extraction model to obtain feature information of the text image; inputting the feature information of the text image into a regression branch model to determine the position of the text in the text image; inputting the feature information of the text image into a classification branch model to determine the attributes of the text.

[0210] In the above embodiment of the present application, the method also includes: obtaining a first training sample and a second training sample, wherein the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image; using the first training sample to train the initial model to obtain an initial structure detection model; using the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0211] In the above embodiment of the present application, obtaining the second training sample includes: obtaining a plurality of second text images; and processing the plurality of second text images using a data enhancement algorithm to generate a second training sample.

[0212] In the above embodiment of the present application, the initial model is trained using the first training sample to obtain the initial structure detection model, including: inputting each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; inputting the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; inputting the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; based on the predicted position and annotation position of the first text, as well as the classification result, the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model, and the classification branch model of the initial model are updated to obtain the initial structure detection model.

[0213] In the above embodiment of the present application, the initial structure detection model is trained using the second training sample to obtain the structure detection model, including: inputting each second text image into the feature extraction model of the initial structure detection model to obtain feature information of each second text image; inputting the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; based on the annotation attribute and prediction attribute of the second text, updating the network parameters of the classification branch model of the initial structure detection model to obtain the structure detection model.

[0214] In the above embodiment of the present application, after the text image is recognized using the structure detection model to obtain the recognition result of the text image, the method also includes: determining the confidence level corresponding to the recognition result based on the recognition result of the text image; determining the target labeling method of the recognition result based on the confidence level; and outputting the recognition result according to the target labeling method.

[0215] In the above embodiment of the present application, after outputting the recognition result, the method further includes: receiving response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and updating the structure detection model based on the response data.

[0216] In the above embodiment of the present application, updating the structure detection model based on the response data includes: generating a new second training sample based on the response data; and training the structure detection model using the new second training sample to obtain an updated structure detection model.

[0217] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0218] Example 7

[0219] According to an embodiment of the present application, an image processing device is also provided. Fig.11 As shown, the device 1100 includes: an acquisition module 1102 and an identification module 1104 .

[0220] Among them, the acquisition module 1102 is used to acquire the text image; the recognition module 1104 is used to use the structure detection model to recognize the text image to obtain the recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model may include: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than the preset number.

[0221] It should be noted that the acquisition module 1102 and the identification module 1104 correspond to steps S202 to S204 in Example 1, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0222] In the above embodiment of the present application, the identification module includes: a first acquisition unit, a first determination unit, and a second determination unit.

[0223] Among them, the first acquisition unit is used to input the text image into the feature extraction model to obtain the feature information of the text image; the first determination unit is used to input the feature information of the text image into the regression branch model to determine the position of the text in the text image; the second determination unit is used to input the feature information of the text image into the classification branch model to determine the attributes of the text.

[0224] In the above embodiments of the present application, the device further includes: a first training module and a second training module.

[0225] Among them, the acquisition module is also used to obtain a first training sample and a second training sample, wherein the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image; the first training module is used to use the first training sample to train the initial model to obtain an initial structure detection model; the second training module is used to use the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0226] In the above embodiments of the present application, the acquisition module includes: a second acquisition unit and a processing unit.

[0227] The second acquisition unit is used to acquire a plurality of second text images; and the processing unit is used to process the plurality of second text images using a data enhancement algorithm to generate second training samples.

[0228] In the above embodiment of the present application, the first training module includes: a third acquisition unit, a third determination unit, a fourth determination unit and a first updating unit.

[0229] Among them, the third acquisition unit is used to input each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; the third determination unit is used to input the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; the fourth determination unit is used to input the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; the update unit is used to update the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model and the classification branch model of the initial model based on the predicted position and annotation position of the first text, as well as the classification result, to obtain the initial structure detection model.

[0230] In the above embodiment of the present application, the second training module includes: a fourth acquisition unit, a fifth determination unit, and a second updating unit.

[0231] Among them, the fourth acquisition unit is used to input each second text image into the feature extraction model of the initial structure detection model to obtain the feature information of each second text image; the fifth determination unit is used to input the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; the second update unit is used to update the network parameters of the classification branch model of the initial structure detection model based on the annotation attributes and prediction attributes of the second text to obtain the structure detection model.

[0232] In the above embodiment of the present application, the device further includes: a first determination module, a second determination module and an output module.

[0233] Among them, the first determination module is used to determine the confidence corresponding to the recognition result based on the recognition result of the text image; the second determination module is used to determine the target labeling method of the recognition result based on the confidence; and the output module is used to output the recognition result according to the target labeling method.

[0234] In the above embodiments of the present application, the device further includes: a receiving module and an updating module.

[0235] The receiving module is used to receive response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and the updating module is used to update the structure detection model based on the response data.

[0236] In the above embodiments of the present application, the update module includes: a generation unit and a training unit.

[0237] The generating unit is used to generate a new second training sample based on the response data; the training unit is used to train the structure detection model using the new second training sample to obtain an updated structure detection model.

[0238] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0239] Example 8

[0240] According to an embodiment of the present application, an image processing device is also provided. Fig.12 As shown, the device 1200 includes: a display module 1202 and a marking module 1204 .

[0241] Among them, the display module 1202 is used to display the text image; the marking module 1204 is used to mark the recognition result of the text image on the text image, wherein the recognition result is obtained by using the structure detection model to recognize the text image, and the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than the preset number.

[0242] It should be noted that the display module 1202 and the marking module 1204 correspond to steps S602 to S604 in Example 2, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 2. It should be noted that the above-mentioned modules as part of the device can be run in the computer terminal 10 provided in Example 1.

[0243] In the above embodiment of the present application, the marking module includes: a first acquiring unit, a first determining unit, and a second determining unit.

[0244] Among them, the first acquisition unit is used to input the text image into the feature extraction model to obtain the feature information of the text image; the first determination unit is used to input the feature information of the text image into the regression branch model to determine the position of the text in the text image; the second determination unit is used to input the feature information of the text image into the classification branch model to determine the attributes of the text.

[0245] In the above embodiments of the present application, the device further includes: an acquisition module, a first training module and a second training module.

[0246] Among them, the acquisition module is used to obtain a first training sample and a second training sample, wherein the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image; the first training module is used to use the first training sample to train the initial model to obtain an initial structure detection model; the second training module is used to use the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0247] In the above embodiments of the present application, the acquisition module includes: a second acquisition unit and a processing unit.

[0248] The second acquisition unit is used to acquire a plurality of second text images; and the processing unit is used to process the plurality of second text images using a data enhancement algorithm to generate second training samples.

[0249] In the above embodiment of the present application, the first training module includes: a third acquisition unit, a third determination unit, a fourth determination unit and a first updating unit.

[0250] Among them, the third acquisition unit is used to input each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; the third determination unit is used to input the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; the fourth determination unit is used to input the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; the update unit is used to update the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model and the classification branch model of the initial model based on the predicted position and annotation position of the first text, as well as the classification result, to obtain the initial structure detection model.

[0251] In the above embodiment of the present application, the second training module includes: a fourth acquisition unit, a fifth determination unit, and a second updating unit.

[0252] Among them, the fourth acquisition unit is used to input each second text image into the feature extraction model of the initial structure detection model to obtain the feature information of each second text image; the fifth determination unit is used to input the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; the second update unit is used to update the network parameters of the classification branch model of the initial structure detection model based on the annotation attributes and prediction attributes of the second text to obtain the structure detection model.

[0253] In the above embodiment of the present application, the marking module includes: a sixth determination unit, a seventh determination unit and a marking unit.

[0254] Among them, the sixth determination unit is used to determine the confidence corresponding to the recognition result based on the recognition result of the text image; the seventh determination unit is used to determine the target labeling method of the recognition result based on the confidence; and the marking unit is used to mark the recognition result of the text image on the text image according to the target labeling method.

[0255] In the above embodiments of the present application, the device further includes: a receiving module and an updating module.

[0256] The receiving module is used to receive response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and the updating module is used to update the structure detection model based on the response data.

[0257] In the above embodiments of the present application, the update module includes: a generation unit and a training unit.

[0258] The generating unit is used to generate a new second training sample based on the response data; the training unit is used to train the structure detection model using the new second training sample to obtain an updated structure detection model.

[0259] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0260] Example 9

[0261] According to an embodiment of the present application, an image processing device is also provided. Fig.13 As shown, the device 1300 includes: an acquisition module 1302 , a first training module 1304 , and a second training module 1306 .

[0262] Among them, the acquisition module 1302 is used to obtain a first training sample and a second training sample, wherein the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than a preset number; the first training module 1304 is used to train the initial model using the first training sample to obtain an initial structure detection model; the second training module 1306 is used to train the initial structure detection model using the second training sample to obtain a structure detection model, wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text contained in the text image in the text image, and the second branch model is used to identify the text image and obtain the attributes of the text.

[0263] It should be noted that the acquisition module 1302, the first training module 1304 and the second training module 1306 correspond to steps S702 to S706 in Example 3, and the three modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above-mentioned Example 3. It should be noted that the above-mentioned modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0264] The first training sample in the above embodiment of the present application includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image.

[0265] In the above embodiments of the present application, the acquisition module includes: a first acquisition unit and a processing unit.

[0266] The first acquisition unit is used to acquire a plurality of second text images; and the processing unit is used to process the plurality of second text images using a data enhancement algorithm to generate second training samples.

[0267] In the above embodiment of the present application, the first training module includes: a second acquisition unit, a first determination unit, a second determination unit and a first update unit.

[0268] Among them, the second acquisition unit is used to input each first text image into the feature extraction model of the initial model in the structure detection model to obtain the feature information of each first text image; the first determination unit is used to input the feature information of each first text image into the regression branch model of the initial model in the structure detection model to determine the predicted position of the first text in each first text image; the second determination unit is used to input the feature information of each first text image into the classification branch model of the initial model in the structure detection model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; the first update unit is used to update the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model and the classification branch model of the initial model based on the predicted position and annotation position of the first text, as well as the classification result.

[0269] In the above embodiment of the present application, during the training process using the second training sample, the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged.

[0270] In the above embodiment of the present application, the second training module includes: a third acquisition unit, a third determination unit, and a second updating unit.

[0271] Among them, the third acquisition unit is used to input each second text image into the feature extraction model of the initial structure detection model to obtain the feature information of each second text image; the third determination unit is used to input the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; the second update unit is used to update the network parameters of the classification branch model of the initial structure detection model based on the annotation attributes and prediction attributes of the second text.

[0272] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0273] Example 10

[0274] According to an embodiment of the present application, an image processing device is also provided. Fig.14 As shown, the device 1400 includes: an acquisition module 1402 and an identification module 1404 .

[0275] Among them, the acquisition module 1402 is used to acquire the certificate image; the recognition module 1404 is used to use the structure detection model to recognize the certificate image to obtain the recognition result of the certificate image, wherein the recognition result includes: the attributes of the text contained in the certificate image, and the position of the text in the certificate image; wherein the structure detection model may include: a first branch model and a second branch model, the first branch model is used to recognize the certificate image and obtain the position of the text in the certificate image, the second branch model is used to recognize the certificate image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than the preset number.

[0276] It should be noted that the acquisition module 1402 and the identification module 1404 correspond to steps S802 to S804 in Example 4, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 4. It should be noted that the above-mentioned modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0277] In the above embodiments of the present application, the device further includes: a first determination module and a generation module.

[0278] The first determination module is used to determine the target format of the certificate image; the generation module is used to generate text data corresponding to the certificate image based on the target format and the recognition result.

[0279] In the above embodiment of the present application, the identification module includes: a first acquisition unit, a first determination unit, and a second determination unit.

[0280] Among them, the first acquisition unit is used to input the ID image into the feature extraction model to obtain the feature information of the ID image; the first determination unit is used to input the feature information of the ID image into the regression branch model to determine the position of the text in the ID image; the second determination unit is used to input the feature information of the ID image into the classification branch model to determine the attributes of the text.

[0281] In the above embodiments of the present application, the device further includes: a first training module and a second training module.

[0282] Among them, the acquisition module is also used to obtain a first training sample and a second training sample, wherein the first training sample includes: multiple first ID images, and the annotation position of the first text contained in each first ID image, and the second training sample includes: multiple second ID images, and the annotation attributes of the second text contained in each second ID image; the first training module is used to use the first training sample to train the initial model to obtain an initial structure detection model; the second training module is used to use the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0283] In the above embodiments of the present application, the acquisition module includes: a second acquisition unit and a processing unit.

[0284] The second acquisition unit is used to acquire a plurality of second certificate images; and the processing unit is used to process the plurality of second certificate images using a data enhancement algorithm to generate second training samples.

[0285] In the above embodiment of the present application, the first training module includes: a third acquisition unit, a third determination unit, a fourth determination unit and a first updating unit.

[0286] Among them, the third acquisition unit is used to input each first ID document image into the feature extraction model of the initial model to obtain the feature information of each first ID document image; the third determination unit is used to input the feature information of each first ID document image into the regression branch model of the initial model to determine the predicted position of the first text in each first ID document image; the fourth determination unit is used to input the feature information of each first ID document image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; the update unit is used to update the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model and the classification branch model of the initial model based on the predicted position and annotation position of the first text, as well as the classification result, to obtain the initial structure detection model.

[0287] In the above embodiment of the present application, the second training module includes: a fourth acquisition unit, a fifth determination unit, and a second updating unit.

[0288] Among them, the fourth acquisition unit is used to input each second ID document image into the feature extraction model of the initial structure detection model to obtain the feature information of each second ID document image; the fifth determination unit is used to input the feature information of each second ID document image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; the second update unit is used to update the network parameters of the classification branch model of the initial structure detection model based on the annotation attributes and prediction attributes of the second text to obtain the structure detection model.

[0289] In the above embodiment of the present application, the device further includes: a second determination module, a third determination module and an output module.

[0290] Among them, the second determination module is used to determine the confidence corresponding to the recognition result based on the recognition result of the certificate image; the third determination module is used to determine the target labeling method of the recognition result based on the confidence; and the output module is used to output the recognition result according to the target labeling method.

[0291] In the above embodiments of the present application, the device further includes: a receiving module and an updating module.

[0292] The receiving module is used to receive response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and the updating module is used to update the structure detection model based on the response data.

[0293] In the above embodiments of the present application, the update module includes: a generation unit and a training unit.

[0294] The generating unit is used to generate a new second training sample based on the response data; the training unit is used to train the structure detection model using the new second training sample to obtain an updated structure detection model.

[0295] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0296] Embodiment 11

[0297] According to an embodiment of the present application, an image processing device is also provided. Fig.15 As shown, the device 1500 includes: a receiving module 1502 , an identification module 1504 and an output module 1506 .

[0298] Among them, the receiving module 1502 is used to receive the text image uploaded by the client; the recognition module 1504 is used to use the structure detection model to recognize the text image to obtain the recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; the output module 1506 is used to output the recognition result to the client; wherein the structure detection model may include: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than the preset number.

[0299] It should be noted that the receiving module 1502, the identifying module 1504 and the output module 1506 correspond to steps S902 to S906 in Embodiment 5, and the three modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Embodiment 5. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Embodiment 1.

[0300] In the above embodiment of the present application, the identification module includes: a first acquisition unit, a first determination unit, and a second determination unit.

[0301] Among them, the first acquisition unit is used to input the text image into the feature extraction model to obtain the feature information of the text image; the first determination unit is used to input the feature information of the text image into the regression branch model to determine the position of the text in the text image; the second determination unit is used to input the feature information of the text image into the classification branch model to determine the attributes of the text.

[0302] In the above embodiments of the present application, the device further includes: a first training module and a second training module.

[0303] Among them, the acquisition module is also used to obtain a first training sample and a second training sample, wherein the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image; the first training module is used to use the first training sample to train the initial model to obtain an initial structure detection model; the second training module is used to use the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0304] In the above embodiments of the present application, the acquisition module includes: a second acquisition unit and a processing unit.

[0305] The second acquisition unit is used to acquire a plurality of second text images; and the processing unit is used to process the plurality of second text images using a data enhancement algorithm to generate second training samples.

[0306] In the above embodiment of the present application, the first training module includes: a third acquisition unit, a third determination unit, a fourth determination unit and a first updating unit.

[0307] Among them, the third acquisition unit is used to input each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; the third determination unit is used to input the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; the fourth determination unit is used to input the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; the update unit is used to update the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model and the classification branch model of the initial model based on the predicted position and annotation position of the first text, as well as the classification result, to obtain the initial structure detection model.

[0308] In the above embodiment of the present application, the second training module includes: a fourth acquisition unit, a fifth determination unit, and a second updating unit.

[0309] Among them, the fourth acquisition unit is used to input each second text image into the feature extraction model of the initial structure detection model to obtain the feature information of each second text image; the fifth determination unit is used to input the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; the second update unit is used to update the network parameters of the classification branch model of the initial structure detection model based on the annotation attributes and prediction attributes of the second text to obtain the structure detection model.

[0310] In the above embodiment of the present application, the device further includes: a first determination module and a second determination module.

[0311] Among them, the first determination module is used to determine the confidence corresponding to the recognition result based on the recognition result of the text image; the second determination module is used to determine the target labeling method of the recognition result based on the confidence; and the output module is also used to output the recognition result according to the target labeling method.

[0312] In the above embodiments of the present application, the device further includes: an update module.

[0313] The receiving module is further used to receive response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and the updating module is used to update the structure detection model based on the response data.

[0314] In the above embodiments of the present application, the update module includes: a generation unit and a training unit.

[0315] The generating unit is used to generate a new second training sample based on the response data; the training unit is used to train the structure detection model using the new second training sample to obtain an updated structure detection model.

[0316] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0317] Example 12

[0318] According to an embodiment of the present application, an image processing device is also provided. Fig.16 As shown, the device 1600 includes: a first calling module 1602 , a second calling module 1604 and a third calling module 1606 .

[0319] Among them, the first calling module 1602 is used to receive a text image by calling a first interface, wherein the first interface includes: a first parameter and a second parameter, the parameter value of the first parameter is the text image, and the parameter value of the second parameter is the target type corresponding to the text image; the second calling module 1604 is used to call a structure detection model based on the target type, and use the structure detection model to identify the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; the third calling module 1606 is used to output the recognition result by calling the second interface, wherein the second interface includes: a third parameter, and the parameter value of the third parameter is the recognition result; wherein the structure detection model may include: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample includes unstructured annotation data, the second training sample includes structured annotation data, and the number of the second training samples is less than the preset number.

[0320] It should be noted that the first calling module 1602, the second calling module 1604 and the third calling module 1606 correspond to steps S1002 to S1006 in Example 6, and the three modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above-mentioned Example 6. It should be noted that the above-mentioned modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0321] In the above embodiment of the present application, the second calling module includes: a first obtaining unit, a first determining unit, and a second determining unit.

[0322] Among them, the first acquisition unit is used to input the text image into the feature extraction model to obtain the feature information of the text image; the first determination unit is used to input the feature information of the text image into the regression branch model to determine the position of the text in the text image; the second determination unit is used to input the feature information of the text image into the classification branch model to determine the attributes of the text.

[0323] In the above embodiments of the present application, the device further includes: an acquisition module, a first training module and a second training module.

[0324] Among them, the acquisition module is used to obtain a first training sample and a second training sample, wherein the first training sample includes: multiple first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: multiple second text images, and the annotation attributes of the second text contained in each second text image; the first training module is used to use the first training sample to train the initial model to obtain an initial structure detection model; the second training module is used to use the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0325] In the above embodiments of the present application, the acquisition module includes: a second acquisition unit and a processing unit.

[0326] The second acquisition unit is used to acquire a plurality of second text images; and the processing unit is used to process the plurality of second text images using a data enhancement algorithm to generate second training samples.

[0327] In the above embodiment of the present application, the first training module includes: a third acquisition unit, a third determination unit, a fourth determination unit and a first updating unit.

[0328] Among them, the third acquisition unit is used to input each first text image into the feature extraction model of the initial model to obtain the feature information of each first text image; the third determination unit is used to input the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; the fourth determination unit is used to input the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; the update unit is used to update the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model and the classification branch model of the initial model based on the predicted position and annotation position of the first text, as well as the classification result, to obtain the initial structure detection model.

[0329] In the above embodiment of the present application, the second training module includes: a fourth acquisition unit, a fifth determination unit, and a second updating unit.

[0330] Among them, the fourth acquisition unit is used to input each second text image into the feature extraction model of the initial structure detection model to obtain the feature information of each second text image; the fifth determination unit is used to input the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; the second update unit is used to update the network parameters of the classification branch model of the initial structure detection model based on the annotation attributes and prediction attributes of the second text to obtain the structure detection model.

[0331] In the above embodiment of the present application, the device further includes: a first determination module, a second determination module and an output module.

[0332] Among them, the first determination module is used to determine the confidence corresponding to the recognition result based on the recognition result of the text image; the second determination module is used to determine the target labeling method of the recognition result based on the confidence; and the output module is used to output the recognition result according to the target labeling method.

[0333] In the above embodiments of the present application, the device further includes: a receiving module and an updating module.

[0334] The receiving module is used to receive response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and the updating module is used to update the structure detection model based on the response data.

[0335] In the above embodiments of the present application, the update module includes: a generation unit and a training unit.

[0336] The generating unit is used to generate a new second training sample based on the response data; the training unit is used to train the structure detection model using the new second training sample to obtain an updated structure detection model.

[0337] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0338] Example 13

[0339] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the image processing method provided by the above embodiment.

[0340] Optionally, in this embodiment, the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.

[0341] Optionally, in this embodiment, a computer-readable storage medium is configured to store program code for executing the following steps: acquiring a text image; identifying the text image using a structure detection model to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attributes of the text, and the structure detection model is trained using a first training sample and a second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than a preset number.

[0342] Optionally, the storage medium is also configured to store program codes for executing the following steps: inputting the text image into a feature extraction model to obtain feature information of the text image; inputting the feature information of the text image into a regression branch model to determine the position of the text in the text image; inputting the feature information of the text image into a classification branch model to determine the attributes of the text.

[0343] Optionally, the storage medium is also configured to store program codes for executing the following steps: obtaining a first training sample and a second training sample, wherein the first training sample includes: a plurality of first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: a plurality of second text images, and the annotation attributes of the second text contained in each second text image; using the first training sample to train the initial model to obtain an initial structure detection model; using the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0344] Optionally, the storage medium is further configured to store program codes for executing the following steps: acquiring a plurality of second text images; and processing the plurality of second text images using a data enhancement algorithm to generate second training samples.

[0345] Optionally, the storage medium is also configured to store program codes for executing the following steps: inputting each first text image into a feature extraction model of an initial model to obtain feature information of each first text image; inputting the feature information of each first text image into a regression branch model of an initial model to determine a predicted position of the first text in each first text image; inputting the feature information of each first text image into a classification branch model of an initial model to determine a classification result, wherein the classification result is used to characterize whether the current position is text; based on the predicted position and annotated position of the first text, as well as the classification result, updating the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model, and the classification branch model of the initial model to obtain an initial structure detection model.

[0346] Optionally, the storage medium is also configured to store program codes for executing the following steps: inputting each second text image into a feature extraction model of an initial structure detection model to obtain feature information of each second text image; inputting the feature information of each second text image into a classification branch model of an initial structure detection model to determine prediction attributes of the second text; and updating network parameters of the classification branch model of the initial structure detection model based on the annotation attributes and prediction attributes of the second text to obtain a structure detection model.

[0347] Optionally, the above-mentioned storage medium is also configured to store program codes for executing the following steps: determining the confidence level corresponding to the recognition result based on the recognition result of the text image; determining the target labeling method of the recognition result based on the confidence level; and marking the recognition result of the text image on the text image according to the target labeling method.

[0348] Optionally, the storage medium is further configured to store program codes for executing the following steps: receiving response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and updating the structure detection model based on the response data.

[0349] Optionally, the storage medium is further configured to store program codes for executing the following steps: generating a new second training sample based on the response data; and training the structure detection model using the new second training sample to obtain an updated structure detection model.

[0350] As an optional example, the storage medium is also configured to store program code for executing the following steps: displaying a text image; marking a recognition result of the text image on the text image, wherein the recognition result is obtained by recognizing the text image using a structure detection model, and the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using a first training sample and a second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than a preset number.

[0351] As an optional example, the storage medium is also configured to store program code for executing the following steps: obtaining a first training sample and a second training sample, wherein the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than a preset number; using the first training sample to train the initial model to obtain an initial structure detection model; using the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify text images and obtain the position of text in the text image, and the second branch model is used to identify text images and obtain attributes of text.

[0352] As an optional example, the storage medium is also configured to store program code for executing the following steps: acquiring an ID image; identifying the ID image using a structural detection model to obtain an identification result of the ID image, wherein the identification result includes: the attributes of the text contained in the ID image, and the position of the text in the ID image; wherein the structural detection model includes: a first branch model and a second branch model, the first branch model is used to identify the ID image and obtain the position of the text in the ID image, the second branch model is used to identify the ID image and obtain the attributes of the text, and the structural detection model is trained using the first training sample and the second training sample in sequence, and the number of the second training samples is less than a preset number.

[0353] Optionally, the storage medium is further configured to store program codes for executing the following steps: determining a target format of the document image; and generating text data corresponding to the document image based on the target format and the recognition result.

[0354] As an optional example, the storage medium is also configured to store program code for executing the following steps: receiving a text image uploaded by a client; using a structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; outputting the recognition result to the client; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence, and the number of the second training samples is less than a preset number.

[0355] As an optional example, the storage medium is also configured to store program code for executing the following steps: receiving a text image by calling a first interface, wherein the first interface includes: a first parameter and a second parameter, the parameter value of the first parameter is the text image, and the parameter value of the second parameter is the target type corresponding to the text image; calling a structure detection model based on the target type, and using the structure detection model to identify the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; outputting the recognition result by calling a second interface, wherein the second interface includes: a third parameter, the parameter value of the third parameter is the recognition result; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence, and the number of the second training samples is less than the preset number.

[0356] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0357] Embodiment 14

[0358] The embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.

[0359] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.

[0360] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the image processing method: obtaining a text image; using a structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using a first training sample and a second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than a preset number.

[0361] Optionally, Fig.17 is a structural block diagram of a computer terminal according to an embodiment of the present application. Fig.17 As shown, the computer terminal 10 may include: one or more (only one is shown in the figure) processors 1702 and a memory 1704.

[0362] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the image processing method and device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned image processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0363] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a text image; use the structure detection model to recognize the text image to obtain the recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than the preset number.

[0364] Optionally, the processor may also execute the program code of the following steps: inputting the text image into a feature extraction model to obtain feature information of the text image; inputting the feature information of the text image into a regression branch model to determine the position of the text in the text image; inputting the feature information of the text image into a classification branch model to determine the attributes of the text.

[0365] Optionally, the processor may also execute the program code of the following steps: obtaining a first training sample and a second training sample, wherein the first training sample includes: a plurality of first text images, and the annotation position of the first text contained in each first text image, and the second training sample includes: a plurality of second text images, and the annotation attributes of the second text contained in each second text image; using the first training sample to train the initial model to obtain an initial structure detection model; using the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

[0366] Optionally, the processor may also execute program codes of the following steps: acquiring a plurality of second text images; and processing the plurality of second text images using a data enhancement algorithm to generate second training samples.

[0367] Optionally, the processor may also execute the following program code: input each first text image into the feature extraction model of the initial model to obtain feature information of each first text image; input the feature information of each first text image into the regression branch model of the initial model to determine the predicted position of the first text in each first text image; input the feature information of each first text image into the classification branch model of the initial model to determine the classification result, wherein the classification result is used to characterize whether the current position is text; based on the predicted position and annotation position of the first text, as well as the classification result, update the network parameters of the feature extraction model of the initial model, the regression branch model of the initial model, and the classification branch model of the initial model to obtain an initial structure detection model.

[0368] Optionally, the processor may also execute the program code of the following steps: input each second text image into the feature extraction model of the initial structure detection model to obtain feature information of each second text image; input the feature information of each second text image into the classification branch model of the initial structure detection model to determine the prediction attribute of the second text; based on the annotation attribute and prediction attribute of the second text, update the network parameters of the classification branch model of the initial structure detection model to obtain the structure detection model.

[0369] Optionally, the processor may also execute the program code of the following steps: based on the recognition result of the text image, determine the confidence level corresponding to the recognition result; based on the confidence level, determine the target labeling method of the recognition result; and mark the recognition result of the text image on the text image according to the target labeling method.

[0370] Optionally, the processor may also execute program codes of the following steps: receiving response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; and updating the structure detection model based on the response data.

[0371] Optionally, the processor may also execute program codes of the following steps: generating a new second training sample based on the response data; and training the structure detection model using the new second training sample to obtain an updated structure detection model.

[0372] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: display a text image; mark the recognition result of the text image on the text image, wherein the recognition result is obtained by using a structure detection model to recognize the text image, and the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using a first training sample and a second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than a preset number.

[0373] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a first training sample and a second training sample, wherein the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than a preset number; use the first training sample to train the initial model to obtain an initial structure detection model; use the second training sample to train the initial structure detection model to obtain a structure detection model, wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, and the second branch model is used to identify the text image and obtain the attributes of the text.

[0374] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain an ID image; use the structure detection model to identify the ID image to obtain the identification result of the ID image, wherein the identification result includes: the attributes of the text contained in the ID image, and the position of the text in the ID image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the ID image and obtain the position of the text in the ID image, the second branch model is used to identify the ID image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence, and the number of the second training samples is less than the preset number.

[0375] Optionally, the processor may further execute program codes of the following steps: determining a target format of the document image; and generating text data corresponding to the document image based on the target format and the recognition result.

[0376] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receive the text image uploaded by the client; use the structure detection model to recognize the text image to obtain the recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; output the recognition result to the client; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence, and the number of the second training samples is less than the preset number.

[0377] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receive the text image by calling the first interface, wherein the first interface includes: a first parameter and a second parameter, the parameter value of the first parameter is the text image, and the parameter value of the second parameter is the target type corresponding to the text image; call the structure detection model based on the target type, and use the structure detection model to identify the text image to obtain the recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; output the recognition result by calling the second interface, wherein the second interface includes: a third parameter, and the parameter value of the third parameter is the recognition result; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attributes of the text, and the structure detection model is trained using the first training sample and the second training sample in sequence, and the number of the second training samples is less than the preset number.

[0378] According to the application embodiment, the structure detection model is first trained by using a large number of first training samples containing unstructured annotation data, which can ensure that the structure detection model can detect and identify different types of images. Then, the structure detection model is trained using a small number of second training samples containing structured annotation data, which can fine-tune the structure detection model and improve the accuracy of the structure detection model. The purpose of training the structure detection model can be achieved through a small amount of structured processing. Therefore, after acquiring the text image, the text image can be recognized with high accuracy through the structure detection model, so that the recognition result of the obtained text image is more accurate, achieving the technical effect of reducing the annotation cost of training samples and improving the recognition accuracy of the structure detection model, thereby solving the technical problem of the high training cost of the structure detection model in the related technology.

[0379] It can be understood by those skilled in the art that Fig.17The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig.17 It does not limit the structure of the above electronic device. For example, the computer terminal A may also include Fig.17 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig.17 Different configurations shown.

[0380] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0381] Embodiment 15

[0382] According to an embodiment of the present application, an image processing system is also provided, including:

[0383] Processor; and

[0384] The memory is connected to the processor and is used to provide the processor with instructions for processing the following processing steps: obtaining a text image; using a structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using a first training sample and a second training sample in sequence, the first training sample includes unstructured annotation data, the second training sample includes structured annotation data, and the number of the second training samples is less than a preset number.

[0385] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0386] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0387] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0388] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0389] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0390] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0391] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0392] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. An image processing method, characterized in that: include: Get text image; Recognize the text image using a structure detection model to obtain a recognition result of the text image, wherein the recognition result includes: attributes of the text contained in the text image, and positions of the text in the text image; The structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attribute of the text, and the structure detection model is trained by using the first training sample and the second training sample in sequence; The first training sample includes unstructured labeled data, and the second training sample includes structured labeled data.

2. The method according to claim 1, characterized in that The first branch model and the second branch model use the same feature extraction model, and the feature extraction model is used to receive an input image and extract features of the image; The first branch model further includes: a regression branch model, the input layer of the regression branch model is connected to the output layer of the feature extraction model, and the regression branch model is used to process the features of the image output by the feature extraction model to obtain the position of the text contained in the image; The second branch model also includes: a classification branch model, the input layer of the classification branch model is connected to the output layer of the feature extraction model, and the classification branch model is used to process the features of the image output by the feature extraction model to obtain the attribute that the image contains text.

3. The method according to claim 2, characterized in that The text image is recognized by using a structure detection model, and the recognition result of the text image is obtained, including: Inputting the text image into the feature extraction model to obtain feature information of the text image; Inputting feature information of the text image into the regression branch model to determine the position of the text in the text image; The feature information of the text image is input into the classification branch model to determine the attributes of the text.

4. The method according to claim 2, characterized in that: The method further comprises: Acquire the first training sample and the second training sample, wherein the first training sample includes: a plurality of first text images, and a marked position of a first character contained in each first text image, and the second training sample includes: a plurality of second text images, and a marked attribute of a second character contained in each second text image, and the number of the second training samples is less than a preset number; Using the first training sample to train the initial model to obtain an initial structure detection model; The initial structure detection model is trained using the second training sample to obtain the structure detection model, wherein network parameters of the feature extraction model and the regression branch model of the initial structure detection model remain unchanged during the training process.

5. The method according to claim 4, characterized in that Acquiring the second training sample includes: Acquire the plurality of second text images; The plurality of second text images are processed using a data enhancement algorithm to generate the second training samples.

6. The method according to claim 4, characterized in that Training the initial structure detection model using the second training sample includes: Inputting each second text image into the feature extraction model of the initial structure detection model to obtain feature information of each second text image; Inputting the feature information of each second text image into the classification branch model of the initial structure detection model to determine the predicted attribute of the second text; Based on the annotation attribute and the prediction attribute of the second character, the network parameters of the classification branch model of the initial structure detection model are updated.

7. An image processing method, characterized in that: include: Display text images; Marking a recognition result of the text image on the text image, wherein the recognition result is obtained by recognizing the text image using a structure detection model, and the recognition result includes: attributes of the text contained in the text image, and positions of the text in the text image; The structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attribute of the text, and the structure detection model is trained by using the first training sample and the second training sample in sequence; The first training sample includes unstructured labeled data, and the second training sample includes structured labeled data.

8. The method according to claim 7, characterized in that Marking the recognition result of the text image on the text image includes: Based on the recognition result of the text image, determining a confidence level corresponding to the recognition result; Determining a target labeling method for the recognition result based on the confidence level; The recognition result is marked on the text image according to the target marking method.

9. The method according to claim 7, characterized in that: After marking the recognition result of the text image on the text image, the method further includes: Receiving response data corresponding to the recognition result, wherein the response data is obtained by modifying the recognition result; The structure detection model is updated based on the response data.

10. The method according to claim 9, characterized in that Updating the structure detection model based on the response data includes: Based on the response data, generating a new second training sample; The structure detection model is trained using the new second training sample to obtain an updated structure detection model.

11. An image processing method, characterized in that: include: Obtain a first training sample and a second training sample; Using the first training sample to train the initial model to obtain an initial structure detection model; The initial structure detection model is trained using the second training sample to obtain a structure detection model, wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text contained in the text image in the text image, and the second branch model is used to identify the text image and obtain the attribute of the text; The first training sample includes unstructured labeled data, and the second training sample includes structured labeled data.

12. The method according to claim 11, characterized in that The first training samples include: multiple first text images and the marked position of the first text contained in each first text image, and the second training samples include: multiple second text images and the marked attributes of the second text contained in each second text image, and the number of the second training samples is less than the preset number.

13. The method according to claim 12, characterized in that Acquiring the second training sample includes: Acquire the plurality of second text images; The plurality of second text images are processed using a data enhancement algorithm to generate the second training samples.

14. An image processing method, characterized in that: include: Get the document image; Using the structure detection model to identify the certificate image, and obtain a recognition result of the certificate image, wherein the recognition result includes: attributes of the text contained in the certificate image, and the position of the text in the certificate image; Among them, the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the document image and obtain the position of the text in the document image, the second branch model is used to identify the document image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, and the second training sample contains structured annotation data.

15. The method according to claim 14, characterized in that After the document image is recognized by using the structure detection model to obtain a recognition result of the document image, the method further includes: Determining a target format of the document image; Based on the target format and the recognition result, text data corresponding to the document image is generated.

16. The method according to claim 14, characterized in that The first branch model and the second branch model use the same feature extraction model, and the feature extraction model is used to receive an input image and extract features of the image; The first branch model further includes: a regression branch model, the input layer of the regression branch model is connected to the output layer of the feature extraction model, and the regression branch model is used to process the features of the image output by the feature extraction model to obtain the position of the text contained in the image; The second branch model also includes: a classification branch model, the input layer of the classification branch model is connected to the output layer of the feature extraction model, and the classification branch model is used to process the features of the image output by the feature extraction model to obtain the attribute that the image contains text.

17. The method according to claim 16, characterized in that The method further comprises: Obtaining the first training sample and the second training sample, wherein the first training sample includes: a plurality of first document images, and a marked position of a first text contained in each first document image, and the second training sample includes: a plurality of second document images, and a marked attribute of a second text contained in each second document image, and the number of the second training samples is less than a preset number; Using the first training sample to train the initial model to obtain an initial structure detection model; The initial structure detection model is trained using the second training sample to obtain the structure detection model, wherein network parameters of the feature extraction model and the regression branch model in the initial structure detection model remain unchanged during the training process.

18. The method according to claim 17, characterized in that Acquiring the second training sample includes: Acquire the plurality of second document images; The plurality of second document images are processed using a data enhancement algorithm to generate the second training samples.

19. An image processing method, characterized in that: include: Receive text images uploaded by the client; Recognize the text image using a structure detection model to obtain a recognition result of the text image, wherein the recognition result includes: attributes of the text contained in the text image, and positions of the text in the text image; Outputting the recognition result to the client; Among them, the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, the second training sample contains structured annotation data, and the number of the second training samples is less than the preset number.

20. An image processing method, characterized in that: include: The text image is received by calling a first interface, wherein the first interface includes: a first parameter and a second parameter, the parameter value of the first parameter is the text image, and the parameter value of the second parameter is the target type corresponding to the text image; Calling a structure detection model based on the target type, and using the structure detection model to recognize the text image to obtain a recognition result of the text image, wherein the recognition result includes: attributes of the text contained in the text image, and positions of the text in the text image; Outputting the recognition result by calling a second interface, wherein the second interface includes: a third parameter, a parameter value of the third parameter being the recognition result; Among them, the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, and the second training sample contains structured annotation data.

21. An image processing device, characterized in that: include: An acquisition module, used for acquiring text images; A recognition module, used to recognize the text image using a structure detection model to obtain a recognition result of the text image, wherein the recognition result includes: attributes of the text contained in the text image, and positions of the text in the text image; Among them, the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, and the second training sample contains structured annotation data.

22. An image processing device, characterized in that: include: A display module, used for displaying text images; A marking module, used for marking the recognition result of the text image on the text image, wherein the recognition result is obtained by recognizing the text image using a structure detection model, and the recognition result includes: attributes of the text contained in the text image, and the position of the text in the text image; Among them, the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the text image and obtain the position of the text in the text image, the second branch model is used to identify the text image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, and the second training sample contains structured annotation data.

23. An image processing device, characterized in that: include: An acquisition module, used to acquire a document image; A recognition module, used to recognize the document image using a structure detection model to obtain a recognition result of the document image, wherein the recognition result includes: attributes of the text contained in the document image, and the position of the text in the document image; Among them, the structure detection model includes: a first branch model and a second branch model, the first branch model is used to identify the document image and obtain the position of the text in the document image, the second branch model is used to identify the document image and obtain the attributes of the text, the structure detection model is trained using the first training sample and the second training sample in sequence, the first training sample contains unstructured annotation data, and the second training sample contains structured annotation data.

24. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the image processing method according to any one of claims 1 to 20.

25. A computer terminal, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the image processing method according to any one of claims 1 to 20 when running.

26. An image processing system, characterized in that: include: processor; as well as A memory, connected to the processor, for providing instructions for the processor to process the following processing steps: acquiring a text image; The text image is recognized by using a structure detection model to obtain a recognition result of the text image, wherein the recognition result includes: the attributes of the text contained in the text image, and the position of the text in the text image; wherein the structure detection model includes: a first branch model and a second branch model, the first branch model is used to recognize the text image and obtain the position of the text in the text image, the second branch model is used to recognize the text image and obtain the attributes of the text, and the structure detection model is trained using a first training sample and a second training sample in sequence, the first training sample includes unstructured annotation data, and the second training sample includes structured annotation data.

Citation Information

Patent Citations

  • Structured processing method and device of image, storage medium and electronic equipment

    CN111144210A

  • Method for extracting structural data from image, apparatus and device

    WO2020113561A1