Character Detection Method, Device, Equipment and Medium Based on Federal OCR Model
Through iterative training of the federal OCR model, the problem of low accuracy of OCR recognition in the financial field is solved, and more efficient identity verification and data security are achieved.
Patent Information
- Application Number
- CN202010202677.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-03-20
AI Technical Summary
The existing OCR technology has low accuracy in identity verification in the financial field, resulting in serious waste of human resources.
The character detection method based on the federated OCR model is adopted, and the local initial OCR model is iteratively trained through the joint gradient generated by the coordination end. Combined with the model gradient of multi-party nodes, a federated OCR model is built to perform character detection of image information.
Without leaking user privacy data, the accuracy of OCR identification is improved, the waste of human resources is reduced, and more efficient identity verification is achieved.
Smart Images

Figure CN111401367B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial technology (Fintech), and particularly to a character detection method, device, equipment and medium based on a federated OCR model. Background Art
[0002] In recent years, with the rapid development of Internet financial technology (Fintech), more and more technologies (big data, distributed, blockchain, artificial intelligence, etc.) have been applied in the financial field.
[0003] In order to ensure the security of financial business operations in the financial field, users are required to upload ID photo information for financial business personnel to conduct identity verification. Currently, identity verification is mainly carried out by manually checking ID photo information, which seriously wastes human resources; some financial institutions in the financial field adopt OCR (Optical Character Recognition, that is, directly converting the text content on pictures and photos into editable text) technology for identity verification. The introduction of OCR technology has reduced the waste of human resources, but the recognition models in current OCR technology have not been fully learned, resulting in low OCR recognition accuracy. Summary of the Invention
[0004] The main purpose of the present invention is to propose a character detection method, device, equipment and medium based on a federated OCR model, aiming to solve the technical problem of low current OCR recognition accuracy.
[0005] To achieve the above object, the present invention provides a character detection method based on a federated OCR model. The character detection method based on a federated OCR model includes the following steps:
[0006] When receiving an OCR recognition request, obtain the image information to be recognized associated with the OCR recognition request;
[0007] Call the federated OCR model to perform character detection on the image information, obtain the OCR recognition result and output it. Among them, the federated OCR model is obtained by iteratively training the local initial OCR model based on the joint gradient sent by the coordination end, and the joint gradient is generated by the coordination end based on the model gradients of multiple parties.
[0008] Optionally, before obtaining the image information to be recognized associated with the OCR recognition request when receiving the OCR recognition request, the method further includes:
[0009] Mark the image information in the local storage to form a training sample set, and extract a preset proportion of training samples from the training sample set;
[0010] Train an initial OCR model with the training samples to obtain model gradients, and send the model gradients to the coordination end so that the coordination end generates joint gradients based on the model gradients fed back by multiple parties' nodes;
[0011] Receive the joint gradients sent by the coordination end, update the initial OCR model according to the joint gradients to obtain a trained OCR model, and obtain the OCR feature vector of the trained OCR model;
[0012] Process the OCR feature vector through a preset loss function to obtain a loss value, and send the loss value to the coordination end to determine whether the OCR model is trained through the coordination end's analysis of the loss value;
[0013] When receiving the training completion prompt sent by the coordination end, use the trained OCR model as the federated OCR model.
[0014] Optionally, the step of calling the federated OCR model to perform character detection on the image information, obtaining an OCR recognition result and outputting it includes:
[0015] Call the federated OCR model to perform text detection on the image information and extract the text area in the image information;
[0016] Perform character recognition on the text area through the federated OCR model to obtain the character information contained in the text area, and use the character information as the OCR recognition result and output it.
[0017] Optionally, the step of performing character recognition on the text area through the federated OCR model to obtain the character information contained in the text area, and using the character information as the OCR recognition result and outputting it includes:
[0018] Perform character recognition on the text area through the federated OCR model to determine the character type of the characters in the text area;
[0019] Obtain the character detection sub-model corresponding to the character type in the federated OCR model, perform character recognition on the text area through the character detection sub-model to obtain the character information contained in the text area, and use the character information as the OCR recognition result and output it.
[0020] Optionally, after the step of calling the federated OCR model to perform character detection on the image information, obtaining an OCR recognition result and outputting it, it includes:
[0021] When the OCR recognition result is incorrect, output a marking prompt to prompt the user to mark the image information;
[0022] Use the labeled image information as training samples, train the federated OCR model according to the training samples to obtain model gradients, and send the model gradients to the coordination end so that the coordination end generates joint gradients based on the model gradients fed back by multiple parties;
[0023] Receive the joint gradients sent by the coordination end and update the federated OCR model according to the joint gradients.
[0024] Optionally, the step of outputting a model training prompt to prompt the user to label the image information when the OCR recognition result is incorrect includes:
[0025] When the OCR recognition result is incorrect, determine the error type of the error;
[0026] When the error type is region detection error, output a region annotation prompt to prompt the user to annotate the text region in the image information;
[0027] When the error type is character recognition error, output a character annotation prompt to prompt the user to input the character information contained in the image information.
[0028] Optionally, after the step of obtaining the image information to be recognized associated with the OCR recognition request when receiving the OCR recognition request, the method further includes:
[0029] When the image information is a document image, obtain the document type of the document image;
[0030] The step of calling the federated OCR model to perform character detection on the image information, obtaining an OCR recognition result and outputting it includes:
[0031] Call the federated OCR model corresponding to the document type to perform character detection on the image information, obtain an OCR recognition result and output it.
[0032] In addition, to achieve the above object, the present invention also provides a character detection device based on a federated OCR model. The character detection device based on a federated OCR model includes:
[0033] A request receiving module, configured to obtain the image information to be recognized associated with the OCR recognition request when receiving the OCR recognition request;
[0034] A call detection module, configured to call the federated OCR model to perform character detection on the image information, obtain an OCR recognition result and output it, where the federated OCR model is obtained by iteratively training the local initial OCR model based on the joint gradients sent by the coordination end, and the joint gradients are processed and generated by the coordination end based on the model gradients of multiple parties.
[0035] In addition, to achieve the above object, the present invention also provides a character detection device based on a federated OCR model. The character detection device based on the federated OCR model includes: a memory, a processor, and a computer program corresponding to OCR recognition based on the federated OCR model stored on the memory and executable on the processor. When the computer program corresponding to the OCR recognition based on the federated OCR model is executed by the processor, the steps of the character detection method based on the federated OCR model as described above are implemented.
[0036] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium. A computer program corresponding to OCR recognition based on the federated OCR model is stored on the computer-readable storage medium. When the computer program corresponding to the OCR recognition based on the federated OCR model is executed by a processor, the steps of the character detection method based on the federated OCR model as described above are implemented.
[0037] The present invention provides a character detection method, device, equipment, and medium based on a federated OCR model. In an embodiment of the present invention, when an OCR recognition request is received, image information to be recognized associated with the OCR recognition request is obtained; a federated OCR model is called to perform character detection on the image information, and an OCR recognition result is obtained and output. Among them, the federated OCR model is obtained by iteratively training the local initial OCR model based on the joint gradient sent by the coordination end, and the joint gradient is generated by the coordination end based on the model gradients of multiple parties. In an embodiment of the present invention, a federated OCR model is pre-constructed. The federated OCR model determines the joint gradient based on the model gradients of multiple nodes in the alliance chain and is jointly trained according to the joint gradient. The federated OCR model enables multiple parties to fully learn without disclosing their data privacy. In this embodiment, character detection is performed on the image information through the federated OCR model, improving the accuracy of OCR recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention;
[0039] Figure 2 is a schematic flowchart of the first embodiment of the character detection method based on the federated OCR model of the present invention;
[0040] Figure 3 is a schematic flowchart of the second embodiment of the character detection method based on the federated OCR model of the present invention;
[0041] Figure 4 is a schematic flowchart of the third embodiment of the character detection method based on the federated OCR model of the present invention;
[0042] Figure 5 This is a schematic diagram of the functional modules of an embodiment of the character detection device based on the federated OCR model of the present invention.
[0043] The realization of the objectives, functional features, and advantages of the present invention will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Detailed implementation manners
[0044] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0045] As Figure 1 shown, Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment solution of the present invention.
[0046] The character detection device based on the federated OCR model in the embodiment of the present invention may be a coordination terminal device. As Figure 1 shown, the character detection device based on the federated OCR model may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0047] Those skilled in the art can understand that Figure 1 the device structure shown in
[0048] does not constitute a limitation on the device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 1 As
[0049] In Figure 1In the device shown, the network interface 1004 is mainly used to connect to the background coordination end and communicate data with the background coordination end; the user interface 1003 is mainly used to connect to the client (user end) and communicate data with the client; and the processor 1001 can be used to call the computer program stored in the memory 1005 corresponding to the OCR recognition based on the federated OCR model and execute the operations in the following character detection method based on the federated OCR model.
[0050] Based on the above hardware structure, an embodiment of the character detection method based on the federated OCR model of the present invention is proposed.
[0051] When an OCR recognition request is received, obtain the image information to be recognized associated with the OCR recognition request;
[0052] Call the federated OCR model to perform character detection on the image information, obtain the OCR recognition result and output it. Among them, the federated OCR model is obtained by iteratively training the local initial OCR model based on the joint gradient sent by the coordination end, and the joint gradient is generated by the coordination end based on the model gradients of multiple parties.
[0053] In this embodiment, the character detection method based on the federated OCR model is applied to the character detection device in financial institutions (such as banking institutions, insurance institutions, and securities institutions) in the financial industry.
[0054] The character detection device is a node in the alliance chain. The federated OCR model is pre-stored in the character detection device. The federated OCR model is obtained by iteratively training the local initial OCR model in the character detection device based on the joint gradient sent by the coordination end (the joint gradient is generated by the coordination end based on the model gradients of multiple parties). The training steps of the federated OCR model in this embodiment are given as follows:
[0055] When an OCR model update instruction is received, obtain the initial OCR model;
[0056] Mark the image information stored locally to form a training sample set, and extract a preset proportion of training samples from the training sample set;
[0057] Train the initial OCR model with the training samples to obtain the model gradient, and send the model gradient to the coordination end, so that the coordination end generates a joint gradient based on the model gradients fed back by multiple parties;
[0058] Receive the joint gradient sent by the coordination end, update the initial OCR model according to the joint gradient to obtain the trained OCR model, and obtain the OCR feature vector of the trained OCR model;
[0059] Process the OCR feature vectors through a preset loss function to obtain a loss value, and send the loss value to a coordination end to determine whether the OCR model is trained completely by analyzing the loss value through the coordination end;
[0060] When receiving a training completion prompt sent by the coordination end, use the trained OCR model as a federated OCR model.
[0061] Specifically, it includes:
[0062] 1. First, the coordination end sends an initial OCR model to be trained to each character detection device in the consortium chain. The initial OCR model here can be partially trained or completely untrained. Usually, data transmission in the consortium chain requires encryption. In this embodiment, the process of the coordination end sending the initial OCR model may not require encryption, and the sent model can also be an unencrypted model.
[0063] 2. After each character detection device receives the initial OCR model sent by the coordination end, the character detection device prepares its own training data. Each piece of training data is a picture and the corresponding annotation result. Since the data is at the character detection device, the annotation of the data also needs to be completed by the character detection device itself. Then the character detection device trains the initial OCR model with its own training data. In this embodiment, when the character detection device trains the OCR model, it will not send it to the coordination end for training as in the traditional OCR model training. At the same time, since the sensitive data is only at the character detection device, the annotation of the data is also completed by the character detection device, thus avoiding the leakage of sensitive data and user privacy.
[0064] 3. During the training process of each character detection device, gradients for updating the model will be obtained. The character detection device encrypts the gradients (homomorphic encryption algorithm, differential privacy algorithm, etc.) and then sends the encrypted gradient information back to the coordination end. Since the gradient information is encrypted, others (including the coordination end) except the character detection device itself cannot reverse the original training data, ensuring the security of user data.
[0065] 4. The coordination end uses the method of secure aggregation to integrate the encrypted training results together to obtain complete updated information on the model gradients, and then sends this complete updated information to each character detection device. In this embodiment, the coordination end can obtain complete gradient update information without knowing any original data of each character detection device, while ensuring that the data of all character detection devices is used for model training.
[0066] 5. After each character detection device receives the complete gradient update information, it updates its respective model. In this way, the local OCR models in each character detection device are synchronized, and the training data used is the data of all people. Iterate the above training process. The character detection device obtains the trained OCR model and acquires the OCR feature vector of the trained OCR model. The character detection device processes the OCR feature vector through a preset loss function (the preset loss function is set according to the OCR model) to obtain a loss value, and sends the loss value to the coordination end to determine whether the OCR model is trained completely by analyzing the loss value at the coordination end. Until the loss function converges, when receiving the training completion prompt sent by the coordination end, the character detection device uses the trained OCR model as the federated OCR model.
[0067] In this embodiment, when the character detection device performs federated OCR model training, it can avoid the leakage of user privacy data while achieving sufficient training of the federated OCR model, so as to improve the character detection accuracy through the federated OCR model with sufficient learning.
[0068] Refer to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the character detection method based on the federated OCR model of the present invention. In this embodiment, the character detection method based on the federated OCR model includes:
[0069] Step S10, when receiving an OCR recognition request, obtain the image information to be recognized associated with the OCR recognition request.
[0070] The character detection device receives the OCR recognition request. The triggering method of the OCR recognition request is not specifically limited. That is, the OCR recognition request can be actively triggered by the user. For example, the user clicks the "OCR recognition" button on the display page of the character detection device to actively trigger the OCR recognition request. In addition, the OCR recognition request can also be automatically triggered. For example, the character detection device is pre-set to automatically trigger the OCR recognition request when receiving new image information.
[0071] The character detection device receives the OCR recognition request, and the character detection device obtains the image information to be recognized associated with the OCR recognition request. The image information can be an image of text in a paper document converted into a black and white dot matrix, such as an advertisement image, a license plate number image, an ID card scan information, etc.
[0072] Step S20, call the federated OCR model to perform character detection on the image information, obtain the OCR recognition result and output it.
[0073] The character detection device calls the federated OCR model to perform character detection on the image information, obtains the OCR recognition result and outputs it. Specifically, it includes:
[0074] Step a1, call the federated OCR model to perform text detection on the image information and extract the text regions in the image information;
[0075] Step a2, perform character recognition on the text regions through the federated OCR model, obtain the character information contained in the text regions, and output the character information as the OCR recognition result.
[0076] That is, the character detection device calls the federated OCR model stored locally to perform text detection on the image information. The federated OCR model preprocesses the image to obtain a grayscale image, extracts feature points in the grayscale image through the federated OCR model, and divides the grayscale image into text regions and non-text regions according to the feature points, and the federated OCR model extracts the text regions in the image information; the character detection device performs character recognition on the text regions through the federated OCR model, obtains the character information contained in the text regions, and outputs the character information as the OCR recognition result. In this embodiment, the image is divided into text regions and non-text regions through the federated OCR model, and character detection is performed on the text regions, improving the accuracy of character recognition.
[0077] In addition, in order to improve the accuracy of character detection, a character detection sub-model is set in this embodiment. This is a refinement of step a1 and includes:
[0078] Perform character recognition on the text regions through the federated OCR model to determine the character types of the characters in the text regions;
[0079] Obtain the character detection sub-model corresponding to the character type in the federated OCR model, perform character recognition on the text regions through the character detection sub-model, obtain the character information contained in the text regions, and output the character information as the OCR recognition result.
[0080] That is, the character detection device performs character recognition on the text regions through the federated OCR model to determine the character types of the characters in the text regions; the character detection device obtains the character detection sub-model corresponding to the character type in the federated OCR model, the character detection device performs character recognition on the text regions through the character detection sub-model, obtains the character information contained in the text regions, and the character detection device outputs the character information as the OCR recognition result.
[0081] In this embodiment, since there are many types of characters. For example, the types of characters can be English characters (such as license plate numbers), or Chinese characters (such as scanned ID card images). In order to simplify the federated OCR model and ensure the accuracy of character detection, the federated OCR model in this embodiment includes multiple character detection sub-models. The character detection sub-models are set according to the character types, which can improve the accuracy of character detection without making the federated OCR model complex.
[0082] In the embodiment of the present invention, a federated OCR model is pre-constructed. The federated OCR model determines the joint gradient based on the model gradients of multiple nodes in the consortium chain and is jointly trained according to the joint gradient. The federated OCR model enables multiple parties to fully learn without disclosing their data privacy. In this embodiment, the federated OCR model is used to perform character detection on image information, improving the accuracy of OCR recognition.
[0083] Further, refer to Figure 3 , Figure 3 which is a schematic flowchart of the second embodiment of the character detection method based on the federated OCR model of the present invention.
[0084] Based on the first embodiment of the character detection method based on the federated OCR model of the present invention, the second embodiment of the character detection method based on the federated OCR model of the present invention is proposed.
[0085] This embodiment is the step after step S20 in the first embodiment. The difference between this embodiment and the above embodiment is that:
[0086] Step S30, when the OCR recognition result is incorrect, output a marking prompt to prompt the user to mark the image information.
[0087] The character detection device outputs the OCR recognition result, and the character detection device determines whether the OCR recognition result is incorrect. That is, the implementation manner of the character detection device for determining whether the OCR recognition result is incorrect is not specifically limited. For example, the character detection device determines that the OCR recognition result is incorrect when the confirmation instruction input by the user is an incorrect instruction; conversely, when the confirmation instruction is a correct instruction, the character detection device determines that the OCR recognition result is correct. Or, the character detection device obtains the character information in the OCR recognition result and the standard character information corresponding to the image information. When the character information in the OCR recognition result is different from the standard character information corresponding to the image information, the character detection device determines that the OCR recognition result is incorrect; conversely, when the character information in the OCR recognition result is the same as the standard character information corresponding to the image information, the character detection device determines that the OCR recognition result is correct.
[0088] Step b1, when the OCR recognition result is incorrect, determine the error type of the error;
[0089] Step b2, if the error type is region detection error, output a region annotation prompt to prompt the user to annotate the text region in the image information;
[0090] Step b3, if the error type is character recognition error, output a character annotation prompt to prompt the user to input the character information included in the image information.
[0091] When the character detection device has an error in the OCR recognition result, determine the error type; if the error type is region detection error, the character detection device outputs a region annotation prompt to prompt the user to annotate the text region in the image information; if the error type is character recognition error, the character detection device outputs a character annotation prompt to prompt the input of the character information included in the image information. In this embodiment, the character detection device prompts the user to perform annotation to perform model training based on the annotated image information. Specifically:
[0092] Step S40, use the annotated image information as a training sample, train the federated OCR model according to the training sample to obtain a model gradient, and send the model gradient to the coordination end so that the coordination end generates a joint gradient based on the model gradients fed back by multiple parties.
[0093] The character detection device uses the annotated image information as a training sample. The character detection device trains the federated OCR model according to the training sample to obtain the model gradient of this training. The character detection device sends the model gradient to the coordination end so that the coordination end processes the model gradient to obtain a joint gradient, that is, the coordination end combines the received model gradient and the model gradients of other nodes in the alliance chain except the character detection device, and the coordination end combines each model gradient to obtain a joint gradient, and the coordination end sends the joint gradient to the character detection device.
[0094] Step S50, receive the joint gradient sent by the coordination end, and update the federated OCR model according to the joint gradient.
[0095] The character detection device receives the joint gradient sent by the coordination end, and the character detection device updates the federated OCR model according to the joint gradient. In this embodiment, when there is an error in the character detection of the OCR model, the OCR model can be effectively updated to facilitate improving the accuracy of character detection of the OCR model in the later stage.
[0096] Further, refer to Figure 4 , Figure 4 which is a schematic flowchart of the third embodiment of the character detection method based on the federated OCR model of the present invention.
[0097] Based on the above embodiments of the character detection method based on the federated OCR model of the present invention, the third embodiment of the character detection method based on the federated OCR model of the present invention is proposed.
[0098] This embodiment is the step after step S50 in the third embodiment. The difference between this embodiment and the above embodiments is as follows:
[0099] Step S60, calculate a loss value through a preset loss function, and send the loss value to the coordination end to analyze the loss value through the coordination end to determine whether the federated OCR model is updated and completed;
[0100] Step S70, if the update completion prompt sent by the coordination end is not received within a preset time interval, input a new training sample, and execute the step of training the federated OCR model according to the training sample to obtain a model gradient, and send the model gradient to the coordination end, so that the coordination end generates a joint gradient based on the model gradients fed back by multiple nodes;
[0101] Step S80, when the update completion prompt sent by the coordination end is received, output the update completion prompt.
[0102] The character detection device calculates a loss value through a preset loss function (the preset loss function refers to a function determined in advance according to the federated OCR model, which will not be elaborated in this embodiment). The character detection device sends the loss value to the coordination end to analyze the loss value through the coordination end to determine whether the federated OCR model is updated and completed; that is, the coordination end receives the loss value, the coordination end obtains the loss values of other nodes in the alliance chain except the character detection device, the coordination end combines each loss value to obtain an accumulated loss value, and the coordination end determines whether the accumulated loss value is less than a preset loss value (the preset loss value is set according to the specific scenario). When the accumulated loss value is greater than or equal to the preset loss value, the coordination end determines that the federated OCR model does not converge; when the accumulated loss value is less than the preset loss value, the coordination end determines that the federated OCR model converges. When the coordination end determines that the federated OCR model does not converge, the coordination end does not perform any operation.
[0103] If the update completion prompt sent by the coordination end is not received within a preset time interval (the preset time interval can be set according to the specific scenario, for example, set to 1 minute), the character detection device inputs a new training sample, and executes the step of training the federated OCR model according to the training sample to obtain a model gradient, and sends the model gradient to the coordination end, so that the coordination end processes the model gradient to obtain a joint gradient; when the character detection device receives the update completion prompt sent by the coordination end, the character detection device outputs the update completion prompt.
[0104] In this embodiment, the training steps of the federated OCR model are specifically described. Through sufficient learning of the federated OCR model, the accuracy of character detection is ensured.
[0105] Furthermore, based on the above-mentioned embodiment of the character detection method based on the federated OCR model of the present invention, a fourth embodiment of the character detection method based on the federated OCR model of the present invention is proposed.
[0106] This embodiment is the step after step S10 in the first embodiment. The difference between this embodiment and the above-mentioned embodiment is as follows:
[0107] When the image information is a certificate image, obtain the certificate type of the certificate image;
[0108] In step S20 of this embodiment, the step of calling the federated OCR model to perform character detection on the image information, obtaining the OCR recognition result and outputting it includes:
[0109] Call the federated OCR model corresponding to the certificate type to perform character detection on the image information, obtain the OCR recognition result and output it.
[0110] In this embodiment, the character detection model obtains the attributes of the image information. When the image information is a certificate image, the character detection model obtains the certificate type of the certificate image. The character detection model calls the federated OCR model corresponding to the certificate type to perform character detection on the image information according to the type of the character detection model, obtains the OCR recognition result and outputs it. In this embodiment, there are multiple types of federated OCR models preset in the character detection device. The character detection device can select the federated OCR model according to the type of the image information. In this way, by setting different types of federated OCR models, while reducing the complexity of the federated OCR model, the accuracy of character detection is effectively ensured.
[0111] Refer to Figure 5 , Figure 5 is a schematic diagram of the functional modules of an embodiment of the character detection device based on the federated OCR model of the present invention; The present invention also provides a character detection device based on the federated OCR model. The character detection device based on the federated OCR model includes:
[0112] A request receiving module 10, configured to obtain the image information to be recognized associated with the OCR recognition request when receiving an OCR recognition request;
[0113] A calling detection module 20, configured to call the federated OCR model to perform character detection on the image information, obtain the OCR recognition result and output it, where the federated OCR model is obtained by iteratively training the local initial OCR model based on the joint gradient sent by the coordination end, and the joint gradient is generated by the coordination end based on the model gradients of multiple parties.
[0114] In one embodiment, the character detection device based on the federated OCR model includes:
[0115] A sample marking module, configured to mark the image information in the local storage to form a training sample set, and extract a preset proportion of training samples from the training sample set;
[0116] A gradient generation module, configured to train an initial OCR model through the training samples to obtain model gradients, and send the model gradients to a coordination end, so that the coordination end generates joint gradients based on the model gradients fed back by multiple nodes;
[0117] A model update module, configured to receive the joint gradients sent by the coordination end, update the initial OCR model according to the joint gradients to obtain a trained OCR model, and obtain the OCR feature vector of the trained OCR model;
[0118] An information sending module, configured to process the OCR feature vector through a preset loss function to obtain a loss value, and send the loss value to the coordination end, so that the coordination end analyzes the loss value to determine whether the OCR model is trained;
[0119] A receiving determination module, configured to use the trained OCR model as a federated OCR model when receiving a training completion prompt sent by the coordination end.
[0120] In one embodiment, the calling detection module 20 includes:
[0121] A calling extraction sub-module, configured to call the federated OCR model to perform text detection on the image information and extract the text area in the image information;
[0122] An identification output sub-module, configured to perform character recognition on the text area through the federated OCR model to obtain the character information included in the text area, and output the character information as an OCR recognition result.
[0123] In one embodiment, the identification output sub-module includes:
[0124] An identification determination unit, configured to perform character recognition on the text area through the federated OCR model to determine the character type of the characters in the text area;
[0125] An acquisition output unit, configured to acquire the character detection sub-model corresponding to the character type in the federated OCR model, perform character recognition on the text area through the character detection sub-model to obtain the character information included in the text area, and output the character information as an OCR recognition result.
[0126] In one embodiment, the character detection device based on the federated OCR model includes:
[0127] A prompt output module, configured to output a labeling prompt to prompt the user to label the image information when the OCR recognition result is incorrect;
[0128] A labeling sending module, configured to use the labeled image information as a training sample, obtain model gradients by training the federated OCR model according to the training sample, and send the model gradients to a coordination end, so that the coordination end generates joint gradients based on the model gradients fed back by multiple nodes;
[0129] A receiving and updating module, configured to receive the joint gradients sent by the coordination end and update the federated OCR model according to the joint gradients.
[0130] In one embodiment, the character detection device based on the federated OCR model includes:
[0131] A calculation and sending module, configured to calculate a loss value through a preset loss function and send the loss value to a coordination end to analyze the loss value through the coordination end to determine whether the update of the federated OCR model is completed;
[0132] A sample input module, configured to input a new training sample if an update completion prompt sent by the coordination end is not received within a preset time interval, and execute the steps of obtaining model gradients by training the federated OCR model according to the training sample and sending the model gradients to the coordination end, so that the coordination end generates joint gradients based on the model gradients fed back by multiple nodes;
[0133] A prompt output module, configured to output the update completion prompt when receiving the update completion prompt sent by the coordination end.
[0134] In one embodiment, the prompt output module includes:
[0135] An error type determination unit, configured to determine the error type of the error when the OCR recognition result is incorrect;
[0136] A first output unit, configured to output a region labeling prompt to prompt the user to label the text region in the image information when the error type is region detection error;
[0137] A second output unit, configured to output a character labeling prompt to prompt the user to input the character information included in the image information when the error type is character recognition error.
[0138] In one embodiment, the character detection device based on the federated OCR model includes:
[0139] A type determination module, configured to obtain the document type of the document image when the image information is a document image;
[0140] The calling detection module 20 is further configured to: call the federated OCR model corresponding to the document type to perform character detection on the image information, obtain an OCR recognition result and output it.
[0141] Wherein, the method implemented when the character detection device based on the federated OCR model is executed at the above can refer to the various embodiments of the character detection method based on the federated OCR model of the present invention, which will not be elaborated here.
[0142] In the embodiments of the present invention, a federated OCR model is pre-constructed. The federated OCR model determines a joint gradient based on the model gradients of multiple nodes in the consortium chain and is jointly trained according to the joint gradient. The federated OCR model enables multiple parties to fully learn without disclosing their data privacy. In this embodiment, the federated OCR model is used to perform character detection on the image information, improving the accuracy of OCR recognition.
[0143] The present invention also provides a computer-readable storage medium.
[0144] The computer-readable storage medium of the present invention stores a computer program corresponding to OCR recognition based on the federated OCR model. When the computer program corresponding to OCR recognition based on the federated OCR model is executed by a processor, the steps of the character detection method based on the federated OCR model as described above are implemented.
[0145] Wherein, the method implemented when the computer program corresponding to OCR recognition based on the federated OCR model running on the processor is executed can refer to the various embodiments of the character detection method based on the federated OCR model of the present invention, which will not be elaborated here.
[0146] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or system including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.
[0147] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a coordination terminal, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0149] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A character detection method based on a federated OCR model, characterized in that, The character detection method based on the federated OCR model includes the following steps: Mark the image information in the local storage to form a training sample set, and extract a preset proportion of training samples from the training sample set; Train the initial OCR model with the training samples to obtain model gradients, and send the model gradients to the coordination end so that the coordination end generates joint gradients based on the model gradients fed back by multiple nodes; Receive the joint gradients sent by the coordination end, update the initial OCR model according to the joint gradients to obtain a trained OCR model, and obtain the OCR feature vector of the trained OCR model; Process the OCR feature vector through a preset loss function to obtain a loss value, and send the loss value to the coordination end to determine whether the OCR model is trained by analyzing the loss value through the coordination end; When receiving the training completion prompt sent by the coordination end, use the trained OCR model as the federated OCR model; When receiving an OCR recognition request, obtain the image information to be recognized associated with the OCR recognition request; Call the federated OCR model to perform character detection on the image information, obtain an OCR recognition result and output it, where the federated OCR model is obtained by iteratively training the local initial OCR model based on the joint gradients sent by the coordination end, and the joint gradients are processed and generated by the coordination end based on the model gradients of multiple nodes.
2. The character detection method based on the federated OCR model according to claim 1, wherein The step of calling the federated OCR model to perform character detection on the image information, obtain an OCR recognition result and output it includes: Call the federated OCR model to perform text detection on the image information, and extract the text area in the image information; Perform character recognition on the text area through the federated OCR model to obtain the character information contained in the text area, and use the character information as the OCR recognition result and output it.
3. The character detection method based on the federated OCR model according to claim 2, characterized in that, The step of performing character recognition on the text area through the federated OCR model to obtain the character information contained in the text area, and using the character information as the OCR recognition result and output it includes: Perform character recognition on the text area through the federated OCR model to determine the character type of the characters in the text area; Obtain the character detection sub-model corresponding to the character type in the federated OCR model, perform character recognition on the text area through the character detection sub-model to obtain the character information contained in the text area, and use the character information as the OCR recognition result and output it.
4. The character detection method based on the federated OCR model according to claim 1, characterized in that, After the step of calling the federated OCR model to perform character detection on the image information, obtain an OCR recognition result and output it, it includes: When the OCR recognition result is incorrect, output a marking prompt to prompt the user to mark the image information; Use the marked image information as a training sample, train the federated OCR model according to the training sample to obtain model gradients, and send the model gradients to the coordination end so that the coordination end generates joint gradients based on the model gradients fed back by multiple nodes; Receive the joint gradient sent by the coordination end, and update the federated OCR model according to the joint gradient.
5. The character detection method based on the federated OCR model according to claim 4, wherein The step of outputting a model training prompt to prompt the user to annotate the image information when the OCR recognition result is incorrect includes: When the OCR recognition result is incorrect, determine the error type of the error; If the error type is region detection error, output a region annotation prompt to prompt the user to annotate the text region in the image information; If the error type is character recognition error, output a character annotation prompt to prompt the user to input the character information contained in the image information.
6. The character detection method based on the federal OCR model according to any one of claims 1 to 5, characterized in that, After the step of obtaining the image information to be recognized associated with the OCR recognition request when receiving the OCR recognition request, the method further includes: When the image information is a certificate image, obtain the certificate type of the certificate image; The step of calling the federated OCR model to perform character detection on the image information, obtaining an OCR recognition result and outputting it includes: Call the federated OCR model corresponding to the certificate type to perform character detection on the image information, obtain an OCR recognition result and output it.
7. A character detection device based on a federated OCR model, characterized in that, The character detection device based on the federated OCR model includes: A request receiving module, configured to obtain the image information to be recognized associated with the OCR recognition request when receiving the OCR recognition request; A call detection module, configured to call the federated OCR model to perform character detection on the image information, obtain an OCR recognition result and output it, wherein the federated OCR model is obtained by iteratively training the local initial OCR model based on the joint gradient sent by the coordination end, and the joint gradient is processed and generated by the coordination end based on the model gradients of multiple parties; The character detection device based on the federated OCR model further includes: Mark the image information in the local storage to form a training sample set, and extract a preset proportion of training samples from the training sample set; Train the initial OCR model with the training samples to obtain a model gradient, and send the model gradient to the coordination end, so that the coordination end generates a joint gradient based on the model gradients fed back by multiple parties; Receive the joint gradient sent by the coordination end, update the initial OCR model according to the joint gradient to obtain a trained OCR model, and obtain the OCR feature vector of the trained OCR model; Process the OCR feature vector through a preset loss function to obtain a loss value, and send the loss value to the coordination end to determine whether the OCR model is trained by analyzing the loss value through the coordination end; When receiving the training completion prompt sent by the coordination end, use the trained OCR model as the federated OCR model.
8. A character detection device based on a federated OCR model, characterized in that The character detection device based on the federated OCR model includes: a memory, a processor, and a computer program corresponding to OCR recognition based on the federated OCR model stored on the memory and executable on the processor. When the computer program corresponding to OCR recognition based on the federated OCR model is executed by the processor, the steps of the character detection method based on the federated OCR model according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that, A computer program corresponding to OCR recognition based on the federated OCR model is stored on the computer-readable storage medium. When the computer program corresponding to OCR recognition based on the federated OCR model is executed by a processor, the steps of the character detection method based on the federated OCR model according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Federated learning method, system, and readable storage medium
CN109299728A
OCR recognition model training method and device based on crowdsourcing technology and computer equipment
CN110503089A