Model training method and apparatus, text information detection method and apparatus, and device and medium

By combining bill image samples and their annotation information, as well as bill type information, and bill type information, the detection model training problem under the requirements of privacy and compliance of bill image data is solved, and the accurate detection of bill text information is achieved.

WO2025092288A1PCT designated stage expired Publication Date: 2025-05-08MASHANG CONSUMER FINANCE CO LTD

Patent Information

Application Number
PCT/CN2024/120145
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-09-20
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively train detection models to accurately detect text information in the ticket under the requirements of privacy and compliance of the ticket image data, especially in scenarios where the number of ticket image samples is small and the type is not fixed.

Method used

By obtaining the bill image sample and its corresponding annotation information, the detection model is trained in combination with the bill type information. During the model training process, the relative weights of the first model loss and the second model loss are dynamically adjusted according to the current training iteration number to optimize the training effect of the model.

Benefits of technology

In the scenario where the number of bill image samples is small and the type is not fixed, the detection model is more accurate in the detection of bill text information, ensuring data privacy and compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024120145_08052025_PF_FP_ABST
    Figure CN2024120145_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical fields of artificial intelligence and image recognition, and relates to a model training method and apparatus, a text information detection method and apparatus, and a device and a medium. The present application can increase the accuracy of bill text information detection. The model training method comprises: acquiring a bill image sample and labeling information thereof; acquiring a bill type corresponding to the bill image sample; inputting the bill image sample and the bill type into a detection model to be trained; obtaining a first model loss on the basis of first prediction information outputted from the model and first labeling information, and obtaining a second model loss on the basis of second prediction information outputted from the model and second labeling information; determining relative weights of the first model loss and the second model loss on the basis of a current number of training iterations; and on the basis of a total model loss determined on the basis of the first model loss, the second model loss and the relative weights, training said detection model until a preset model training termination condition is met.
Need to check novelty before this filing date? Find Prior Art

Description

Model training method and text information detection method, device, equipment and medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on October 30, 2023, with application number 202311424426.5, and invention name “Bill text detection model training and detection method, device, equipment and medium”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence and image recognition technology, and in particular to a model training method, text information detection method, device, equipment and medium. Background Art

[0003] With the development of OCR (Optical Character Recognition) technology, OCR technology can be automatically applied to business processes, and many projects have been implemented in many scenarios.

[0004] Among them, due to the privacy and compliance characteristics of bill image data, it is difficult for the entity providing model algorithm services to obtain a large amount of bill image data to iteratively optimize its model algorithm, so the business entity needs to perform self-training. Self-training refers to deploying a self-training system on the business entity's private server, passing in the bill image data and annotation information that need to be detected and identified, and then automatically training to obtain the detection model. The bill image data does not need to be transmitted externally, which is safe.

[0005] Summary of the Invention

[0006] Based on this, it is necessary to provide a model training method, text information detection method, device, computer equipment, storage medium and computer program product to address the above technical problems.

[0007] In the first aspect, the present application provides a model training method. The method includes: obtaining a bill image sample and corresponding annotation information; the annotation information includes first annotation information and second annotation information, the first annotation information is text area annotation information, and the second annotation information is text box annotation information; obtaining the bill type corresponding to the bill image sample; inputting the bill image sample and the corresponding bill type into the detection model to be trained, and obtaining first prediction information and second prediction information output by the detection model, wherein the first prediction information is information corresponding to the text area, and the second prediction information is information corresponding to the text box; obtaining a first model loss based on the first prediction information and the first annotation information, and obtaining a second model loss based on the second prediction information and the second annotation information; determining the relative weight of the first model loss and the second model loss based on the current number of training iterations; and training the detection model to be trained based on the total model loss determined based on the first model loss, the second model loss, and the relative weight until the preset model training end condition is met.

[0008] In one embodiment, determining the relative weight of the first model loss and the second model loss based on the current number of training iterations includes: obtaining a preset iteration number threshold; when the current number of training iterations does not reach the preset iteration number threshold, making the relative weight of the first model loss and the second model loss greater than the preset relative weight; wherein the preset relative weight is used to represent the weight value when the weights of the first model loss and the second model loss are the same.

[0009] In one embodiment, making the relative weight of the first model loss and the second model loss greater than a preset relative weight includes: determining the relative weight corresponding to the current number of training iterations within a preset relative weight selection range based on the current number of training iterations; wherein the current number of training iterations is negatively correlated with the corresponding relative weight.

[0010] In one embodiment, the method further includes: when the current number of training iterations reaches the preset iteration number threshold, determining the relative weight of the first model loss and the second model loss to be the preset relative weight.

[0011] In a second aspect, the present application provides a method for detecting text information. The method comprises: obtaining a bill image to be detected and determining the bill type corresponding to the bill image; inputting the bill image and the corresponding bill type into a trained detection model; wherein the trained detection model is trained according to the model training method described in any of the above embodiments; and obtaining bill text information of the bill image based on text area information and text box information of the bill image output by the trained detection model.

[0012] In a third aspect, the present application further provides a detection model training device. The device includes: a sample acquisition module for acquiring a bill image sample and corresponding annotation information; the annotation information includes first annotation information and second annotation information, wherein the first annotation information is text area annotation information and the second annotation information is text box annotation information; a type acquisition module for acquiring the bill type corresponding to the bill image sample; a sample input module for inputting the bill image sample and the corresponding bill type into a detection model to be trained, and acquiring first prediction information and second prediction information output by the detection model, wherein the first prediction information is information corresponding to the text area and the second prediction information is information corresponding to the text box; a loss acquisition module for obtaining a first model loss based on the first prediction information and the first annotation information, and a second model loss based on the second prediction information and the second annotation information; a weight determination module for determining a relative weight between the first model loss and the second model loss based on the current number of training iterations; and a model training module for training the detection model to be trained according to the total model loss based on the first model loss, the second model loss, and a total model loss determined by the relative weights, until a preset model training end condition is met.

[0013] In a fourth aspect, the present application further provides a text information detection device. The device comprises: an image acquisition module for acquiring a bill image to be detected and determining the bill type corresponding to the bill image; an image input module for inputting the bill image and the corresponding bill type into a trained detection model; wherein the trained detection model is trained according to the model training method described in any of the above embodiments; and an information acquisition module for obtaining the bill text information of the bill image based on the text area information and text box information of the bill image output by the trained detection model.

[0014] In a fifth aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program: obtaining a bill image sample and corresponding annotation information; the annotation information includes first annotation information and second annotation information, the first annotation information is text area annotation information, and the second annotation information is text box annotation information; obtaining the bill type corresponding to the bill image sample; inputting the bill image sample and the corresponding bill type into a detection model to be trained, and obtaining first prediction information and second prediction information output by the detection model, wherein the first prediction information is information corresponding to the text area, and the second prediction information is information corresponding to the text box; obtaining a first model loss based on the first prediction information and the first annotation information, and obtaining a second model loss based on the second prediction information and the second annotation information; determining the relative weight of the first model loss and the second model loss based on the current number of training iterations; determining a total model loss based on the first model loss, the second model loss, and the relative weight, and training the detection model to be trained according to the total model loss until a preset model training end condition is met.

[0015] In a sixth aspect, the present application further provides a computer device. The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the computer device performs the following steps: obtaining a bill image to be detected and determining the bill type corresponding to the bill image; inputting the bill image and the corresponding bill type into a trained detection model; wherein the trained detection model is trained according to the model training method described in any of the above embodiments; and obtaining bill text information of the bill image based on the text area information and text box information of the bill image output by the trained detection model.

[0016] In a seventh aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented: obtaining a bill image sample and corresponding annotation information; the annotation information includes first annotation information and second annotation information, the first annotation information is text area annotation information, and the second annotation information is text box annotation information; obtaining the bill type corresponding to the bill image sample; inputting the bill image sample and the corresponding bill type into a detection model to be trained, and obtaining first prediction information and second prediction information output by the detection model, wherein the first prediction information is information corresponding to the text area, and the second prediction information is information corresponding to the text box; obtaining a first model loss based on the first prediction information and the first annotation information, and obtaining a second model loss based on the second prediction information and the second annotation information; determining a relative weight of the first model loss and the second model loss based on the current number of training iterations; determining a total model loss based on the first model loss, the second model loss, and the relative weight, and training the detection model to be trained according to the total model loss until a preset model training end condition is met.

[0017] In an eighth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps: acquiring a bill image to be detected and determining the bill type corresponding to the bill image; inputting the bill image and the corresponding bill type into a trained detection model; wherein the trained detection model is trained according to the model training method described in any of the above embodiments; and obtaining bill text information of the bill image based on the text area information and text box information of the bill image output by the trained detection model.

[0018] In a ninth aspect, the present application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps: obtaining a bill image sample and corresponding annotation information; the annotation information includes first annotation information and second annotation information, the first annotation information being text area annotation information, and the second annotation information being text box annotation information; obtaining the bill type corresponding to the bill image sample; inputting the bill image sample and the corresponding bill type into a detection model to be trained, and obtaining first prediction information and second prediction information output by the detection model, wherein the first prediction information is information corresponding to the text area, and the second prediction information is information corresponding to the text box; obtaining a first model loss based on the first prediction information and the first annotation information, and obtaining a second model loss based on the second prediction information and the second annotation information; determining a relative weight between the first model loss and the second model loss based on the current number of training iterations; determining a total model loss based on the first model loss, the second model loss, and the relative weight, and training the detection model to be trained according to the total model loss until a preset model training end condition is met.

[0019] In a tenth aspect, the present application further provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps: acquiring a bill image to be detected and determining the bill type corresponding to the bill image; inputting the bill image and the corresponding bill type into a trained detection model; wherein the trained detection model is trained according to the model training method described in any of the above embodiments; and obtaining bill text information of the bill image based on the text area information and text box information of the bill image output by the trained detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] FIG1 is a diagram of an application environment of a related method according to an embodiment of the present application;

[0021] FIG2 is a schematic diagram of a processing method for training a detection model in the current technology;

[0022] FIG3 is a flow chart of a model training method in an embodiment of the present application;

[0023] FIG4 is a schematic diagram of relevant data processing of the detection model in an embodiment of the present application;

[0024] FIG5 is a flow chart showing the steps of determining weights in an embodiment of the present application;

[0025] FIG6 is a flow chart of the steps for obtaining a preset threshold number of iterations in an embodiment of the present application;

[0026] FIG7 is a flow chart showing the steps of the text information detection method according to an embodiment of the present application;

[0027] FIG8 is a structural block diagram of a model training device according to an embodiment of the present application;

[0028] FIG9 is a structural block diagram of a text information detection device according to an embodiment of the present application;

[0029] FIG10 is a diagram showing the internal structure of a computer device according to an embodiment of the present application;

[0030] FIG11 is a diagram showing the internal structure of a computer device in another embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0032] The model training method and text information detection method provided in the embodiments of the present application can be applied in an application environment as shown in FIG1 . The application environment may include a terminal 110 and a server 120. Terminal 110 communicates with server 120 via a network. Server 120 can be used to execute the model training method of the present application to obtain a trained detection model. Server 120 can then transmit the trained detection model to terminal 110 for deployment. Terminal 110 can then execute the text information detection method of the present application and detect text information in a receipt image provided by a user based on the detection model trained by server 120. This application environment may also include a data storage system that can store data to be processed by server 120. The data storage system can be integrated with server 120 or deployed in the cloud or on another network server. Terminal 110 may include, but is not limited to, a personal computer, a laptop, a smartphone, a tablet computer, an Internet of Things device, and a portable wearable device. An Internet of Things device may include a smart speaker, a smart TV, a smart air conditioner, a smart car device, etc. A portable wearable device may include a smart watch, a smart bracelet, a head-mounted device, etc. The server 120 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0033] The model training method and text information detection method of the present application are described in sequence below in combination with various embodiments and corresponding drawings.

[0034] Current model training methods typically train the model by inputting labeled bill image samples. However, the number of bill image samples is typically very small, resulting in low accuracy for the trained detection model and difficulty in accurately detecting text information in bill images. This application addresses the privacy, regulatory compliance, small sample size, and non-fixed type characteristics of bill images by providing a model training method and a text information detection method to optimize the detection model's effectiveness in detecting text information in bills, making the trained detection model more accurate in detecting text information in bill images. In the context of bill image detection, current training involves inputting user-labeled bill image samples into the model. Figure 2 illustrates a detection model training method. By inputting bill image samples and labeled bill image samples into the model, calculating the corresponding model loss, and then training, a trained model is obtained. This model is used to detect text information in bills. Because the bill types corresponding to the bill image samples are non-fixed and the number of labeled bill image samples is small, the aforementioned detection model training method becomes a few-shot learning method, which significantly impacts the model training results. Furthermore, the model uses the same parameters for training all types of bill image samples, which can lead to suboptimal training results. Therefore, the detection model trained using the above training method is unable to accurately detect text information in bill images.

[0035] The model training method provided in this application can provide the corresponding bill type while inputting bill image samples into the detection model to be trained during training, so that the detection model can obtain the corresponding type hint feature based on the bill type, thereby better combining the image features of the bill image sample to detect the bill image sample, and output the corresponding text area prediction information and text box prediction information. During model training, the weights of the first model loss and the second model loss are self-adjusted according to the current number of training iterations to further improve the learning effect of few samples, thereby optimizing the model's effect on bill text information detection training. Even in scenarios where the number of bill image samples is small and the bill type is not fixed, the trained model can more accurately detect the bill text information in the bill image.

[0036] In one embodiment, as shown in FIG3 , a model training method is provided. The method may be applied to the server 120 shown in FIG1 . The method may include the following steps:

[0037] Step S301: Obtain a bill image sample and corresponding annotation information.

[0038] In the embodiments of the present application, the receipt image samples serve as training samples for the detection model. The receipt image samples can be obtained by capturing images of receipts retained by the business entity using a camera, and then transmitted to server 120. The receipt image samples can also be obtained by capturing images of receipts using other electronic devices, which is not limited in this application.

[0039] In an embodiment of the present application, the bill text corresponding to the bill image sample includes a text area and a text box. The text area can be the area of ​​each character in the bill text in the bill image sample. The text box can be a box in the bill image sample used to indicate the location of the entire bill text. In order to improve the performance of the detection model, the bill image sample can be annotated to obtain annotation information. The annotation information includes first annotation information and second annotation information. Specifically, the text area can be annotated to obtain the first annotation information, and the text box can be annotated to obtain the second annotation information. Step S302, obtain the bill type corresponding to the bill image sample.

[0040] In the embodiments of the present application, the bill image samples include various types of bills, such as checks, bills of exchange, and promissory notes. Bills of exchange include sight bills and post-dated bills; promissory notes include ordinary invoices and value-added tax invoices; and checks include transfer checks, certified checks, registered checks, and bearer checks. The bill type corresponding to the bill image sample can be the type name of each bill image, such as Type A and Type B, or the character code corresponding to the type name of each bill image, such as character code 0 corresponding to Type A, character code 1 corresponding to Type B, and so on.

[0041] Step S303: Input the bill image sample and the corresponding bill type into the detection model to be trained, and obtain the first prediction information and the second prediction information output by the detection model, wherein the first prediction information is the information corresponding to the text area, and the second prediction information is the information corresponding to the text box.

[0042] In this step, the server 120 can input the bill image sample and the corresponding bill type into the detection model to be trained, and the detection model to be trained outputs its predicted text area information and text box information based on the bill image sample and the corresponding bill type. The predicted text area information is the first prediction information, and the predicted text box information is the second prediction information. Since the bill type corresponding to the bill image sample is input during the training of the detection model, the detection model to be trained can be fine-tuned for different bill types, and the prediction effect when predicting the text area information and text box information of the bill image is better. In a specific implementation, in order to further improve the model training effect, the server 120 can perform data enhancement processing on the bill image sample. For example, the bill image sample is randomly scaled, the brightness is enhanced, etc., and then the bill image sample is randomly cropped according to a preset size to obtain a cropped bill image, and the cropped bill image is input into the detection model to be trained.

[0043] Specifically, it is explained in conjunction with Figure 4, which shows a schematic diagram of the relevant data processing of the detection model in this application. The bill image sample can be first labeled to obtain the first labeling information and the second labeling information, and then the bill image sample can be data enhanced and randomly cropped to a specified size to obtain the processed bill image sample. The bill type corresponding to the bill image sample can also be determined, and the processed bill image sample and the corresponding bill type are input into the detection model to be trained. Based on the input bill type, the type hint feature is obtained using the bill type constructor in the detection model. For example, when the bill type corresponding to the input bill image sample is a value-added tax invoice, the type hint feature of the value-added tax invoice can be obtained through the bill type constructor. The type hint feature includes that the pixels of the bill image mainly include black, red, and white. The image encoding module in the detection model processes the bill image sample to obtain image features. Based on the type hint feature and the image feature, the text hint module in the detection model processes the bill image sample to obtain text encoding. Based on the text encoding and image features, the visual cue module in the detection model processes the bill image sample to obtain a first visual feature; the first visual feature is fused with the image feature to obtain a second visual feature; the second visual feature is combined with the text encoding through Einstein summation to obtain a text feature; the text feature is processed to obtain first prediction information; and the first prediction information and the second visual feature are combined to obtain second prediction information. In a specific implementation, the image encoding module can use ResNet50; the text cue module includes two regularization layers (linear layers) and an activation function layer, with the activation function layer located between the two regularization layers; and the visual cue module can use a transformer decoder.

[0044] Step S304: Obtain a first model loss based on the first prediction information and the first annotation information, and obtain a second model loss based on the second prediction information and the second annotation information.

[0045] Step S305: Determine the relative weight of the first model loss and the second model loss according to the current number of training iterations.

[0046] Step S306: Determine the total model loss based on the first model loss, the second model loss, and the relative weight, and train the detection model to be trained according to the total model loss until the preset model training end condition is met.

[0047] Steps S304 to S306 are steps related to calculating the model loss and training the detection model to be trained. Specifically, referring to Figure 4, the first model loss and the second model loss are calculated based on the first prediction information and the second prediction information output by the detection model, and the detection model is iteratively trained according to the first model loss and the second model loss. In step S304, the first model loss is calculated based on the first prediction information and the first annotation information. The first model loss can be obtained by binarized cross entropy calculation. The second model loss is calculated based on the second prediction information and the second annotation information. In an embodiment of the present application, the second model loss can use L1 loss, and the second model loss can be directly obtained by subtracting the text box prediction information and the text box annotation information. It should be noted that in step S305, the present application determines the relative weights of the first model loss and the second model loss based on the current number of training iterations of the detection model to be trained. The weight of the second model loss (recorded as the second weight) can be set to 1, and then step S305 can be used to determine the weight of the first model loss (recorded as the first weight). Therefore, in step S306, a weighted summation can be performed based on the first model loss, the second model loss and their corresponding first weight and second weight to obtain the total model loss, and the detection model to be trained can be trained based on the total model loss until the preset model training end condition is met. In a specific implementation, the preset model training end condition can be reaching a preset maximum number of iterations, the total model loss converges, etc. Therefore, the relative weights of the first model loss and the second model loss can be self-adjusted according to the current number of training iterations during the training process, so as to further learn the knowledge in the text and improve the few-sample learning effect, so that the detection model trained in this way can have a more accurate detection effect on the text information in the bill image.

[0048] The model training method of this embodiment obtains a bill image sample and its corresponding annotation information, which includes first annotation information and second annotation information, obtains the bill type corresponding to the bill image sample, inputs the bill image sample and the corresponding bill type into the detection model to be trained, obtains the first prediction information and the second prediction information output by the model, obtains the first model loss based on the first prediction information and the first annotation information, and obtains the second model loss based on the second prediction information and the second annotation information, determines the relative weight of the first model loss and the second model loss based on the current number of training iterations, determines the total model loss based on the first model loss, the second model loss, and the relative weight, and trains the detection model to be trained based on the total model loss until a preset model training end condition is met. This solution can determine the bill type corresponding to the bill image sample during training, and provide the corresponding bill type while inputting the bill image sample into the detection model to be trained, so that the detection model can obtain the corresponding type hint feature based on the bill type, thereby better combining the image features of the bill image sample to detect the bill image sample and output the corresponding first prediction information and second prediction information. During model training, the relative weights of the first model loss and the second model loss are self-adjusted according to the current number of training iterations to further improve the few-sample learning effect, thereby optimizing the model's training effect on bill text information detection. Therefore, even in scenarios where the number of bill image samples is small and the bill types are not fixed, the trained model can have a more accurate detection effect on the bill text information in the bill image.

[0049] In some embodiments, as shown in FIG5 , determining the relative weight of the first model loss and the second model loss according to the current number of training iterations in step S305 may include:

[0050] Step S501: Obtain a preset iteration number threshold.

[0051] Step S502: If the current number of training iterations does not reach a preset iteration threshold, the relative weight of the first model loss and the second model loss is set to be greater than a preset relative weight. The preset relative weight is used to represent the weight value when the weights of the first model loss and the second model loss are the same.

[0052] Step S503: When the current number of training iterations reaches a preset iteration threshold, the relative weight of the first model loss and the second model loss is determined to be a preset relative weight.

[0053] In this embodiment, the relative weight of the first model loss and the second model loss can be dynamically changed with the current number of training iterations, shifting the model's initial focus from text region prediction to both text region prediction and text box prediction. This improves the learning effect of text region detection in scenarios with few samples of bill images and balances the learning effect of text box detection. Specifically, in step S501, a preset iteration threshold can be first determined, and then a determination can be made as to whether the current number of training iterations has reached the preset iteration threshold. Specifically, in step S502, if the current number of training iterations has not reached the preset iteration threshold, the relative weight of the first model loss and the second model loss is set to be greater than a preset relative weight, where the preset relative weight represents the weight value when the weights of the first model loss and the second model loss are the same. For example, by setting this relative weight greater than the preset relative weight, the model initially focuses on text region prediction. In a specific implementation, the preset relative weight can be 1. Additionally, in step S503, if the current number of training iterations has reached the preset iteration threshold, the relative weight of the first model loss and the second model loss can be determined to be the preset relative weight. This allows the model to focus on both text region prediction and text box prediction in the later stages of training, thereby balancing the learning effects of text region detection and text box detection.

[0054] Furthermore, in one embodiment, as shown in FIG6 , obtaining a preset iteration number threshold in step S501 may include:

[0055] Step S601: determine the preset maximum number of iterations for detection model training.

[0056] Step S602: determining the attention stage division parameter corresponding to the detection model training, wherein the attention stage division parameter is used to divide the stages in which the detection model focuses on the text area within a preset maximum number of iterations.

[0057] Step S603: Determine a preset iteration number threshold according to the preset maximum iteration number and the focus phase division parameter.

[0058] In this embodiment, the preset iteration number threshold can be calculated based on the preset maximum number of iterations and the attention stage division parameter. Among them, the preset maximum number of iterations can be set by the user. In step S601, the preset maximum number of iterations for the detection model training set by the user can be obtained and recorded as Emax. Among them, the attention stage division parameter refers to the stage used to divide the detection model's attention to the text area within the preset maximum number of iterations. The parameter can take a value between 0 and 1, such as 4 / 5. In step S603, the preset iteration number threshold 4 / 5*Emax can be obtained based on the product of the preset maximum number of iterations and the attention stage division parameter. This allows users to flexibly set a reasonable preset iteration number threshold according to the actual training needs of the model, so that the detection model can obtain a better model training effect.

[0059] Furthermore, in one embodiment, making the relative weight of the first model loss and the second model loss greater than a preset relative weight in step S502 may include:

[0060] According to the current number of training iterations, the relative weight corresponding to the current number of training iterations is determined within a preset relative weight selection range; wherein the current number of training iterations is negatively correlated with the corresponding relative weight.

[0061] In this embodiment, the relative weight of the first model loss and the second model loss can be made to gradually decrease as the current number of training iterations increases, until the current number of training iterations reaches a preset iteration threshold, and the relative weight of the first model loss and the second model loss obtains the above-mentioned preset relative weight. In this way, the model can gradually and smoothly transition from focusing on text area prediction to paying equal attention to text area prediction and text box prediction as the current number of training iterations increases, further optimizing the model training effect. Specifically, when the current number of training iterations does not reach the preset iteration threshold, the relative weight of the first model loss and the second model loss needs to be greater than the preset relative weight. In the process of selecting the preset relative weight, this embodiment can determine the relative weight corresponding to the current number of training iterations within the preset relative weight selection range. And the current number of training iterations is negatively correlated with the corresponding relative weight. For example, at the beginning of training, a larger relative weight can be selected from the relative weight selection range, and then as the current training iteration number increases, smaller and smaller relative weights are selected until the current training iteration number reaches the preset iteration number threshold, and the above-mentioned preset relative weight is obtained. In a specific implementation, the relative weight selection range can be limited by a maximum value and a minimum value, the maximum value can be 2, and the minimum value can be a preset relative weight, and the preset relative weight is equal to 1. In this way, during the model training process, as the current training iteration number increases, the weight transitions from the maximum value of 2 to the minimum value of 1, thereby optimizing the model training effect and avoiding complex design parameters.

[0062] As a specific implementation, a preset function can be used to determine the corresponding relative weight α(e) for each current training iteration number. The preset function is as follows:

[0063] Among them, e represents the number of current training iterations, E max Indicates the preset maximum number of iterations, Indicates the preset iteration threshold.

[0064] In one embodiment, as shown in FIG7 , a text information detection method is provided. The method can be applied to the terminal 110 shown in FIG1 . The method can include the following steps:

[0065] Step S701: Acquire a bill image to be detected and determine the bill type corresponding to the bill image.

[0066] In this step, the terminal 110 may obtain the bill image to be detected provided by the user and obtain the bill type corresponding to the bill image selected by the user.

[0067] Step S702: Input the bill image and the corresponding bill type into the trained detection model.

[0068] In this step, terminal 110 may first obtain a trained detection model from server 120. This detection model may be trained by server 120 according to the model training method described in any of the above embodiments and then transmitted to terminal 110, thereby enabling terminal 110 to obtain the trained detection model. Then, after obtaining a bill image to be detected and its corresponding bill type, terminal 110 may input the bill image and corresponding bill type into the trained detection model, which then outputs the text area information and text box information of the bill image.

[0069] Step S703: Obtain the bill text information of the bill image according to the text area information and text box information of the bill image output by the trained detection model.

[0070] In this step, terminal 110 can obtain the bill text information based on the text area information and text box information output by the trained detection model. Specifically, the text area information and text box information of the bill image can be returned to the user as the bill text information of the bill image. Furthermore, each bill character in the corresponding bill text can be identified based on the text area information and text box information, and the bill text can be returned to the user.

[0071] The solution of this embodiment can apply the detection model trained by the model training method of this application to the detection and recognition of bill images. In scenarios where the number of bill image samples is small and the bill types are not fixed, the bill text information can be accurately detected based on the trained detection model, so as to provide users with accurate bill text information.

[0072] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0073] Based on the same inventive concept, the embodiments of the present application also provide a model training device for implementing the aforementioned model training method, and a text information detection device for the text information detection method. The implementation solution provided by the device is similar to the implementation solution described in the aforementioned method, so the specific limitations of one or more related device embodiments provided below can be found in the above-mentioned limitations of the related method, and will not be repeated here.

[0074] In one embodiment, as shown in FIG8 , a model training device is provided. The device 800 may include:

[0075] The sample acquisition module 801 is used to acquire a bill image sample and corresponding annotation information; the annotation information includes first annotation information and second annotation information, the first annotation information is text area annotation information, and the second annotation information is text box annotation information;

[0076] A type acquisition module 802 is used to acquire the bill type corresponding to the bill image sample;

[0077] The sample input module 803 is used to input the bill image sample and the corresponding bill type into the detection model to be trained, and obtain the first prediction information and the second prediction information output by the detection model, wherein the first prediction information is the information corresponding to the text area, and the second prediction information is the information corresponding to the text box;

[0078] A loss acquisition module 804 is configured to obtain a first model loss based on the first prediction information and the first annotation information, and to obtain a second model loss based on the second prediction information and the second annotation information;

[0079] A weight determination module 805 is configured to determine a relative weight between the first model loss and the second model loss based on the current number of training iterations;

[0080] The model training module 806 is used to determine the total model loss based on the first model loss, the second model loss, and the relative weight, and train the detection model to be trained according to the total model loss until the preset model training end condition is met.

[0081] In one embodiment, the weight determination module 805 is used to obtain a preset iteration number threshold; when the current training iteration number does not reach the preset iteration number threshold, the relative weight of the first model loss and the second model loss is made greater than the preset relative weight; wherein the preset relative weight is used to represent the weight value when the weights of the first model loss and the second model loss are the same.

[0082] In one embodiment, the weight determination module 805 is used to determine the relative weight corresponding to the current number of training iterations within a preset relative weight selection range based on the current number of training iterations; wherein the current number of training iterations is negatively correlated with the corresponding relative weight.

[0083] In one embodiment, the weight determination module 805 is further used to determine the relative weight of the first model loss and the second model loss as the preset relative weight when the current number of training iterations reaches the preset iteration threshold.

[0084] In one embodiment, the weight determination module 805 is used to determine the preset maximum number of iterations for the detection model training; determine the attention stage division parameter corresponding to the detection model training; wherein the attention stage division parameter is used to divide the stage in which the detection model focuses on the text area within the preset maximum number of iterations; and determine the preset iteration number threshold based on the preset maximum number of iterations and the attention stage division parameter.

[0085] In one embodiment, as shown in FIG9 , a text information detection device is provided. The device 900 may include:

[0086] The image acquisition module 901 is used to acquire a bill image to be detected and determine the bill type corresponding to the bill image;

[0087] An image input module 902 is configured to input the bill image and the corresponding bill type into a trained detection model; wherein the trained detection model is trained according to the model training method described in any of the above embodiments;

[0088] The information acquisition module 903 is used to obtain the bill text information of the bill image according to the text area information and text box information of the bill image output by the trained detection model.

[0089] Each module in the above-mentioned apparatus may be implemented in whole or in part by software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.

[0090] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be shown in FIG10 . The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as bill image samples. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a model training method is implemented.

[0091] In one embodiment, a computer device is provided, which may be a terminal. Its internal structure diagram may be as shown in Figure 11. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, where the wireless means may be implemented via Wi-Fi, a mobile cellular network, NFC (near field communication), or other technologies. When executed by the processor, the computer program implements a text information detection method. The display unit of the computer device is used to form a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0092] Those skilled in the art will understand that the structures shown in Figures 10 and 11 are merely block diagrams of partial structures related to the solution of the present application, and do not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figures, or combine certain components, or have a different component arrangement.

[0093] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0094] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0095] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0097] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0098] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0099] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A model training method, the method comprising: Obtain bill image samples and corresponding annotation information; The annotation information includes first annotation information and second annotation information, the first annotation information is text area annotation information, and the second annotation information is text box annotation information; Obtaining the bill type corresponding to the bill image sample; Input the bill image sample and the corresponding bill type into the detection model to be trained, and obtain the first prediction information and the second prediction information output by the detection model, wherein the first prediction information is the information corresponding to the text area, and the second prediction information is the information corresponding to the text box; Obtain a first model loss according to the first prediction information and the first labeling information, and obtain a second model loss according to the second prediction information and the second labeling information; Determine the relative weight of the first model loss and the second model loss according to the current number of training iterations; Based on the total model loss determined by the first model loss, the second model loss and the relative weight, the detection model to be trained is trained according to the total model loss until a preset model training end condition is met.

2. According to the method of claim 1, determining the relative weight of the first model loss and the second model loss according to the current number of training iterations comprises: Get the preset iteration number threshold; When the current number of training iterations does not reach the preset iteration number threshold, the relative weight of the first model loss and the second model loss is made greater than the preset relative weight; wherein the preset relative weight is used to represent the weight value when the weights of the first model loss and the second model loss are the same.

3. The method according to claim 2, wherein the step of making the relative weight of the first model loss and the second model loss greater than a preset relative weight comprises: According to the current number of training iterations, a relative weight corresponding to the current number of training iterations is determined within a preset relative weight selection range; wherein the current number of training iterations is negatively correlated with the corresponding relative weight.

4. The method according to claim 2, further comprising: When the current number of training iterations reaches the preset iteration number threshold, the relative weight of the first model loss and the second model loss is determined to be the preset relative weight.

5. A text information detection method, the method comprising: Acquire a bill image to be detected, and determine the bill type corresponding to the bill image; Inputting the bill image and the corresponding bill type into a trained detection model; wherein the trained detection model is trained according to any one of claims 1 to 4; The bill text information of the bill image is obtained according to the text area information and text box information of the bill image output by the trained detection model.

6. A model training device, comprising: A sample acquisition module is used to obtain bill image samples and corresponding annotation information; The annotation information includes first annotation information and second annotation information, the first annotation information is text area annotation information, and the second annotation information is text box annotation information; A type acquisition module, used to acquire the bill type corresponding to the bill image sample; A sample input module, used to input the bill image sample and the corresponding bill type into the detection model to be trained, and obtain the first prediction information and the second prediction information output by the detection model, wherein the first prediction information is the information corresponding to the text area, and the second prediction information is the information corresponding to the text box; a loss acquisition module, configured to obtain a first model loss according to the first prediction information and the first annotation information, and obtain a second model loss according to the second prediction information and the second annotation information; A weight determination module, used to determine the relative weight of the first model loss and the second model loss according to the current number of training iterations; A model training module is used to determine the loss of the first model, the loss of the second model and the relative weight. A total model loss is determined, and the detection model to be trained is trained according to the total model loss until a preset model training end condition is met.

7. A text information detection device, the device comprising: An image acquisition module, used to acquire a bill image to be detected and determine the bill type corresponding to the bill image; An image input module, used for inputting the bill image and the corresponding bill type into a trained detection model; wherein the trained detection model is trained according to any one of claims 1 to 4; The information acquisition module is used to obtain the bill text information of the bill image according to the text area information and text box information of the bill image output by the trained detection model.

8. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of claims 1 to 4 or the steps of the method according to claim 5 when executing the computer program.

9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 or the steps of the method according to claim 5 are implemented.

Citation Information

Patent Citations

  • OCR template learning method and device based on small sample, electronic equipment and medium

    CN110874618A

  • Mixed-pasting bill image processing method, device, computer equipment and storage medium

    CN111931664A

  • Image processing method, device and system

    CN114973218A

  • Bill text detection model training and detection method and device, equipment and medium

    CN117975473A

  • Method, apparatus, device and storage medium for recognizing bill image

    US20210383107A1

Cited By

  • Foreign language bill image translation method and device based on artificial intelligence and medium

    CN120913229A