Model training method and device based on multi-directional text dataset

Through the training method of multi-directional text dataset, polygonal areas are converted into rectangular text images, which solves the deformation and ambiguity problems of the text recognition model and improves the accuracy and efficiency of text recognition.

CN115240204BActive Publication Date: 2025-10-10GUANGZHOU YOUMI INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210723662.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-10-10
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

In the existing technology, during the text recognition process, simple alignment, stretching or rotation processing causes text image deformation or character ambiguity, which increases the difficulty of text recognition and reduces the accuracy of text recognition.

Method used

By obtaining a multi-directional text dataset, using the dataset to create algorithms and graphic transformation models, converting polygonal areas into easily recognizable rectangular text images, calculating the loss value and training the target transformation model to ensure that the image meets the preset standards.

Benefits of technology

The recognition accuracy and efficiency of the text recognition model are improved, the recognition difficulty is reduced, and the reliability and training efficiency of the model are improved through intelligent data set creation and loss value calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240204B_ABST
    Figure CN115240204B_ABST
Patent Text Reader

Abstract

The application discloses a kind of model training method and device based on multi-direction text data set, the method includes: according to data set production algorithm, text processing operation is executed to the image to be processed, obtain target data set, according to the graph transformation model to be trained and target data set, the graph transformation operation is executed to the target image obtained, and transformation image is obtained;Again according to feature analysis network, the first loss value of the target image and transformation image is calculated, the second loss value of the predetermined contrast image and transformation image, and the first and second loss values are determined as target loss value;When target loss value is less than preset loss value, determine the graph transformation model to be trained as the target transformation model of completion training, and the target transformation model is used to transform the input image into corrected image meeting preset image standard.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text recognition technology, and in particular to a model training method and device based on a multi-directional text dataset. Background Art

[0002] In the prior art, when designing algorithms for text recognition, the conventional approach is to perform a simple alignment and stretching process on the text image before inputting it into the recognition model. However, this approach easily causes serious deformation of the aligned and stretched text image. Another approach is to perform a simple rotation (such as a 90-degree rotation) on the text image before inputting it into the recognition model. This approach has the technical drawback that some characters will be ambiguous after rotation. For example, "一" will become "|" after rotation, and other common characters will become characters that are difficult for the human eye to distinguish after being rotated 90 degrees. This will also increase the difficulty of text recognition by the text model, that is, reduce the text recognition accuracy of the text model. It can be seen that it is particularly important to provide a method for improving the text recognition accuracy of text images. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a model training method and device based on a multi-directional text dataset, which can convert an input image containing a polygonal area into an easily recognizable rectangular text image, thereby reducing the text recognition difficulty of the text recognition model and thus improving the recognition accuracy of the text recognition model in recognizing text in the image area.

[0004] In order to solve the above technical problems, the first aspect of the present invention discloses a model training method based on a multi-directional text dataset, the method comprising:

[0005] Acquire a first image and a second image, wherein the first image includes at least one annotated text and a first region corresponding to the annotated text, and the second image is an image obtained after image preprocessing of the first image, and includes the annotated text and a second region corresponding to the annotated text;

[0006] performing a preset text processing operation on the determined image to be processed according to the determined data set production algorithm to obtain a target data set, wherein the target data set includes text data in at least two text extension directions;

[0007] performing a graphic transformation operation on the second image according to the determined graphic transformation model to be trained and the target data set to obtain a transformed image corresponding to the second image;

[0008] Calculating a first loss value between the second image and the transformed image, and calculating a second loss value between the first image and the transformed image according to a preset feature analysis network, and determining the first loss value and the second loss value as target loss values;

[0009] When it is determined that the target loss value is less than the preset loss value, the graphic transformation model to be trained is determined to be the target transformation model that has completed training. The target transformation model is used to transform the input image into a corrected image that meets the preset image standard.

[0010] As an optional embodiment, in the first aspect of the present invention, performing a preset text processing operation on the determined image to be processed according to the determined dataset production algorithm to obtain a target dataset includes:

[0011] Acquire a predetermined data set to be processed, the data set to be processed comprising a plurality of images to be processed, each of the images to be processed comprising text data and non-text data, the text data of each image to be processed comprising a text to be processed and a text format corresponding to the text to be processed, and each text to be processed comprising a plurality of subtexts;

[0012] Performing an image segmentation processing operation on all the images to be processed according to a preset image segmentation processing algorithm to obtain text data and non-text data of each of the images to be processed;

[0013] According to a preset text reconstruction algorithm and the non-text data of each of the images to be processed, a text reconstruction operation is performed on the text data of each of the images to be processed to obtain a reconstruction result of the text data of each of the images to be processed, and a target data set is generated based on all the reconstruction results and all the images to be processed.

[0014] As an optional embodiment, in the first aspect of the present invention, performing a text reconstruction operation on the text data of each of the images to be processed based on a preset text reconstruction algorithm and the non-text data of each of the images to be processed to obtain a reconstruction result of the text data of each of the images to be processed includes:

[0015] Performing a first reconstruction operation on the text to be processed of each of the images to be processed according to a preset first reconstruction algorithm to obtain a first reconstruction result for each of the images to be processed, wherein the first reconstruction result for each of the images to be processed is all subtexts included in the text to be processed of the image to be processed;

[0016] Determining a transformed text format of the text to be processed in each of the images to be processed according to the text format corresponding to the text to be processed in each of the images to be processed;

[0017] performing a second reconstruction operation on the first reconstruction result of each of the to-be-processed images according to the preset second reconstruction algorithm and the transformed text format of the to-be-processed text of each of the to-be-processed images, to obtain a second reconstruction result of each of the to-be-processed images, wherein the text format of the to-be-processed text of each of the to-be-processed images in the second reconstruction result of each of the to-be-processed images is the transformed text format corresponding to the to-be-processed text;

[0018] generating target reconstruction data of each of the to-be-processed images as the reconstruction result of the text data of each of the to-be-processed images according to the second reconstruction result of each of the to-be-processed images and the non-text data of each of the to-be-processed images.

[0019] As an optional implementation, in the first aspect of the present application, the method further comprises:

[0020] extracting the labeled text corresponding to the first image, and performing a text matching operation on the labeled text according to a preset text library to obtain at least one target text with a matching degree less than a preset matching degree threshold with the labeled text;

[0021] selecting a preset number of target backgrounds from the determined background library;

[0022] for each of the target texts, covering the target text on all the target backgrounds to obtain a text data set corresponding to the target text, wherein the text data set corresponding to the target text comprises the preset number of target images, and each of the target images is generated by the target text and any of the target backgrounds;

[0023] determining all the text data sets corresponding to the target texts as the target data set.

[0024] As an optional implementation, in the first aspect of the present application, after the step of selecting a preset number of target backgrounds from the determined background library, the method further comprises:

[0025] performing a preset text enhancement operation on each of the target texts to update the target text;

[0026] analyzing each of the updated target texts and analyzing each of the target backgrounds to obtain a text length of each of the target texts and a background length of each of the target backgrounds;

[0027] For each of the target backgrounds, determining whether the length difference is less than or equal to a preset length threshold; when it is determined that the length difference is greater than the preset length threshold, performing the operation of overlaying the target text on all the target backgrounds to obtain a text data set corresponding to the target text for each of the target texts, wherein the length difference is the difference between the background length of the target background and the text length of each of the target texts;

[0028] When it is determined that the length difference is less than or equal to the preset length threshold, performing a background transformation operation on the target background according to the length difference to update the target background, and performing the operation of overlaying the target text on all the target backgrounds for each target text to obtain a text data set corresponding to the target text;

[0029] The difference between the updated background length of the target background and the text length of the target text is greater than the preset length threshold.

[0030] As an optional implementation manner, in the first aspect of the present invention, when it is determined that the target loss value is greater than or equal to the preset loss value, the method further includes:

[0031] Repeat the steps of updating the training times corresponding to the graphic transformation model to be trained and updating the graphic transformation model to be trained, and perform the image transformation operation on the second image according to the graphic transformation model to be trained and the target data set to obtain a transformed image corresponding to the second image, calculate the first loss value and the second loss value according to the feature analysis network, and determine the first loss value and the second loss value as the target loss value corresponding operations, until it is determined that the target loss value is less than the preset loss value, and the graphic transformation model to be trained is determined to be the target transformation model that has completed training.

[0032] As an optional embodiment, in the first aspect of the present invention, the method further comprises:

[0033] performing an image transformation operation on an input image to be corrected according to the target transformation model to obtain a corresponding image transformation result, wherein the image to be corrected is an image including annotated text;

[0034] performing a correction and comparison operation on the image transformation result according to the preset image standard to obtain a correction and comparison result, and determining that the target transformation model meets the preset image standard when the correction and comparison result indicates that the degree of matching between the image transformation result and the preset image standard is greater than or equal to a preset correction degree threshold;

[0035] When the correction comparison result indicates that the matching degree between the image transformation result and the preset image standard is less than a preset correction degree threshold, correction information is generated based on the correction comparison result and the image transformation result. The correction information is provided to the processing personnel responsible for the target transformation model so that the processing personnel adjust the training parameters of the target transformation model according to the correction information.

[0036] A second aspect of the present invention discloses a model training device based on a multi-directional text dataset, the device comprising:

[0037] an acquisition module, configured to acquire a first image and a second image, wherein the first image includes at least one annotated text and a first region corresponding to the annotated text, and the second image is an image obtained after image preprocessing of the first image, and includes the annotated text and a second region corresponding to the annotated text;

[0038] a text processing module, configured to prepare an algorithm based on the determined data set, perform a preset text processing operation on the determined image to be processed, and obtain a target data set, wherein the target data set includes text data in at least two text extension directions;

[0039] a graphics transformation module, configured to perform a graphics transformation operation on the second image according to the determined graphics transformation model to be trained and the target data set, to obtain a transformed image corresponding to the second image;

[0040] a calculation module, configured to calculate a first loss value between the second image and the transformed image, and calculate a second loss value between the first image and the transformed image according to a preset feature analysis network;

[0041] a determination module, configured to determine the first loss value and the second loss value calculated by the calculation module as target loss values;

[0042] The determination module is also used to determine that the graphic transformation model to be trained is the target transformation model that has completed training when it is determined that the target loss value is less than the preset loss value. The target transformation model is used to transform the input image into a corrected image that meets the preset image standard.

[0043] As an optional implementation, in the second aspect of the present invention, the text processing module includes:

[0044] an acquisition submodule, configured to acquire a predetermined data set to be processed, wherein the data set to be processed includes a plurality of images to be processed, each of the images to be processed includes text data and non-text data, the text data of each image to be processed includes a text to be processed and a text format corresponding to the text to be processed, and each text to be processed includes a plurality of subtexts;

[0045] An image segmentation submodule is configured to perform an image segmentation processing operation on all the images to be processed according to a preset image segmentation processing algorithm to obtain text data and non-text data of each of the images to be processed;

[0046] The reconstruction submodule is used to perform a text reconstruction operation on the text data of each image to be processed according to a preset text reconstruction algorithm and the non-text data of each image to be processed, obtain a reconstruction result of the text data of each image to be processed, and generate a target data set based on all the reconstruction results and all the images to be processed.

[0047] As an optional embodiment, in the second aspect of the present invention, the reconstruction submodule performs a text reconstruction operation on the text data of each of the images to be processed according to a preset text reconstruction algorithm and the non-text data of each of the images to be processed, and a method for obtaining a reconstruction result of the text data of each of the images to be processed specifically includes:

[0048] Performing a first reconstruction operation on the text to be processed of each of the images to be processed according to a preset first reconstruction algorithm to obtain a first reconstruction result for each of the images to be processed, wherein the first reconstruction result for each of the images to be processed is all subtexts included in the text to be processed of the image to be processed;

[0049] Determining a transformed text format of the text to be processed in each of the images to be processed according to the text format corresponding to the text to be processed in each of the images to be processed;

[0050] performing a second reconstruction operation on the first reconstruction result of each of the images to be processed according to a preset second reconstruction algorithm and the transformed text format of the text to be processed of each of the images to be processed, thereby obtaining a second reconstruction result of each of the images to be processed, wherein the text format corresponding to the text to be processed of each of the images to be processed in the second reconstruction result is the transformed text format corresponding to the text to be processed;

[0051] According to the second reconstruction result of each of the images to be processed and the non-text data of each of the images to be processed, target reconstruction data of each of the images to be processed is generated as a reconstruction result of the text data of each of the images to be processed.

[0052] As an optional implementation, in the second aspect of the present invention, the text processing module includes:

[0053] a first processing submodule, configured to extract annotated text corresponding to the first image, and perform a text matching operation on the annotated text according to a preset text library to obtain at least one target text whose matching degree with the annotated text is less than a preset matching degree threshold;

[0054] The first processing submodule is further configured to select a preset number of target backgrounds from the determined background image library;

[0055] The first processing submodule is further configured to, for each target text, overlay the target text on all the target backgrounds to obtain a text dataset corresponding to the target text, wherein the text dataset corresponding to the target text includes the preset number of target images, each target image being generated from the target text and any one of the target backgrounds;

[0056] The determination submodule is configured to determine the text data sets corresponding to all the target texts as the target data sets.

[0057] As an optional embodiment, in the second aspect of the present invention, the first processing submodule is further configured to, after selecting a preset number of target backgrounds from the determined background library, perform a preset text enhancement operation on each target text to update the target text;

[0058] The text processing module further includes:

[0059] A second processing submodule is configured to analyze each of the updated target texts and each of the target backgrounds to obtain a text length of each of the target texts and a background length of each of the target backgrounds;

[0060] a judging submodule, configured to judge, for each target background, whether the length difference is less than or equal to a preset length threshold, and, when it is judged that the length difference is greater than the preset length threshold, trigger the first processing submodule to perform, for each target text, an operation of overlaying the target text on all the target backgrounds to obtain a text data set corresponding to the target text, wherein the length difference is the difference between the background length of the target background and the text length of each target text;

[0061] The second processing submodule is further configured to, when it is determined that the length difference is less than or equal to the preset length threshold, perform a background transformation operation on the target background according to the length difference to update the target background, and trigger the first processing submodule to perform the operation of overlaying the target text on all the target backgrounds for each target text to obtain a text dataset corresponding to the target text;

[0062] The difference between the updated background length of the target background and the text length of the target text is greater than the preset length threshold.

[0063] As an optional embodiment, in the second aspect of the present invention, the device further includes:

[0064] An iterative training module is used to repeatedly execute the updating of the training times corresponding to the graphic transformation model to be trained and the updating of the graphic transformation model to be trained when it is determined that the target loss value is greater than or equal to the preset loss value, and trigger the graphic transformation module to execute the image transformation operation on the second image according to the graphic transformation model to be trained and the target data set to obtain a transformed image corresponding to the second image, trigger the calculation module to execute the calculation of the first loss value and the calculation of the second loss value according to the feature analysis network, and trigger the determination module to execute the operation of determining the first loss value and the second loss value as the target loss value corresponding to the target loss value, until it is determined that the target loss value is less than the preset loss value, and the graphic transformation model to be trained is determined to be the target transformation model that has completed training.

[0065] As an optional embodiment, in the second aspect of the present invention, the image transformation module is further configured to perform an image transformation operation on an input image to be corrected according to the target transformation model to obtain a corresponding image transformation result, wherein the image to be corrected is an image including annotated text;

[0066] The device further comprises:

[0067] a correction and comparison module, configured to perform a correction and comparison operation on the image transformation result according to the preset image standard to obtain a correction and comparison result, and determine that the target transformation model meets the preset image standard when the correction and comparison result indicates that the degree of match between the image transformation result and the preset image standard is greater than or equal to a preset correction degree threshold;

[0068] A generation module is used to generate correction information based on the correction comparison result and the image transformation result when the correction comparison result indicates that the matching degree between the image transformation result and the preset image standard is less than a preset correction degree threshold. The correction information is used to be provided to the processing personnel responsible for the target transformation model so that the processing personnel can adjust the training parameters of the target transformation model according to the correction information.

[0069] A third aspect of the present invention discloses another model training device based on a multi-directional text dataset, the device comprising:

[0070] a memory storing executable program code;

[0071] a processor coupled to the memory;

[0072] The processor calls the executable program code stored in the memory to execute the model training method based on a multi-directional text dataset disclosed in the first aspect of the present invention.

[0073] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the model training method based on a multi-directional text dataset disclosed in the first aspect of the present invention.

[0074] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0075] In an embodiment of the present invention, a model training method based on a multi-directional text dataset is provided, the method comprising: obtaining a first image and a second image, the first image comprising at least one annotated text and a first area corresponding to the annotated text, the second image being an image obtained after image preprocessing of the first image, and the second image comprising the annotated text and a second area corresponding to the annotated text; performing a preset text processing operation on the determined image to be processed according to a determined dataset production algorithm to obtain a target dataset, the target dataset comprising text data in at least two text extension directions; performing a graphic transformation operation on the second image according to a determined graphic transformation model to be trained and the target dataset to obtain a transformed image corresponding to the second image; calculating a first loss value between the second image and the transformed image, and calculating a second loss value between the first image and the transformed image according to a preset feature analysis network, and determining the first loss value and the second loss value as target loss values; when it is determined that the target loss value is less than the preset loss value, determining that the graphic transformation model to be trained is a target transformation model that has completed training, the target transformation model being used to transform the input image into a corrected image that meets a preset image standard. It can be seen that the implementation of the present invention can produce an algorithm based on a preset data set and intelligently produce a target data set. The target data set is used as a training sample for subsequent training of the graphic transformation model, which is beneficial to improving the training efficiency; it can also intelligently calculate the first loss value and the second loss value as evaluation indicators of the graphic transformation model to be trained, thereby improving the reliability of the obtained target transformation model; in addition, the trained target transformation model is used to transform the input image into a corrected image of a preset image standard. The corrected image is used as the recognition object of the subsequent text recognition model, which reduces the recognition difficulty of the text recognition model, that is, it is beneficial to improve the recognition efficiency of the text recognition model and improve the accuracy of the recognition results obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0077] Figure 1 This is a flow chart of a model training method based on a multi-directional text dataset disclosed in an embodiment of the present invention;

[0078] Figure 2 This is a flowchart of another model training method based on a multi-directional text dataset disclosed in an embodiment of the present invention;

[0079] Figure 3 is a structural schematic view of a model training device based on a multi-directional text data set according to an embodiment of the present application;

[0080] Figure 4 is a structural schematic view of another model training device based on a multi-directional text data set according to an embodiment of the present application;

[0081] Figure 5 is a structural schematic view of still another model training device based on a multi-directional text data set according to an embodiment of the present application;

[0082] Figure 6 is a structural schematic view of another model training device based on a multi-directional text data set according to an embodiment of the present application;

[0083] Figure 7 is a structural schematic view of still another model training device based on a multi-directional text data set according to an embodiment of the present application;

[0084] Figure 8 is a data effect view of a multi-directional text data set generation method according to an embodiment of the present application. DETAILED DESCRIPTION

[0085] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor fall within the scope of protection of the present application.

[0086] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish different objects, and are not used to describe a particular order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or the like that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed or can optionally include other steps or units inherent to the process, method, product, or the like.

[0087] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0088] The present invention discloses a model training method and device based on a multi-directional text dataset. The method can intelligently produce a target dataset based on a preset dataset production algorithm. The target dataset is used as a training sample for subsequent training of a graphic transformation model, which is beneficial to improving training efficiency. The method can also intelligently calculate a first loss value and a second loss value as evaluation indicators for the graphic transformation model to be trained, thereby improving the reliability of the obtained target transformation model. In addition, the trained target transformation model is used to transform an input image into a corrected image of a preset image standard. The corrected image serves as the recognition object of a subsequent text recognition model, reducing the recognition difficulty of the text recognition model, that is, it is beneficial to improving the recognition efficiency of the text recognition model and improving the accuracy of the obtained recognition results. Detailed descriptions are given below.

[0089] Example 1

[0090] See also Figure 1 , Figure 1 This is a flow chart of a model training method based on a multi-directional text dataset disclosed in an embodiment of the present invention. Figure 1 The model training method based on a multi-directional text dataset described above can be applied to a model training device based on a multi-directional text dataset, and the embodiment of the present invention does not limit this. Figure 1 As shown, the model training method based on the multi-directional text dataset may include the following operations:

[0091] 101. Acquire a first image and a second image, where the first image includes at least one annotated text and a first area corresponding to the annotated text, and the second image includes the annotated text and a second area corresponding to the annotated text.

[0092] In the embodiment of the present invention, it should be noted that the first area is an area corresponding to the rectangular annotation method, and the second area is an area corresponding to the polygonal annotation method. In actual applications, the first image and the second image are consistent in text content and background template, and only the annotation method is different, that is, the annotation area where the text content in the first image is located is a rectangular area, and the area where the text content in the second image is located is a polygonal area.

[0093] In the embodiment of the present invention, the second image is an image obtained by performing image preprocessing on the first image. The method of performing image preprocessing on the first image to obtain the second image may specifically include the following operations:

[0094] Performing an image preprocessing operation on the first image using the determined image preprocessing function to obtain a preprocessed image corresponding to the first image;

[0095] Selecting at least one pre-background image matching the pre-processed image from the determined background database;

[0096] An image synthesis operation is performed on the pre-processed image and each pre-background image to obtain a synthesis result as the second image.

[0097] In an embodiment of the present invention, it is further explained as follows: in actual application, the image to be processed (first image) is slightly deformed by stretching methods such as affine transformation or fan transformation, and the text image including the rectangular mark is transformed into a batch of polygonal text images; and a polygonal art template image (the background pixel values ​​outside the polygonal area are all 0) is selected to manually synthesize a batch of polygonal text images (second images) for a section of text. Finally, a batch of polygonal text images are constructed (each image has a corresponding original rectangular image). Among them, the synthesized polygonal text image and the corresponding original rectangular image need to be consistent in background template color and text style. This is to avoid destroying the content and style of the original image.

[0098] 102. According to the determined data set production algorithm, a preset text processing operation is performed on the determined image to be processed to obtain a target data set.

[0099] In an embodiment of the present invention, the target data set includes text data in at least two text extension directions. Further, the text data in the text extension directions may include horizontal text data and vertical text data.

[0100] 103. Perform a graphic transformation operation on the second image according to the determined graphic transformation model to be trained and the target data set to obtain a transformed image corresponding to the second image.

[0101] In the embodiment of the present invention, when the area where the marked text is located in the second image is a polygonal area, the transformed image corresponding to the second image is an image in which the polygonal area where the marked text is located is converted into a rectangular area.

[0102] 104. Calculate a first loss value between the second image and the transformed image, and calculate a second loss value between the first image and the transformed image according to a preset feature analysis network, and determine the first loss value and the second loss value as target loss values.

[0103] In an embodiment of the present invention, when calculating the first loss value, the second image and the transformed image are input into a feature extraction network to extract an input feature vector of the classification layer, wherein the feature extraction network may be a pre-trained model based on ImageNet, such as ResNet18. Specifically, in practical applications, the Gram loss of the second image and the transformed image may be calculated using a preset style loss function, and the formula of the style loss function is as follows:

[0104] is the output of the style image at position (i, j, k) of the CNN layer l, where (i, j, k) corresponds to height, width, and channel;

[0105]

[0106]

[0107]

[0108] Among them, G is a k*k size gram matrix, k is the number of channels;

[0109] G kk′ is the value of the position (k, k') in the gram matrix, which is obtained by accumulating the product of the values ​​of the corresponding positions of the kth and k'th channels; the second image and the transformed image are input into the style loss function respectively to obtain two Gram matrices. Substitute the two Gram matrices into J style In (S, G), the difference between the two matrices is calculated.

[0110] In an embodiment of the present invention, the calculation function used to calculate the second loss value is a predetermined regression loss function, which may be a mean absolute error function MAE. The image feature data corresponding to the first image and the image feature data corresponding to the transformed image are input into the loss regression function to calculate the loss of the two images. Specifically, the formula of the regression loss function is as follows:

[0111]

[0112] 105. When it is determined that the target loss value is less than the preset loss value, the graphic transformation model to be trained is determined to be the target transformation model that has completed training.

[0113] In an embodiment of the present invention, after determining that the graphic transformation model to be trained is the target transformation model that has completed training, the target transformation model is used to transform the input image into a corrected image that meets the preset image standard. In actual applications, the trained target transformation model can transform the input image including polygonal text content into an image including rectangular text content.

[0114] It can be seen that implementation Figure 1 The described model training method based on a multi-directional text dataset can intelligently produce a target dataset based on a preset dataset production algorithm. The target dataset is used as a training sample for subsequent training of a graphic transformation model, which is beneficial to improving training efficiency. It can also intelligently calculate the first loss value and the second loss value as evaluation indicators for the graphic transformation model to be trained, thereby improving the reliability of the obtained target transformation model. In addition, the trained target transformation model is used to transform the input image into a corrected image of a preset image standard. The corrected image is used as the recognition object of a subsequent text recognition model, which reduces the recognition difficulty of the text recognition model, and is beneficial to improving the recognition efficiency of the text recognition model and improving the accuracy of the recognition results.

[0115] In an optional embodiment, the method of performing a preset text processing operation on the determined image to be processed according to the determined dataset production algorithm to obtain the target dataset may specifically include the following operations:

[0116] Obtaining a predetermined data set to be processed, the data set to be processed including a plurality of images to be processed, each image to be processed including text data and non-text data, the text data of each image to be processed including a text to be processed and a text format corresponding to the text to be processed, and each text to be processed including a plurality of subtexts;

[0117] Perform image segmentation processing on all images to be processed according to a preset image segmentation processing algorithm to obtain text data and non-text data of each image to be processed;

[0118] According to the preset text reconstruction algorithm and the non-text data of each image to be processed, a text reconstruction operation is performed on the text data of each image to be processed to obtain the reconstruction result of the text data of each image to be processed, and a target data set is generated based on all the reconstruction results and all the images to be processed.

[0119] In this optional embodiment, it should be noted that the non-text data of each image to be processed can be image background data other than text; wherein, the preset image segmentation processing algorithm can be an adaptive grayscale image threshold segmentation algorithm, which is not limited in the embodiment of the present invention.

[0120] In this optional embodiment, optionally, the above-mentioned method of performing a text reconstruction operation on the text data of each image to be processed based on a preset text reconstruction algorithm and the non-text data of each image to be processed to obtain a reconstruction result of the text data of each image to be processed may specifically include the following operations:

[0121] Performing a first reconstruction operation on the to-be-processed text of each to-be-processed image according to a preset first reconstruction algorithm to obtain a first reconstruction result for each to-be-processed image, wherein the first reconstruction result for each to-be-processed image is all subtexts included in the to-be-processed text of the to-be-processed image;

[0122] Determining a transformed text format of the text to be processed in each image to be processed according to the text format corresponding to the text to be processed in each image to be processed;

[0123] performing a second reconstruction operation on the first reconstruction result of each image to be processed according to a preset second reconstruction algorithm and a transformed text format of the text to be processed of each image to be processed, thereby obtaining a second reconstruction result of each image to be processed, wherein the text format corresponding to the text to be processed of each image to be processed in the second reconstruction result of each image to be processed is the transformed text format corresponding to the text to be processed;

[0124] According to the second reconstruction result of each image to be processed and the non-text data of each image to be processed, target reconstruction data of each image to be processed is generated as the reconstruction result of the text data of each image to be processed.

[0125] In this optional embodiment, the first reconstruction algorithm can be the dilation algorithm in image morphology processing, and the corresponding first reconstruction result is the dilation map of each image to be processed, so as to connect the various parts of a character together (such as the dot in the jade character will be connected to the character "王"); the second reconstruction algorithm can be the kmeans center clustering algorithm, that is, after obtaining the dilation map, each character is isolated by the center clustering algorithm; the processing effect in actual application can be seen in the following figure. Figure 8 , Figure 8 From left to right, the unprocessed original image, the binary image processed by the adaptive grayscale image threshold segmentation algorithm, the expanded image processed by the dilation algorithm, and the clustered image processed by the clustering algorithm are shown. Each character in the binary image is relatively loose, and after dilation, they are connected together. However, there is also a certain connection between different characters. After clustering, each character is separated. At this time, the minimum distance between each connected domain in the image can be calculated. Based on this distance, each character can be accurately separated. If the text is vertical, the separated characters are spliced ​​in the horizontal direction; if the text is horizontal, they are spliced ​​in the vertical direction. After the splicing is completed, the horizontal and vertical characters are uniformly filled with the same content template to synthesize the complete data.

[0126] It can be seen that in this optional embodiment, an algorithm for producing a multi-directional text dataset is provided. By transforming the existing dataset, a target dataset is obtained. The target dataset is used to calculate the loss value of the image to determine the training progress of the graphic transformation model to be trained. The provided dataset production algorithm solves the problem of text interpretation distortion that occurs after a simple horizontal and vertical rotation of the text in the existing text dataset, thereby improving the recognition efficiency and accuracy of text data.

[0127] In another optional embodiment, when it is determined that the target loss value is greater than or equal to the preset loss value, the following steps of updating the training times corresponding to the graphic transformation model to be trained and updating the graphic transformation model to be trained are repeated; according to the updated graphic transformation model to be trained and the target data set determined above, an image transformation operation is performed on the second image to obtain a transformed image corresponding to the second image; according to the feature analysis network, the first loss value and the second loss value are calculated, and the two loss values ​​are determined as the target loss values, until it is determined that the target loss value is less than the preset loss value, the graphic transformation model to be trained is determined to be the target transformation model that has completed training.

[0128] It can be seen that in this optional embodiment, for the case where the target loss value is greater than or equal to the preset loss value, an iterative training scheme is provided to ensure that the target transformation model finally obtained is a model with a target loss value less than the preset loss value, thereby improving the reliability of the target transformation model finally obtained.

[0129] In yet another optional embodiment, before the above step 104 calculates the first loss value between the second image and the transformed image according to the preset feature analysis network, and before calculating the second loss value between the first image and the transformed image, the method may further include the following operations:

[0130] According to the determined graphic transformation model and the target data set, a preset graphic inverse transformation operation is performed on the transformed image to obtain an inverse transformed image;

[0131] According to the preset feature analysis network, the third loss value of the second image and the inverse transformed image is calculated:

[0132] The above-mentioned method of determining the first loss value and the second loss value as the target loss value specifically includes:

[0133] The first loss value, the second loss value, and the third loss value are determined as target loss values.

[0134] In this optional embodiment, for the calculation method of the third loss value, please refer to the corresponding specific instructions when calculating the second loss data in Example 1, and the embodiment of the present invention will not be repeated.

[0135] In this optional embodiment, it should be noted that the model used for performing the preset graphic inverse transformation operation on the transformed image has the same structure as the graphic transformation model to be trained, that is, when processing the transformed image and the second image, two transformation models will be determined in advance, and the weight parameters corresponding to each transformation model will match the image to be processed by the transformation model, and then the transformed image and the second image will be subjected to graphic transformation operations respectively through their respective corresponding transformation models.

[0136] It can be seen that in this optional embodiment, the first loss value, the second loss value, and the third loss value can be intelligently calculated as evaluation indicators of the graphic transformation model to be trained, thereby improving the reliability and accuracy of the obtained target transformation model.

[0137] Example 2

[0138] See also Figure 2 , Figure 2 This is a flow chart of another model training method based on a multi-directional text dataset disclosed in an embodiment of the present invention. Figure 2 The model training method based on a multi-directional text dataset described above can be applied to a model training device based on a multi-directional text dataset, and the embodiment of the present invention does not limit this. Figure 2 As shown, the model training method based on the multi-directional text dataset may include the following operations:

[0139] 201. Acquire a first image and a second image, where the first image includes at least one annotated text and a first area corresponding to the annotated text, and the second image includes the annotated text and a second area corresponding to the annotated text.

[0140] 202. According to the determined data set production algorithm, a preset text processing operation is performed on the determined image to be processed to obtain a target data set.

[0141] 203. Perform a graphic transformation operation on the second image according to the determined graphic transformation model to be trained and the target data set to obtain a transformed image corresponding to the second image.

[0142] 204. Calculate a first loss value between the second image and the transformed image, and calculate a second loss value between the first image and the transformed image according to a preset feature analysis network, and determine the first loss value and the second loss value as target loss values.

[0143] 205. When it is determined that the target loss value is less than the preset loss value, the graphic transformation model to be trained is determined to be the target transformation model that has completed training.

[0144] In the embodiment of the present invention, for other descriptions of steps 201 to 205, please refer to other specific descriptions of steps 101 to 105 in the first embodiment, which will not be repeated in the embodiment of the present invention.

[0145] 206. Perform an image transformation operation on the input image to be corrected according to the target transformation model to obtain a corresponding image transformation result.

[0146] In the embodiment of the present invention, the image to be corrected is an image including annotated text.

[0147] 207. Perform a correction and comparison operation on the image transformation result according to a preset image standard to obtain a correction and comparison result.

[0148] 208. When the correction comparison result indicates that the matching degree between the image transformation result and the preset image standard is greater than or equal to a preset correction degree threshold, it is determined that the target transformation model meets the preset image standard.

[0149] 209. When the correction comparison result indicates that the matching degree between the image transformation result and the preset image standard is less than a preset correction degree threshold, correction information is generated according to the correction comparison result and the image transformation result.

[0150] In the embodiment of the present invention, the correction information is used to provide to a processing personnel responsible for the target transformation model, so that the processing personnel can adjust the training parameters of the target transformation model according to the correction information.

[0151] It can be seen that implementation Figure 2 The described model training method based on a multi-directional text dataset can, after training the target transformation model, intelligently input the image to be corrected into the target transformation model, and compare the image transformation results according to preset image standards, and further correct and compare the trained target transformation model, which is beneficial to improving the reliability and accuracy of the target transformation model; in addition, when the correction comparison result indicates that the matching degree between the image transformation result and the preset image standard is less than the preset correction degree threshold, correction information can also be adaptively generated, thereby improving the processing efficiency of the processing personnel in handling the correction information.

[0152] In an optional embodiment, the method of performing a preset text processing operation on the determined image to be processed according to the determined dataset production algorithm to obtain the target dataset may include the following operations:

[0153] Extracting the annotated text corresponding to the first image, and performing a text matching operation on the annotated text according to a preset text library to obtain at least one target text whose matching degree with the annotated text is less than a preset matching degree threshold;

[0154] Selecting a preset number of target backgrounds from the determined background gallery;

[0155] For each target text, the target text is overlaid on all target backgrounds to obtain a text dataset corresponding to the target text. The text dataset corresponding to the target text includes a preset number of target images, each target image is generated by the target text and any target background;

[0156] The text dataset corresponding to all target texts is determined as the target dataset.

[0157] In this optional embodiment, the texts included in the preset text library are all texts of uniform image size and do not include a text background; similarly, the images included in the background library include image backgrounds but no text.

[0158] In this optional embodiment, optionally, after selecting a preset number of target backgrounds from the determined background library, the method may further include the following operations:

[0159] Performing a preset text enhancement operation on each target text to update the target text;

[0160] Analyze each updated target text and each target background to obtain the text length of each target text and the background length of each target background;

[0161] For each target background, determine whether the length difference is less than or equal to a preset length threshold. When it is determined that the length difference is greater than the preset length threshold, perform the above operation for each target text, overlay the target text on all target backgrounds, and obtain a text data set corresponding to the target text. The length difference is the difference between the background length of the target background and the text length of each target text.

[0162] When it is determined that the length difference is less than or equal to the preset length threshold, a background transformation operation is performed on the target background according to the length difference to update the target background, and the above-mentioned operation of overlaying the target text on all target backgrounds to obtain a text dataset corresponding to the target text is performed for each target text;

[0163] The difference between the background length of the updated target background and the text length of the target text is greater than a preset length threshold.

[0164] In the optional embodiment, further illustrated as follows, a way of re-making a text data set is provided: the template library and the character image library are determined in advance, wherein each character image in the character image library is equal in size and does not contain a background; each template library does not contain a text, and the template library can be various simple color backgrounds, image backgrounds of actual scenes, etc., and the actual required horizontal and vertical background templates can be directly obtained by rotation. After receiving an input text, the corresponding text image is matched from the character image library, and then the corresponding background template image is randomly taken from the template library, the character image is overlaid on the background template image, and before the overlaying, the length of the spliced character image and the length of the background template are counted first. If the length of the background template is too small, the length of the background template is automatically lengthened, which can be obtained by scaling or splicing a copy of the background template itself, until the length of the background template is equal to the length of the spliced character image. A large number of effective horizontal and vertical arrangement data sets can be obtained by this method. Additionally, in order to enrich the characteristics of the text in the data set, before the character image is overlaid on the template image, N texts therein are randomly subjected to data enhancement (including color transformation, affine transformation, etc.), the color transformation is used to realize texts of different colors, and the affine transformation is used for text deformation, so as to enhance the robustness of model learning, wherein N is less than or equal to the length of the text.

[0165] It can be seen that in the optional embodiment, another way of making a text data set is provided, the generated text data set can be targeted to fit the input text into the matching background template, solving the text interpretation distortion and other situations after the text is simply rotated horizontally and vertically in the existing text data set, and the more abundant background templates improve the reliability and accuracy of text matching, and to some extent, also improve the recognition efficiency and accuracy of the text data.

[0166] Embodiment three

[0167] Please refer to Figure 3 , Figure 3 is a structural schematic diagram of a model training device based on a multi-directional text data set disclosed by the embodiment of the application. The model training device based on the multi-directional text data set can be a model training terminal based on the multi-directional text data set, a model training equipment based on the multi-directional text data set, a model training system based on the multi-directional text data set, or a model training server based on the multi-directional text data set. The model training server based on the multi-directional text data set can be a local server, a remote server, or a cloud server (also known as a cloud server). When the model training server based on the multi-directional text data set is a non-cloud server, the non-cloud server can be in communication connection with the cloud server, and the embodiment of the application does not make any limitation. Figure 3As shown, the model training device based on the multi-directional text dataset may include an acquisition module 301, a text processing module 302, a graphic transformation module 303, a calculation module 304 and a determination module 305, wherein:

[0168] The acquisition module 301 is used to acquire a first image and a second image, wherein the first image includes at least one annotated text and a first area corresponding to the annotated text; the second image is an image obtained after image preprocessing of the first image, and the second image includes the annotated text and a second area corresponding to the annotated text.

[0169] The text processing module 302 is configured to create an algorithm based on the determined data set, perform a preset text processing operation on the determined image to be processed, and obtain a target data set, wherein the target data set includes text data in at least two text extension directions.

[0170] The graphic transformation module 303 is configured to perform a graphic transformation operation on the second image obtained by the acquisition module 301 according to the determined graphic transformation model to be trained and the target data set obtained by the text processing module 302 to obtain a transformed image corresponding to the second image.

[0171] The calculation module 304 is configured to calculate a first loss value between the second image and the transformed image, and calculate a second loss value between the first image and the transformed image according to a preset feature analysis network.

[0172] The determination module 305 is configured to determine the first loss value and the second loss value calculated by the calculation module 304 as target loss values.

[0173] The determination module 305 is also used to determine that the graphic transformation model to be trained is the target transformation model that has completed training when it is determined that the target loss value is less than the preset loss value. The target transformation model is used to transform the input image into a corrected image that meets the preset image standard.

[0174] It can be seen that implementation Figure 3 The described model training device based on a multi-directional text dataset can intelligently produce a target dataset based on a preset dataset production algorithm. The target dataset is used as a training sample for subsequent training of a graphic transformation model, which is beneficial to improving training efficiency. It can also intelligently calculate the first loss value and the second loss value as evaluation indicators for the graphic transformation model to be trained, thereby improving the reliability of the obtained target transformation model. In addition, the trained target transformation model is used to transform the input image into a corrected image of a preset image standard. The corrected image is used as the recognition object of a subsequent text recognition model, which reduces the recognition difficulty of the text recognition model, and is beneficial to improving the recognition efficiency of the text recognition model and improving the accuracy of the recognition results.

[0175] In an optional embodiment, if Figure 4 As shown, the text processing module 302 may include an acquisition submodule 3021, an image segmentation submodule 3022, and a reconstruction submodule 3023, wherein:

[0176] The acquisition submodule 3021 is used to obtain a predetermined data set to be processed, where the data set to be processed includes several images to be processed, each image to be processed includes text data and non-text data, the text data of each image to be processed includes the text to be processed and the text format corresponding to the text to be processed, and each text to be processed includes several sub-texts.

[0177] The image segmentation submodule 3022 is used to perform image segmentation processing operations on all images to be processed according to a preset image segmentation processing algorithm to obtain text data and non-text data of each image to be processed.

[0178] The reconstruction submodule 3023 is used to perform a text reconstruction operation on the text data of each image to be processed based on the preset text reconstruction algorithm and the non-text data of each image to be processed obtained by the image segmentation submodule 3022, obtain the reconstruction result of the text data of each image to be processed, and generate a target data set based on all the reconstruction results and all the images to be processed obtained by the acquisition submodule 3021.

[0179] In this optional embodiment, the reconstruction submodule 3023 optionally performs a text reconstruction operation on the text data of each image to be processed according to a preset text reconstruction algorithm and the non-text data of each image to be processed, and obtains the reconstruction result of the text data of each image to be processed in a manner specifically including:

[0180] Performing a first reconstruction operation on the to-be-processed text of each to-be-processed image according to a preset first reconstruction algorithm to obtain a first reconstruction result for each to-be-processed image, wherein the first reconstruction result for each to-be-processed image is all subtexts included in the to-be-processed text of the to-be-processed image;

[0181] Determining a transformed text format of the text to be processed in each image to be processed according to the text format corresponding to the text to be processed in each image to be processed;

[0182] performing a second reconstruction operation on the first reconstruction result of each image to be processed according to a preset second reconstruction algorithm and a transformed text format of the text to be processed of each image to be processed, thereby obtaining a second reconstruction result of each image to be processed, wherein the text format corresponding to the text to be processed of each image to be processed in the second reconstruction result of each image to be processed is the transformed text format corresponding to the text to be processed;

[0183] According to the second reconstruction result of each image to be processed and the non-text data of each image to be processed, target reconstruction data of each image to be processed is generated as the reconstruction result of the text data of each image to be processed.

[0184] It can be seen that implementation Figure 4 The described model training device based on a multi-directional text dataset provides an algorithm for producing a multi-directional text dataset. By transforming an existing dataset, a target dataset is obtained. The target dataset is used to calculate the loss value of the image to determine the training progress of the graphic transformation model to be trained. The provided dataset production algorithm solves the problem of text interpretation distortion that occurs after a simple horizontal or vertical rotation of the text in the existing text dataset, thereby improving the recognition efficiency and accuracy of text data.

[0185] In another optional embodiment, Figure 5 As shown, the text processing module 302 may include a first processing submodule 3024 and a determination submodule 3025, wherein:

[0186] The first processing submodule 3024 is used to extract the annotated text corresponding to the first image, and perform a text matching operation on the annotated text according to a preset text library to obtain at least one target text whose matching degree with the annotated text is less than a preset matching degree threshold.

[0187] The first processing submodule 3024 is further configured to select a preset number of target backgrounds from the determined background image library.

[0188] The first processing submodule 3024 is further used to cover each target text on all target backgrounds to obtain a text dataset corresponding to the target text. The text dataset corresponding to the target text includes a preset number of target images, and each target image is generated by the target text and any target background.

[0189] The determination submodule 3025 is configured to determine the text data sets corresponding to all target texts as the target data sets.

[0190] In this optional embodiment, optionally, the first processing submodule 3024 is further configured to perform a preset text enhancement operation on each target text after selecting a preset number of target backgrounds from the determined background gallery to update the target text.

[0191] As well as Figure 5 As shown, the text processing module 302 may further include a second processing submodule 3026 and a judgment submodule 3027, wherein:

[0192] The second processing submodule 3026 is configured to analyze each updated target text and each target background to obtain the text length of each target text and the background length of each target background.

[0193] The judgment submodule 3027 is used to judge whether the length difference is less than or equal to the preset length threshold for each target background. When it is judged that the length difference is greater than the preset length threshold, the first processing submodule 3024 is triggered to perform the above-mentioned operation for each target text, covering the target text on all target backgrounds to obtain the text data set corresponding to the target text. The length difference is the difference between the background length of the target background and the text length of each target text.

[0194] The second processing sub-module 3026 is also used to perform a background transformation operation on the target background according to the length difference to update the target background when it is determined that the length difference is less than or equal to the preset length threshold, and trigger the first processing sub-module 3024 to perform the above-mentioned operation of covering the target text on all target backgrounds for each target text to obtain a text data set corresponding to the target text; wherein the difference between the background length of the updated target background and the text length of the target text is greater than the preset length threshold.

[0195] It can be seen that implementation Figure 5 The described model training device based on a multi-directional text dataset provides another way to produce a text dataset. The generated text dataset can specifically fit the input text into a matching background template, solving the problem of text interpretation distortion that occurs after a simple horizontal or vertical rotation of the text in the existing text dataset. The richer background templates improve the reliability and accuracy of text matching when performing text matching, and to a certain extent also improve the recognition efficiency and accuracy of text data.

[0196] In another optional embodiment, Figure 6 As shown, the apparatus may further include an iterative training module 306, wherein:

[0197] The iterative training module 306 is used to repeatedly execute the training times corresponding to the updating of the graphic transformation model to be trained and the updating of the graphic transformation model to be trained when it is determined that the target loss value is greater than or equal to the preset loss value, and trigger the graphic transformation module 303 to perform an image transformation operation on the second image according to the graphic transformation model to be trained and the target data set to obtain a transformed image corresponding to the second image, and trigger the calculation module 304 to calculate the first loss value and the second loss value according to the feature analysis network, and trigger the determination module 305 to perform the operation of determining the first loss value and the second loss value as corresponding to the target loss value, until it is determined that the target loss value is less than the preset loss value, and the graphic transformation model to be trained is determined to be the target transformation model that has completed training.

[0198] It can be seen that implementation Figure 6 The described model training device based on a multi-directional text dataset provides an iterative training scheme for the case where the target loss value is greater than or equal to the preset loss value, so as to ensure that the final target transformation model is a model with a target loss value less than the preset loss value, thereby improving the reliability of the final target transformation model.

[0199] In another optional embodiment, the graphic transformation module 303 is further configured to perform an image transformation operation on the input image to be corrected according to a target transformation model to obtain a corresponding image transformation result, where the image to be corrected is an image including annotated text.

[0200] As well as Figure 6 As shown, the device may further include a correction and comparison module 307 and a generation module 308, wherein:

[0201] The correction and comparison module 307 is used to perform a correction and comparison operation on the image transformation result obtained by the graphic transformation module 303 according to the preset image standard to obtain a correction and comparison result. When the correction and comparison result indicates that the matching degree between the image transformation result and the preset image standard is greater than or equal to the preset correction degree threshold, it is determined that the target transformation model meets the preset image standard.

[0202] The generation module 308 is used to generate correction information based on the correction comparison result and the image transformation result obtained by the graphic transformation module 303 when the correction comparison result obtained by the correction comparison module 307 indicates that the matching degree between the image transformation result and the preset image standard is less than the preset correction degree threshold. The correction information is used to be provided to the processing personnel responsible for the target transformation model so that the processing personnel can adjust the training parameters of the target transformation model according to the correction information.

[0203] It can be seen that implementation Figure 6The described model training device based on a multi-directional text dataset can intelligently input the image to be corrected into the target transformation model, and compare the image transformation results according to the preset image standard, and further correct and compare the trained target transformation model, which is beneficial to improving the reliability and accuracy of the target transformation model; in addition, when the correction comparison result indicates that the matching degree between the image transformation result and the preset image standard is less than the preset correction degree threshold, it can also adaptively generate correction information, thereby improving the processing efficiency of the processing personnel in handling the correction information.

[0204] Example 4

[0205] See also Figure 7 , Figure 7 This is a structural diagram of another model training device based on a multi-directional text dataset disclosed in an embodiment of the present invention. Figure 7 As shown, the model training device based on the multi-directional text dataset may include:

[0206] A memory 401 storing executable program code;

[0207] a processor 402 coupled to the memory 401;

[0208] The processor 402 calls the executable program code stored in the memory 401 to execute the steps of the model training method based on the multi-directional text dataset described in the first embodiment of the present invention or the second embodiment of the present invention.

[0209] Example 5

[0210] An embodiment of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the steps in the model training method based on a multi-directional text dataset described in Example 1 or Example 2 of the present invention.

[0211] Example 6

[0212] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps in the model training method based on a multi-directional text dataset described in Example 1 or Example 2.

[0213] The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art can understand and implement the present invention without inventive effort.

[0214] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by means of hardware. Based on this understanding, the above technical solution, in essence, or the portion that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0215] Finally, it should be noted that the model training method and device based on a multi-directional text dataset disclosed in the embodiment of the present invention are only preferred embodiments of the present invention, and are only used to illustrate the technical solution of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A model training method based on a multi-directional text dataset, characterized in that: The method comprises: Acquire a first image and a second image, wherein the first image includes at least one annotated text and a first region corresponding to the annotated text, and the second image is an image obtained after image preprocessing of the first image, and includes the annotated text and a second region corresponding to the annotated text; performing a preset text processing operation on the determined image to be processed according to the determined data set production algorithm to obtain a target data set, wherein the target data set includes text data in at least two text extension directions; performing a graphic transformation operation on the second image according to the determined graphic transformation model to be trained and the target data set to obtain a transformed image corresponding to the second image; Calculating a first loss value between the second image and the transformed image, and calculating a second loss value between the first image and the transformed image according to a preset feature analysis network, and determining the first loss value and the second loss value as target loss values; When it is determined that the target loss value is less than the preset loss value, the graphic transformation model to be trained is determined to be the target transformation model that has completed training. The target transformation model is used to transform the input image into a corrected image that meets the preset image standard.

2. The model training method based on a multi-directional text dataset according to claim 1, characterized in that The step of performing a preset text processing operation on the determined image to be processed according to the determined data set production algorithm to obtain a target data set includes: Acquire a predetermined data set to be processed, the data set to be processed comprising a plurality of images to be processed, each of the images to be processed comprising text data and non-text data, the text data of each image to be processed comprising a text to be processed and a text format corresponding to the text to be processed, and each text to be processed comprising a plurality of subtexts; Performing an image segmentation processing operation on all the images to be processed according to a preset image segmentation processing algorithm to obtain text data and non-text data of each of the images to be processed; According to a preset text reconstruction algorithm and the non-text data of each of the images to be processed, a text reconstruction operation is performed on the text data of each of the images to be processed to obtain a reconstruction result of the text data of each of the images to be processed, and a target data set is generated based on all the reconstruction results and all the images to be processed.

3. The model training method based on a multi-directional text dataset according to claim 2, characterized in that: The step of performing a text reconstruction operation on the text data of each of the images to be processed according to a preset text reconstruction algorithm and the non-text data of each of the images to be processed to obtain a reconstruction result of the text data of each of the images to be processed includes: Performing a first reconstruction operation on the text to be processed of each of the images to be processed according to a preset first reconstruction algorithm to obtain a first reconstruction result for each of the images to be processed, wherein the first reconstruction result for each of the images to be processed is all subtexts included in the text to be processed of the image to be processed; Determining a transformed text format of the text to be processed in each of the images to be processed according to the text format corresponding to the text to be processed in each of the images to be processed; performing a second reconstruction operation on the first reconstruction result of each of the images to be processed according to a preset second reconstruction algorithm and the transformed text format of the text to be processed of each of the images to be processed, thereby obtaining a second reconstruction result of each of the images to be processed, wherein the text format corresponding to the text to be processed of each of the images to be processed in the second reconstruction result is the transformed text format corresponding to the text to be processed; According to the second reconstruction result of each of the images to be processed and the non-text data of each of the images to be processed, target reconstruction data of each of the images to be processed is generated as a reconstruction result of the text data of each of the images to be processed.

4. The model training method based on a multi-directional text dataset according to claim 1, characterized in that The step of performing a preset text processing operation on the determined image to be processed according to the determined data set production algorithm to obtain a target data set includes: Extracting the annotated text corresponding to the first image, and performing a text matching operation on the annotated text according to a preset text library to obtain at least one target text whose matching degree with the annotated text is less than a preset matching degree threshold; Selecting a preset number of target backgrounds from the determined background gallery; For each target text, overlay the target text on all the target backgrounds to obtain a text dataset corresponding to the target text, wherein the text dataset corresponding to the target text includes the preset number of target images, each target image being generated by the target text and any one of the target backgrounds; The text data sets corresponding to all the target texts are determined as the target data sets.

5. The model training method based on a multi-directional text dataset according to claim 4 is characterized in that: After selecting a preset number of target backgrounds from the determined background library, the method further includes: Performing a preset text enhancement operation on each target text to update the target text; Analyzing each of the updated target texts and each of the target backgrounds to obtain a text length of each of the target texts and a background length of each of the target backgrounds; For each of the target backgrounds, determining whether the length difference is less than or equal to a preset length threshold; when it is determined that the length difference is greater than the preset length threshold, performing the operation of overlaying the target text on all the target backgrounds to obtain a text data set corresponding to the target text for each of the target texts, wherein the length difference is the difference between the background length of the target background and the text length of each of the target texts; When it is determined that the length difference is less than or equal to the preset length threshold, performing a background transformation operation on the target background according to the length difference to update the target background, and performing the operation of overlaying the target text on all the target backgrounds for each target text to obtain a text data set corresponding to the target text; The difference between the updated background length of the target background and the text length of the target text is greater than the preset length threshold.

6. The model training method based on a multi-directional text dataset according to any one of claims 1 to 5, characterized in that: When it is determined that the target loss value is greater than or equal to the preset loss value, the method further includes: Repeat the steps of updating the training times corresponding to the graphic transformation model to be trained and updating the graphic transformation model to be trained, and perform the image transformation operation on the second image according to the graphic transformation model to be trained and the target data set to obtain a transformed image corresponding to the second image, calculate the first loss value and the second loss value according to the feature analysis network, and determine the first loss value and the second loss value as the target loss value corresponding operations, until it is determined that the target loss value is less than the preset loss value, and the graphic transformation model to be trained is determined to be the target transformation model that has completed training.

7. The model training method based on a multi-directional text dataset according to claim 6, characterized in that: The method further comprises: performing an image transformation operation on an input image to be corrected according to the target transformation model to obtain a corresponding image transformation result, wherein the image to be corrected is an image including annotated text; performing a correction and comparison operation on the image transformation result according to the preset image standard to obtain a correction and comparison result, and determining that the target transformation model meets the preset image standard when the correction and comparison result indicates that the degree of matching between the image transformation result and the preset image standard is greater than or equal to a preset correction degree threshold; When the correction comparison result indicates that the matching degree between the image transformation result and the preset image standard is less than a preset correction degree threshold, correction information is generated based on the correction comparison result and the image transformation result. The correction information is provided to the processing personnel responsible for the target transformation model so that the processing personnel adjust the training parameters of the target transformation model according to the correction information.

8. A model training device based on a multi-directional text dataset, characterized in that: The device comprises: an acquisition module, configured to acquire a first image and a second image, wherein the first image includes at least one annotated text and a first region corresponding to the annotated text, and the second image is an image obtained after image preprocessing of the first image, and includes the annotated text and a second region corresponding to the annotated text; a text processing module, configured to prepare an algorithm based on the determined data set, perform a preset text processing operation on the determined image to be processed, and obtain a target data set, wherein the target data set includes text data in at least two text extension directions; a graphics transformation module, configured to perform a graphics transformation operation on the second image according to the determined graphics transformation model to be trained and the target data set, to obtain a transformed image corresponding to the second image; a calculation module, configured to calculate a first loss value between the second image and the transformed image, and calculate a second loss value between the first image and the transformed image according to a preset feature analysis network; a determination module, configured to determine the first loss value and the second loss value calculated by the calculation module as target loss values; The determination module is also used to determine that the graphic transformation model to be trained is the target transformation model that has completed training when it is determined that the target loss value is less than the preset loss value. The target transformation model is used to transform the input image into a corrected image that meets the preset image standard.

9. A model training device based on a multi-directional text dataset, characterized in that: The device comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the model training method based on a multi-directional text dataset as described in any one of claims 1-7.

10. A computer storage medium, characterized in that The computer storage medium stores computer instructions, which, when called, are used to execute the model training method based on a multi-directional text dataset as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Character recognition model training method, character recognition method and device

    CN114596570A

  • Text recognition method and apparatus

    WO2021115091A1