Image processing model training method, and image processing method and apparatus
By using text generation tools to generate negative sample labels and shared weight networks to train the image processing model, the time-consuming and labor-intensive problem in the training process of image processing model is solved, and faster and more accurate model training and recognition effects are achieved.
Patent Information
- Application Number
- PCT/CN2024/140309
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-12-18
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the training process of image processing models is time-consuming and labor-intensive, and requires a large amount of manual labeling of image data, resulting in a long training cycle.
The text generation tool is used to generate images with negative sample labels, and combined with the positive sample labels of the real image, the loss value adjustment of the text detection model and the text judgment model is adjusted, the image processing model is trained, and the features are extracted using the shared weight network.
The workload of manual annotation is reduced, the speed and accuracy of model training is improved, and the generalization ability and robustness of the model are enhanced.
Smart Images

Figure CN2024140309_03072025_PF_FP_ABST
Abstract
Description
Image processing model training method, image processing method and device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 27, 2023, with application number 202311813675.3 and invention name “Training method of image processing model, image processing method and device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image processing, and in a possible implementation, relates to a training method for an image processing model, an image processing method, a training device for an image processing model, an image processing device, an electronic device, and a storage medium. Background Art
[0003] With the advancement of technology, image processing technology is increasingly being applied to various fields. Because text in images carries more explicit information, processing text in images has always been a hot topic. Among related technologies, some artificial intelligence models have been applied to text processing in images, offering advantages such as high accuracy and speed.
[0004] However, because text can vary in font, shape, size, and other aspects, specialized AI models typically need to be trained for specific application scenarios. The conventional approach to model training relies on manual annotation of large amounts of collected image data, followed by training based on the annotation results. This manual annotation process is time-consuming and labor-intensive, resulting in lengthy training cycles. Summary of the Invention
[0005] The present application has been made in view of the above-mentioned problems.
[0006] According to one aspect of the present application, a training method for an image processing model is provided. The image processing model includes a text detection model and a text judgment model. The text detection model is used to perform text recognition on an image, and the text judgment model is used to detect the authenticity of text in the image. The training method includes:
[0007] Step S110: obtaining a first image and a corresponding positive sample label, wherein the first image includes a first text, and the positive sample label includes authenticity information of the first text;
[0008] Step S120: using a text generation tool to obtain a second image and a corresponding negative sample label, wherein the second image includes a second text generated by the text generation tool, and the negative sample label includes authenticity information and text recognition information of the second text;
[0009] In step S130, both the first image and the second image are input into a text judgment model so that the text judgment model outputs a detection result, and the second image is input into a text detection model so that the text detection model outputs a text recognition result. Based on the positive sample label, the negative sample label, the text recognition result and the detection result, the loss value of the image processing model is calculated, and the corresponding parameters of the image processing model are adjusted using the loss value to train the image processing model.
[0010] In one possible implementation, the loss value includes a first loss value of a text judgment model and a second loss value of a text detection model.
[0011] Step S130 includes: first, inputting both the first image and the second image into a text judgment model, so that the text judgment model outputs a detection result, calculating a first loss value of the text judgment model based on the positive sample label, the negative sample label, and the detection result, and adjusting corresponding parameters of the text judgment model using the first loss value to train the text judgment model;
[0012] Then, the second image is input into the text detection model so that the text detection model outputs the text recognition result. Based on the negative sample label and the text recognition result, the second loss value of the text detection model is calculated, and the corresponding parameters of the text detection model are adjusted using the second loss value to train the text detection model.
[0013] In one possible implementation, the loss value is used to adjust the parameters corresponding to the image processing model to train the image processing model, including: using the loss value to simultaneously adjust the parameters corresponding to the text judgment model and the parameters corresponding to the text detection model to train the image processing model.
[0014] In one possible implementation, the loss value of the image processing model is calculated based on the positive sample label, the negative sample label, the text recognition result, and the detection result, including: calculating the first loss value of the text judgment model based on the positive sample label, the negative sample label, and the detection result; calculating the second loss value of the text detection model based on the negative sample label and the text recognition result; and calculating the loss value of the image processing model based on the first loss value and the second loss value.
[0015] In one possible implementation, the loss value of the image processing model is calculated based on the first loss value and the second loss value, including: performing a weighted sum of the first loss value and the second loss value to determine the calculated sum as the loss value of the image processing model.
[0016] In one possible implementation, the text detection model and the text judgment model have a shared weight network.
[0017] In one possible implementation, obtaining the second image and the corresponding negative sample label includes: obtaining a background image; using a text generation tool to generate a second text and obtain a negative sample label corresponding to the second text; and mapping the second text to the background image to generate the second image.
[0018] According to another aspect of the present application, an image processing method is provided, which includes: obtaining an image to be processed, which includes text; inputting the image to be processed into an image processing model trained by the above-mentioned training method to output a text recognition result of the image to be processed and / or an authenticity detection result of the text in the image to be processed.
[0019] According to another aspect of the present application, a training device for an image processing model is provided. The image processing model includes a text detection model and a text judgment model. The text detection model is used to perform text recognition on an image, and the text judgment model is used to detect the authenticity of text in the image. The training device includes:
[0020] A first acquisition module is configured to acquire a first image and a corresponding positive sample label, wherein the first image includes a first text, and the positive sample label includes authenticity information of the first text;
[0021] a second acquisition module, configured to obtain a second image and a corresponding negative sample label using a text generation tool, wherein the second image includes a second text generated using the text generation tool, and the negative sample label includes authenticity information and text recognition information of the second text;
[0022] A training module is used to input both the first image and the second image into the text judgment model so that the text judgment model outputs a detection result, input the second image into the text detection model so that the text detection model outputs a text recognition result, calculate the loss value of the image processing model based on the positive sample label, the negative sample label, the text recognition result and the detection result, and use the loss value to adjust the corresponding parameters of the image processing model to train the image processing model.
[0023] According to another aspect of the present application, an image processing device is provided, which includes: a third acquisition module for acquiring an image to be processed, wherein the image to be processed includes text; a processing module for inputting the image to be processed into an image processing model trained by the above-mentioned training method to output a text recognition result of the image to be processed and / or an authenticity detection result of the text in the image to be processed.
[0024] According to another aspect of the present application, an electronic device is provided, including a processor and a memory, wherein the memory stores computer program instructions, and the computer program instructions are used by the processor to execute the above-mentioned image processing model training method and / or image processing method when the processor is running.
[0025] According to another aspect of the present application, a storage medium is provided, on which program instructions are stored. The program instructions are used to execute the above-mentioned image processing model training method and / or image processing method when running.
[0026] In the process of training the image processing model, not only the second image including artificial characters is used, but also the first image including real characters. If only the second image is used for model training, although it has a certain recognition ability for real images, the recognition accuracy is low. In the above embodiment, a text generation tool is used to obtain the second image and the corresponding negative sample label, and the image processing model is trained based on the first image and its corresponding positive sample label and the second image and its corresponding negative sample label. Thus, the image processing model is trained based on the semantic labels generated by the text generation tool. This solution effectively saves the energy of the staff and speeds up the model training.
[0027] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0029] FIG1 shows a schematic flow chart of a method for training an image processing model according to an embodiment of the present application;
[0030] FIG2 shows a second image according to one embodiment of the present application;
[0031] FIG3 is a schematic diagram showing image processing using an image processing model according to an embodiment of the present application;
[0032] FIG4 shows a second image according to another embodiment of the present application;
[0033] FIG5 shows a schematic block diagram of a training device for an image processing model according to an embodiment of the present application;
[0034] FIG6 shows a schematic block diagram of an image processing apparatus according to an embodiment of the present application;
[0035] FIG7 shows a schematic block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solutions and advantages of the present application more apparent, the following is a detailed description of example embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application described in this application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this application.
[0037] To at least partially address the above technical issues, according to one aspect of the present application, a method for training an image processing model is provided. The image processing model is used to process images containing text. The image processing model includes a text recognition model, a text detection model, and a text detection model.
[0038] The text detection model is used to identify text in an image. In one possible implementation, the text detection model can locate text in an image and identify the located text. The text detection model can include a convolutional recurrent neural network (CRNN) and a linked temporal classification model.
[0039] The text judgment model is used to detect the authenticity of text in an image. It can be used to distinguish whether an input image is a synthetic fake. Text judgment models can include network models such as the Pointer Generation Network (PGNet), Deep Bilateral Network (DBNet), and Character Region Aware Text Detection (CRAFT). They can also be object detection network models such as Fully Convolutional Single-Stage Object Detection (FCOS) and Yolo Object Detection (YOLO). Text judgment models can assist text detection models in achieving more accurate text recognition.
[0040] Figure 1 shows a schematic flow chart of a method for training an image processing model according to an embodiment of the present application. As shown in Figure 1 , the method includes steps S110, S120, and S130.
[0041] Step S110: Obtain a first image and a corresponding positive sample label. The first image includes a first text. The positive sample label includes information about the authenticity of the first text.
[0042] The first image can be an image originally captured or obtained by processing an originally captured image, and includes a limited set of characters in the industrial field, for example, an image with a production date or an image with a product number, etc. The first image can be an RGB image or a grayscale image. The first image can be an original image directly captured by an image capture device, or an image after a preprocessing operation is performed on the original image. The preprocessing operation can include all operations to improve the visual effect of the image, increase its clarity, or highlight the text in the image. By way of example and not limitation, the preprocessing operation can include digitization, geometric transformation, normalization, filtering, and other operations on the original image. The first image includes the first text. Because the first image is originally captured or obtained by processing an originally captured image, the first text therein includes real characters. The real characters can include numbers and letters, etc.
[0043] The first image has a corresponding positive sample label, which can indicate the authenticity of the first text. As mentioned above, since the first images are all real images, the first texts are all real texts.
[0044] Step S120: Using a text generation tool, obtain a second image and a corresponding negative sample label, wherein the second image includes a second text generated by the text generation tool. The negative sample label includes authenticity information and text recognition information of the second text.
[0045] In an embodiment of the present application, a small amount of images with clean backgrounds that do not include any text can be collected in an industrial scene. Then, a text generation tool can be used to automatically generate various types of text content at specified locations in the image. Figure 2 shows a second image according to an embodiment of the present application. The text in the second image shown in Figure 2 is characters generated by a text generation tool. For the convenience of description, the characters in the second text generated by the text generation tool are referred to as artificial characters hereinafter. The text generation tool may include but is not limited to a TTF text generation tool. In the embodiment of the present application, there is no limitation on the type of text generation tool. Any tool that can automatically generate text content is within the scope of protection of this application.
[0046] The artificial characters in the second image can correspond to real characters in the first text in the first image. For example, the artificial characters in the second image have the same font, shape, or size as the real characters in the first image. This helps the text judgment model and the text detection model perform text detection and text recognition, respectively.
[0047] The second image also has its corresponding negative sample label. The negative sample label can indicate the authenticity of the second text and text recognition information. Among them, because the second image is generated using a text generation tool, the second text therein can be considered to be false, that is, the second image is also false. Text recognition information may include text position information, text content information, etc. in the image. It can be understood that when the text generation tool generates artificial characters in the second image, it can also generate corresponding negative sample labels at the same time. In this way, the process of manually collecting a large number of images containing text and annotating the images is avoided.
[0048] In step S130, both the first image and the second image are input into a text judgment model so that the text judgment model outputs a detection result, and the second image is input into a text detection model so that the text detection model outputs a text recognition result. Based on the positive sample label, the negative sample label, the text recognition result and the detection result, the loss value of the image processing model is calculated, and the corresponding parameters of the image processing model are adjusted using the loss value to train the image processing model.
[0049] The model's loss value measures the difference between the model's predictions and the true labels. A smaller loss value indicates a closer match between the model's predictions and the true labels, which translates to better model performance. If the model's loss value is high, you can adjust the model's parameters to improve its accuracy and generalization.
[0050] The detection result output by the text judgment model can indicate whether the input image is a true image or a false image. A true image is an image including real characters, and a false image is an image including artificial characters. The positive sample label corresponding to the first image can indicate that the first image is a true image, and the negative sample label corresponding to the second image can indicate that the second image is a false image. Therefore, the loss value of the loss function of the text judgment model can be calculated based on the positive sample label, the negative sample label, and the detection result output by the text judgment model. The parameters of the text judgment model can be adjusted based on the loss value until the detection result output by the text judgment model is consistent with the label of the input image with a high probability, that is, the text judgment model can accurately detect the authenticity of the text in the image.
[0051] The text detection model outputs the text recognition result of the image, which may specifically include the position of the text in the image and the specific content of the text. The negative sample label corresponding to the second image can represent the position and specific content of the artificial characters in the second image. The loss value of the loss function of the text detection model can be calculated based on the negative sample label and the text recognition result output by the text detection model. The parameters of the text detection model can be adjusted based on the loss value until the text recognition result output by the text detection model is consistent with the negative sample label of the second image with a high probability, that is, the text detection model can accurately recognize the text in the image.
[0052] The trained text judgment model and the trained text detection model constitute the trained image processing model.
[0053] In the process of training the image processing model, not only the second image including artificial characters is used, but also the first image including real characters. If only the second image is used for model training, although it has a certain recognition ability for real images, the recognition accuracy is low. In the above embodiment, a text generation tool is used to obtain the second image and the corresponding negative sample label, and the image processing model is trained based on the first image and its corresponding positive sample label and the second image and its corresponding negative sample label. Thus, the image processing model is trained based on the semantic labels generated by the text generation tool. This solution effectively saves the energy of the staff and speeds up the model training.
[0054] In some embodiments, the text detection model and the text judgment model in the image processing model have a shared weight network.
[0055] Both the text detection model and the text judgment model are neural network models. A neural network is formed by a large number of interconnected neurons. After receiving input, each neuron performs a linear weighted process. This linear weighted process can be represented by weights. When a neuron has two inputs, each input is multiplied by an associated weight and then added together to produce the output. These weights can be randomly initialized during neural network training and updated during model training.
[0056] The text detection model and text judgment model include a shared weight network. In other words, the weights of these shared weight networks remain consistent. During training, if the weight of a neuron in one shared weight network changes, the weight of the corresponding neuron in the other network also changes accordingly. For example, this shared weight network can be called a backbone network and can be used to extract image features.
[0057] Figure 3 shows a schematic diagram of image processing using an image processing model according to an embodiment of the present application. In Figure 3, the first shared weight network and the first sub-network constitute a text detection model. It can be understood that the first sub-network can be a downstream task network of the text detection model. The second shared weight network and the second sub-network constitute a text judgment model. Because the text judgment model is used to detect the authenticity of the text in the image, the second sub-network can be implemented using a discriminator. As shown in Figure 3, the second image and the first image are input into the first shared weight network and the second shared weight network respectively. For the second image calculated by the first shared weight network, it will be input into both the first sub-network and the second sub-network. For the first image calculated by the second shared weight network, it will only be input into the second sub-network. From the perspective of the text judgment model, it receives the first image and the second image, and outputs the authenticity of the two respectively. From the perspective of the text detection model, it only receives the second image and outputs the text recognition result in the second image.
[0058] In one possible implementation, the text recognition result may include the position information of the text in the image, the position information of each character in the text, and the text content information. It can be understood that the position information of the text can be represented by the circumscribed rectangle of the text. For the convenience of description, the rectangular box is referred to as the text box below. In one possible implementation, the coordinates of the vertices of the opposite corners of the text box can be used to represent it, such as the coordinates of the upper left vertex and the coordinates of the lower right vertex. Alternatively, the text box can also be represented by its center point and its length and width. The position information of the characters can be represented by the coordinates of the center position of the characters. The text content information indicates which characters the text can include, such as numerical characters such as 1, 2, 3, and English characters such as a, b, c, etc. Based on the output results of the above-mentioned image processing model and the above-mentioned positive sample labels and negative sample labels, the image processing model can be trained.
[0059] In the above technical solution, the image processing model includes not only a text detection model but also a text judgment model, and both models also have a shared weight network. This allows the shared weight network to extract features not only from the first image but also from the second image. This ensures the image processing accuracy of the image processing model. Furthermore, the shared weight network can effectively reduce the number of parameters required for the image processing model, improving the model's generalization capabilities and enabling the application of knowledge learned from one training task to another related task, further improving the training efficiency of the image processing model.
[0060] In some embodiments, obtaining the second image and the corresponding negative sample label in step S120 includes: obtaining a background image; using a text generation tool to generate a second text and obtain a negative sample label corresponding to the second text; and mapping the second text to the background image to generate the second image.
[0061] The background image can be an image without text printed on it, and the second text can be at any position in the background image. By using a text generation tool, the generated artificial characters are randomly mapped to the background image to form a second image, and corresponding negative sample labels can also be generated. The second image may include a strong interference signal and also have features similar to images in real scenes. Figure 4 shows a second image according to another embodiment of the present application. As shown in Figure 4, the text in the second image is generated by using a text generation tool to generate artificial characters and then mapping them to the background image. The second image has both strong randomness and the features of a real image. The second image generated in this way effectively improves the robustness of the image processing model.
[0062] In the above embodiment, since the second image is highly random and has its own label information, the manual collection of a large number of images with text and manual labeling operations are avoided, which effectively improves the speed of second image collection and further improves the training efficiency of the image processing model.
[0063] In some embodiments, the loss value includes a first loss value of the text judgment model and a second loss value of the text detection model. Step S130 includes the following steps. First, the first image and the second image are both input into the text judgment model so that the text judgment model outputs a detection result, and based on the positive sample label, the negative sample label and the detection result, the first loss value of the text judgment model is calculated, and the first loss value is used to adjust the corresponding parameters of the text judgment model to train the text judgment model. Then, the second image is input into the text detection model so that the text detection model outputs a text recognition result, and based on the negative sample label and the text recognition result, the second loss value of the text detection model is calculated, and the corresponding parameters of the text detection model are adjusted using the second loss value to train the text detection model.
[0064] The first image and the second image are respectively input into the text recognition model, and the text recognition model outputs respective detection results. A first loss value of the text recognition model is calculated based on the detection results output by the text recognition model, the positive sample label corresponding to the first image, and the negative sample label corresponding to the second image. The parameters of the text recognition model can be updated based on this first loss value. The text recognition model is trained by repeatedly performing calculations and updates.
[0065] After the text recognition model is trained, only the second image is fed into the text detection model, which then outputs a text recognition result. Based on the text recognition result and the negative sample label output by the text detection model, a second loss value is calculated for the text detection model, and this second loss value is used to complete the training of the text detection model.
[0066] In the above embodiment, the text judgment model is trained first, followed by the text detection model, ensuring the effectiveness of both training. In particular, for embodiments where the text judgment model and the text detection model share a weight network, after the text judgment model is trained, the text detection model can directly utilize some of the parameters of the text judgment model for training, thereby accelerating the training of the image processing model.
[0067] In some embodiments, in step S130, the loss value is used to adjust the parameters corresponding to the image processing model to train the image processing model, including: using the loss value to simultaneously adjust the parameters corresponding to the text judgment model and the parameters corresponding to the text detection model to train the image processing model.
[0068] In the above embodiment, the text judgment model and the text detection model can be trained simultaneously on different computing units. In other words, the image processing model is trained as a whole. Training the text judgment model and the text detection model simultaneously can effectively simplify the image processing model training process, speed up the image processing model training process, and improve training efficiency.
[0069] In some embodiments, calculating the loss value of the image processing model based on the positive sample labels, negative sample labels, text recognition results, and detection results in step S130 includes the following steps S131 to S133. In step S131, a first loss value of the text judgment model is calculated based on the positive sample labels, negative sample labels, and detection results. In step S132, a second loss value of the text detection model is calculated based on the negative sample labels and text recognition results. In step S133, a loss value of the image processing model is calculated based on the first loss value and the second loss value.
[0070] The first loss value represents the difference between the detection result and the positive sample label and the negative sample label, thereby representing the detection accuracy of the text judgment model. In one possible implementation, the first loss value can be calculated using a cross entropy loss function.
[0071] The second loss value represents the difference between the text recognition result and the negative sample label, thereby indicating the text recognition accuracy of the text detection model. As mentioned above, the text recognition result may include a text box, which is used to indicate the location of the text in the image. Alternatively, the text box can be represented by its center point and its length and width. The second loss value may include text box classification loss, text box regression loss, and / or text content loss. The text box classification loss is primarily used to determine whether a text box contains text. The text box regression loss is used to optimize the position and size of the text box. For each text box, its offset relative to the true text box needs to be calculated. The text box regression loss can be calculated using a regression loss function. The text content loss is used to determine the text detection model's ability to recognize text. In one possible implementation, the second loss value of the text detection model can be obtained by taking a weighted sum of the text box classification loss, the text box regression loss, and the text content loss. The loss function of the text detection model can be determined based on the model. In one possible implementation, if PGNet is used as the text detection model, the loss function in PGNet can be used.
[0072] In the above embodiment, by respectively calculating the first loss value and the second loss value, and using the above two loss values to calculate the loss value of the image processing model, the loss value of the image processing model can be accurately calculated, thereby effectively improving the stability and robustness of the image processing model.
[0073] In some embodiments, step S133, based on the first loss value and the second loss value, calculates the loss value of the image processing model, including: performing weighted summation on the first loss value and the second loss value to determine the calculated sum as the loss value of the image processing model.
[0074] The relative importance of the first and second loss values in the image processing model can be comprehensively considered to determine their respective weights. If a loss value contributes more significantly to the performance of the image processing model, it can be given a higher weight, giving it greater weight. This helps balance the relationships between the different loss values, ensuring the desired performance of the image processing model.
[0075] In the above embodiment, a weighted sum is performed according to the weights of the first loss value and the second loss value, and the sum result is used as the loss value of the image processing model to train the image processing model, which can make the image processing model have stronger reliability and robustness.
[0076] According to another aspect of the present application, an image processing method is provided, comprising the following steps: First, obtaining an image to be processed, wherein the image to be processed includes text. Then, the image to be processed is input into the image processing model trained by the above-described training method to output a text recognition result for the image to be processed and / or an authenticity detection result for the text in the image to be processed.
[0077] It is understood that the image to be processed can be an RGB image or a grayscale image. The image to be processed can be the original image directly captured by the image acquisition device, or it can be an image that has undergone preprocessing operations on the original image. The preprocessing operations can include all operations to improve the visual effect of the image, enhance its clarity, or highlight text in the image. By way of example and not limitation, the preprocessing operations can include digitization, geometric transformation, normalization, filtering, and other operations on the original image.
[0078] In some embodiments of the present application, the above-described image processing method can be used to determine text recognition results in a processed image. In this embodiment, the presence of a text detection model within the image processing model facilitates training a more accurate text detection model. Consequently, when the image processing model is used to process the processed image, more accurate text recognition results can be obtained.
[0079] In some embodiments of the present application, the aforementioned image processing method can be used to determine the authenticity of text in an image to be processed. In this embodiment, the presence of a text detection model within the image processing model facilitates training a more accurate text determination model. Consequently, when the image processing model is used to process the image to be processed, a more accurate text determination result can be obtained.
[0080] In the above technical solution, the image processing model trained by the above training method outputs the text recognition result of the image to be processed and / or the authenticity detection result of the text in the image to be processed. This solution ensures the accuracy of the image processing result.
[0081] According to another aspect of the present application, a training device for an image processing model is also provided. The image processing model is used to process images containing text. The image processing model includes a text detection model and a text judgment model. The text detection model is used to perform text recognition on an image, and the text judgment model is used to detect the authenticity of text in the image. Figure 5 shows a schematic block diagram of a training device for an image processing model according to one embodiment of the present application. As shown in Figure 5, the training device 500 includes a first acquisition module 510, a second acquisition module 520, and a training module 530.
[0082] The first acquisition module 510 is used to obtain a first image and a corresponding positive sample label, wherein the first image includes a first text, and the positive sample label includes information about the authenticity of the first text. The second acquisition module 520 is used to obtain a second image and a corresponding negative sample label using a text generation tool, wherein the second image includes a second text generated using the text generation tool, and the negative sample label includes information about the authenticity of the second text and text recognition information. The training module 530 is used to input both the first image and the second image into a text judgment model so that the text judgment model outputs a detection result, and input the second image into a text detection model so that the text detection model outputs a text recognition result, and calculate the loss value of the image processing model based on the positive sample label, the negative sample label, the text recognition result, and the detection result, and use the loss value to adjust the corresponding parameters of the image processing model to train the image processing model.
[0083] According to another aspect of the present application, an image processing device is also provided. FIG6 shows a schematic block diagram of an image processing device according to an embodiment of the present application. As shown in FIG6 , the image processing device 600 includes a third acquisition module 610 and a processing module 620 .
[0084] The third acquisition module 610 is used to acquire an image to be processed, wherein the image to be processed includes text. The processing module 620 is used to input the image to be processed into the image processing model trained by the above-mentioned training method to output a text recognition result of the image to be processed and / or an authenticity detection result of the text in the image to be processed.
[0085] According to another aspect of the present application, an electronic device 700 is also provided. Figure 7 shows a schematic block diagram of electronic device 700 according to one embodiment of the present application. As shown in Figure 7, electronic device 700 includes a processor 710 and a memory 720. The memory 720 stores computer program instructions, which, when executed by the processor, are used to execute the above-described image processing model training method and / or image processing method.
[0086] In addition, according to another aspect of the present application, a storage medium is also provided, on which program instructions are stored, and when the program instructions are executed by a computer or a processor, the computer or the processor executes the training method of the above-mentioned image processing model and / or the corresponding steps of the image processing method according to the embodiment of the present application, and is used to implement the training device of the above-mentioned image processing model and / or the corresponding module of the image processing device according to the embodiment of the present application or the corresponding module in the above-mentioned electronic device. The storage medium may, for example, include a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above-mentioned storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0087] A person skilled in the art can understand the specific implementation and beneficial effects of the image processing model training device, electronic device and storage medium by reading the specific description of the image processing model training method and / or image processing method. For the sake of brevity, they will not be repeated here.
[0088] Example 1: A training method for an image processing model, the image processing model comprising a text detection model and a text judgment model, the text detection model being used to perform text recognition on an image, and the text judgment model being used to detect the authenticity of text in the image;
[0089] The training method comprises:
[0090] Step S110: Acquire a first image and a corresponding positive sample label, wherein the first image includes a first text, and the positive sample label includes authenticity information of the first text;
[0091] Step S120: using a text generation tool to obtain a second image and a corresponding negative sample label, wherein the second image includes a second text generated by the text generation tool, and the negative sample label includes authenticity information and text recognition information of the second text;
[0092] Step S130: Input the first image and the second image into the text judgment model so that the text judgment model outputs a detection result; input the second image into the text detection model so that the text recognition model outputs a text recognition result; calculate the loss value of the image processing model based on the positive sample label, the negative sample label, the text recognition result and the detection result; and use the loss value to adjust the corresponding parameters of the image processing model to train the image processing model.
[0093] Example 2: According to the training method of the image processing model introduced in Example 1, the loss value includes the first loss value of the text judgment model and the second loss value of the text detection model.
[0094] The step S130 includes:
[0095] First, the first image and the second image are both input into the text judgment model, so that the text judgment model outputs the detection result, and a first loss value of the text judgment model is calculated based on the positive sample label, the negative sample label, and the detection result. The first loss value is used to adjust the corresponding parameters of the text judgment model to train the text judgment model;
[0096] Then, the second image is input into the text detection model so that the text detection model outputs the text recognition result. Based on the negative sample label and the text recognition result, a second loss value of the text detection model is calculated, and the second loss value is used to adjust the corresponding parameters of the text detection model to train the text detection model.
[0097] Embodiment 3: According to the image processing model training method described in any one of Embodiments 1-2, the method of using the loss value to adjust the parameters corresponding to the image processing model to train the image processing model includes:
[0098] The loss value is used to simultaneously adjust parameters corresponding to the text judgment model and parameters corresponding to the text detection model to train the image processing model.
[0099] Example 4: The training method of the image processing model according to any one of Examples 1-3,
[0100] The calculating the loss value of the image processing model based on the positive sample label, the negative sample label, the text recognition result, and the detection result includes:
[0101] Calculating a first loss value of the text judgment model based on the positive sample label, the negative sample label, and the detection result;
[0102] Calculating a second loss value of the text detection model based on the negative sample label and the text recognition result;
[0103] Based on the first loss value and the second loss value, a loss value of the image processing model is calculated.
[0104] Embodiment 5: According to the training method of the image processing model described in any one of Embodiments 1-4, calculating the loss value of the image processing model based on the first loss value and the second loss value includes:
[0105] A weighted sum is performed on the first loss value and the second loss value to determine the calculated sum as the loss value of the image processing model.
[0106] Example 6: According to the training method of the image processing model introduced in any one of Examples 1-5, the text detection model and the text judgment model have a shared weight network.
[0107] Example 7: According to the training method of the image processing model described in any one of Examples 1-6, obtaining the second image and the corresponding negative sample label includes:
[0108] Get the background image;
[0109] Using a text generation tool, generate the second text and obtain a negative sample label corresponding to the second text;
[0110] The second text is mapped to the background image to generate the second image.
[0111] Embodiment 8: An image processing method, the processing method comprising:
[0112] Acquire an image to be processed, wherein the image to be processed includes text;
[0113] The image to be processed is input into the image processing model trained by the image processing model training method described in any one of Examples 1-7 to output the text recognition result of the image to be processed and / or the authenticity detection result of the text in the image to be processed.
[0114] Embodiment 9: A training device for an image processing model, the image processing model comprising a text detection model and a text judgment model, the text detection model being used to perform text recognition on an image, and the text judgment model being used to detect the authenticity of text in the image;
[0115] The training device comprises:
[0116] A first acquisition module is configured to acquire a first image and a corresponding positive sample label, wherein the first image includes a first text, and the positive sample label includes authenticity information of the first text;
[0117] A second acquisition module is configured to obtain a second image and a corresponding negative sample label using a text generation tool, wherein the second image includes a second text generated using the text generation tool, and the negative sample label includes authenticity information and text recognition information of the second text;
[0118] A training module is used to input both the first image and the second image into the text judgment model so that the text judgment model outputs a detection result, input the second image into the text detection model so that the text detection model outputs a text recognition result, calculate the loss value of the image processing model based on the positive sample label, the negative sample label, the text recognition result and the detection result, and use the loss value to adjust the corresponding parameters of the image processing model to train the image processing model.
[0119] Embodiment 10: An image processing device, comprising:
[0120] A third acquisition module is used to acquire an image to be processed, wherein the image to be processed includes text;
[0121] A processing module is used to input the image to be processed into an image processing model trained by the image processing model training method introduced in any of Examples 1-7 to output the text recognition result of the image to be processed and / or the authenticity detection result of the text in the image to be processed.
[0122] Example 11: An electronic device comprising a processor and a memory, wherein the memory stores computer program instructions, and the computer program instructions are used by the processor to execute the training method of the image processing model introduced in any of Examples 1-7 and / or the image processing method described in Example 8 when the processor is running.
[0123] Example 12: A storage medium having program instructions stored thereon, wherein the program instructions are used to execute the training method of the image processing model as described in any one of Examples 1-7 and / or the image processing method as described in Example 8 during runtime.
[0124] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present application. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as required by the appended claims.
[0125] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0126] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units described is merely a logical function division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not performing some features.
[0127] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0128] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this approach of the present application should not be interpreted as reflecting the intention that the application claimed for protection requires more features than those explicitly recited in each claim. More precisely, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.
[0129] It will be understood by those skilled in the art that, except where mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus disclosed herein may be combined in any combination. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0130] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.
[0131] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules in the training device and image processing device of the image processing model according to the embodiment of the present application. The present application can also be implemented as a device program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0132] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbols placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0133] The above description is merely a specific embodiment or illustration of a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. The scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A training method for an image processing model, characterized in that, The image processing model includes a text detection model and a text judgment model. The text detection model is used to perform text recognition on an image, and the text judgment model is used to detect the authenticity of the text in the image; The training method includes: Step S110: Obtain a first image and a corresponding positive sample label. Among them, the first image includes first text, and the positive sample label includes information about the authenticity of the first text; Step S120: Use a text generation tool to obtain a second image and a corresponding negative sample label. Among them, the second image includes second text generated by the text generation tool, and the negative sample label includes information about the authenticity of the second text and text recognition information; Step S130: Input both the first image and the second image into the text judgment model so that the text judgment model outputs a detection result. Input the second image into the text detection model so that the text recognition model outputs a text recognition result. Based on the positive sample label, the negative sample label, the text recognition result, and the detection result, calculate the loss value of the image processing model, and use the loss value to adjust the corresponding parameters of the image processing model to train the image processing model.
2. The training method of the image processing model according to claim 1, wherein, The loss value includes a first loss value of the text judgment model and a second loss value of the text detection model, Step S130 includes: First, input both the first image and the second image into the text judgment model so that the text judgment model outputs the detection result. Based on the positive sample label, the negative sample label, and the detection result, calculate the first loss value of the text judgment model, and use the first loss value to adjust the corresponding parameters of the text judgment model to train the text judgment model; Then, input the second image into the text detection model so that the text detection model outputs the text recognition result. Based on the negative sample label and the text recognition result, calculate the second loss value of the text detection model, and use the second loss value to adjust the corresponding parameters of the text detection model to train the text detection model.
3. The training method of the image processing model according to claim 1, wherein The using the loss value to adjust the corresponding parameters of the image processing model to train the image processing model includes: Using the loss value to simultaneously adjust the corresponding parameters of the text judgment model and the text detection model to train the image processing model.
4. The training method of the image processing model according to claim 3, wherein The calculating the loss value of the image processing model based on the positive sample label, the negative sample label, the text recognition result, and the detection result includes: Calculating the first loss value of the text judgment model based on the positive sample label, the negative sample label, and the detection result; Calculating the second loss value of the text detection model based on the negative sample label and the text recognition result; Calculating the loss value of the image processing model based on the first loss value and the second loss value.
5. The training method of the image processing model according to claim 4, characterized in that Calculating the loss value of the image processing model based on the first loss value and the second loss value includes: Performing weighted summation on the first loss value and the second loss value to determine the calculated sum as the loss value of the image processing model.
6. The training method of the image processing model according to any one of claims 1 to 5, characterized in that, The text detection model and the text judgment model have a shared weight network.
7. The training method of the image processing model according to any one of claims 1 to 5, characterized in that, Obtaining the second image and the corresponding negative sample label includes: Obtaining a background image; Using a text generation tool to generate the second text and obtaining the negative sample label corresponding to the second text; Mapping the second text to the background image to generate the second image.
8. An image processing method, characterized in that, The processing method includes: Obtaining an image to be processed, where the image to be processed includes text; Inputting the image to be processed into the image processing model trained by the training method according to any one of claims 1 to 7 to output the text recognition result of the image to be processed and / or the authenticity detection result of the text in the image to be processed.
9. A training device for an image processing model, characterized in that The image processing model includes a text detection model and a text judgment model. The text detection model is used for text recognition of an image, and the text judgment model is used for detecting the authenticity of the text in the image; The training device includes: A first acquisition module, configured to acquire a first image and a corresponding positive sample label, where the first image includes a first text, and the positive sample label includes information about the authenticity of the first text; A second acquisition module, configured to use a text generation tool to obtain a second image and a corresponding negative sample label, where the second image includes a second text generated by the text generation tool, and the negative sample label includes information about the authenticity of the second text and text recognition information; A training module, configured to input both the first image and the second image into the text judgment model to output a detection result by the text judgment model, input the second image into the text detection model to output a text recognition result by the text detection model, calculate the loss value of the image processing model based on the positive sample label, the negative sample label, the text recognition result, and the detection result, and use the loss value to adjust the corresponding parameters of the image processing model to train the image processing model.
10. An image processing apparatus, characterized in that, The processing device includes: A third acquisition module, configured to acquire an image to be processed, where the image to be processed includes text; A processing module, configured to input the image to be processed into the image processing model trained by the training method according to any one of claims 1 to 7 to output the text recognition result of the image to be processed and / or the authenticity detection result of the text in the image to be processed.
11. An electronic device, comprising a processor and a memory, characterized in that, The memory stores computer program instructions, and when the computer program instructions are run by the processor, they are used to execute the training method of the image processing model according to any one of claims 1 to 7 and / or the image processing method according to claim 8.
12. A storage medium, on which program instructions are stored, characterized in that, The program instructions are used to execute the training method of the image processing model according to any one of claims 1 to 7 and / or the image processing method according to claim 8 when running.
Citation Information
Patent Citations
Text erasing method, text erasing model training method and device and storage medium
CN113469878A
Neural network training method and device and image segmentation method
CN113822428A
Text detection model training method and device, equipment and storage medium
CN114067321A
Text image synthesis model training method and device, text image synthesis method and device, equipment and medium
CN115619903A
Training method of image processing model, and image processing method and device
CN117475448A