Detection Model Training Method and Device, Image Detection Method and Device
Through the semi-supervised training method, the features of labeled and unlabeled sample images are used to solve the problem that existing text detection methods rely on manual annotation and low accuracy in complex scene detection, achieving more efficient and accurate text detection.
Patent Information
- Application Number
- CN202111570174.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-12-21
AI Technical Summary
Existing text detection methods rely on high-cost sample images manually annotated, and the detection accuracy in complex scenarios is not high, resulting in low image processing efficiency.
By acquiring labeled sample images and labelless sample images, using a Generative Adversarial Network (GAN) for semi-supervised training, the generation module learns the features of labeled and labelless sample images, and improves the training efficiency and accuracy of the detection model.
Reduce dependence on high-cost annotation samples, improve text detection accuracy in complex scenarios, and improve image processing efficiency.
Smart Images

Figure CN114266308B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for training a detection model. Background Art
[0002] With the rapid development of computer technology, the field of image processing has also developed rapidly. Among them, text detection is also a very important branch in the field of image processing. Most of the existing text detections are based on manually labeled text images as the training sample images of the model. The training sample images require a large amount of manpower and material resources to label them, or cost a high price to purchase labeled sample images, with high costs. At the same time, when detecting images with relatively complex scenes, the accuracy of the text content detection results is not high, and thus the efficiency of image processing is relatively low. Summary of the Invention
[0003] In view of this, the embodiments of this specification provide a method for training a detection model. One or more embodiments of this specification also relate to an image detection method, a detection model training device, an image detection device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0004] According to the first aspect of the embodiments of this specification, a method for training a detection model is provided, including:
[0005] Obtain a labeled sample image and an unlabeled sample image, input the labeled sample image into a generation module to obtain a first detection result, and input the unlabeled sample image into the generation module to obtain a second detection result, where the labeled sample image includes an initial sample image and a label corresponding to the initial sample image;
[0006] Process the labeled sample image to obtain a first fused image and a first identifier, and process the unlabeled sample image and the second detection result to obtain a second fused image and a second identifier;
[0007] Input the first fused image and the first identifier, and the second fused image and the second identifier into a discrimination module to determine the discrimination result of the second fused image;
[0008] When the discrimination result meets a preset processing condition, determine a first loss function based on the first fused image and the second fused image, and determine a second loss function based on the first detection result and the label corresponding to the initial sample image;
[0009] Training the generation module based on the first loss function and the second loss function, and determining the generation module obtained when the training stop condition is reached as the target generation module.
[0010] According to a second aspect of the embodiments of the present specification, there is provided an image detection method, including:
[0011] Obtaining an image to be processed;
[0012] Inputting the image to be processed into the target generation module to obtain a detection result of the image to be processed, where the target generation module is trained by the above detection model training method.
[0013] According to a third aspect of the embodiments of the present specification, there is provided a detection model training device, including:
[0014] An image acquisition module, configured to acquire a labeled sample image and an unlabeled sample image, input the labeled sample image into the generation module to obtain a first detection result, and input the unlabeled sample image into the generation module to obtain a second detection result, where the labeled sample image includes an initial sample image and a label corresponding to the initial sample image;
[0015] An image processing module, configured to process the labeled sample image to obtain a first fused image and a first identifier, and process the unlabeled sample image and the second detection result to obtain a second fused image and a second identifier;
[0016] An image input module, configured to input the first fused image and the first identifier, and the second fused image and the second identifier into the discrimination module to determine a discrimination result of the second fused image;
[0017] A loss function determination module, configured to, when the discrimination result meets a preset processing condition, determine a first loss function based on the first fused image and the second fused image, and determine a second loss function based on the first detection result and the label corresponding to the initial sample image;
[0018] A training module, configured to train the generation module based on the first loss function and the second loss function, and determine the generation module obtained when the training stop condition is reached as the target generation module.
[0019] According to a fourth aspect of the embodiments of the present specification, there is provided an image detection device, including:
[0020] An acquisition module, configured to acquire an image to be processed;
[0021] A detection module, configured to input the image to be processed into a target generation module, and obtain a detection result of the image to be processed, where the target generation module is trained by the above detection model training method.
[0022] According to a fifth aspect of the embodiments of the present specification, a computing device is provided, including:
[0023] A memory and a processor;
[0024] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the processor executes the computer-executable instructions, the steps of the above detection model training method or image detection method are implemented.
[0025] According to a sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above detection model training method or image detection method are implemented.
[0026] According to a seventh aspect of the embodiments of the present specification, a computer program is provided. When the computer program is executed on a computer, the computer is made to execute the steps of the above detection model training method or image detection method.
[0027] In an embodiment of the present specification, by obtaining a labeled sample image and an unlabeled sample image, inputting the labeled sample image into a generation module to obtain a first detection result, and inputting the unlabeled sample image into the generation module to obtain a second detection result, where the labeled sample image includes an initial sample image and a label corresponding to the initial sample image; processing the labeled sample image to obtain a first fused image and a first identifier, processing the unlabeled sample image and the second detection result to obtain a second fused image and a second identifier; inputting the first fused image and the first identifier and the second fused image and the second identifier into a discrimination module to determine a discrimination result of the second fused image; when the discrimination result meets a preset processing condition, determining a first loss function based on the first fused image and the second fused image, and determining a second loss function based on the first detection result and the label corresponding to the initial sample image; training the generation module based on the first loss function and the second loss function, and determining the generation module obtained when a training stop condition is reached as the target generation module.
[0028] Specifically, by obtaining labeled sample images and unlabeled sample images, and inputting them into the generation module respectively, the first detection result and the second detection result are obtained. At the same time, the labeled sample images and the unlabeled sample images are processed respectively to obtain the first fusion image and the first identifier, the second fusion image and the second identifier, and the first fusion image and the second fusion image are input into the discrimination module to generate a discrimination result. When it is determined that the discrimination result meets the preset processing conditions, the first loss function and the second loss function are determined, and then the generation module is trained. This method uses a semi-supervised training method with labeled sample images and unlabeled sample images, and the generation module and the discrimination module are used to process the sample images respectively. The generation module can also learn the information of the unlabeled sample images. In addition, for the input of the discrimination module, the pre-processed fusion images are also used, so that the discrimination module has clear learning objectives and judgment criteria. Finally, according to the discrimination result, two loss functions are determined to train the generation module, so that the generation module can learn the features of the unlabeled sample images. The generation module can be applied not only to the scenario of labeled images, but also to the scenario of unlabeled images, thereby improving the accuracy of the detection result of the generation module and enhancing the image processing efficiency. Description of the Drawings
[0029] Figure 1 is a flowchart of a method for training a detection model provided by an embodiment of this specification;
[0030] Figure 2 is a schematic diagram of the processing process of a method for training a detection model provided by an embodiment of this specification;
[0031] Figure 3 is a flowchart of an image detection method provided by an embodiment of this specification;
[0032] Figure 4 is a schematic diagram of the structure of a detection model training device provided by an embodiment of this specification;
[0033] Figure 5 is a schematic diagram of the structure of an image detection device provided by an embodiment of this specification;
[0034] Figure 6 is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed Embodiments
[0035] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.
[0036] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0037] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0038] First, the noun terms related to one or more embodiments of this specification are explained.
[0039] Semi-supervised: Using both unlabeled data and labeled data for model learning.
[0040] OCR: Optical Character Recognition, which refers to the process by which an electronic device (such as a scanner or digital camera) examines printed characters on paper and translates the shapes into computer text using character recognition methods.
[0041] Text detection: Detecting the regions of each line of text boxes.
[0042] Labeled data: The original image and the label.
[0043] Unlabeled data: The original image without a corresponding label.
[0044] Generative Adversarial Networks (GAN): A deep learning model that can be widely applied to algorithms such as image generation, image translation, and image conversion.
[0045] Generator: A deep learning model.
[0046] Discriminator: A deep learning model.
[0047] Channel dimension: The channel dimension of an image.
[0048] Batch: In deep learning, one-step learning involves multiple images being learned together. Batch is the general term for these images.
[0049] In OCR applications, the text detection model for images needs to handle different scenarios. If the model is not trained on data from a specific scenario, the accuracy of the results output by the model in this scenario will drop significantly. However, the annotation workload for text detection in images is huge. If a supervised training method is used, it means that a large amount of manpower and material resources are required to complete the training process of the model, which is not only inefficient but also prone to errors. Then, without manual annotation, useful information can be learned with a small amount of annotated data in a specific scenario, and training based on this pre-trained model can achieve a high accuracy.
[0050] Moreover, in current models for text detection in images, most do not consider the connection between image channels. When detecting text regions in complex backgrounds (such as complex colors, textures, etc.), there are often omissions, and the final determined text detection positions are often inaccurate, and there will also be misjudgment situations. Based on this, the embodiments of this specification provide a method for training a detection model, which constructs a scheme for pre-training a text detection model using unlabeled data, and can improve the accuracy of the model while greatly reducing the annotation time and cost of OCR.
[0051] In this specification, a method for training a detection model is provided. This specification also relates to an image detection method, a detection model training device, an image detection device, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail one by one in the following embodiments.
[0052] Figure 1 The flowchart of a method for training a detection model provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0053] It should be noted that the method for training a detection model provided in the embodiments of this specification uses a generative adversarial network (GAN), which includes two parts: a generator and a discriminator. The generator is a segmentation-based text detection network, such as a Differential Binarization network, and the discriminator can be a lightweight binary classification network. However, the types of the generator and the discriminator used in this embodiment are not limited in any way.
[0054] Step 102: Obtain labeled sample images and unlabeled sample images, input the labeled sample images into the generation module to obtain a first detection result, and input the unlabeled sample images into the generation module to obtain a second detection result.
[0055] Among them, the labeled sample image includes an initial sample image and the label corresponding to the initial sample image.
[0056] Among them, the initial sample image can be understood as the original image of the supervised training sample, and the label corresponding to the initial sample image can be understood as the text box corresponding to the original image of the supervised training sample; the unlabeled sample image can be understood as the original image of the unsupervised training sample.
[0057] In specific implementation, an initial sample image, the label corresponding to the initial sample image, and an unlabeled sample image are obtained, and the initial sample image and the label corresponding to the initial sample image are input into the generation module. The generation module can output a first detection result predicted for the initial sample image. Among them, the first detection result can be understood as the detection box of the area where the corresponding text in the initial sample image is detected by the generator; then, the unlabeled sample image is input into the generation module, and the generation module can also detect the unlabeled sample original image and output a second detection result predicted for the unlabeled sample original image. Among them, the second detection result can be understood as the detection box of the area where the corresponding text in the unlabeled sample image is detected by the generator.
[0058] In practical applications, each time the model starts training, a batch of pre-training images can be obtained. For example, in the first round of iteration, 10 initial sample images can be obtained, and the 10 initial sample images carry corresponding labels, and the label can be understood as the text box that has been manually labeled; in addition, 10 unlabeled sample images need to be obtained. Among them, the unlabeled sample image is an image without a labeled text box. Then, the 10 initial sample images and the corresponding labels carried by the 10 initial sample images are input into the generation module (generator), and the generator makes predictions and outputs 10 first detection results (the detection results are text boxes for the area where the text in each image is located); then, the 10 unlabeled sample images are input into the generator, and the generator makes predictions and outputs 10 second detection results (the detection results are text boxes for the area where the text in each image is located).
[0059] It should be emphasized that in order to enable the trained model to be applied to various types of images and accurately output corresponding detection results, then, in a large number of sample images, the sample images need to be divided into labeled sample images and unlabeled sample images according to different scenarios; specifically, before obtaining the labeled sample images and unlabeled sample images, it further includes:
[0060] Determine the first type of image as the labeled sample image and the second type of image as the unlabeled sample image, where the first type of image is an image in which the area of the text connected region is greater than a preset area threshold, and the second type of image is an image other than the first type of image.
[0061] Among them, the connected region can be understood as the region where the text regions in the picture can be connected.
[0062] In practical applications, the first type of image can be understood as an image in a general scenario. For this type of image, the results output by the model are generally more accurate, and the cost of manual annotation is not high; while the second type of image can be understood as an image in a specific scenario. For this type of image in the model, the output results may have large errors. Moreover, for this type of image, manual annotation will also consume a large amount of manpower and material resources. Therefore, it is generally not recommended to perform manual annotation on this type of image in a specific scenario.
[0063] It should be noted that the general scenario can be understood as a scenario where the area of the connected region where the text in the image is located is greater than the preset area threshold, and the characteristic scenario can be understood as a scenario where the area of the connected region where the text in the image is located is less than or equal to the preset area threshold; for example, the image in the general scenario can be an image of an electronic document. When the areas of the regions where the text is located are connected together, the total area of the formed connected region is greater than the preset area threshold; the image in the specific scenario can be a hospital test report. The regions where the text is located are irregularly distributed, and the multiple regions where the corresponding text is located cannot be connected together to form a large connected area. Therefore, the total area of the connected region is not greater than the preset area threshold.
[0064] Furthermore, due to the differences in the dataset distributions of the general scenario and the specific scenario. In addition to determining the image type by the area of the connected region of the text in the picture, the pattern type can also be determined in advance by whether it is labeled data. For example, the sample images in the general scenario can be publicly available labeled datasets or datasets created for training the model; the sample images in the specific scenario can be datasets with no or few labeled data. At the same time, due to problems such as complex backgrounds (for example, there are both handwritten fonts and printed fonts, which are difficult to distinguish), and it is also difficult to create such datasets; the embodiments of this specification do not make any limitations on the distinction between the general scenario and the specific scenario data here.
[0065] The detection model training method provided by the embodiments of this specification classifies the sample images by using the images with the area of the text connected region greater than the preset area threshold as labeled sample images and the images with the area of the text connected region not greater than the preset area threshold as unlabeled sample images, and trains the model in a semi-supervised manner. During the training process of the model, data in a specific scenario is added, so that the model will also have a high detection accuracy for images in the characteristic scenario during the later application process.
[0066] Step 104: Process the labeled sample images to obtain a first fused image and a first label, and process the unlabeled sample images and the second detection result to obtain a second fused image and a second label.
[0067] Among them, the first fused image can be understood as an image obtained by fusing the original image in the labeled sample image with the label corresponding to the original image, and the second fused image can be understood as an image obtained by fusing the original image of the unlabeled sample image with the result predicted by the generator for the original image.
[0068] The first label can be understood as the label after labeling the first fused image through manual or machine processing methods, and the second label can be understood as the label after labeling the second fused image through manual or machine processing methods.
[0069] In practical applications, first process the multiple labeled sample images input in this iteration to obtain multiple first fused images corresponding to the labeled sample images and the corresponding first labels; then process the multiple unlabeled sample images and the detection results output by the generator for the unlabeled sample images to obtain the second fused images of the unlabeled sample images and the corresponding second labels.
[0070] Further, it is necessary to first perform splicing processing on the labeled sample images to obtain the first fused image; specifically, the processing of the labeled sample images to obtain the first fused image and the first label includes:
[0071] Splice the initial sample image in the labeled sample image and the label corresponding to the initial sample image to obtain the first fused image;
[0072] Label the first fused image to determine the first label corresponding to the first fused image.
[0073] In specific implementation, for the labeled sample image, first splice the initial sample image and the label corresponding to the initial sample image in the image channel dimension to obtain the first fused image; in order to better train the discriminator in the subsequent process, it is necessary to verify whether the result output by the discriminator is accurate in this method. Therefore, label the first fused image to determine the first label corresponding to the first fused image, where the first label can be a binary value, for example, 1 represents that the fused image is a labeled sample image.
[0074] In practical applications, in this round of iteration, it is necessary to perform splicing processing on each initial sample image and the label corresponding to each initial sample image in a pairwise manner, so as to obtain multiple first fusion images, and then perform labeling on each first fusion image to determine the first identifier corresponding to each fusion image. For example, if there are 5 original images in this round, then each original image and the label corresponding to the original image are spliced to obtain 5 first fusion images, and the 5 first fusion images are respectively labeled, so that each first fusion image has an identifier indicating whether the fusion image is labeled data or unlabeled data.
[0075] The detection model training method provided by the embodiments of this specification determines the first fusion image by splicing the original image and the label in the labeled sample image, and performs labeling on the first fusion image to determine the first identifier, which is convenient for subsequently inputting the first fusion image into the discriminator to enable the discriminator to determine whether the fusion image is labeled data, so as to test the discrimination ability of the discriminator.
[0076] Furthermore, for the splicing process of unlabeled sample data, the unlabeled original image is spliced with the label of the text box predicted by the generator for the unlabeled original image, so as to obtain a second fusion image; specifically, the processing of the unlabeled sample image and the second detection result to obtain a second fusion image and a second identifier includes:
[0077] Perform splicing processing on the unlabeled sample image and the second detection result to obtain a second fusion image;
[0078] Perform labeling on the second fusion image to determine the second identifier corresponding to the second fusion image.
[0079] It should be noted that since there is no corresponding label indicating the position of the text area in the image when the unlabeled sample data is input, after the unlabeled original image is input into the generator, the generator can output a predicted detection result, which indicates the position where the text in the unlabeled original image exists, that is, the text box.
[0080] Furthermore, in practical applications, multiple unlabeled sample original images in this round of iteration can be spliced with the detection results corresponding to each original image output by the generator to obtain multiple second fusion images; similarly, it is also necessary to perform labeling on each second fusion image to determine whether the fusion image is unlabeled data, so as to facilitate subsequent verification of the discrimination result of the discriminator.
[0081] The detection model training method provided by the embodiments of this specification determines a second fusion image by splicing the original image of the unlabeled sample image and the detection result predicted by the generator, and labels the second fusion image to determine a second identifier, which is convenient for subsequently inputting the second fusion image into a discriminator to enable the discriminator to determine whether the fusion image is unlabeled data, so as to test the discrimination ability of the discriminator.
[0082] Based on this, the splicing process in the detection model training method provided by the embodiments of this specification includes:
[0083] Determine the first splicing information in the sample image to be processed, where the first splicing information includes the first dimension value, the second dimension value, and the color channel identifier of the pixel point;
[0084] Determine the second splicing information of the label corresponding to the sample image to be processed, where the second splicing information includes the first dimension value, the second dimension value, and the label identifier of the pixel point;
[0085] Perform splicing processing on the image channel dimension based on the first splicing information and the second splicing information.
[0086] In practical applications, for the splicing processing of images, the first splicing information in the sample image to be processed can be obtained, that is, the first dimension value, the second dimension value, and the color channel identifier of each pixel point in the sample image to be processed. For example, the first splicing information is (800, 600, 3). It should be noted that compared with grayscale images, each pixel point in an RGB color image has 3 channels. Each channel occupies 8 bits of memory space. In memory, RGB images are stored in the form of a two-dimensional array; then the second splicing information in the sample image to be processed can also be obtained, that is, the first dimension value, the second dimension value, and the label identifier of each pixel point in the label of the sample image to be processed. For example, the second splicing information is (800, 600, 1). Among them, the label identifier can be understood as identifying whether the pixel point is a text area. For example, when the label identifier is 1, it means that the pixel point is a text area, and when the label is 0, it means that the pixel point is a non-text area; finally, perform splicing processing on the image channel dimension based on the first splicing information and the second splicing information. Finally, the spliced information is (800, 600, 4). It should be noted that the label of the above sample image to be processed and the detection result output by the generator need to be multiplied by 255 before splicing.
[0087] The detection model training method provided by the embodiments of this specification obtains a fusion image by performing splicing processing on the pixel points in the sample image to be processed, so that the features corresponding to the text boxes in the image can be fused in the image, which is convenient for the discriminator to judge whether the fusion image is labeled data.
[0088] Step 106: Input the first fused image, the first identifier, the second fused image, and the second identifier into the discrimination module to determine the discrimination result of the second fused image.
[0089] Among them, the discrimination module can be understood as the discriminator mentioned in the above embodiments, which is used to output a probability value to judge true or false. For example, if the probability value is greater than or equal to 0.5, it is true; if it is less than 0.5, it is false. Currently, the definition of true and false can be understood as a concept defined artificially, and this specification embodiment will not elaborate too much on this.
[0090] In practical applications, input the first fused image, the first identifier, the second fused image, and the second identifier into the discriminator together, and then the discrimination result of the second fused image can be obtained. Among them, this discrimination result can be understood as the probability value that the second fused image is labeled data. If the discrimination result is 0.8 (greater than 0.5), then it is determined that the second fused image is labeled data (positive sample). If the discrimination result is 0.3 (less than 0.5), then it is determined that the second fused image is unlabeled data (negative sample).
[0091] Step 108: When the discrimination result meets the preset processing condition, determine the first loss function based on the first fused image and the second fused image, and determine the second loss function based on the first detection result and the label corresponding to the initial sample image.
[0092] Among them, the preset processing condition can be understood as a situation where there is a misjudgment in the discrimination result output by the discriminator. If it is determined that the discriminator has a misjudgment, then it can be determined that the current discrimination result meets the preset processing condition. If it is determined that there is no misjudgment in this discrimination result, then it can be determined that the current discrimination result does not meet the preset processing condition.
[0093] In practical applications, the condition for determining that the discriminator has a misjudgment is to determine whether the discrimination result output by the discriminator matches the identifier determined on the fused image. For example, the discriminator outputs the discrimination results of 10 second fused images. Among them, the discrimination results of 8 second fused images are less than 0.5, and the discrimination results of 2 second fused images are greater than 0.5. Then, it can be determined that among the 10 second fused images in this round, 2 fused images are displayed as having labeled data; since the second fused image has been labeled before being input into the discriminator, and it is determined that the second fused image is unlabeled data, then among the results output by the discriminator, if there are 2 second fused images as having labeled data, it can be determined that there is a misjudgment phenomenon in the discrimination process of this round of iteration of the discriminator.
[0094] Further, after it is determined that the discriminator has misjudged, that is, the preset processing condition is satisfied, then the first loss function can be determined based on the first fusion image and the second fusion image; at the same time, the second loss function can also be determined according to the first detection result output by the generator and the label corresponding to the initial sample image (the label of the labeled sample image has been manually labeled in advance).
[0095] According to the discrimination result output by the discriminator, it is necessary to determine the loss function based on the first fusion image and the second fusion image. In order to make the determined difference smaller, the calculation method provided by the detection model training method in this specification embodiment is to adopt the mean square deviation of variance; specifically, the determining the first loss function based on the first fusion image and the second fusion image includes:
[0096] Extract the features of the first fusion image to obtain a first feature map;
[0097] Extract the features of the second fusion image to obtain a second feature map;
[0098] Process the first feature map to determine the first processing value of the first fusion image, and process the second feature map to obtain the second processing value of the second fusion image;
[0099] Take the difference between the first processing value and the second processing value as the first loss function.
[0100] Specifically, after inputting the first fusion image and the second fusion image into the discriminator, the discriminator can respectively extract the features of the first fusion image and the second fusion image to obtain a first feature map and a second feature map; then process the first feature map to obtain the first processing value of the first fusion image, and also process the second feature map to obtain the second processing value of the second fusion image. Finally, the result of the square of the difference between the first processing value and the second processing value is determined as the first loss function.
[0101] In practical applications, the first fusion image input into the discriminator can be multiple fusion images in this round of iteration, and the second fusion image can be multiple fusion images in this round of iteration. Then, feature extraction can be performed on each fusion image to obtain corresponding feature maps; then calculate the mean square deviation for each feature map of the first fusion image, and then calculate the mean square deviation for each feature map of the second fusion image. Finally, continue to calculate the mean square deviation of the mean square deviation of each feature map to obtain the first processing value and the second processing value, and take the result of the square of the difference between the first processing value and the second processing value as the first loss function.
[0102] For example, there are 10 first fusion images and 10 second fusion images input to the discriminator. In this round of iteration, first, feature extraction is performed on each first fusion image and each second fusion image, respectively obtaining 10 first feature maps and 10 second feature maps; then, the mean square error is calculated for the 10 first feature maps to obtain 10 first values to be processed, and the mean square error is calculated for the 10 second feature maps to obtain 10 second values to be processed; further, the mean square error is calculated for the 10 first values to be processed to obtain a first processed value, and the mean square error is calculated for the 10 second values to be processed to obtain a second processed value. Finally, the difference between the first processed value and the second processed value is calculated, and the result of the square of the difference is determined as the first loss function.
[0103] It should be noted that the calculation method of the loss function between the first fusion image and the second fusion image in the detection model training method provided in this specification embodiment is for the purpose of guiding the learning of the generator to make the prediction result of the unlabeled original image approach the label corresponding to the labeled original image. The present specification does not make specific limitations on the above-provided calculation method.
[0104] The detection model training method provided in this specification embodiment extracts feature maps from the first fusion image and the second fusion image, and then performs two mean square error calculations, making the difference after the mean square error smaller, being able to accurately determine the loss function, which is convenient for subsequent adjustment of the parameters in the generator.
[0105] Step 110: Train the generation module based on the first loss function and the second loss function, and determine the generation module obtained when the training stop condition is reached as the target generation module.
[0106] In practical applications, during the iterative process of this model training, the determined first loss function and second loss function are used to train the generator. Further, in subsequent multiple rounds of iteration, the parameters of the generator are continuously adjusted until the training stop condition is reached, and the obtained generator can be used as the target generator.
[0107] It should be noted that the training stop condition can be reaching a preset number of times, or other preset thresholds determined according to the application scenario. The present specification does not make specific limitations on this training stop condition.
[0108] In addition, the discriminator can not only output the discrimination result of the second fusion image, but also output the discrimination result of the first fusion image, thereby realizing the training of the discriminator; specifically, after inputting the first fusion image and the first identifier into the discrimination module, it further includes:
[0109] Determine the discrimination result of the first fused image. When the discrimination result meets the preset processing conditions, determine the fused image to be processed from the first fused image based on the discrimination result;
[0110] Determine the label corresponding to the labeled sample image to be processed and the first detection result based on the fused image to be processed. Determine the third loss function based on the label corresponding to the labeled sample image to be processed and the first detection result of the labeled sample image to be processed;
[0111] Train the discrimination module based on the first loss function and the third loss function until the training stop condition is reached.
[0112] Among them, the fused image to be processed can be understood as the first fused image corresponding to the case where there is a misjudgment in the discrimination, determined according to the discrimination result and the first identifier.
[0113] In specific implementation, after the first fused image and the first identifier are input into the discriminator, the discriminator can also output the discrimination result of the first fused image. At the same time, when it is determined that the discrimination result meets the preset processing conditions, the fused image to be processed can be determined from the first fused image according to the discrimination result; after determining which first fused images have misjudgments, the label corresponding to the labeled sample image to be processed (that is, the label corresponding to the original labeled image) and the first detection result (the detection result predicted by the generator for the labeled data) can be determined according to the fused image to be processed, and the third loss function can be calculated according to the label corresponding to the original labeled image and the detection result predicted by the generator for the labeled data; finally, the discriminator is trained according to the calculated first loss function and the third loss function until the training stop condition is reached.
[0114] In practical applications, in the case of misjudgment in the discrimination result of the discriminator for the first fused image, there is no need to use labeled data and unlabeled data to determine the loss function, but only calculate the loss function for the sample images with misjudgments in the labeled sample images. Then, for the training process of the discriminator, the parameters in the discriminator still need to be adjusted based on the first loss function and the third loss function to continuously train the discriminator until the training stop condition is reached. Among them, the training stop condition for the discriminator is the same as that for the generator, and the specific conditions will not be elaborated here.
[0115] For example, during one iteration, there are 10 first fusion images input to the discriminator. The discriminator's discrimination results for the 10 first fusion images are as follows: the discrimination results for 8 first fusion images are greater than or equal to 0.5, and the discrimination results for 2 first fusion images are less than 0.5. Then, it can be determined that there are misjudgments in the 2 first fusion images. Thus, the 2 first fusion images with misjudgment situations can be obtained as the fusion images to be processed. Furthermore, the labels of the labeled sample images corresponding to each fusion image to be processed can be obtained, as well as the detection results predicted by the labeled sample images in the generator. Then, the third loss function can be calculated respectively. Finally, the discriminator can be trained based on the first loss function and the third loss function until the training stop condition is reached.
[0116] The detection model training method provided in the embodiments of this specification requires continuous training of the discriminator while continuously training the generator, so that the discrimination results of the discriminator become more and more accurate, and the phenomenon of misjudgment becomes less and less. Furthermore, when the generator and the discriminator are alternately updated, the generator can learn the information of unlabeled data.
[0117] It should be noted that the detection model training method provided in the above embodiments details the steps of how to train the supervised data and unsupervised data in one round of iteration. However, during the model training process, a large amount of sample data is required to train the generator and the discriminator, and continuously adjust the parameters in the generator and the discriminator, so that the capabilities of the generator and the discriminator become stronger and stronger. Through the discrimination of the discriminator on the output results of the generator, the text detection results of the generator for images become more and more accurate.
[0118] In summary, through the design of the above detection model training method process, when the generator and the discriminator models are alternately updated, the generator can learn the information of unlabeled data; the design of the discriminator input enables the discriminator to have clear learning objectives and judgment criteria; the design of the loss function in the discriminator enables the generator to minimize the gap between labeled data and unlabeled data; the generator uses the cross-entropy loss function, which enables the input of the label term in the loss function to be non-integer, and the obtained effect is to pull the prediction result of each pixel point of the discriminator towards 0 or 1 (i.e., the text box area and the non-text box area), so that the discriminator can learn the features of unlabeled data.
[0119] The following combination of attached Figure 2 , Figure 2 shows a schematic diagram of the processing process of a detection model training method provided in an embodiment of this specification.
[0120] Figure 2 There are two types of data in Figure 1 and the original Figure 1The corresponding label, and the second is unlabeled data, which includes the original Figure 2 ; Figure 2 There are two modules, namely the generator and the discriminator. The generator can be a segmentation-based text detection network, such as the DifferentialBinarization network, and the discriminator is a lightweight binary classification network.
[0121] In practical applications, in each round of iteration, the network parameters of the discriminator are first frozen. The training process is as follows: 1) Input the original Figure 1 and the label in the labeled data into the generator, and the generator can output the segmentation result of the labeled data; 2) Input the original Figure 2 in the unlabeled data into the generator, and the generator can output the segmentation result of the unlabeled data; 3) Concatenate the original Figure 1 in the labeled data and the label in the channel dimension of the image to obtain the labeled data fusion map; 4) Concatenate the original Figure 2 in the unlabeled data and the segmentation result of the unlabeled data in the channel dimension of the image to obtain the unlabeled data fusion map; 5) Input the labeled data fusion map and the unlabeled data fusion map into the discriminator, and the discriminator discriminates the labeled data fusion map and the unlabeled data fusion map to determine whether the size of the probability value is greater than or equal to 0.5. The discrimination result with the probability value determined to be greater than or equal to 0.5 is used as the positive sample, and the discrimination result with the probability value determined to be less than 0.5 is used as the negative sample; 6) Determine the first loss function according to the first discrimination result corresponding to the labeled data fusion map output by the discriminator; 7) Determine the second loss function according to the second discrimination result corresponding to the unlabeled data fusion map output by the discriminator; among them, the calculation method of the second loss function can calculate the mean square error of the variance based on the feature maps corresponding to the labeled data fusion map and the unlabeled data fusion map, and then determine the loss function; 8) Determine the third loss function according to the segmentation result of the labeled data and the original Figure 1 label in the labeled data; 9) Train the discriminator by backpropagating the first loss function and the second loss function. The positive sample is the labeled data, and the negative sample is the unlabeled data; 10) Train the generator by backpropagating the second loss function and the third loss function; 11) After multiple rounds of the above iteration, the obtained generator is the pre-trained model to be obtained.
[0122] It should be noted that when training specific scenario data with this pre-trained model, the accuracy will be significantly improved. The input of the discriminator is the combination of the original labeled data image and the label, and the combination of the original unlabeled data image and the discriminator's predicted image. The effect is to establish the connection between the original image and the predicted image of the generator, making the discrimination target of the discriminator clear. Further calculation of the second loss function has the effect of guiding the learning for the prediction of the generator to approach the label. At the same time, the unlabeled data used by the generator is the entire original image, rather than just some pixel points, and the effect is that there is a lot of available information, and the generator can utilize the global information of the unlabeled data.
[0123] The detection model training method provided in the embodiments of this specification, through the combination of a generator and a discriminator, in the data input into the discriminator, uses the splicing processing method to establish the connection between the original image and the predicted image of the generator. At the same time, based on the spliced labeled data fusion graph and unlabeled data fusion graph, the loss function is calculated, which increases the guidance for the learning of the generator. Furthermore, the unlabeled data used is the entire original image, using all available pixel points, mastering the entire global information, which is beneficial for the generator to learn and train an efficient generator.
[0124] See Figure 3 , Figure 3 shows a flowchart of an image detection method provided according to an embodiment of this specification, specifically including the following steps.
[0125] Step 302: Obtain the image to be processed.
[0126] Among them, the image to be processed can be understood as an image for which text detection needs to be performed on the text in the image. For example, a hospital report form, where the connected regions of the text in the report form are relatively small and are mainly distributed in various regions of the image. Therefore, it can be determined that the hospital report form is an image in a specific scenario.
[0127] Step 304: Input the image to be processed into the target generation module to obtain the detection result of the image to be processed, where the target generation module is trained by the detection model training method provided in the above embodiments.
[0128] Among them, the target generation module can be understood as an image detection model. By inputting the image to be processed into the image detection model, all the text detection frames in the image to be processed can be output.
[0129] In practical applications, the image to be processed is input into the image detection model. The image detection model is a target generation module trained by the detection model training method provided in the above embodiments. When a hospital report form is input into the image detection model, text detection frames for all the text in the hospital report form can be input, which facilitates subsequent text extraction and other processing based on the text detection frames. The image detection method provided in the embodiments of this specification does not make specific limitations in this regard.
[0130] The image detection method provided in the embodiments of this specification uses the target generation module determined by the training method of the detection model mentioned in the above embodiments. For the image to be processed in any scenario, it can accurately and efficiently output the text detection frame corresponding to the image to be processed, which facilitates subsequent feature extraction of the text based on the text detection frame.
[0131] Corresponding to the above method embodiments, this specification also provides embodiments of a detection model training device. Figure 4 The structural schematic diagram of a detection model training device provided by an embodiment of this specification is shown. As Figure 4 shown, the device includes:
[0132] An image acquisition module 402, configured to acquire a labeled sample image and an unlabeled sample image, input the labeled sample image into the generation module to obtain a first detection result, and input the unlabeled sample image into the generation module to obtain a second detection result, where the labeled sample image includes an initial sample image and the label corresponding to the initial sample image;
[0133] An image processing module 404, configured to process the labeled sample image to obtain a first fused image and a first identifier, and process the unlabeled sample image and the second detection result to obtain a second fused image and a second identifier;
[0134] An image input module 406, configured to input the first fused image, the first identifier, the second fused image, and the second identifier into the discrimination module to determine the discrimination result of the second fused image;
[0135] A loss function determination module 408, configured to, when the discrimination result meets a preset processing condition, determine a first loss function based on the first fused image and the second fused image, and determine a second loss function based on the first detection result and the label corresponding to the initial sample image;
[0136] A training module 410, configured to train the generation module based on the first loss function and the second loss function, and determine the generation module obtained when the training stop condition is reached as the target generation module.
[0137] Optionally, the device further includes:
[0138] A determination module, configured to determine a discrimination result of a first fused image, and based on the discrimination result, determine a to-be-processed fused image from the first fused image when the discrimination result meets a preset processing condition;
[0139] Based on the to-be-processed fused image, determine a label corresponding to a to-be-processed labeled sample image and a first detection result, and determine a third loss function based on the label corresponding to the to-be-processed labeled sample image and the first detection result of the to-be-processed labeled sample image;
[0140] Train the discrimination module based on the first loss function and the third loss function until a training stop condition is reached.
[0141] Optionally, the image processing module 404 is further configured to:
[0142] Perform splicing processing on the initial sample image in the labeled sample image and the label corresponding to the initial sample image to obtain a first fused image;
[0143] Label the first fused image to determine a first identifier corresponding to the first fused image.
[0144] Optionally, the image processing module 404 is further configured to:
[0145] Perform splicing processing on the unlabeled sample image and the second detection result to obtain a second fused image;
[0146] Label the second fused image to determine a second identifier corresponding to the second fused image.
[0147] Optionally, the device further includes:
[0148] A splicing module, configured to determine first splicing information in a to-be-processed sample image, where the first splicing information includes a first dimension value, a second dimension value, and a color channel identifier of a pixel point;
[0149] Determine second splicing information of a label corresponding to the to-be-processed sample image, where the second splicing information includes a first dimension value, a second dimension value, and a label identifier of a pixel point;
[0150] Perform splicing processing on the first splicing information and the second splicing information in the image channel dimension.
[0151] Optionally, the loss function determination module 408 is further configured to:
[0152] Extract the features of the first fused image to obtain a first feature map;
[0153] Extract the features of the second fused image to obtain a second feature map;
[0154] Process the first feature map to determine a first processing value of the first fused image, and process the second feature map to obtain a second processing value of the second fused image;
[0155] Use the difference between the first processing value and the second processing value as a first loss function.
[0156] Optionally, the device further includes:
[0157] An image determination module configured to determine a first type of image as a labeled sample image and a second type of image as an unlabeled sample image, where the first type of image is an image in which the area of the text connected region is greater than a preset area threshold, and the second type of image is an image other than the first type of image.
[0158] The detection model training device provided in this specification obtains a labeled sample image and an unlabeled sample image, and inputs them into a generation module respectively to obtain a first detection result and a second detection result. At the same time, the labeled sample image and the unlabeled sample image are processed respectively to obtain a first fused image and a first identifier, a second fused image and a second identifier, and the first fused image and the second fused image are input into a discrimination module to generate a discrimination result. When it is determined that the discrimination result meets the preset processing conditions, a first loss function and a second loss function are determined, and then the generation module is trained. This method uses a semi-supervised training method with labeled sample images and unlabeled sample images, and the generation module and the discrimination module are used to process the sample images respectively. The generation module can also learn the information of the unlabeled sample images. In addition, for the input of the discrimination module, the pre-processed fused images are also used, so that the discrimination module has a clear learning objective and judgment criterion; finally, according to the discrimination result, two loss functions are determined to train the generation module, so that the generation module can learn the features of the unlabeled sample images. This generation module can be applied not only to the scenario of labeled images, but also to the scenario of unlabeled images, thereby improving the accuracy of the detection results of the generation module and enhancing the image processing efficiency.
[0159] The above is a schematic solution of a detection model training device according to this embodiment. It should be noted that the technical solution of this detection model training device and the technical solution of the above detection model training method belong to the same concept. For the details not described in the technical solution of this detection model training device, reference can be made to the description of the technical solution of the above detection model training method.
[0160] Corresponding to the above method embodiments, this specification also provides embodiments of an image detection device. Figure 5 FIG. shows a schematic structural diagram of an image detection device provided by an embodiment of this specification. As Figure 5 shown, the device includes:
[0161] An acquisition module 502, configured to acquire an image to be processed;
[0162] A detection module 504, configured to input the image to be processed into a target generation module to obtain a detection result of the image to be processed, where the target generation module is trained by the above detection model training method.
[0163] The image detection device provided by this specification, through the target generation module determined by the training method of the detection model mentioned in the above embodiment, can accurately and efficiently output a text detection frame corresponding to the image to be processed for any image to be processed, facilitating subsequent feature extraction of the text based on the text detection frame, etc.
[0164] The above is a schematic solution of an image detection device according to this embodiment. It should be noted that the technical solution of this image detection device and the technical solution of the above image detection method belong to the same concept. For the details not described in detail in the technical solution of the image detection device, reference can be made to the description of the technical solution of the above image detection method.
[0165] Figure 6 FIG. shows a block diagram of the structure of a computing device 600 provided by an embodiment of this specification. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.
[0166] The computing device 600 further includes an access device 640, and the access device 640 enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interfaces (e.g., a Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0167] In an embodiment of this specification, the above components of the computing device 600 andFigure 6 Other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0168] The computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a PC. The computing device 600 can also be a mobile or stationary server.
[0169] Wherein, the processor 620 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above detection model training method or image detection method are implemented.
[0170] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above detection model training method or image detection belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above detection model training method or image detection method.
[0171] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the above detection model training method or image detection method are implemented.
[0172] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above detection model training method or image detection belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above detection model training method or image detection method.
[0173] An embodiment of this specification also provides a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the above detection model training method or image detection method.
[0174] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solutions of the above detection model training method or image detection method belong to the same concept. For the details not described in the technical solution of the computer program, reference can be made to the descriptions of the technical solutions of the above detection model training method or image detection method.
[0175] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0176] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0177] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0178] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0179] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of the present specification. These embodiments are selected and specifically described in the present specification to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification. The present specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for training a detection model, comprising: obtaining labeled sample images and unlabeled sample images, inputting the labeled sample images into a generation module to obtain a first detection result, and inputting the unlabeled sample images into the generation module to obtain a second detection result, wherein the labeled sample images include initial sample images and labels corresponding to the initial sample images; processing the labeled sample images to obtain a first fused image and a first identifier, processing the unlabeled sample images and the second detection result to obtain a second fused image and a second identifier; inputting the first fused image, the first identifier, the second fused image, and the second identifier into a discrimination module to determine a discrimination result of the second fused image; when the discrimination result meets a preset processing condition, determining a first loss function based on the first fused image and the second fused image, and determining a second loss function based on the first detection result and the label corresponding to the initial sample image; training the generation module based on the first loss function and the second loss function, and determining the generation module obtained when the training stop condition is reached as the target generation module.
2. The method for training a detection model according to claim 1, after inputting the first fused image and the first identifier into the discrimination module, further comprising: determining a discrimination result of the first fused image, and when the discrimination result meets a preset processing condition, determining a to-be-processed fused image from the first fused image based on the discrimination result; determining a label corresponding to the to-be-processed labeled sample image and a first detection result based on the to-be-processed fused image, and determining a third loss function based on the label corresponding to the to-be-processed labeled sample image and the first detection result of the to-be-processed labeled sample image; training the discrimination module based on the first loss function and the third loss function until the training stop condition is reached.
3. The method for training a detection model according to claim 1, wherein the processing the labeled sample images to obtain a first fused image and a first identifier comprises: performing a splicing process on the initial sample image and the label corresponding to the initial sample image in the labeled sample image to obtain a first fused image; marking the first fused image to determine a first identifier corresponding to the first fused image.
4. The method for training a detection model according to claim 3, wherein the processing the unlabeled sample images and the second detection result to obtain a second fused image and a second identifier comprises: performing a splicing process on the unlabeled sample images and the second detection result to obtain a second fused image; marking the second fused image to determine a second identifier corresponding to the second fused image.
5. The method for training a detection model according to claim 3 or 4, wherein the splicing process comprises: determining first splicing information in the to-be-processed sample image, wherein the first splicing information includes a first dimension value, a second dimension value, and a color channel identifier of a pixel point; Determine the second splicing information corresponding to the sample image to be processed, where the second splicing information includes the first dimension value, the second dimension value of the pixel point, and the label identifier; Perform splicing processing on the image channel dimension based on the first splicing information and the second splicing information.
6. The detection model training method according to claim 1, wherein determining the first loss function based on the first fused image and the second fused image, comprises: Extract the features of the first fused image to obtain a first feature map; Extract the features of the second fused image to obtain a second feature map; Process the first feature map to determine the first processing value of the first fused image, and process the second feature map to obtain the second processing value of the second fused image; Use the difference between the first processing value and the second processing value as the first loss function.
7. The detection model training method according to claim 1, before obtaining the labeled sample image and the unlabeled sample image, further comprises: Determine the first type of image as the labeled sample image, and determine the second type of image as the unlabeled sample image, where the first type of image is an image in which the area of the text connected region is greater than a preset area threshold, and the second type of image is an image other than the first type of image.
8. An image detection method, comprises: Obtain the image to be processed; Input the image to be processed into the target generation module to obtain the detection result of the image to be processed, where the target generation module is trained by the detection model training method according to any one of claims 1-7.
9. A detection model training device, comprises: An image acquisition module, configured to acquire a labeled sample image and an unlabeled sample image, input the labeled sample image into the generation module to obtain a first detection result, and input the unlabeled sample image into the generation module to obtain a second detection result, where the labeled sample image includes the initial sample image and the label corresponding to the initial sample image; An image processing module, configured to process the labeled sample image to obtain a first fused image and a first identifier, and process the unlabeled sample image and the second detection result to obtain a second fused image and a second identifier; An image input module, configured to input the first fused image and the first identifier, and the second fused image and the second identifier into the discrimination module to determine the discrimination result of the second fused image; A loss function determination module, configured to, when the discrimination result meets a preset processing condition, determine a first loss function based on the first fused image and the second fused image, and determine a second loss function based on the first detection result and the label corresponding to the initial sample image; A training module, configured to train the generation module based on the first loss function and the second loss function, and determine the generation module obtained when the training stop condition is reached as the target generation module.
10. An image detection device, comprises: An acquisition module, configured to acquire the image to be processed; A detection module, configured to input the image to be processed into a target generation module to obtain a detection result of the image to be processed, wherein the target generation module is trained by the detection model training method according to any one of claims 1-7.
11. A computing device, comprising: a memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the detection model training method according to any one of claims 1-7 are implemented.
12. A computer-readable storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the detection model training method according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Recognition model training method and device and risk website recognition method and device
CN110807197A
Target recognition model training method and device and electronic equipment
CN112990432A