This invention discloses a training method for a scene text erasure model. A scene
text detection dataset is used as the
training set for the scene text erasure model. The last classification layer of the
baseline model is replaced with two parallel classification
layers, thus dividing the entire model into a background restoration
branch and a text erasure
branch, resulting in the scene text erasure model. The background restoration
branch is trained by taking a partially occluded
background image as input and predicting the background filling content of text regions and randomly occluded regions. During training, the
background image is used as the
label for this branch to supervise its learning process. The text erasure branch is trained by taking an input image as input and predicting the background filling content after text regions are erased and restored. During training, the replaced image is used as the pseudo-
label for this branch to supervise its learning process. This invention trains the scene text erasure model using only a
text detection dataset in a weakly supervised manner.