This application discloses a
text recognition method, apparatus, electronic device, and storage medium. This invention can be applied to various scenarios such as cloud technology,
artificial intelligence, intelligent transportation, and assisted driving. This application can acquire a text image; perform
convolution processing on the text image to obtain a first feature; perform dilated
convolution processing on the text image to obtain a second feature; fuse the first feature and the second feature to obtain a fused feature; and recognize the text content in the text
image based on the fused feature. In this application,
convolution can extract the spatially denser (smaller scale) first feature, while dilated convolution can extract the spatially sparser (larger scale) second feature. By fusing these two features, a fused feature that balances spatial continuity can be obtained. Therefore, recognizing the text content in a text image using this fused feature can improve recognition accuracy.