The invention discloses an
intelligent character recognition and structured
processing method based on multi-
modal deep learning, and belongs to the technical field of
character recognition, and the method comprises the following steps: 1, receiving an image uploaded by a user, carrying out the detection of a character main
body region, and extracting a main
body region image containing characters; step 2, preprocessing the image; 3, performing character target detection on the preprocessed image by adopting a YOLOv8 model, and positioning a text region; step 4, identifying the detected character area by using a CRNN model, and outputting text content; 5, performing
layout analysis on the recognition result, and determining a reading sequence and a
logic structure of the text; and step 6, using a
natural language model, combining the text content and the coordinate information, and carrying out structured output on an identification result. Multiple tasks such as business licenses and licenses are integrated into one model, the problem of
multiple models in the prior art is solved, the
resource utilization rate is high, and the recognition precision is high.