The invention relates to the technical field of text
data processing, and discloses a text
data extraction method,
system, equipment and medium based on multi-
modal fusion and self-evolution learning, which comprises the following steps of: performing
feature extraction and spatial alignment on a printed text, a handwritten
annotation and a dynamic table of a mixed format document to obtain a
semantic feature of an image-text table, and inputting the
semantic feature into a dynamic analysis layer; analyzing metaphor expressions and synonymous heterogeneous fields through field extraction and a context semantic reasoning mechanism, and outputting structured data; performing grammar compliance
verification by adopting a regularization engine, and performing comparison
verification through a federal learning mechanism; and inputting the verified data into the
reinforcement learning model, updating the analysis rule and the
model parameters through strategy iteration, and feeding back the updated analysis rule and
model parameters to the dynamic analysis layer to complete closed-
loop optimization. According to the method, the
processing precision and efficiency of the complex document are greatly improved, the manual intervention requirement is remarkably reduced, and meanwhile, the
privacy protection and compliance requirements are met.