The application discloses a multi-
modal large model driven geometric image reading
analysis method and
system, relates to the technical field of
data processing, and comprises the following steps: after receiving image data uploaded by a user, a candidate set of geometric elements is extracted, and an initial confidence is set for each candidate element; a geometric labeling and recognition channel is started in parallel, visual encoders are used to extract the candidate set of geometric elements, visual embedding is generated, an OCR channel is activated at the same time for
text recognition, and text semantic embedding is generated; through confidence weighted cross-
modal alignment fusion, structured
geometric representation is generated, which is input into a mixed reasoning and identification model, and an analysis result is output. The application solves the technical problems that the existing geometric
image analysis method cannot effectively process image and text information at the same time, and the analysis accuracy and efficiency of complex geometric structures are low, and achieves the technical effects of accurately extracting geometric information and text labeling in an image through the fusion of multi-
modal large models, and improving the analysis accuracy and efficiency.