The invention discloses a universal streetscape multi-dimensional
perception method based on a multi-
modal large
language model, and relates to the technical field of
computer vision, and the method comprises the steps: obtaining a multi-source streetscape image, carrying out the preprocessing, generating a streetscape text description result, carrying out the analysis
processing of the streetscape text description result based on a preset rule, and generating a streetscape text description
data set; based on the streetscape text description
data set, training a pre-configured multi-
modal large
language model through a low-rank
adaptation mechanism to obtain a streetscape
perception model; and carrying out secondary training on the streetscape
perception model by utilizing an improved thinking chain mechanism to obtain a second-order streetscape
perception model which is used for perceiving multi-dimensional complex information such as visual sense, emotion, sound sense and the like. According to the method, the multi-
modal large
language model is utilized, the cross-modal representation capability, the context reasoning mechanism and the zero sample migration capability are achieved, and large-scale pre-training data and an advanced
deep learning architecture are utilized, so that deep fusion understanding of vision and language information and intelligent analysis of complex scenes are achieved.