A method for cross-
modal semantic understanding and generation of IMU signals for LLM (Limited
Language Model) is proposed. In the offline multimodal alignment stage, a pre-trained visual
language model is used to generate motion-related text descriptions from videos synchronized with the IMU signals. Multi-granular features are extracted through temporal fusion, and an IMU
encoder is trained using coarse-grained inter-sample alignment and fine-grained intra-sample alignment strategies to achieve initial alignment of IMU signals and text in a shared
semantic space. In the online retrieval and enhancement generation stage, features are extracted from newly input IMU signals using the pre-trained IMU
encoder, and similar
signal-text pairs are retrieved from a pre-constructed IMU-text
pairing RAG
database. A
structural similarity graph between IMU signals is constructed, and a
central node and its neighbors are selected to form an
anchor point-cluster structure. This structure, along with the retrieved text descriptions, is input into a large
language model to generate high-quality text descriptions consistent with the
semantics of the IMU signals. This invention enables a large
language model to generate high-quality text descriptions consistent with motion
semantics, even under the constraint of imprecise alignment between the IMU and text, thus improving its perceptual performance in tasks such as
activity recognition, behavior analysis, and human-computer interaction.