The invention relates to the technical field of content transmission methods, in particular to a content transmission method based on a multi-
modal streaming media fusion
technology system, which comprises the following steps: S1, multi-
source data acquisition and time-space alignment: adopting an IEEE 1588
PTP protocol, deploying a hardware
clock synchronization module at a camera, a
microphone and IMU equipment, improving the
timestamp precision to a mu s level, and performing time-space alignment on the camera, the
microphone and the IMU equipment; carrying out sampling rate normalization on the sensor data; s2, feature level cross-
modal fusion: extracting each
modal feature by using a lightweight Transform model, executing cross-modal attention calculation in an
edge server, and generating a fusion feature
tensor Ffusion belonging to RN * D; the
system content transmission method based on the multi-modal streaming media fusion technology solves the problems that a traditional multi-modal
transmission system usually adopts a separated transmission architecture, video streams are transmitted through an RTMP protocol, audio streams are encoded through Opus and are transmitted through
WebRTC, sensor data are sent through an
MQTT protocol in a
JSON format, ABR only adjusts the video
code rate, and the transmission efficiency is low. And joint optimization of multi-
modal data is not coordinated.