This invention discloses a
multimodal data fusion coding method for constructing a unified underlying language for national
big data, belonging to the fields of national
big data architecture, multimodal governance, and domestically produced unified coding. Based on the mathematical fundamental feature theory and the STE domestic coding
system, coupled with the GZ-BigData-RISC-V domestic dedicated
chip and a multimodal AI fusion model, it achieves normalized access to
multimodal data, unified extraction of fundamental features, unified underlying language coding, lossless fusion, and secure storage. Quantitative thresholds are set for cross-
modal fusion efficiency ≥85%,
data consistency ≥99.5%, coding latency ≤1ms, and lossless
fusion rate 100%, completing the unified coding and
semantic association of text, image, audio, video, time-series, and geospatial data. This method breaks down
data heterogeneity barriers, constructs the only nationally interoperable unified underlying language for
big data, and is 100% domestically produced and controllable throughout the entire process, supporting the construction of a national integrated big
data center and promoting the market-oriented circulation of data elements.