Model pre-training optimization method and system for fusing multi-modal data

By iteratively aligning and correcting speech vectors at different resolution scales and dynamically adjusting the training loss in conjunction with uncertainty metrics, the robustness problem of fusing speech modalities with other modal features in noisy environments is solved, achieving more efficient noise suppression and feature extraction.

CN120579141BActive Publication Date: 2026-07-03NANJING TORTOISE & HARE RACE SOFTWARE RES INST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING TORTOISE & HARE RACE SOFTWARE RES INST CO LTD
Filing Date
2025-06-05
Publication Date
2026-07-03

Smart Images

  • Figure CN120579141B_ABST
    Figure CN120579141B_ABST
Patent Text Reader

Abstract

The application provides a model pre-training optimization method and system for fusing multi-modal data, and relates to the technical field of multi-modal technology. The method comprises the following steps: after obtaining voice modal data and other modal data, sampling at low resolution and high resolution two scales respectively, generating corresponding voice vector sets and calculating uncertainty metrics; globally aligning the voice and other modal data at the first resolution scale, and transferring the global alignment result to the second resolution scale to perform refinement correction, and finally returning to the first resolution scale, thereby forming a multiple mesh format back-and-forth iteration; by monitoring and observing noise and state noise in real time and integrating them into the uncertainty metric, the weighted coefficient can be dynamically reduced during training for high-noise sections, and a higher weight is given to low-noise or stable sections to strengthen effective features; the application can adaptively suppress noise interference and retain voice mutation details in a multi-modal scene, and has better robustness and generalization performance.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Anti-noise multi-modal characterization method based on coarse-to-fine progressive cross-modal attention

    CN115309950A

  • Multi-modal medical data fusion evaluation method and apparatus, device, and storage medium

    WO2023098524A1