Model pre-training optimization method and system for fusing multi-modal data
By iteratively aligning and correcting speech vectors at different resolution scales and dynamically adjusting the training loss in conjunction with uncertainty metrics, the robustness problem of fusing speech modalities with other modal features in noisy environments is solved, achieving more efficient noise suppression and feature extraction.
CN120579141BActive Publication Date: 2026-07-03NANJING TORTOISE & HARE RACE SOFTWARE RES INST CO LTD +1
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING TORTOISE & HARE RACE SOFTWARE RES INST CO LTD
- Filing Date
- 2025-06-05
- Publication Date
- 2026-07-03
Smart Images

Figure CN120579141B_ABST
Abstract
The application provides a model pre-training optimization method and system for fusing multi-modal data, and relates to the technical field of multi-modal technology. The method comprises the following steps: after obtaining voice modal data and other modal data, sampling at low resolution and high resolution two scales respectively, generating corresponding voice vector sets and calculating uncertainty metrics; globally aligning the voice and other modal data at the first resolution scale, and transferring the global alignment result to the second resolution scale to perform refinement correction, and finally returning to the first resolution scale, thereby forming a multiple mesh format back-and-forth iteration; by monitoring and observing noise and state noise in real time and integrating them into the uncertainty metric, the weighted coefficient can be dynamically reduced during training for high-noise sections, and a higher weight is given to low-noise or stable sections to strengthen effective features; the application can adaptively suppress noise interference and retain voice mutation details in a multi-modal scene, and has better robustness and generalization performance.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Anti-noise multi-modal characterization method based on coarse-to-fine progressive cross-modal attention
CN115309950A
Multi-modal medical data fusion evaluation method and apparatus, device, and storage medium
WO2023098524A1