Efficient generation of multimodal sequences using between-frame and within-frame machine-learned models
Patent Information
- Application Number
- EP2024837241
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-12-09
- Publication Date
- 2025-09-03
AI Technical Summary
Existing methods for generating multimodal sequences, such as autoregressive models, face high computational costs, especially when dealing with many tokens per frame, while non-autoregressive methods often produce lower quality outputs.
The system employs a multiscale machine-learned architecture that uses a first machine-learned sequence processing model to generate a frame token, and a second model to generate frame-aligned tokens within each frame, reducing computational costs while maintaining output quality.
This approach enables the efficient generation of multimodal sequences with similar or better quality than autoregressive methods, at a lower computational cost, and achieves better performance and energy efficiency compared to prior methods.
Smart Images

Figure US2024059142_19062025_PF_FP_ABST