Efficient generation of multimodal sequences using between-frame and within-frame machine-learned models

EP4609319A1Pending Publication Date: 2025-09-03GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024837241
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-09
Publication Date
2025-09-03

AI Technical Summary

Technical Problem

Existing methods for generating multimodal sequences, such as autoregressive models, face high computational costs, especially when dealing with many tokens per frame, while non-autoregressive methods often produce lower quality outputs.

Method used

The system employs a multiscale machine-learned architecture that uses a first machine-learned sequence processing model to generate a frame token, and a second model to generate frame-aligned tokens within each frame, reducing computational costs while maintaining output quality.

Benefits of technology

This approach enables the efficient generation of multimodal sequences with similar or better quality than autoregressive methods, at a lower computational cost, and achieves better performance and energy efficiency compared to prior methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059142_19062025_PF_FP_ABST
    Figure US2024059142_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A multimodal sequence can be generated on a frame-by-frame basis, where each frame can be associated with a multi-token portion of the sequence. A first machine-learned model (a "frame model") can generate a frame token representative of an entire frame (e.g. a time frame), and one or more additional models ("depth models") can generate individual tokens associated with the frame based on the frame token.
Need to check novelty before this filing date? Find Prior Art