基于向量量化变分自编码器的离散特征建模方法和视觉信息生成系统
By employing a two-stage training and K-means clustering initialization method, the feature latent space of the vector quantization variational autoencoder is optimized, solving the problem of insufficient discrete code representation capability in existing technologies and achieving high-quality visual information generation.
Patent Information
- Application Number
- CN202511327500.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-07-17
- Estimated Expiration
- 2045-09-17
AI Technical Summary
In existing technologies, vector quantization variational autoencoders lack robustness when optimizing the latent space of image features and have limited discrete code representation capabilities, resulting in poor performance of image generation models.
A two-stage training approach is adopted. First, continuous features are learned through pre-trained variational autoencoders. The codebook is initialized using the K-means clustering algorithm and vector quantization variational autoencoders are constructed by combining linear projection layers. Then, discrete codebooks and encoding feature spaces are trained. Finally, an autoregressive image generation model is used for visual generation.
By optimizing the feature latent space and codebook, the expressive power of discrete visual features was improved, the model performance was stabilized, and high-quality visual information understanding and generation were achieved.
Smart Images

Figure CN121259517B_ABST
Abstract
Citation Information
Patent Citations
Discrete speech coding method, electronic equipment and storage medium
CN119296556A
Image classification employing image vectors compressed using vector quantization
US20120076401A1