基于向量量化变分自编码器的离散特征建模方法和视觉信息生成系统

By employing a two-stage training and K-means clustering initialization method, the feature latent space of the vector quantization variational autoencoder is optimized, solving the problem of insufficient discrete code representation capability in existing technologies and achieving high-quality visual information generation.

CN121259517BActive Publication Date: 2026-07-17ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511327500.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2026-07-17
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

In existing technologies, vector quantization variational autoencoders lack robustness when optimizing the latent space of image features and have limited discrete code representation capabilities, resulting in poor performance of image generation models.

Method used

A two-stage training approach is adopted. First, continuous features are learned through pre-trained variational autoencoders. The codebook is initialized using the K-means clustering algorithm and vector quantization variational autoencoders are constructed by combining linear projection layers. Then, discrete codebooks and encoding feature spaces are trained. Finally, an autoregressive image generation model is used for visual generation.

Benefits of technology

By optimizing the feature latent space and codebook, the expressive power of discrete visual features was improved, the model performance was stabilized, and high-quality visual information understanding and generation were achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121259517B_ABST
    Figure CN121259517B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于向量量化变分自编码器的离散特征建模方法和视觉信息生成系统,属于视觉信息理解与生成领域。所述方法采用两阶段向量量化变分自编码器训练策略:第一阶段利用大规模真实图像数据进行连续特征学习,构建高质量特征潜在空间;第二阶段在引入K‑means聚类初始化的码本和可学习投影层后,学习表达能力更强的离散码本与编码特征空间。基于预训练的向量量化变分自编码器,提取图像的离散编码并用于训练自回归生成模型,从而完成高质量图像重构与生成任务。本发明通过联合特征潜在空间与码字潜在空间的优化,显著提升了视觉生成的感知质量、像素保真度及多样性,在多种数据集和任务中表现出优异的泛化性与鲁棒性。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Discrete speech coding method, electronic equipment and storage medium

    CN119296556A

  • Image classification employing image vectors compressed using vector quantization

    US20120076401A1