Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

168 results about "Vector quantisation" patented technology

Multi-Scale Temporal Attention Processing System for Multimodal Deep Learning with Vector-Quantized Variational Autoencoder

A system and method for multi-scale temporal attention processing in multimodal technology deep learning systems. This system processes time-series, textual, sentiment, and structured tabular data across three hierarchically-organized temporal streams—quarterly, weekly, and intraday levels—with bidirectional cross-temporal information flow. Scale-specific attention mechanisms are optimized for respective temporal granularities, while an adaptive controller dynamically weights each temporal level based on real-time market volatility indicators. A multi-scale fusion processor integrates attention-weighted representations to generate temporally unified representations preserving both short-term market dynamics and long-term trends. This approach enables superior forecasting and risk assessment by leveraging temporal correlations across multiple time scales while automatically adapting to changing market conditions. The system facilitates interpretable AI analysis through attention visualization and enables synthetic scenario generation for model testing.
Owner:ATOMBEAM TECH INC

Adversarial-robust vector quantized variational autoencoder with secure latent space for time-series data

A system and methods for implementing adversarial-robust compression and reconstruction using a vector quantized variational autoencoder (VQ-VAE) with secure latent space management. The system provides comprehensive protection against adversarial attacks through multi-channel threat detection, adaptive defensive parameters, and coordinated response mechanisms. Input data is continuously monitored for potential threats, and defensive parameters are dynamically adjusted based on detected threat levels. The system implements bounded constraints and hierarchical projections to maintain latent space security while preserving compression efficiency. Multi-stage reconstruction with progressive validation ensures reliable data recovery even under adversarial conditions. The system coordinates defensive responses across all compression and reconstruction processes, implementing various recovery mechanisms when security violations are detected. This approach enables robust compression and reconstruction of time-series data while maintaining protection against various forms of adversarial manipulation.
Owner:ATOMBEAM TECH INC

Multimodal data processing and generation system using VQ-VAE and latent transformer

A system and method for processing and generating multimodal data using a combination of Vector Quantized Variational Autoencoder (VQ-VAE) and Latent Transformer architectures. The system efficiently handles diverse data types including time-series, textual, sentiment, and structured tabular data through specialized encoding modules. A novel fusion module integrates these encodings, capturing cross-modal relationships. The fused representation is compressed into a discrete latent space, processed by a latent transformer, and then reconstructed and enhanced. This approach enables superior data compression, reconstruction, analysis, and generation of synthetic scenarios.
Owner:ATOMBEAM TECH INC

Audio coding and decoding method, device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an audio coding and decoding method, device, equipment and medium, and the method comprises the steps: carrying out the sliding window segmentation processing of an input audio signal, and generating signal segments; processing the signal segments through an encoder containing a multi-layer self-attention mechanism to generate continuous potential representations; performing decomposition vector quantization processing on the continuous potential representation to generate a discrete code; the discrete codes are processed through a decoder comprising a multi-layer self-attention mechanism, and reconstructed signal segments are generated; and splicing the reconstructed signal segments to generate a complete audio signal. According to the method, a traditional convolution structure is replaced by a multi-layer self-attention mechanism, global time sequence dependence modeling is carried out on the audio signals after sliding window segmentation, effective compression of potential representation is realized in combination with decomposition vector quantization, and continuity and fidelity of audio reconstruction are improved on the premise that calculation complexity is not increased.
Owner:PING AN TECH (SHENZHEN) CO LTD

Semantic communication method based on semantic perception mask and vector quantization codebook

The invention discloses a semantic communication method based on a semantic perception mask and a vector quantization codebook, and belongs to the technical field of wireless communication. Specifically, the invention provides a semantic perception compression strategy, semantic perception masking is performed on an input image according to the semantic importance size represented by the attention vector of the last self-attention layer of an image encoder of a CLIP model, image blocks with relatively low semantic importance can be selectively masked instead of random masking, and therefore, the semantic perception compression strategy can be applied to image processing. Therefore, the semantic information learning capability of the model is enhanced, and the training efficiency is improved. In addition, an improved robust vector quantization codebook shared at a transmitting end and a receiving end is designed. According to the codebook, coding features are represented by orthogonal and trainable base vector indexes in transmission, so that the robustness of a system is enhanced, and the transmission data volume is reduced. According to the method provided by the invention, the semantic information processing efficiency can be remarkably improved, and better performance is shown under the challenging low signal-to-noise ratio condition.
Owner:SOUTHEAST UNIV

Neural network image compression using representation adaptive VQ encoder and base encoder

A vector quantization (VQ) neural network-based image compression method includes encoding, by a first encoder, a first image to obtain a first latent feature corresponding to the first image; generating, based on the first latent feature and using a leading codebook, a leading codebook indices map and a first codeword map corresponding to the leading codebook indices map, wherein the leading codebook comprises codebook indices that are assigned to identical vectors in multiple codebooks; encoding the leading codebook indices map to generate an encoded leading codebook indices map; and transmitting the encoded leading codebook indices map to a decoder.
Owner:FUTUREWEI TECHNOLOGIES INC

Generation of latent representations of images using a machine learning model

The present disclosure describes techniques for generating latent representations of images using a machine learning model. An image is split and flattened into a series of patches. The series of patches is concatenated with a sequence of latent tokens. The concatenated patches and latent tokens are input into an encoder of the machine learning model. A one-dimensional (1D) latent representation of the image is generated by the encoder. Vector quantization is performed on the 1D latent representation of the image by a vector quantizer of the machine learning model to generate quantized latent tokens. The image is reconstructed based on the quantized latent tokens by a decoder of the machine learning model.
Owner:LEMON INC(GB)

Dynamic environment adaptive semantic communication system and method

The invention relates to a dynamic environment adaptive semantic communication system and method, which are used for realizing the whole task-oriented semantic communication, the system comprises a transmitting terminal network, a codebook space, a wireless channel and a receiving terminal network, the transmitting terminal network comprises a feature extractor and a joint source channel encoder, namely a JSC encoder; the feature extractor is used for identifying and extracting features related to a specific task from the original input data; the JSC encoder is used for: mapping the continuous feature vector output by the feature extractor to a discrete symbol representation suitable for wireless transmission; the codebook space is used for mapping a continuous feature vector output by the feature extractor into a discrete codeword index through a vector quantization mechanism; the wireless channel is used for converting the discrete code word index into an electromagnetic signal capable of being physically transmitted and completing a transmission process of the electromagnetic signal in a noise and fading environment; according to the method, the comprehensive performance optimization of task-oriented semantic communication is realized.
Owner:SHANDONG UNIV

Emotion prediction system, method, device and medium

The invention provides an emotion prediction system, method and device and a medium, a vector quantization variational auto-encoder extracts time-continuous differential entropy features from electroencephalogram signals, maps the features to a submerged space to obtain feature vectors, performs vector quantization to obtain codebook features, splices the codebook features and the feature vectors to obtain dynamic emotion depth features, and predicts the emotion according to the dynamic emotion depth features. The flexible actor-commentator model captures a time sequence dependency relationship of dynamic emotion depth features to obtain depth features for prediction, the depth features for prediction are processed to obtain an emotion prediction value, and the flexible actor-commentator model not only can consider prediction accuracy, but also can obtain a prediction result based on a set reward function. And whether the change trend of the predicted value of the previous and later time steps is consistent with the change trend of the real value can be considered, so that the modeling of the dynamic emotion which continuously changes in time is realized, the regression prediction of the continuous emotion state is realized, and the application of emotion recognition in a brain-computer interface and human-computer interaction is promoted.
Owner:SHENZHEN UNIV

Signal feature compression method based on codebook discrete quantization and multi-task learning

The invention discloses a signal feature compression method based on codebook discrete quantization and multi-task learning, and belongs to the technical field of crossing of signal processing and artificial intelligence. In order to solve the problem that in the prior art, compression efficiency, semantic retention and calculation complexity are difficult to consider in high-dimensional IQ signal compression, high-dimensional continuous features of an original IQ signal are extracted through an encoder; carrying out vector quantization by utilizing the learnable codebook to generate a discrete index and calculating quantization loss; inputting the discrete features into a decoder branch reconstruction signal to calculate reconstruction loss, and inputting a classifier branch prediction category to calculate classification loss; processing traditional features and coding features by adopting a multi-layer perceptron, and calculating comparison loss through cosine similarity; combining optimization quantization loss, reconstruction loss, classification loss and comparison loss to train a model; and finally outputting a discrete index as a compression feature. The method realizes efficient compression and classification semantic reservation, and is suitable for wireless communication, Internet of Things and other scenes.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Distortion and perceptual code rate adaptive variable code rate depth image compression method

The invention provides a distortion and perceptual code rate self-adaptive variable code rate depth image compression method, which is used for constructing an encoder model, and comprises the following steps of: dividing an image to be encoded into image blocks with different granularities, extracting hierarchical features, and filtering by using an image mask to obtain non-repetitively represented hierarchical features; vector quantization is carried out on the hierarchical features based on the hierarchical vector quantization codebook, and discrete index representation is generated and coded into a code stream file; constructing a decoder model, namely decoding a code stream file into discrete index representation, recovering layered features through a layered vector quantization codebook, fusing and decoding by using a mixed condition decoder, and generating a reconstructed image; carrying out joint training on the encoder model, the decoder model and the layered vector quantization codebook; and encoding an input image into a code stream file through the trained encoder model, and decoding the code stream file into a reconstructed image through the trained decoder model. According to the invention, fine bit rate control and adaptive switching between perception and distortion optimization can be realized.
Owner:WUHAN UNIV

Visual target detection method and device based on vector quantization and uncertainty perception

The invention discloses a visual target detection method and device based on vector quantization uncertainty perception, and the method comprises the steps: collecting image data in an open scene as original data, marking the collected original data according to the category, and taking the marked original data as an initial task data set; training a target detection model on the initial task data set, wherein the target detection model comprises a target detection module, a vector quantization module and an uncertainty perception label distribution module; a new category of interest is screened from the explored unknown objects, and image data is collected and marked to serve as a new task data set; meanwhile, selecting a part of samples from the old task data set as a playback sample set; and finely adjusting the target detection model on the new task training set and the old task playback set to realize continuous expansion and evolution of visual target detection. The device comprises a processor and a memory.
Owner:TIANJIN UNIV

Semantic information transmission communication method and system based on VQGAN optimization

The invention provides a semantic information transmission communication method and system based on VQGAN optimization. The method comprises the following steps: establishing a VQGAN model for an image; inputting an image to be transmitted into an encoder of the VQGAN model, and extracting image features corresponding to the image by the encoder; searching the index represented by the discrete potential space with the nearest distance according to the feature vector of the image feature; the index of the embedded vector is recorded, channel coding and modulation are carried out, and transmission is carried out through a channel; after channel transmission, the index is recovered through demodulation and channel decoding; and finding the spatial representation in the codebook according to the index, and performing image reconstruction by using a decoder of a receiving end to obtain a reconstructed image corresponding to the to-be-transmitted image. According to the method, a discrete codebook is established in a vector quantization mode, transmission content is reduced by only transmitting an index of the codebook, and high-resolution image reconstruction is performed based on a generative adversarial network and dynamic change of a confrontation channel and compatibility among different devices based on traditional channel coding.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Self-supervised emotion recognition method based on heart-brain joint codebook and related equipment

The embodiment of the invention provides a self-supervised emotion recognition method based on a heart and brain combined codebook and related equipment, and belongs to the technical field of physiological signal processing and artificial intelligence. The method comprises the following steps: respectively defining heart beats of electrocardiosignals and electroencephalogram signal segments with equal lengths as words, and constructing sentences; a shared heart and brain joint codebook is created and trained, and electrocardio and electroencephalogram words are mapped to a unified discrete semantic space through vector quantization so as to learn cross-subject general characterization; then, discretizing a signal sentence by using the codebook, and combining space and position embedding and inputting a Transform encoder to carry out mask pre-training so as to learn context semantics of the signal; and finally, finely tuning the pre-training model for an emotion recognition task. According to the method, deep semantic fusion of heart and brain signals is realized through signal structuring and codebook sharing, dependence on labeled data is effectively overcome, and emotion recognition accuracy and cross-subject generalization ability are remarkably improved.
Owner:SOUTH CHINA UNIV OF TECH

Multi-modal time sequence fusion voice drive gesture generation method

The invention discloses a multi-modal time sequence fusion voice-driven gesture generation method, which comprises the following steps of: firstly, learning compact discrete representation of gesture motion through a vector quantization variational auto-encoder model, and constructing a quantized potential space for a subsequent generation task; extracting audio features of the voice audio through an audio encoder; performing low-dimensional embedding on the identity of the speaker to obtain an identity feature; performing time sequence alignment on the audio features and the historical gesture sequence and then splicing the audio features and the historical gesture sequence along feature dimensions to form multi-modal initial representation; deep feature fusion is carried out through a multi-modal time sequence fusion module integrating a self-attention mechanism, a cross attention mechanism taking identity features as conditions and a Mamba module; and finally, reconstructing a target gesture sequence through a pre-trained decoder. According to the method, the technical problems of insufficient multi-modal fusion, low calculation efficiency and single generated action are solved, and the gesture animation which is natural, smooth and personalized and meets the real-time interaction requirement can be generated.
Owner:JIANGXI NORMAL UNIV

System and Method for Machine Learning Based CSI Codebook Generation and CSI Reporting

A system and method for a wireless device to derive a channel state information (CSI) codebook based on a decoder and a vector quantization codebook, receive a reference signal from a base station, derive an estimated channel from the reference signal, select an entry from the CSI codebook based on the estimated channel and a selection criterion, and report an index of the selected entry to the base station. The wireless device may receive a subset indication, derive therefrom a second CSI codebook, and select the CSI codebook entry from the second CSI codebook. A CSI compression machine learning (ML) system and vector quantization codebook may be obtained by a network controller. The decoder may be a part of the CSI compression ML system, the vector quantization codebook may be based on an encoder of the CSI compression ML system, and they may be sent to the wireless device.
Owner:HUAWEI TECH CO LTD

Vector quantization methods for UE-driven multi-vendor sequential training

A UE-associated entity may train an encoder to encode uplink control information. The UE-associated entity may determine a quantization codebook to be applied to the encoded uplink control information. The UE-associated entity may share a sequential training dataset with a base station-associated entity, the sequential training dataset including: one of an input vector set or an output vector set; and one of an encoded and unquantized intermediate vector set or an encoded and quantized intermediate vector set. The base station-associated entity may train a decoder based on a quantization codebook and at least the sequential training dataset from the UE-associated entity. When the base station-associated entity receives multiple sequential training datasets for different vendors, the base station-associated entity may train a multi-vendor decoder based on the multiple sequential training datasets.
Owner:QUALCOMM INC

Robust semantic communication method, device and equipment based on sparse vector coding

The invention relates to the technical field of wireless communication, and discloses a robust semantic communication method, device and equipment based on sparse vector coding, and the method comprises the steps: extracting continuous semantic features of original data through a semantic encoder model; discretizing the continuous semantic features into index bits through a vector quantization method by using a learnable feature dictionary; the index bit is mapped into a sparse vector by adopting sparse vector coding, and the sparse vector is sent after being expanded by a codebook; and a receiving end identifies the index through a multipath matching pursuit algorithm and reconstructs the original data. According to the method, a semantic feature extraction model, a vector quantization technology and sparse vector coding are fused, and gradient approximation, joint loss function optimization and a transmission parameter dynamic adjustment mechanism are designed, so that the problems of poor digital transmission compatibility and weak channel adaptability of traditional semantic communication are effectively solved; and the semantic transmission reliability and the system resource utilization efficiency under the dynamic channel condition are improved.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Procedural multi-stage vector quantization for low precision integer neural network weights

An example computing device may include memory and one or more processors. The one or more processors are configured to determine a Cartesian grid in a space. The Cartesian grid may correspond to a precision of an arithmetic processing unit. The one or more processors are configured to, for each of a plurality of neural network weights, determine a respective vector in the space, such that each respective vector belongs to a corresponding predefined vector on the Cartesian grid and represents a respective compressed neural network weight of a plurality of compressed neural network weights. The one or more processors are configured to execute, based on the plurality of compressed network weights, a neural network. The one or more processors are configured to generate, based on executing the neural network, an output.
Owner:QUALCOMM INC

Voice call real-time transcription system and method

The invention provides a voice call real-time transcription system and method, and relates to the technical field of computers, and the system comprises a network element module which is used for obtaining corresponding audio data when a call request of a user side is detected; the voice streaming engine is used for carrying out hierarchical compression on the audio data based on a preset perceptual weighted vector quantization algorithm to obtain audio compressed data, and carrying out format conversion processing on the audio compressed data to obtain temporary audio data; the voice engine is used for performing feature extraction on the temporary audio data to obtain multi-modal feature data, and processing the multi-modal feature data based on a preset voice recognition model to obtain text information; and the analysis and optimization module is used for obtaining corresponding real-time transliteration text data according to the text information and a preset vocabulary library based on a preset large model. According to the method, the voice information is comprehensively represented by using the multi-modal feature data, so that the voice recognition model can more accurately convert the voice into the text.
Owner:CHINA UNICOM WO MUSIC & CULTURE CO LTD

Codebook compression for vector quantized neural networks

Systems and techniques are described herein for quantizing a codebook used in the context of quantizing post-training parameters (e.g., vectors of weights) of a pre-trained model. For example, a device can perform rank reduction on a tensor of a codebook associated with parameters of a layer of a pre-trained machine learning model to generate a first tensor factor having a first shape and a second tensor factor having a second shape. The device can perform an optimization technique on the first tensor factor and the second tensor factor to minimize an output reconstruction error of the layer. The device can quantize the first tensor factor to generate a reduced size codebook.
Owner:QUALCOMM INC

Radar radiation source identification method

The invention relates to the technical field of communication, and discloses a radar radiation source identification method, which comprises the following steps: acquiring an original radiation source signal, and carrying out signal processing operation on the original radiation source to obtain a clear radiation source signal; performing vector quantization operation according to the clear radiation source signal to obtain a quantized radiation source signal; inputting the quantized radiation source signal into a pre-trained signal screening model to obtain a to-be-identified radiation source signal; according to the to-be-identified radiation source signal, performing a nearest neighbor node extraction operation to obtain a nearest neighbor node set; according to the nearest neighbor node set and a preset target node, performing calculation to obtain a signal distance; and according to the signal distance, carrying out maximum value extraction operation to obtain a maximum signal distance, and determining the radiation source corresponding to the maximum signal distance as a final radiation source. The method can improve the recognition precision of radar radiation source signals in a complex electronic countermeasure environment.
Owner:NANJING WEJOY TECH CO LTD

Compression of machine-learned models by vector quantization

A computing system can include one or more processors and one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the computing system to perform operations including obtaining model structure data indicative of a plurality of parameters of a machine-learned model; determining a codebook comprising a plurality of centroids, the plurality of centroids having a respective index of a plurality of indices indicative of an ordering of the codebook; determining a plurality of codes respective to the plurality of parameters, the plurality of codes respectively comprising a code index of the plurality of indices corresponding to a closest centroid of the plurality of centroids to a respective parameter of the plurality of parameters; and providing encoded data as an encoded representation of the plurality of parameters of the machine-learned model, the encoded data comprising the codebook and the plurality of codes.
Owner:AURORA OPERATIONS INC

Communication method and apparatus

A communication method and apparatus. A terminal device determines channel state information based on S indexes of S vectors. Each of the S vectors is included in a first vector quantization dictionary. The first vector quantization dictionary includes N1 vectors, where N1 and S are both positive integers. The terminal device sends the channel state information to a network device. Because a vector included in the vector quantization dictionary usually has a relatively large dimension, quantizing channel information by using the first vector quantization dictionary is equivalent to performing dimension expansion on the channel information or maintaining a relatively high dimension. High-precision feedback is implemented by using relatively low signaling overheads.
Owner:HUAWEI TECH CO LTD

Efficient post-training vector quantization for deep neural network weights

Systems and techniques are described for quantizing parameters (e.g., post-training vectors) associated with a pre-trained model. For example, a device can obtain a codebook for a group of weights of a pre-trained machine learning model. The device can determine a compression ratio based on the codebook and at least one of a vector quantization dimensionality, a group size, a codebook bit-width, or a scale group size. The device can quantize, via a vector quantization engine, the group of weights of the pre-trained machine learning model a plurality of columns at a time according to the compression ratio to generate a quantized pre-trained model.
Owner:QUALCOMM INC

Dense monitoring Internet of Things remote estimation method based on multi-cell semantic enhancement

The invention discloses a dense monitoring Internet of Things remote estimation method based on multi-cell semantic enhancement, and belongs to the field of wireless communication. The method comprises the following steps: firstly, an edge device observes a reasoning target, encodes an observation value according to a semantic codebook, and transmits the observation value to an edge server; secondly, the edge server carries out vector quantization on the received signal and sends a quantized code word index to a cloud server through a forward link; and finally, the cloud server performs dequantization and joint decoding on the received quantization index to obtain an estimated value of the reasoning target in the multiple cells, and a remote estimation task is completed. In addition, a loss function is constructed based on an information bottleneck theory and a straight-through estimation gradient approximation method, and the system is trained to be optimal by adopting a two-stage training strategy. According to the method, semantic information extraction of the remote estimation task in a multi-cell scene is realized, the performance of the remote estimation task can be improved while code word redundancy is inhibited, and the utilization efficiency of spectrum resources is remarkably improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Mutual alignment vector quantization

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating training an encoder neural network to generate discrete latent representations of data items by performing both a forward and a backward function during training.
Owner:GDM HOLDING LLC

Speech synthesis method and device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to the fields of financial science and technology and medical health, and discloses a speech synthesis method, device, equipment and medium, and the method comprises the steps: obtaining original speech, and extracting high-dimensional features from the original speech to obtain high-dimensional speech features; inputting the high-dimensional language features into a pre-trained vector quantizer for discretization to obtain a plurality of discrete Tokens; a prediction Token sequence is generated through a TTS generator according to text information corresponding to the original voice and the multiple discrete Token, and the TTS generator is obtained by adopting a sample set to train and verify a large language model; and inputting the predicted Token sequence into a voice decoder for voice synthesis to obtain target voice. And the quality and the accuracy of the synthesized speech are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Audio missing fragment repairing method

The invention provides an audio missing fragment repairing method, which comprises the following steps of S1, performing short-time Fourier transform on a damaged audio signal to obtain an STFT amplitude spectrum, and converting the STFT amplitude spectrum into a Mel scale through a frequency filter bank to obtain a Mel spectrum of the damaged signal; s2, using the trained vector quantization variational auto-encoder to encode the Mel spectrum of the damaged signal into a potential spatial feature with a relatively low dimension; s3, applying a diffusion model method, taking the damaged potential spatial features as conditions, splicing sampled Gaussian noise, performing T-step denoising through the trained denoising network, and outputting predicted complete potential spatial features; and S4, decoding the predicted complete potential features through an auto-encoder to obtain a predicted complete Mel spectrum, repairing a signal phase by using a sensing HiFiGAN vocoder, and outputting a repaired voice signal. According to the invention, natural and coherent restoration of long voice missing segments is realized, and the robustness of the method in a noise environment is improved.
Owner:DALIAN MARITIME UNIVERSITY

Contrastive image captioning neural networks with vector quantization

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training an image processing neural network to generate discrete representations of input images. The discrete representations can then be used for any of a variety of downstream tasks.
Owner:GOOGLE LLC