Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

185 results about "Vector quantization" patented technology

Vector quantization (VQ) is a classical quantization technique from signal processing that allows the modeling of probability density functions by the distribution of prototype vectors. It was originally used for data compression. It works by dividing a large set of points (vectors) into groups having approximately the same number of points closest to them. Each group is represented by its centroid point, as in k-means and some other clustering algorithms.

Mapping latent space of vector quantized variational autoencoders to functional basis vectors for enhanced data representation and manipulation

A method is provided for mapping the latent space of a Vector Quantized Variational AutoEncoder (VQ-VAE) to polynomial basis vectors. The method includes training a VQ-VAE model on a dataset to obtain a set of codebook vectors representing the latent space; defining a polynomial basis for the latent space, the polynomial basis containing terms up to a predetermined order; mapping each codebook vector to the polynomial basis by determining polynomial coefficients that represent each codebook vector in terms of the polynomial basis; and using the polynomial coefficients to reconstruct and manipulate latent space representations.
Owner:LEPTUDE INC

Method and system for generating 3D (three-dimensional) human motion under text driving by using 2D (two-dimensional) video

The invention discloses a method and a system for generating 3D (three-dimensional) human motion under text driving by utilizing a 2D (two-dimensional) video. The method comprises the following steps of: acquiring the video and preprocessing to obtain a two-dimensional key point sequence and text description; the two-dimensional key point sequence passes through a spatiotemporal feature adapter to obtain a potential spatiotemporal feature sequence, a residual vector quantizer quantizes and outputs a three-dimensional SMP L parameter sequence, and meanwhile, potential spatiotemporal features and a discrete Token sequence are mapped; preprocessing a text to extract a semantic vector, partially covering a Token sequence of a basic quantization layer, reconstructing a prediction sequence through a predictor in combination with the semantic vector, and obtaining a complete sequence through a refiner; constructing a total loss function and a text-to-action loss function to train the module; and inputting the text description and the basic quantization layer Token to a trained module, outputting a three-dimensional SMPL parameter sequence, and rendering to generate a three-dimensional human body grid and animation. According to the method, the end-to-end generation from the text to the three-dimensional SMPL action is realized only by two-dimensional key points and text description.
Owner:ZHEJIANG UNIV

Quantization Error Compensation for Vector Computing

A method for performing a computing task includes: extracting one or more features from a user content; converting the features to a floating point query vector; quantizing the floating point query vector; obtaining a database vector including one or more floating point feature vectors; determining a compensation vector based on a data distribution of the floating point query vector; quantizing the floating point feature vectors; determining an error function based on a difference between data distributions of i) the quantized query vector compensated with the compensation vector, and ii) the floating point query vector; determining, based on the error function, values of the compensation vector corresponding to the quantized feature vectors; combining the quantized query vectors and the values of the compensation vector to obtain one or more compensated query vectors; and performing the computing task using the compensated query vectors and the quantized feature vectors to obtain an output.
Owner:MACRONIX INTERNATIONAL CO LTD

Semantic information transmission method and device, equipment and storage medium

The invention relates to the technical field of wireless semantic communication, and discloses a semantic information transmission method, device and equipment and a storage medium, a sending end extracts features of to-be-transmitted data, and matches the features with vectors in a preset discrete vector codebook to obtain discrete feature vectors; and determining an orthogonal frequency division multiplexing time-frequency resource distribution grid according to the feature importance weights of the discrete feature vector and the initial feature vector, generating a sending signal, and transmitting the sending signal to a receiving end. And the receiving end inputs the semantic feature vector obtained after processing the received signal into the vector quantization variational self-decoder, and the semantic feature vector is matched with the vector in the discrete vector codebook again to correct the transmission error, thereby obtaining the reconstructed data to be transmitted. Discretization processing of the features is achieved through a discrete vector codebook, important semantic feature bits are quantized according to the importance weight of each feature and then distributed to orthogonal frequency division multiplexing for transmission, so that key semantic information is protected, and anti-interference transmission of the semantic information is achieved.
Owner:PENG CHENG LAB

Method and system for motion generation from input text

A method for training a model for generating a representation of long-term motion from a text input comprises: training a motion encoder of an autoencoder to compress and map an input motion into a latent representation comprising a sequence of latent vectors in a discrete latent space, each latent vector representing a fixed length of motion; training a quantization module to quantize the latent vectors to a sequence of quantized latent vectors in quantized latent space; and training a motion decoder to reconstruct the quantized sequence as a sequence of single-frame pose representations. A text encoder is trained to predict a latent sequence conditioned on a text input and a duration using the mapped latent representation as a target.
Owner:NAVER CORP

Speech reconstruction method and system based on entropy coding residual quantization and spectrum repair

ActiveCN121096348ASpeech recognitionFrequency spectrumSpeech reconstruction
The invention provides a voice reconstruction method and system based on entropy coding residual quantization and frequency spectrum restoration, and relates to the technical field of artificial intelligence voice signal processing, and the method comprises the steps: obtaining an original voice waveform, inputting the voice waveform into a neural voice coding and decoding model, firstly entering a coder to map the input voice waveform into acoustic potential representation, and then entering a frequency spectrum restoration model; performing residual quantization on the acoustic potential characterization layer by layer through a residual vector quantization module, introducing a gating-based dynamic layer number selection mechanism and entropy regularization constraint, enabling bits to be adaptively distributed among different voice segments, reconstructing reconstructed acoustic features of the potential characterization, inputting the reconstructed acoustic features into a decoder, restoring the reconstructed acoustic features into a time domain waveform, and outputting the time domain waveform. And mapping to a logarithmic magnitude spectrum domain through a spectrum repairing module, predicting a residual error in the logarithmic magnitude spectrum domain and performing confidence gating fusion to obtain a complex spectrum, and outputting after time domain synthesis to obtain reconstructed speech. According to the invention, high fidelity, intelligibility and transmission reliability of the voice can be considered at an extremely low bit rate.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

High-fidelity three-dimensional Gaussian sputtering lightweight method for resource-constrained equipment

PendingCN121095405A3D-image renderingColor-codingGaussian units
The invention discloses a high-fidelity three-dimensional Gaussian sputtering (3DGS) lightweight method for resource-constrained equipment. The method aims at solving the problems of high storage and computing resource consumption caused by the fact that a large number of parameters are stored in an existing 3DGS technology, and geometric distortion possibly occurring when details of a scene center are processed is overcome. The core of the method lies in a multi-stage progressive optimization framework, and the framework cooperatively applies four key technologies of Gaussian cutting and opacity regularization, dynamic spherical harmonic function adjustment, entropy constraint vector quantization and coordinate space shrinkage. Wherein in Gaussian clipping, redundant gauss are eliminated by quantifying the contribution degree of a Gaussian unit; the dynamic spherical harmonic function adjustment adaptively adjusts the order of color coding according to the scene complexity; the entropy constraint vector quantization is used for compressing a plurality of Gaussian attributes so as to realize more compact representation; and the coordinate space shrinkage is realized through nonlinear transformation, so that the rendering precision of details of the center of the scene is remarkably improved. According to the method, while the rendering precision and quality are kept, remarkable storage compression is realized, and the method is particularly suitable for deployment of resource-limited platforms such as mobile equipment.
Owner:HENAN UNIVERSITY OF TECHNOLOGY

Semantic-driven agent capability discovery method and device for agent internet

The invention discloses a semantic-driven intelligent agent capability discovery method and device oriented to the intelligent agent Internet, and the method comprises the steps: firstly generating an intelligent agent structured portrait which covers the three-dimensional information of skills, roles and states; semantic coding is carried out on the portrait, and the portrait is converted into a high-dimensional semantic vector; then, the vector is partitioned, a codebook is generated in each subspace in a clustering mode, and the codebook comprises representative vectors and serial numbers of the representative vectors; and quantizing the sub-vectors into numbers based on a codebook, and splicing the numbers to form discrete identification codes of the agent index. And when the index is updated, incremental maintenance is executed according to the distance between the new agent sub-vector and the vector in the codebook. Meanwhile, a generative retrieval model is trained, task query is directly mapped into discrete identification codes, historical and new task samples are mixed for continuous learning in training, and stability constraints are introduced. According to the method, through semantic portraits, a quantitative index mechanism and memory enhancement continuous learning, an end-to-end retrieval and rapid updating capability discovery scheme is constructed.
Owner:XI AN JIAOTONG UNIV

Multi-modal driven human body action generation method based on large language model

A multi-modal human body action generation method based on a large language model comprises the steps that structural features are extracted based on 3D action data, body part-level atomic semantic description is generated in combination with the large language model, and a multi-modal alignment data set containing texts, voices and music is constructed; whole-body actions are decoupled according to parts, an independent vector quantization encoder is adopted to perform residual quantization, and an atomic action token strongly associated with a fine-grained text is generated; splicing the text description and the action token into a mixed action sentence containing a special mark according to a body structure; multi-modal input joint modeling is realized through a large language model, a fine-grained text and an action token sequence are synchronously generated, 3D actions conforming to semantics are output through decoding, and zero sample generation and part-level accurate control are supported. According to the method, the limitation of coarse-grained alignment of a traditional method is broken through, and the semantic consistency, multi-modal adaptability and local controllability of action generation are remarkably improved through fine-grained semantic mapping, decoupling action coding and unified sequence modeling.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Semi-supervised signal modulation identification method based on multi-codebook discrete virtual adversarial training

The invention provides a semi-supervised signal modulation identification method based on multi-codebook discrete virtual adversarial training, and the method comprises the steps: carrying out the high-dimensional potential representation coding of input labeled and unlabeled sample signals through employing a vector quantization generative adversarial network, so as to obtain a continuous potential representation; mapping the continuous potential representation into a discrete vector sequence through a quantizer; virtual confrontation disturbance meeting the condition that the L2 norm does not exceed epsilon is applied to the unlabeled sample at the input end, and a disturbance sample is generated; respectively inputting a labeled sample, a label-free sample and a disturbance sample into a classifier, calculating cross entropy loss and KL divergence consistency loss, and jointly optimizing parameters of an encoder, a codebook and the classifier in a weighted sum form; therefore, the discrete potential representation is obtained through the vector quantization generative adversarial network, disturbance is constructed for the unlabeled sample at the input end, and joint optimization is performed in combination with the cross entropy and the KL divergence loss, so that the recognition accuracy and robustness under the conditions of few labels and low signal-to-noise ratio are improved.
Owner:XIAMEN UNIV

CSI feedback method based on Transform and entropy constraint vector quantization

The embodiment of the invention provides a CSI (Channel State Information) feedback method based on Transform and entropy constraint vector quantization. The method is applied to the technical field of wireless communication. The method comprises the following steps: receiving an input angle-time delay domain channel matrix, and performing feature extraction on the channel matrix through a convolutional layer to obtain a first feature tensor; the first feature tensor is input into a trained Transform encoder for analysis processing, and a latent vector is obtained; performing quantization processing on the latent vector through the trained vector quantization variational auto-encoder to obtain discretization representation of the latent vector; performing dequantization processing on the discretized representation of the latent vector to obtain a dequantized latent vector; performing feature extraction on the dequantized latent vector through a convolutional layer to generate an initial channel feature; the initial channel characteristics are input into a trained Transform decoder for analysis processing, and final channel characteristics are obtained; and the final channel feature is mapped to the target resolution to obtain a reconstructed channel matrix, so that the reconstruction precision of the channel matrix is improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Zero-sample semantic segmentation map construction method combining road visual language model and dynamic object

The invention discloses a zero-sample semantic segmentation map construction method combining a road visual language model and a dynamic object, and belongs to the technical field of zero-sample semantic segmentation map construction methods.According to the method, feature recovery is carried out on a degraded look-around image in severe weather through a vector quantized image restoration model, and the feature of the degraded look-around image is restored; therefore, the image features under the normal weather condition are recovered. Then, analyzing the restored image by using a road visual language model, generating semantic description, and extracting information of key elements on the road;
Owner:BEIHANG UNIV

Hidden space confrontation sample generation method and system based on multi-scale feature separation

The invention discloses a hidden space adversarial sample generation method and system based on multi-scale feature separation, and the method comprises the steps: employing a neural network quantization training method based on straight-through estimation, and training a hierarchical vector quantization variational auto-encoder; carrying out differentiable Haar wavelet transformation on the input image by adopting a wavelet packet transformation algorithm, decomposing the input image into a low-frequency component and a high-frequency component, and realizing multi-scale feature separation; inputting the high-frequency component into a hierarchical vector quantization variational auto-encoder, and extracting and quantizing global high-frequency features and local high-frequency detail features; in the potential space, a learnable disturbance variable is introduced, a potential vector after disturbance is constructed, and the potential vector is reconstructed into an adversarial sample through a decoder; and based on a preset disturbance target, carrying out iterative optimization on the disturbance vector until a confrontation sample which satisfies an attack success condition and is optimized in visual quality is generated. According to the method, a wavelet domain variational auto-encoder and a hidden space iterative attack algorithm are fused, and an adversarial sample with high fidelity and clear interpretation is generated.
Owner:XINJIANG UNIVERSITY

Transferable vector quantization alignment method and device based on unsupervised domain adaptation

The invention relates to a transferable vector quantization alignment method and device based on unsupervised domain adaptation, and the method comprises the steps: extracting the features of a source domain and a target domain of a source domain data set and a target domain data set through a feature extraction unit, and calculating an overall feature distribution loss function through searching a feature item closest to the features in a codebook; calculating a local alignment loss function through a bottleneck layer, performing classification through a classifier to obtain a source domain pseudo-label and a target domain pseudo-label, calculating a cross entropy classification loss function according to the source domain pseudo-label and a corresponding truth value label, performing normalization processing on the target domain pseudo-label, introducing mutual information to obtain a sample weight in each target domain data set, and obtaining a sample weight in each target domain data set; and calculating mutual information weighted maximization confusion matrix loss functions, and updating parameters in each module by using the loss functions until convergence to obtain a target classification architecture with cross-domain extraction, feature alignment and time sequence signal classification capabilities. By adopting the method, the identification performance of the label-free target domain can be improved.
Owner:NAT UNIV OF DEFENSE TECH

Aircraft control audio coding method and system based on dynamic acoustic masking

ActiveCN120636422ASpeech recognitionNoiseIntelligibility (communication)
The invention belongs to the technical field of audio coding, and particularly discloses an air traffic control audio coding method and system based on dynamic acoustic masking, and the method comprises the steps: collecting an air traffic control audio signal, and carrying out the complex short-time Fourier transform to generate an audio time-frequency diagram; inputting the time-frequency graph into a personalized auditory feature extraction model, extracting a physiological feature graph and a dynamic weight matrix, and fusing the physiological feature graph and the dynamic weight matrix to generate a final masking matrix; performing Hadamard product on the masking matrix and the time-frequency map to obtain a perception saliency map; extracting audio features based on the perceptual saliency map, constructing a basic codebook and a tree residual codebook to perform hierarchical vector quantization compression, and outputting compression representation; and converting the compressed representation into a coding triple consisting of a basic index, a path index and a termination flag bit, and taking the coding triple as a final coding result. The method adapts to individual hearing differences, improves instruction intelligibility and compression efficiency, and is suitable for high-noise aviation communication scenes.
Owner:NAVAL AVIATION UNIV

EEG characterization method based on wavelet neural quantization training and semantic alignment

The invention provides an EEG characterization method based on wavelet neural quantization training and semantic alignment, and the training method comprises the steps: S1, obtaining training data which comprises a plurality of EEG signal samples, and each sample comprises a plurality of EEG sub-signals; s2, using the training data to train a wavelet quantization nerve marker and a nerve codebook thereof: using a vector quantization encoder of the wavelet quantization nerve marker to carry out discrete wavelet transform on input samples, generating a feature block obtained by discrete wavelet transform and corresponding to each sub-signal of each sample, and extracting a potential representation of each feature block; acquiring a closest discrete vector matched from a plurality of discrete vectors in the neural codebook according to the potential representation, and reconstructing the EEG sub-signal corresponding to the feature block according to the discrete vector matched with each feature block by using a neural decoder based on inverse wavelet transform; and training and updating the parameters of the nerve marker and discrete vectors of the nerve codebook according to the first total loss function.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A method, device, storage medium and terminal for generating text to image

The present invention discloses a method, device, storage medium and terminal for generating text to image. The method includes: obtaining a text description, tokenizing the text description and generating a text shape sequence; generating at least one first image based on the text shape sequence, a pre-trained image generation model and a vector quantization autoencoder; inputting each first image into a pre-trained scoring model to obtain a probability value for each first image; based on the probability value of each first image, screening first images with a probability value greater than a preset threshold to generate at least one second image; and increasing the resolution of the second image based on a pre-trained resolution enhancement model to generate a target image. Therefore, by adopting the embodiments of the present application, it is possible to ensure that the content of the generated image is consistent with the semantics of the descriptive text, greatly reducing the error between the two, and effectively improving the resolution of the generated image.
Owner:BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

Multi-scale space-time autoregression moving track generation method and device and storage medium

The invention discloses a multi-scale space-time autoregression movement track generation method, which comprises the following steps of: sequentially carrying out spatial scale representation and time scale representation on an original movement track to obtain a plurality of space-time representation sequences; performing residual vector quantization processing on each time-space representation sequence to obtain a discretized quantization index tag sequence corresponding to each time-space representation sequence; performing scale-by-scale autoregression generation on each quantization index tag sequence to obtain a discretized prediction quantization index tag sequence corresponding to each spatio-temporal representation sequence; and restoring the plurality of predictive quantization index tag sequences into a complete moving track. According to the embodiment of the invention, the moving track generation method has obvious advantages in the aspects of generation efficiency, track structure consistency, multi-scale modeling capability, individual behavior expression and the like.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Adaptive semantic joint source-channel coding method, system, electronic device and storage medium

This application provides an adaptive semantic joint source-channel coding method, system, electronic device, and storage medium. The method includes: Step S1: acquiring raw input data and real-time channel state information, wherein the real-time channel state information includes at least the signal-to-noise ratio (SNR); Step S2: using a lightweight semantic coding module to extract multi-scale semantic features from the raw input data, and incorporating the SNR embedding vector in a feature modulation manner during the coding process to obtain channel-adaptive semantic features; Step S3: using an SNR embedding and channel-adaptive attention module to adjust the attention weights of the semantic features to obtain attention features; Step S4: using a dynamic codebook generation and vector quantization module to convert the attention features into a discrete codebook index sequence; Step S5: using a joint source-channel coding module to modulate and map the codebook index sequence into complex channel symbols and transmit them.
Owner:KAIFENG UNIV

Electric power vision large model multi-scale semi-supervised target detection method and system

The invention discloses an electric power vision large model multi-scale semi-supervised target detection method and system, and the method comprises the steps: training a multi-scale vision word segmentation device based on a vector quantization knowledge distillation framework, and constructing an electric power semantic enhanced multi-scale vision codebook, the multi-scale visual codebook comprises a top-layer codebook used for encoding a global structure of the power equipment and a bottom-layer codebook used for encoding defect local features of the equipment; based on the multi-scale visual codebook, adopting a mask image modeling task to pre-train a visual large model on an unmarked power image; performing semi-supervised fine tuning on the visual large model by using the labeled electric power image data and the unlabeled electric power image data; using an SMLS fine tuning strategy to update the visual large model parameters; and inputting a to-be-detected power image into the updated visual large model, and outputting a detection result of the multi-scale target. According to the method, the problems of scarcity of annotation data, model semantic gaps and high calculation complexity in an electric power scene in the prior art are solved.
Owner:STATE GRID ELECTRIC POWER RES INST +1

Encoding device, decoding device, encoding method, and decoding method

An encoding device comprising: a quantization circuit that generates a quantization parameter that includes information about a vector quantization codebook; and a control circuit that sets the number of available bits according to conditions for encoding based on the difference between the number of bits available for encoding of the target sub-vector and the number of bits for the quantization parameter of the target sub-vector.
Owner:PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA

Collaborative distillation and quantitative compression method oriented to image semantic representation

The invention belongs to the technical field of computer vision, and particularly relates to an image semantic representation-oriented collaborative distillation and quantitative compression method and an innovative joint training framework, which deeply couples two processes of vector quantization and knowledge distillation, realizes collaborative optimization through bidirectional information interaction, and can realize high compression ratio and high compression efficiency at the same time. And the quality of semantic representation output by the model is kept to the maximum extent.
Owner:山东齐鲁壹点传媒有限公司 +1

Vector quantization auto-encoder learning method based on Kepler codebook theory

The invention discloses a Kepler codebook theory-based vector quantization auto-encoder learning method, which comprises the following steps of: firstly, constructing vector quantization characteristic representation based on a Kepler codebook theory to obtain Kepler codebook distribution; then the distribution is used as a regularization constraint of a quantization hidden space, the constraint precision is improved by combining a codebook partitioning technology, and the learning process of the vector quantization auto-encoder is optimized; meanwhile, Kepler codebook distribution is applied to regularization constraint of the variational hidden space, and learning of the variational auto-encoder is optimized. The method is characterized in that each codebook vector can be fully trained and utilized through the Kepler codebook theory, and the codebook vector is applied to a quantization auto-encoder and a variational auto-encoder as a hidden space regularization constraint, so that the problem of codebook collapse (low utilization rate) is effectively solved; and the modeling capability of an image feature extraction module on complex textures and detail information is remarkably improved. According to the method, a new theoretical framework and a technical path are provided for self-encoder learning.
Owner:SUN YAT SEN UNIV

Encoding device, decoding device, encoding method and decoding method

This invention reduces the number of encoding bits in vector quantization. The encoding apparatus includes: a quantization circuit that generates quantization parameters, the quantization parameters including first information related to the codebook of the vector quantization and second information related to the code vectors contained in the codebook; and a control circuit that uses a second number of bits to control the encoding of the first information of a sub-vector, the second number of bits being the difference between the first number of bits available for encoding the sub-vector in vector quantization and the number of bits of the quantization parameters of the sub-vector.
Owner:PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA

Vector quantization-based encoding cache methods, systems, electronic devices, and media

This application provides a method, system, electronic device, and medium for caching codebooks based on vector quantization. The method includes: acquiring a codebook to be processed; sorting the codebook to be processed to obtain a target codebook index; obtaining an index boundary based on the target codebook index and a codebook cache; the codebook cache being a cache for storing the codebook to be processed; comparing the index boundary and the target codebook index to obtain a judgment result; and storing the codebook to be processed in the corresponding cache within the codebook cache based on the judgment result. This application improves memory performance and execution efficiency by placing codebook entries at different locations in the GPU's memory hierarchy based on the codebook's usage frequency, thus solving the problems of low efficiency in shared memory and global memory.
Owner:SHANGHAI JIAOTONG UNIV +1

Vehicle-mounted audio equalization method and system based on residual vector quantization

The invention discloses a vehicle-mounted audio equalization method and system based on residual vector quantization, and relates to the technical field of vehicle-mounted electronic equipment.The method comprises the steps that a white noise audio signal is played through a vehicle-mounted loudspeaker, a microphone is used for collecting the audio signal lost through in-vehicle propagation, and time-frequency characteristics are obtained through digital preprocessing; performing discretization processing on the target IIR filter coefficient through a residual vector quantization encoder to generate a plurality of layers of token sequences; inputting the time-frequency feature matrix into a convolutional neural network for feature extraction to obtain a low-dimensional feature vector, and through a multi-layer perceptron prediction head, respectively obtaining classification probability distribution of each layer of token and predicting a token sequence; and inputting the predicted token sequence into a residual vector quantization decoder, reconstructing the token sequence into an IIR filter coefficient, loading the IIR filter coefficient into a graphic equalizer, and performing frequency response compensation on an input audio signal. According to the method, the audio frequency response curve can be quickly and adaptively corrected, and the in-vehicle tone quality and the user hearing comfort are remarkably improved.
Owner:GUANGZHOU CHEERILEE ELECTRONIC TECH CO LTD

Low-power-consumption and low-resource-consumption super-dimensional computing system oriented to edge computing equipment

The invention discloses an edge computing device-oriented super-dimensional computing system with low power consumption and low resource consumption, which is characterized in that a preprocessing unit adopts a binary convolution kernel to carry out convolution operation on input image data and binarizes a convolution result to obtain binary data; the super-vector coding unit maps binary data corresponding to the training data and the reasoning data into a training super-vector and a reasoning super-vector; the training unit generates n class super-vectors according to label clustering of the training super-vectors; a dynamic retraining unit compares the similarity of the training super-vector and the class super-vector to obtain a most similar class label, when the most similar class label is inconsistent with a real label, a retraining updating process is started, and after retraining is completed, n class super-vectors are quantized and then input into a reasoning classification unit; and the reasoning and classifying unit adopts matrix multiplication to compare the similarity of the reasoning super-vector and the quantized class super-vector, and outputs a classifying result. The precision can be improved, the power consumption is reduced, and the hardware overhead is reduced.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Method and apparatus for configurable CSI feedback including vector quantization

Aspects of the present application provide apparatus, devices, and methods for a channel state information (CSI) feedback scheme that can accommodate various scenarios to report channel information to a base station or a gNB. A method includes a receiver receiving CSI configuration information, where the CSI configuration information includes vector quantization configuration information. A receiver receives a reference signal from a transmitter, where the reference signal is used to determine the CSI of a channel in which the reference signal is received. The receiver measures the received reference signal. The receiver determines a CSI parameter of the measured reference signal by performing vector quantization based on the vector quantization configuration information. After determining the CSI parameter, the receiver transmits the CSI parameter to the transmitter. In some embodiments, the receiver may be a user equipment (UE), and the transmitter may be a base station.
Owner:HUAWEI TECH CO LTD

Infrared target identification multi-dimensional complexity characterization method in complex ground scene

The invention discloses a multi-dimensional complexity characterization method for infrared target identification in a complex ground scene, which belongs to the technical field of infrared target identification and tracking, and comprises the following steps: comparing the difference between the current feature of a target and the historical reference feature of the target, and quantifying feature degradation; carrying out feature vector extraction by adopting a module of a Zernike moment; target motion feature vectors are extracted, and target motion complexity is calculated; extracting a target historical reference feature and a target local background feature, quantifying a difference between the features, and calculating a similarity representation local background suspected degree; extracting feature vectors and motion vectors of the real target and the false motion target, and quantifying feature differences and inter-frame motion differences to represent the interference degree of the false motion target; and the target feature degradation degree and the like are fused into comprehensive identification tracking complexity. According to the method, the tracking difficulty of the current frame or sequence can be reflected, and a theoretical basis and decision support are provided for online switching of algorithm strategies, dynamic allocation of system resources and scientific evaluation of cross-scene algorithm performance.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

High-fidelity low-illumination image enhancement method and system based on one-step diffusion

The invention provides a high-fidelity low-illumination image enhancement method and system based on one-step diffusion, and the method comprises the steps: training a codebook based on a vector quantization generative adversarial network through a normal illumination image, obtaining a normal illumination codebook, and obtaining a high-fidelity low-illumination image through a latent variable refinement and alignment model; refining and aligning conditional latent variables of low illumination acquired by a VAE encoder of a pre-trained diffusion model to obtain refined conditional latent variables and aligned conditional latent variables, and quantifying the refined conditional latent variables and the aligned conditional latent variables to obtain quantized latent variables; taking the refined conditional latent variable as a diffusion starting point through a denoising network model based on one-step diffusion, taking the quantized latent variable as a vector for replacing text embedding, and obtaining an enhanced latent variable corresponding to the normal illumination image; and decoding to obtain a final enhancement result by using wavelet-based jump connection. According to the method, the generalization of complex and diversified real images is ensured while the low-illumination image is enhanced in a high-fidelity manner.
Owner:SICHUAN RES INST OF SHANGHAI JIAOTONG UNIV