Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

37 results about "Vector quantization" patented technology

Vector quantization (VQ) is a classical quantization technique from signal processing that allows the modeling of probability density functions by the distribution of prototype vectors. It was originally used for data compression. It works by dividing a large set of points (vectors) into groups having approximately the same number of points closest to them. Each group is represented by its centroid point, as in k-means and some other clustering algorithms.

Adaptive semantic joint source-channel coding method, system, electronic device and storage medium

This application provides an adaptive semantic joint source-channel coding method, system, electronic device, and storage medium. The method includes: Step S1: acquiring raw input data and real-time channel state information, wherein the real-time channel state information includes at least the signal-to-noise ratio (SNR); Step S2: using a lightweight semantic coding module to extract multi-scale semantic features from the raw input data, and incorporating the SNR embedding vector in a feature modulation manner during the coding process to obtain channel-adaptive semantic features; Step S3: using an SNR embedding and channel-adaptive attention module to adjust the attention weights of the semantic features to obtain attention features; Step S4: using a dynamic codebook generation and vector quantization module to convert the attention features into a discrete codebook index sequence; Step S5: using a joint source-channel coding module to modulate and map the codebook index sequence into complex channel symbols and transmit them.
Owner:KAIFENG UNIV

A skin care efficacy vector quantization method and system fusing multi-dimensional constraints and semantic intent

This invention relates to the field of natural language processing, and in particular provides a method and system for quantifying skincare efficacy vectors by integrating multidimensional constraints and semantic intent. The method includes receiving four-dimensional heterogeneous semantic input from users; constructing a multidimensional knowledge base, which includes a semantic weight association library, an intent-vector mapping library, and a constraint-correction matrix library; calculating a dynamic priority scalar for each skin problem by combining the influence of skincare motivation and skincare goals on various major skin problems; calling the intent-vector mapping library to obtain the primitive efficacy vectors corresponding to each skin problem, weighting the primitive efficacy vectors using the dynamic priority scalar, and aggregating them using a max-pooling algorithm to generate a baseline demand vector; calling the constraint-correction matrix library to obtain the corresponding correction matrix according to the constraints, performing nonlinear correction processing on the baseline demand vector, and outputting the final target efficacy representation vector. This invention achieves accurate, safe, and personalized quantification of skincare efficacy.
Owner:SHANGHAI CHAOGUI BIOTECHNOLOGY DEVELOPMENT CO LTD

An implementation method and processing terminal of audio compression bit allocation and quantization

PendingCN122337218ABit allocationAudio frequency
This invention discloses a method for implementing bit allocation and quantization in audio compression, comprising the following steps: Step 1: Obtaining the audio signal, the perceptual importance of the audio signal, and the preset total target number of bits; Step 2: Based on the bit allocation vector corresponding to the minimum value of the calculated objective function, allocating bits to each frequency band of the audio signal, the sum of the bits allocated to each frequency band being the total target number of bits; Step 3: Feeding the audio signal and the bit allocation results for each frequency band into a residual vector quantization network for actual feature discretization, obtaining the quantized residual, thereby completing the quantization. This invention achieves adaptive bit allocation and quantization, thereby ensuring perceptually optimized encoding of audio and avoiding excessive differences in bit allocation between adjacent frequency bands that could lead to inter-band discontinuity artifacts in the reconstructed audio.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Continuous dynamic emotion prediction method, training method and device

ActiveCN122087622BPredictive methodsMood
The application provides a continuous dynamic emotion prediction method, a training method and equipment. A vector quantization variational self-encoding and decoding module learns a space-time codebook, better learns the causality and space-time semantics of dynamic emotions from electroencephalogram data, adopts a Transformer for long time series modeling, better learns time series dependence and context relationships, enables the model to learn a dynamic emotion latent space with stronger generalization and better capture of dynamic trends, further optimizes the time series prediction regression ability of the model through a reward mechanism and reinforcement learning, and the model can not only dynamically adapt to the dynamics of human emotions and accurately predict the current emotions of subjects, and the accuracy and effectiveness are guaranteed.
Owner:SHENZHEN UNIV

A context-based audio adaptive entropy encoding method and processing terminal

PendingCN122337214ASide informationTemporal context
This invention discloses a context-based adaptive entropy coding method for audio, comprising the following steps: First, the quantized codebook index sequence Q and quantization side information S of the audio signal are sequentially fed into a causal convolutional context model, a cross-band attention network, and a cascaded codebook dependency network for processing, to obtain temporal context feature vectors, frequency context feature vectors, and quantization-level context feature vectors for the audio signal in the time domain, respectively. Second, the temporal context feature vectors, frequency context feature vectors, and quantization-level feature vectors are input into a neural probability prediction model for processing, estimating the conditional probability distribution of each codebook index. Third, the predicted conditional probability distributions are input into an adaptive arithmetic encoder to generate the encoded bitstream, and necessary header information is added to the encoded bitstream to form a complete bitstream. This invention can further improve the compression efficiency of audio coding.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

A speech synthesis method based on an implicit continuous consistency model

PendingCN122313940AData setText annotation
This invention relates to the field of speech synthesis technology, specifically to a speech synthesis method based on an implicit continuous consistency model. The method involves constructing a dataset containing audio and its text annotations; building a residual vector quantization variational autoencoder, training it using a joint loss until convergence, and extracting the mean and variance of all audio data mapped to a latent vector distribution; sampling Gaussian noise latent variables based on the mean and variance, sampling Gaussian noise and time steps, and adding noise to obtain speaker features; calculating the continuous consistency loss using the time steps, speaker features, text, audio data, and the consistency model; and optimizing the continuous consistency loss until the consistency model converges. This method utilizes residual vector quantization technology to achieve high-rate audio feature compression and decoupling, and combines the single-step sampling characteristics of the consistency model in the latent space to improve the inference efficiency and training stability of speech synthesis, enabling the model to generate high-fidelity speech in a very small number of iterations.
Owner:HARBIN INST OF TECH AT WEIHAI +1

An agent full-link semantic-behavior collaborative control method

The present application relates to the technical field of rail transit intelligent agent, and discloses an intelligent agent full-link semantic-behavior collaborative control method, comprising: high-frequency capture of semantic elements of track signal equipment, when the newly added semantic type meets the trigger threshold, incremental update is carried out, and a dynamic evolution graph containing semantic nodes and device nodes is constructed; the history state sequence of each semantic node evolving with time is extracted, forming a semantic flow; a time-space model of the semantic flow is established by using a time sequence graph neural network TGNN, and finally a semantic flow state vector containing time-space information is generated; a semantic-behavior double-flow embedding space is constructed, a dynamic mapping relationship between the semantic flow state vector and the behavior primitive vector is established, and the semantic flow state vector is quantized into a specific robot behavior primitive vector; the actual behavior execution value of the robot is collected in real time, the deviation of the actual behavior execution value from the semantic constraint threshold is calculated, when the deviation exceeds the preset threshold, the physical deviation is traced back to the semantic element level through a reverse mapping mechanism for correction, and the full-link semantic coverage and the dynamic adaptation capability are improved.
Owner:XI AN JIAOTONG UNIV +1

Time-Frequency Dual-Mode Alignment Method and Apparatus for Multivariable Time-Series Signals

This invention relates to the field of signal processing technology, and provides a time-frequency dual-modal alignment method and apparatus for multivariable time-series signals. It utilizes a cross-scale time-series feature extraction network to extract discriminative features containing both time and frequency information from the multivariable time-series signal using cross-variable attention and cross-time attention. Then, it employs a codebook sharing time-frequency feature prototypes across domains for dual-modal vector quantization alignment based on spectral consistency, thereby eliminating domain offset through explicit frequency domain constraints. Furthermore, it uses an adaptive pseudo-label optimization strategy based on channel mutual information to suppress noise channel interference and improve the confidence of the target domain pseudo-label. Finally, under an end-to-end unified framework, it jointly optimizes the global alignment loss, local alignment loss, mutual information weighted maximization of the confusion matrix loss, and source domain cross-entropy loss to determine the target domain pseudo-label, deeply adapting to the physical characteristics of the time-series signal and simultaneously solving the problems of multivariable coupling, long-term time dependence, and time-frequency feature collaborative alignment.
Owner:NAT UNIV OF DEFENSE TECH

A Controllable Generative Enhancement Method and System for Breath Sound Categories Based on Discrete Tokens and Transformers

This invention provides a controllable generative enhancement method for respiratory sound categories based on discrete tokens and Transformer, belonging to the interdisciplinary field of medical signal processing and artificial intelligence. The method employs a two-stage framework: In the first stage, a conditional vector quantization autoencoder is trained, injecting pathological category labels during encoding and introducing multi-scale category prototypes based on statistical analysis of similar samples during decoding and reconstruction, thereby constructing a discrete latent space with explicit semantic structure. In the second stage, an autoregressive Transformer model is trained to learn the category conditional distribution of token sequences within this space. In the enhancement stage, for the target minority categories, the Transformer generates new token sequences, which are then fused with the corresponding category's multi-scale prototypes and broadcast decoded to reconstruct high-fidelity, pathologically semantically consistent synthetic respiratory sound samples. Finally, the synthetic samples are added to the training set to balance the data distribution, effectively improving the downstream respiratory sound classification model, especially its performance and overall robustness in identifying a minority of abnormal categories.
Owner:SHANGHAI UNIV

Text-guided scene generation method and device based on scene semantic graph

The application provides a text-guided scene generation method and device based on a scene semantic graph, and is applied to the technical field of text-guided scene generation. The method comprises the following steps: based on a received scene description text, a trained target scene semantic graph construction encoder is called to construct a scene semantic graph corresponding to the scene description text, and the scene semantic graph construction encoder comprises a discrete diffusion model; based on the scene semantic graph, a trained target vector quantization variational autoencoder and a trained target scene layout decoder are called to generate a target scene corresponding to the scene description text, and the scene layout decoder comprises a continuous diffusion model. The application converts the text into an intermediate representation to explicitly express object relationships and global layout structures, the continuous diffusion model performs noise adding and noise removing in a semantic graph latent space, learns a reasonable distribution, guarantees the performance in semantic consistency, spatial reasonableness and functional constraints, and generates a target scene in combination with the semantic graph, so that the practicability and reasonableness of the generated result can be improved.
Owner:TSINGHUA UNIVERSITY +1

Postoperative cognitive quantitative evaluation system and method based on fusion of electroencephalogram and near-infrared spectrum

PendingCN122320488AImprove the problem of easy missed diagnosis of high-risk risksImprove problems that are easily missedDynamic monitoringCognitive status
This invention relates to the field of perioperative neurological monitoring technology, and particularly to a postoperative cognitive quantitative assessment system and method that integrates electroencephalography (EEG) and near-infrared spectroscopy. The system includes: a baseline anchoring module, which extracts preoperative scalp EEG and near-infrared spectral signal features, concatenates them into a first vector, and calculates the covariance to generate a preoperative baseline manifold; a matrix construction module, which generates a cross-modal manifold matrix based on real-time bimodal signals within a time window; a feature extraction module, which calculates the geodesic distance between the cross-modal manifold matrix and the preoperative baseline manifold, projects it onto the tangent space using a logarithmic mapping to generate a tangent matrix, and combines the geodesic distance to generate a topological feature vector; a quantitative assessment module, which inputs the vector into a softmax function layer to output a probability distribution and calculates a cognitive assessment index; and a closed-loop intervention module, which outputs an intervention command when the index meets preset conditions and provides a timestamp to reset the time window. This system achieves closed-loop manifold-based assessment and dynamic monitoring of postoperative cognitive status.
Owner:ZHANJIANG CENT PEOPLES HOSPITAL

Hybrid storage progressive precision writing and retrieval method and apparatus

The application discloses a hybrid storage progressive precision writing and retrieval method and device, wherein the method comprises the following steps: quantizing a current real value vector of a target online model to a target bit width, determining an initial base value vector corresponding to the current real value vector of the target bit width, and writing the initial base value vector into a preset resistive random access memory; based on the current real value vector of the target bit width and the initial base value vector, calculating an initial residual error vector corresponding thereto, and performing a preset residual error quantization encoding operation on the initial residual error vector to obtain a residual error code word; performing a preset bit plane decomposition and block organization processing operation on the residual error code word to obtain a plurality of residual error bit vector blocks, and continuously storing the residual error bit vector blocks into a dynamic random access memory; and based on the initial base value vector and the initial residual error vector, performing a preset dynamic precision retrieval operation on a current query vector to obtain a corresponding retrieval result, so that the dynamic balance of retrieval performance and precision is realized.
Owner:TSINGHUA UNIVERSITY

A vitiligo auxiliary diagnosis method and device for multi-modal medical image collaborative segmentation and classification and a storage medium

PendingCN122265236AImage enhancementImage analysisDisease activityImage pair
The embodiment of the application discloses a kind of vitiligo auxiliary diagnosis methods, devices and storage medium of multi-modal medical image collaborative segmentation and classification, wherein the method comprises: obtaining the multi-modal image pair of the clinical image and Wood lamp image of the same examinee;According to the imaging characteristics of two kinds of modalities, the image pair is preprocessed and modality-specific data enhancement;The multi-modal image pair after processing is input into the feature extraction network to obtain each modality feature, and spatial guidance information for subsequent segmentation is generated;Each modality feature is input into vector quantization fusion module for cross-modal feature fusion to obtain semantic consistent fusion feature;Based on the fusion feature, a segmentation branch and a classification branch are constructed to realize the joint output of lesion region segmentation and disease activity classification, and the collaborative effect of the two tasks is improved through inter-task interaction. By using the present application, end-to-end joint optimization of vitiligo lesion segmentation and disease activity classification can be realized, and the diagnostic performance is improved.
Owner:GUANGZHOU UNIVERSITY

Method and apparatus for constructing text-to-speech model, electronic device, readable medium, and program product

PCT designated stageWO2026138018A1Speech inputAcoustics
A method and apparatus for constructing a text-to-speech model, an electronic device, a readable medium, and a program product. The method comprises: inputting preset training speech into a preset vector quantizer, so as to obtain a training semantic dispersion feature of the training speech, the training semantic dispersion feature comprising a language style of the training speech (101); acquiring training text corresponding to the training speech, and using the training text and the training semantic dispersion feature to train a preset autoregressive speech model, so as to obtain a semantic dispersion feature generation model (102); acquiring a training Mel-frequency spectrogram corresponding to the training semantic dispersion feature (103); using the training semantic dispersion feature and the training Mel-frequency spectrogram to train a preset optimal transport conditional flow matching model, so as to obtain a Mel-frequency spectrogram generation model (104); and on the basis of the Mel-frequency spectrogram generation model and the semantic dispersion feature generation model, constructing a text-to-speech model (105).
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Adaptive working condition aware fuel cell hybrid tram layered management method

ActiveCN120422725BData setEngineering
The application discloses a fuel cell hybrid railcar layered energy management method based on adaptive working condition sensing; in the identification layer, a sliding window mechanism is used to extract load working condition time domain and frequency domain features, feature data is clustered based on a spectral clustering algorithm driven by a deep auto-encoder, a data set with a category label is obtained, and a deep dynamic learning vector quantization neural network classifier is trained; in the strategy layer, a double-delay deep deterministic policy gradient reinforcement learning algorithm is used to construct a reward function, a lithium battery SOC fluctuation penalty term limit parameter in the reward function is adaptively adjusted according to a real-time load working condition category output by the identification layer, and an optimal power distribution scheme between multiple fuel cell power generation systems and lithium batteries is obtained by training a reinforcement learning intelligent agent; according to the performance degradation degree of different fuel cell stacks, a distributed collaborative control strategy considering performance differences is used to distribute the output power of each stack, so that the coordinated control of the operating states of the multiple fuel cell power generation systems is realized.
Owner:SOUTHWEST JIAOTONG UNIV +1

Intelligent model adaptation method and device based on cloud-edge collaboration

This invention provides a cloud-edge collaborative intelligent model adaptation method and apparatus. The method, applied to a cloud server, includes: receiving a quantization vector sent by an edge device; determining the quantization vector based on target information of the edge device, including device status information, computing resources, application scenario information, and requirement indicators; performing multi-dimensional matching and weighted scoring between the quantization vector and feature vectors corresponding to multiple intelligent models, and determining a target intelligent model from the multiple intelligent models based on the scoring results; and sending the target intelligent model to the edge device. The method and apparatus of this invention overcome the computational waste caused by traditional manual fixed deployment schemes, automatically adapting intelligent models for edge devices, thereby improving the real-time performance and accuracy of edge devices when performing complex inference tasks.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Continuous dynamic emotion prediction methods, training methods and equipment

This application proposes a continuous dynamic emotion prediction method, training method, and device. It learns a spatiotemporal codebook through a vector quantization variational self-encoding / decoding module, enabling better learning of the causal and spatiotemporal semantics of dynamic emotions from EEG signal data. Employing Transformer for long-term temporal modeling allows for better learning of temporal dependencies and contextual relationships, enabling the model to learn a more generalizable and dynamic emotion latent space that better captures dynamic trends. Through a reward mechanism and reinforcement learning, the model's temporal prediction and regression capabilities are further optimized. The model not only dynamically adapts to the dynamics of human emotions but also accurately predicts the subject's current emotion, ensuring both accuracy and effectiveness.
Owner:SHENZHEN UNIV

Hardware-aware register-level operator fusion and simd search acceleration system and method

PendingCN122450509AData streamFloating point
The application discloses a kind of hardware perception type vector quantization calculation fusion acceleration instruction parallel processing system and method.For the memory bandwidth bottleneck problem of asymmetric distance calculation in large-scale vector retrieval, the application is inside processor register, 8-bit quantization encoding is loaded, zero extension to 32-bit floating point precision, affine correction based on quantization offset and scale factor, and distance accumulation operation for query vector are fused into single instruction multiple data stream driven in-register execution pipeline.The method eliminates the intermediate memory copy and register overflow operation in the traditional scheme, improves the throughput capacity of vector retrieval and processor resource utilization under the premise of maintaining the calculation accuracy.
Owner:SHANGHAI LINGXIN INTELLIGENT TECHNOLOGY CO LTD

A method for analyzing electroencephalogram data based on reinforcement learning

The application discloses a kind of electroencephalogram data analysis methods based on reinforcement learning, comprising the following steps: S1 gathers each brain area electroencephalogram and is preprocessed to generate fragment;S2 calculates frequency domain features to generate characteristic vector sequence;S3 constructs and initializes growing prototype learning vector quantization network, and outputs nerve state code;S4 executes prototype addition, merging, deletion according to growth criterion and updates state code;S5 constructs reinforcement learning state by state code and performance parameter and determines action execution task;S6 updates action value table according to SARSA and generates group base model;S7 generates individualized action value by copying base and forms individualized model by SARSA update.The present application realizes electroencephalogram training self-adaptive decision and individualized continuous evolution, improves training efficiency and long-term stability.
Owner:ANHUI DUONIANNI INTELLIGENT TECHNOLOGY CO LTD

Gait vision recognition-based cow lameness prediction method and system

This invention belongs to the field of action recognition technology. It discloses a method and system for predicting bovine lameness based on gait visual recognition. The method includes acquiring video data of the bovine movement process using a multi-view camera device, identifying the multi-view two-dimensional coordinates of key bovine body nodes, and fusing these coordinates to reconstruct a three-dimensional trajectory in a global reference frame. Based on the three-dimensional trajectory, the method divides the bovine gait cycle, extracts the velocity, acceleration, and relative motion vectors of key bovine body nodes within each gait cycle, and quantifies the gait cycle characteristics. Based on the gait cycle characteristics, a vertical displacement stagnation criterion is defined. This criterion compares the range of the vertical residual of the hoof relative to the torso within a time window with an adaptive threshold to determine hoof micro-stagnation and output a stagnation index. This significantly improves the accuracy, robustness, and foresight of lameness recognition, while also possessing good individual adaptability and scalable application capabilities.
Owner:河南省种业发展中心

A lightweight encryption communication method between an intelligent fusion terminal and an electric energy meter

The present application relates to the field of digital information transmission, and particularly relates to a lightweight encryption communication method between an intelligent fusion terminal and an electric energy meter, which comprises the following steps: first, collecting background electromagnetic noise energy, equivalent impedance phase and inter-harmonic amplitude of a power grid line to form a multi-dimensional feature vector; then, quantizing the vector into an environment fingerprint, deriving a session key and an initialization vector based on the environment fingerprint and a cloud, and encrypting business data such as electric energy measurement values; then, calculating a comprehensive disturbance index and comparing the comprehensive disturbance index with a dynamic threshold value, and triggering key update when the power grid disturbance exceeds the threshold value; finally, the cloud decrypts and restores the encrypted data packet by querying the stored session key. The present application realizes dynamic binding of the encryption key and the physical environment of the power grid, thereby ensuring communication security and significantly reducing system overhead.
Owner:HANGZHOU XILI INTELLIGENT TECH CO LTD

An intelligent fault diagnosis method for an aero-engine based on a physical constraint spectrum codebook

PendingCN122388359AEngineeringLinear modulation
A physical constraint spectral codebook-based intelligent fault diagnosis method for aero-engines belongs to the fields of aero-engine fault diagnosis and artificial intelligence technology. First, the multi-source vibration signals of the aero-engine are normalized in the order domain and injected with rotational speed information through feature linear modulation. Then, a one-dimensional convolutional neural network and spectral attention mechanism are used to extract fault-sensitive features. Next, a physical constraint spectral codebook is constructed, embedding the order ratio of aero-engine bearing fault features into the vector quantization codebook structure in the form of a sinusoidal basis template, and using hierarchical dual-space quantization to separate sensor features and fault representations. Finally, prototype matching inference is used in the codebook activation histogram space to achieve fault diagnosis across operating conditions and models with few samples. This invention integrates fault physics knowledge into discrete representation learning at the structural level, giving codebook entries interpretable physical semantics, thus solving the problems of missing physical knowledge and insufficient generalization ability in aero-engine fault diagnosis across operating conditions and models.
Owner:XI AN JIAOTONG UNIV

SYSTEM AND METHOD FOR ADAPTIVE VECTOR QUANTIZATION OF DEEP NEURAL NETWORKS USING MACHINE LEARNING

The present disclosure provides a system (106) and a method (400) for the adaptive quantization of deep neural networks using machine learning. The system (106) provides input through a first neural network layer (304-1 to 304-N). The system (106) assigns a set of weight variables (306-1 to 306-N) to the first neural network layer (304-1 to 304-N) and dynamically assigns one or more equal bit-width values ​​to generate a quantized neural network layer (304-1 to 304-N). The system (106) computes a power loss function 318 and a fidelity loss function 320 based on the first neural network layer (304-1 to 304-N) and the quantized neural network layer (326-1 to 326-N).The system (106) optimizes a variation between a quantity 314 and a quantized quantity 316 by calculating a loss function 324 based on the power loss function 318 and the loyalty loss function 320.
Owner:MERCEDES BENZ GROUP AG

An inter-vehicle communication blind filling method and system based on laser radar point cloud target detection and semantic joint source channel coding, and a storage medium

PendingCN122457628ANetwork ConvergenceIntelligent Network
The application discloses a kind of based on laser radar point cloud target detection and semantic joint source channel coding inter-vehicle communication blind filling method, system and storage medium, it is related to intelligent networked vehicle cooperative sensing and semantic communication technical field.Collect the laser radar point cloud data and pose information of target vehicle and cooperative vehicle;Using PointPillars network extracts point cloud feature;Cooperative vehicle carries out joint source channel coding to feature and obtains continuous value semantic feature, is converted into discrete bit sequence by vector quantization and is sent to target vehicle;After receiving, target vehicle carries out semantic recovery and decoding, recovers cooperative vehicle feature;According to pose information, feature level coordinate is aligned;Three-stage self-attention fusion network is used to fuse features;Three-dimensional target detection is carried out to fused feature.The amount of data transmitted can be reduced to 1 / 128 to 1 / 512 of the original point cloud feature, while maintaining high detection accuracy under noisy wireless communication conditions, effectively supplementing the detection of the target vehicle's blind area.
Owner:TONGJI UNIV

A method for decoding an audio signal with perception enhancement and a processing terminal

PendingCN122337220ATime domainSide information
This invention discloses a method for decoding audio signals with enhanced perception, comprising: obtaining a bitstream encoded from the original audio signal, the bitstream including a codebook index and quantization side information, and an embedding vector generated by embedding the quantization side information; performing arithmetic decoding on the bitstream to recover the codebook index and quantization side information in the bitstream; inputting the codebook index into a multi-level inverse vector quantization network for inverse quantization processing to obtain the final quantization feature vector; fusing the quantization feature vector and the embedding vector channel by channel using feature modulation to obtain a fused feature vector; inputting the fused feature vector into a neural waveform synthesis network to generate time-domain audio sampling points from the time-frequency features of the fused feature vector, thereby obtaining a reconstructed audio signal; and inputting the reconstructed audio signal into a perception enhancement post-processing network to obtain the final reconstructed audio signal. This invention improves the quality of the decoded and reconstructed audio signal.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

A transaction governance system and method based on metadata negotiation and dynamic degradation

PendingCN122339988AFeature vectorEngineering
This application discloses a transaction governance system and method based on metadata negotiation and dynamic degradation, relating to the interdisciplinary fields of computer science and artificial intelligence. The system addresses the impedance mismatch problem between probabilistic generative AI and deterministic physical business systems by constructing a Host module with auditing and decision-making capabilities at the MCP protocol layer. It utilizes techniques such as semantic metadata negotiation, multi-dimensional feature vector quantization decision-making, deterministic fingerprint verification, and lock resource materialization mapping. This application achieves adaptive switching of transaction modes through the dynamic degradation mechanism of MCP semantic metadata; constructs causal anchors for AI dialogue flows through deterministic fingerprint and sequence number verification; transforms physical conflicts into proactive avoidance through lock resource materialization mapping; and prevents destructive self-healing through shadow prompts. It achieves controlled consistency risks, execution security, and recoverability guarantees across heterogeneous systems, promoting the evolution of transaction governance paradigms from static hard-coded links to real-time dynamic orchestration of AI.
Owner:SICHUAN HUSHAN ELECTRIC APPLIANCE

Method and apparatus for efficient vector quantization based on pre-quantization and post-correction

This invention proposes a method and apparatus for efficient vector quantization based on pre-quantization and post-correction, relating to the field of image data discrete compression and reconstruction technology. The method includes: encoding input data into a continuous representation using a pre-trained variational autoencoder; dividing the continuous representation into multiple channel groups using a multi-channel quantization strategy, assigning an independent codebook to each channel group and generating corresponding quantization features; constructing an EfficientViT-based post-corrector to optimize the quantization features, generating optimized features; and using a pre-trained variational autoencoder decoder to decode the optimized quantization features to obtain reconstructed data.
Owner:TSINGHUA UNIVERSITY

A sleep-wake automatic detection method based on discrete time-frequency representation learning

PendingCN122250925ABiological modelsSensorsSleep wakingTime–frequency representation
The application discloses a sleep and wakefulness automatic detection method based on discrete time-frequency representation learning, and belongs to the technical field of sleep medical signal processing. The method comprises three core steps of input data generation, representation learning and sleep and wakefulness detection. First, multi-channel physiological signals are preprocessed and divided into sub-blocks, then self-supervised representation learning is completed through convolution embedding, Transformer coding, vector quantization and time-frequency reconstruction, finally continuous and discrete features are fused, and a wakefulness discrimination result is output through a classification network. The application adopts a two-stage training mechanism, reduces the dependence on labeled data, introduces a discrete codebook to improve noise resistance, and through feature fusion, both fine-grained time sequence information and semantic stability are considered, so that high-precision detection can be realized only by using limited physiological signal channels, the application is suitable for home sleep monitoring scenes, and has high popularization value.
Owner:SHENZHEN TECH UNIV