Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

16 results about "Perceptual coding" patented technology

In digital audio perceptual coding is a coding method used to reduce the amount of data needed to produce high-quality sound. Perceptual coding takes advantage of the human ear, screening out a certain amount of sound that is perceived as noise.

Immersive video perception transmission method based on implicit integrated viewport prediction and representation learning code rate decision

The invention discloses an immersive video perception transmission method based on implicit integrated viewport prediction and representation learning code rate decision, and relates to the technical field of immersive video transmission and intelligent content perception coding optimization. In order to solve the problems of insufficient viewport prediction generalization, non-adaptive user preference modeling and poor stability under network disturbance in the prior art, the invention provides a method for establishing an implicit integrated viewport prediction model of a multi-input-output structure by collecting user head motion trail and view field thermodynamic diagram data; constructing a code rate decision engine of mutual information constraint by combining representation learning and reinforcement learning, and realizing tile-level dynamic code rate allocation according to a prediction result and user preference; introducing an adversarial training mechanism of bandwidth jitter and packet loss disturbance to enhance robustness; and end-to-end adaptive optimization is realized by adopting cloud, edge and terminal collaborative deployment. The method is suitable for a high-bandwidth and low-delay scene, the definition of a core visual area can be guaranteed, meanwhile, the bandwidth is remarkably saved, and the user experience continuity is improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Cooperative control system of primary and secondary deep fusion pole-mounted circuit breaker

The invention relates to the technical field of circuit breaker control, in particular to a cooperative control system of a primary and secondary deep fusion pole-mounted circuit breaker, which comprises a perceptual coding module, an intelligent optimization module, a cooperative decision module, a digital twinning module, a causal reasoning module, a learning evolution module, a semantic communication module and an execution feedback module, multi-modal data acquisition is realized through a neuromorphic sensor, a protection strategy is generated in combination with quantum optimization and chaos detection, collaborative decision is realized by using multi-agent reinforcement learning, strategy verification is performed by means of digital twinning and meta-learning, a decision process is optimized based on causal reasoning, and continuous evolution of a model is realized through neural architecture search. And finally, a standardized control instruction is generated through semantic communication, and closed-loop control is formed, so that the fault early warning accuracy, the control response speed and the equipment service life are remarkably improved.
Owner:XI AN BAOGUANG INTELLIGENT ELECTRIC CO LTD +1

Apparatus for determining a lowest integer number of bits required for representing non-differential gain values for compression of HOA data frame representations

To provide an apparatus for determining a minimum integer number of bits required for representing non-differential gain values for compression of a HOA data frame representation.SOLUTION: When compressing the HOA data frame representation, a gain control (15, 151) is applied for each channel signal before it is perceptually encoded (16). The gain values are transferred in a differential manner as side information. However, to start decoding such a streamed compressed HOA data frame representation, an absolute gain value is needed, which should be encoded with a minimum number of bits. In order to determine such a lowest integer number of bits (β e), the HOA data frame representation (C (k)) is rendered in the spatial domain to virtual loudspeaker signals on a unit sphere, followed by a normalization of the directional signal data frame representation (C (k)). Next, the minimum integer bit number is set to β e = Γ log2 (Γ log2 ({{square root} KMAX}·O) ¬ + 1) ¬.SELECTED DRAWING: Figure 1
Owner:DOLBY INTERNATIONAL AB

Automatic detection method of cbct head shadow measurement marker points based on multi-geometry guidance and specific perception coding

PendingCN122335671APattern recognition3d image
An automatic detection method for CBCT cephalometric landmarks based on multi-geometric guidance and specific perceptual coding includes the following steps: Step S1, downsampling the 3D CBCT image and obtaining a preliminary coordinate set through a coarse localization network; Step S2, cropping image blocks centered on the coordinates and extracting local features through a visual encoder containing a shared basic encoder and a low-rank adapter; Step S3, calculating the relative position matrix of the landmarks, encoding spatial relationships using radial basis functions, and constructing a multi-anatomical heterogeneous map; Step S4, inputting visual and edge features into a multi-geometric guidance Transformer, fusing global constraints and updating features using an attention mechanism; Step S5, extracting directional geometric relationships using spherical harmonic functions to construct higher-order update terms, and dynamically updating the heterogeneous map using a gated residual mechanism; Step S6, predicting coordinate offsets through a multi-layer network and performing iterative optimization to output high-precision 3D coordinates. This method significantly improves detection accuracy and robustness.
Owner:ZHEJIANG UNIV OF TECH

Transfer learning-based imaginary voice classification method and system

The invention relates to the technical field of brain-computer interfaces, in particular to an imaginary voice classification method and system based on transfer learning, and the method comprises the steps: training a full-module model through health feature electroencephalogram signals, extracting spatial attention features through a perceptual coding module, capturing cross-brain region correlation features in combination with a sequence correlation module, and classifying the cross-brain region correlation features; semantic conversion from motion to language is realized; multi-modal noise is introduced through the injury simulation module, electroencephalogram signal characteristics in an injury scene are simulated, and the model in the first stage has higher robustness; by freezing the perceptual coding layer parameters of the first-stage model and only updating the subsequent module parameters, the obtained second-stage model improves the classification accuracy and generalization ability in the damage feature space; according to the method, the accuracy of the language intention recognition task is improved, and the method has higher individual adaptability and anti-noise capability.
Owner:GUANGDONG OCEAN UNIVERSITY +1

Light field image perceptual coding method based on intra coding tree unit level code rate allocation

The application discloses a light field image perceptual coding method based on intra coding tree unit level code rate allocation, which firstly selects part of sub-aperture images in a sub-aperture image array and arranges the part of sub-aperture images into a pseudo video sequence; then obtains a depth map and a saliency map of a center sub-aperture image by using a depth estimation network and a saliency detection network; then calculates a code rate allocation weight of each coding tree unit in the selected sub-aperture image by using the center sub-aperture image, the depth map and the saliency map, and performs target code rate allocation by using the code rate allocation weight; finally, synthesizes a sub-aperture image which is not selected by using a light field angle super-resolution reconstruction network, and combines the sub-aperture image with a decoded sub-aperture image to form a complete decoded light field image; the method has the advantages that perceptual redundancy existing in the light field image is effectively removed, and the visual quality and structural consistency of a salient region can be maintained at a lower code rate.
Owner:NINGBO UNIV

Emergency rescue video call system and method for people trapped in a malfunctioning elevator

PendingCN122317227ANoise (video)Packet loss
This invention relates to the field of edge-cloud collaboration technology, specifically to an emergency rescue video call system and method for people trapped in malfunctioning elevators. The system includes: a state-aware coding module that calculates the signal-to-noise ratio slope to compress macroblocks and construct a video stream; a keyframe loss assessment module that calculates packet loss density to obtain truncation loss rate; an asymmetric error correction scheduling module that stretches the step size to construct a hybrid stream; a delay mapping retransmission module that monitors delay parameters to obtain detection results; and a cross-modal time-domain synchronization module that stretches audio and shifts and calibrates time markers. In this invention, spatial redundancy compression is performed based on dynamic image and signal-to-noise ratio patterns to adapt to physical channel attenuation. The load missing density is analyzed and the step distance is dynamically stretched to perform asymmetric error correction scheduling. Combined with the delay detection results, periodic peak search and non-modulation extension are performed on the waveform to drive the retransmission packet marker shift calibration. In a weak network transmission environment, the audio-visual data is kept aligned in the time domain and the screen stuttering and disconnection faults are eliminated.
Owner:GORE ELEVATOR TIANJIN

An immersive video perceptual transmission method based on implicit integrated viewport prediction and representation learning code rate decision

ActiveCN121728261BMulti inputPacket loss
An immersive video perception transmission method based on implicit ensemble viewport prediction and representation learning-based bitrate decision-making is proposed, relating to the fields of immersive video transmission and intelligent content-aware coding optimization. To address the problems of insufficient generalization in viewport prediction, non-adaptive user preference modeling, and poor stability under network disturbances in existing technologies, this method establishes an implicit ensemble viewport prediction model with a multi-input / output structure by collecting user head motion trajectory and field-of-view heatmap data; it constructs a bitrate decision engine with mutual information constraints by combining representation learning and reinforcement learning, and achieves tile-level dynamic bitrate allocation based on prediction results and user preferences; it introduces an adversarial training mechanism against bandwidth jitter and packet loss disturbances to enhance robustness; and it adopts a collaborative deployment of cloud, edge, and terminal to achieve end-to-end adaptive optimization. This method is suitable for high-bandwidth, low-latency scenarios, significantly saving bandwidth and improving user experience continuity while ensuring the clarity of the core view area.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Method and apparatus for decoding compressed HOA signals

PendingJP2026009938ASpeech analysisStereophonic systemsSource encodingPerceptual coding
To provide a method for decompressing a compressed HOA signal and an apparatus for decompressing a compressed HOA signal.SOLUTION: A method for compressing an HOA signal, which is an input HOA representation with input time frames (C (k)) of HOA coefficient sequences, includes a spatial HOA encoding of the input time frames and subsequent perceptual and source encoding. Each input time frame is decomposed (802) into a frame of predominant sound signals (XPS (k-1)) and a frame of ambient HOA component (CAMB (k-1)). The ambient HOA component (CAMB (k - 1)) comprises, in the layered mode, first HOA coefficient sequences of the input HOA representation (c n (k - 1)) in lower positions and second HOA coefficient sequences (CAMB, n (k - 1)) in remaining higher positions. The second HOA coefficient sequences are part of a HOA representation of a residual between the input HOA representation and the HOA representation of the predominant sound signals.SELECTED DRAWING: Figure 5
Owner:DOLBY INTERNATIONAL AB

Methods and apparatus for determining for decoding a compressed HOA sound representation

When compressing an HOA data frame representation, a gain control (15, 151) is applied for each channel signal before it is perceptually encoded (16). The gain values are transferred in a differential manner as side information. However, for starting decoding of such streamed compressed HOA data frame representation absolute gain values are required, which should be coded with a minimum number of bits. For determining such lowest integer number (βe) of bits the HOA data frame representation (C(k)) is rendered in spatial domain to virtual loudspeaker signals lying on a unit sphere, followed by normalisation of the HOA data frame representation (C(k)). Then the lowest integer number of bits is set to βe=┌log2(┌log2(√{square root over (KMAX)}·O)┐+1)┐.
Owner:DOLBY LABORATORIES LICENSING CORP

Voice processing method and electronic device

The present disclosure relates to a speech processing method, apparatus, electronic device, computer readable storage medium and computer program product. The method comprises: obtaining a trained speech translation model, wherein the trained speech translation model comprises an acoustic encoder and a text encoder, wherein the acoustic encoder and / or the text encoder comprises an intermediate CTC module between adjacent first and second layers, the intermediate CTC module being configured to determine an input of the second layer based on an output of the first layer and a word embedding matrix; and inputting a source language speech to be processed into the trained speech translation model to obtain a corresponding target language text. In this way, the speech translation model is able to solve the independent assumption problem inherent in CTC by introducing an intermediate CTC module between two adjacent layers in the acoustic encoder and / or the text encoder, which integrates the predictive perceptual coding into the encoding information, and thus is able to improve the performance of the speech translation processing.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

A memristor-based perceptual neuron circuit and application

The application discloses a kind of perception neuron circuits based on memristor and application, belong to intelligent perception technical field;Integrated perception neuron array is constructed by word line and bit line, and each perception neuron in array is converted into pulse signal and is output in parallel with its perceived external environment information;Each perception neuron in array includes series resistance sensor and threshold transition type memristor;Sensor acts as adjustable resistance, so that memristor works in local active area and normally pulse firing, play the role of sensing environmental information and voltage division, also reduce the difference between devices and devices of multiple perception circuit integration memristor caused by encoding error of perception encoding circuit.In addition, the parasitic capacitance of memristor itself acts as necessary capacitor in memristor neuron circuit;Through the above design, the perception neuron circuit is more compact, greatly improves the integration density, and the reduction of area and hardware saving also reduces energy consumption, improves energy efficiency.
Owner:HUAZHONG UNIV OF SCI & TECH

Audio transmission control method, device and equipment and computer storage medium

PendingCN121938383ASpeech analysisTime-division multiplexPathPingBit allocation
The invention relates to an audio transmission control method, apparatus and device, and a computer storage medium. The method comprises the steps of constructing a multi-level clock source architecture of a master clock, a region-level boundary clock and an end transparent clock in a distributed network environment; a dynamic path delay compensation mechanism and an edge node collaborative calibration strategy are combined to realize time synchronization; a multi-channel parallel processing unit, a low-delay FIR / IIR filter bank and a dynamic bit distribution engine are obtained based on a DSP processor; audio transmission when the network bandwidth is limited is realized by combining a perceptual coding technology, a neural network auxiliary quantization algorithm and a bandwidth self-adaptive dynamic switching strategy; constructing a hybrid prediction model based on a machine learning algorithm; acquiring and collecting multi-dimensional operation data in real time; and dynamically predicting a transmission health state based on the multi-dimensional operation data and the hybrid prediction model. The method and the device have the effects of improving the accuracy of time synchronization and dynamically predicting the transmission state to deal with in advance.
Owner:SHENZHEN SOUNDFIT TECH CO LTD

A multi-level multi-module collaborative video perception coding optimization method and device

The application discloses a kind of multi-level multi-module collaborative video sensing coding optimization method and device, and the derivation of frame level coding distortion prediction and frame level quantization parameter is carried out by original video coding distortion prediction;The image of original video is intraframe / interframe prediction, and the difference calculation is carried out to the predicted image and original image obtained, and residual image is obtained, and the residual image is filtered by the predicted coding distortion, and the residual image after filtering is based on residual block transformation, and then according to the predicted frame level coding distortion and frame level quantization parameter, perceptual quantization is carried out;Rate distortion optimization is carried out based on perceptual quantization parameter, and intraframe / interframe prediction is optimized;Perceptual quality enhancement network is constructed, and is used to optimize intraframe / interframe prediction;After the image of original video is predicted, difference calculation, residual filtering, transformation, perceptual quantization based on optimized intraframe / interframe prediction, entropy coding is carried out.
Owner:HANGZHOU DIANZI UNIV +1

Digital music multi-track intelligent sound mixing method based on perceptual coding

PendingCN121686979AElectrophonic musical instrumentsPerceptual codingSignal processing
The invention discloses a digital music multi-track intelligent sound mixing method based on perceptual coding, and belongs to the field of digital music. The method comprises four steps of music content understanding and analysis, adaptive track grouping mixing, perception importance dynamic adjustment and optimization, and final sound mixing signal generation and output. According to the method, the music content understanding layer based on the deep learning model is introduced, so that structured semantic analysis of multi-track music is realized, and the sound mixing process is improved from the traditional signal processing level to the music semantic understanding level. The system can automatically identify key features such as musical instruments, harmony and rhythm, and understand music function association between tracks, so that accurate context information is provided for subsequent intelligent processing, a sound mixing decision is more in line with the artistic law of music creation, and the music expressive force and specialty of a sound mixing result are remarkably improved.
Owner:XINGAN VOCATIONAL & TECH COLLEGE

Encoded HOA data frame representation comprising non-differential gain values associated with channel signals of individual ones of the data frames of the HOA data frame representation

To provide a method and apparatus for determining a minimum integer number of bits required to represent non-differential gain values for compression of a Higher Order Ambisonics (HOA) data frame representation.SOLUTION: When compressing the HOA data frame representation in the HOA compressor, a gain control 15, 151 is applied for each channel signal before it is perceptually encoded 16. Those gain values should be encoded with a minimum number of bits, and in order to determine such a lowest integer number of bits (β e), the HOA data frame representation (C (K)) is rendered in the spatial domain to virtual loudspeaker signals on a unit sphere, following which the HOA component is normalized to the directional signal data frame representation (C (K)). Then, the lowest integer number of bits is set to β e = [log2 ([log2 ({√ KMAX}·O)] + 1)].SELECTED DRAWING: Figure 1
Owner:DOLBY INTERNATIONAL AB