Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

27 results about "Discrete cosine transform" patented technology

A discrete cosine transform (DCT) expresses a finite sequence of data points in terms of a sum of cosine functions oscillating at different frequencies. This is the standard data compression technique widely used by most digital media standards, for image compression (e.g. JPEG and HEIF, where small high-frequency components can be discarded), video coding (e.g. MPEG and H.26x), digital audio (e.g. MP3 and AAC) and digital television (e.g. SDTV and HDTV). DCTs are also important to numerous applications in science and engineering, such as spectral methods for the numerical solution of partial differential equations.

Hybrid coding processing method, system and equipment based on video engine

The invention relates to the field of hybrid coding of video engines, and provides a hybrid coding processing method, system and equipment based on a video engine, which comprises the step of intelligently switching an H.265 inter-frame predictive coding mode and an MJPEG (Multijoint Joint Photographic Experts Group) intra-frame compression coding mode by dynamically monitoring the proportion of a motion area in a video picture. When a scene is static, discrete cosine transform is adopted to compress space redundancy, motion vector compensation is started to eliminate time redundancy when motion is violent, and meanwhile, key frames are doubly screened by using pixel difference analysis and a perceptual hash algorithm, and repeated frames of visual redundancy are eliminated. And finally, space-time association is established for the optimized double-code-stream data through a timestamp index system, and efficient storage and accurate reconstruction of the mixed code stream are realized. The problems that the coding efficiency is low, no self-adaptive coding mode exists, redundant frames are not optimized, and code stream management is insufficient are solved.
Owner:SICHUAN SILICON MICROELECTRONICS TECHNOLOGY CO LTD

Multi-rate audio mixing

PCT designated stageWO2025250369A1Speech analysisSpecial service for subscribersDiscrete cosine transformAudio frequency
This disclosure provides methods, components, devices and systems for multirate audio mixing. Some aspects more specifically relate to mixing audio streams with different sample rates. In some examples, an audio source device may convert audio streams with different sample rates to the frequency domain using a modified discrete cosine transform (MDCT), and the audio source device may mix the audio streams with different sample rates in the frequency domain. The audio source device may apply a pre-emphasis filter after mixing the audio streams in the frequency domain. An audio stream with a higher sample rate may be down-sampled by dropping frequency bins after converting the audio stream to the frequency domain. Additionally, or alternatively, an audio stream with a lower sample rate may be up-sampled by padding frequency bins of the frequency domain-converted audio stream.
Owner:QUALCOMM INC

An efficient gaussian splatting reconstruction method and device for dynamic endoscopic scenes

The application discloses a kind of high-efficiency Gaussian splashing reconstruction method and device for dynamic endoscope scene in the technical field of medical image processing, comprising: the endoscope video sequence obtained is preprocessed, and initial Gaussian point cloud is generated;The motion trajectory of each Gaussian point in initial Gaussian point cloud is modeled using Gaussian point cloud morphing model based on discrete cosine transform, and dynamic Gaussian scene is generated;The dynamic and static attributes of Gaussian point cloud morphing model based on discrete cosine transform are compressed and stored using residual-aware hybrid precision quantization strategy;Dynamic Gaussian scene is rendered using hardware-aware dense inference strategy, and three-dimensional reconstruction image of endoscope video is generated in real time.The application can effectively improve the expression efficiency of endoscope scene dynamic modeling, while reducing storage / deployment overhead.
Owner:HUBEI UNIV OF ARTS & SCI

Editing tool tracing method and system for multi-tool editing image

The invention discloses an editing tool traceability method for multi-tool editing images, and the method comprises the steps: extracting frequency domain block-level features through sliding window discrete cosine transform based on a dual-domain block feature enhancement mode of a spatial domain and a frequency domain for a to-be-traced image data set, and combining the spatial domain block-level features to obtain a to-be-traced image data set; realizing double-domain feature fusion by adopting a cross-attention mechanism, and outputting fusion features; based on a training data sampling strategy of error feedback, a training sample is divided into a plurality of subsets according to the number of editing steps, the sampling weight is dynamically adjusted based on error feedback of each subset, and model training is optimized; and inputting the fusion features into the optimized model to realize traceability of each editing tool. According to the method provided by the invention, one-time accurate identification of all the editing participating tools in the multi-tool editing image can be realized.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A speech filtering method, device, storage medium and equipment

The present application discloses a speech filtering method, device, storage medium and equipment, which belongs to the field of speech coding and decoding technology. The method mainly includes: encoding the speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without a post-filtering module to obtain speech spectrum coefficients; inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; and according to the remaining decoding steps of the standard decoder without a post-filtering module, inputting the target spectrum coefficients into the low-latency improved inverse discrete cosine transform module of the standard decoder to obtain the target speech signal corresponding to the target spectrum coefficients. The present application omits the complex post-filtering operation in the Bluetooth encoding process, and only uses the pre-trained neural network model for filtering in the Bluetooth decoding process, so that it achieves a sound quality close to that of standard decoding.
Owner:BEIJING BAIRUI INTERNET TECH CO LTD

Self-adaptive robust video watermarking method based on H.264AVC video coding

The invention relates to a self-adaptive robust video watermarking method based on H.264AVC video coding. The self-adaptive robust video watermarking method comprises the steps of preprocessing, watermark embedding and extracting. The method comprises the following steps: preprocessing: graying an image, carrying out integer discrete cosine transform and quantization processing on a gray value, then carrying out Z-shaped coefficient scanning, then carrying out data formula transformation and 4-bit unsigned binary coding conversion, and finally introducing an error correction coding link for improving the watermark data recovery capability. When a watermark is embedded, deterministic mapping from a macro block group to a bit is established in a key frame, and the same information is embedded in a distributed manner at different time points. During watermark extraction, a judgment method based on bit-level majority statistics is introduced so as to perform fusion processing on extraction results of the same watermark line in a plurality of key frames. According to the method, the overall performance of the watermarking system in the aspects of synchronism, robustness and invisibility can be improved.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Method, device, equipment and medium for extracting video watermark based on quick response code

The present application discloses a method, apparatus, device, and medium for extracting video watermarks based on a quick response code, relating to the field of image processing technology. The method comprises: obtaining a first scene change frame and a quick response code matrix of a first video, processing the above contents separately to obtain an encrypted quick response code matrix and multiple image blocks of different sizes; performing a discrete cosine transform on the luminance component of the first scene change frame, combining the encrypted quick response code matrix to obtain DC coefficients and AC coefficients; then performing an inverse discrete cosine transform on the multiple image blocks of different sizes to embed watermark information and obtain a first video with a watermark; processing the first video with a watermark to obtain a target estimated value of the encrypted quick response code matrix; decrypting and decoding the target estimated value in sequence to obtain target watermark information, thereby extracting the watermark information. The method can dynamically adjust the size of the embedded watermark, improve extraction efficiency, and enhance the security of the watermark.
Owner:TIANJIN XITONG ELECTRONICS EQUIP CO LTD +1

Counterfeit positioning method based on attention enhancement and adaptive frequency selection

PendingCN121937453AImage enhancementImage analysisEngineeringDiscrete cosine transform
The invention discloses a forgery positioning method based on attention enhancement and adaptive frequency selection, and belongs to the technical field of image forgery detection. The invention aims to solve the problem of poor positioning precision and robustness caused by inaccurate feature capture of a forged area and insufficient utilization of multi-scale and frequency information in the existing forged positioning technology. The scheme comprises the following steps: extracting multi-scale spatial features of an input image through an encoder; channel and space attention weighting is carried out on the multi-scale features by using an attention enhancement module so as to enhance and counterfeit related features; carrying out discrete cosine transform and adaptive weight learning on the enhanced features by adopting an adaptive frequency selection module, screening effective frequency components and carrying out cross-domain fusion; and finally, generating a positioning mask of the forged region based on the fusion features. The method can effectively improve the positioning accuracy of the forged area in the image and the robustness in a complex scene, and is suitable for digital media evidence obtaining and content security auditing.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Audio understanding method and system, electronic equipment and storage medium

PendingCN121963770AImprove understanding and analysis efficiencyEffective filteringSpeech analysisEngineeringDiscrete cosine transform
The invention discloses an audio understanding method and system, electronic equipment and a storage medium, and the method comprises the steps: obtaining each coding unit of each audio clip of a target audio, and carrying out the compression through discrete cosine transform based on the information entropy of each coding unit, and obtaining a compressed coding unit; polling each audio clip, and calculating the acoustic significance membership degree of each compressed coding unit and the time sequence correlation membership degree of each current fusion coding unit of the previous audio clip; calculating the contribution weight of each compression coding unit and the contribution weight of each current fusion coding unit to the compression coding unit by using the membership degree of each compression coding unit; respectively utilizing the contribution weight corresponding to each compression coding unit to carry out weighted summation on the compression coding unit and each current fusion coding unit to obtain each fusion coding unit of the current audio clip, inputting the fusion coding units into a large language model, and outputting an understanding analysis result; and fusing the understanding analysis results into a result of the target audio.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Dynamic image compression method

The invention provides an image compression system and method. The method includes acquiring an image frame and dividing it into a set of image blocks. Then obtaining a preliminary quantitative configuration file; each image block is subjected to a discrete cosine transform (DCT) to create a frequency dependent image profile, which is then quantized using a preliminary quantization profile and compressed using at least one compression encoding algorithm. And comparing the size of the compressed image frame with a target size range, and updating the preliminary quantitative configuration file according to a comparison result. In addition, the method further comprises the step of dynamically adjusting the target size range according to the environment or network condition of the system, so that the real-time adaptability and efficiency of processing the dynamic environment and the complex scene are improved.
Owner:BEKEN CORP

An image denoising method based on frequency space joint guided dynamic kernel generation network

The application provides an image denoising method based on frequency space joint guided dynamic kernel generation network, which comprises the following steps: inputting a noisy image into an initial convolution layer to obtain a shallow feature representation; adopting an encoder to perform hierarchical down-sampling and context extraction, and recovering spatial details through a jump connection in a decoder; adopting a fixed two-dimensional discrete cosine transform filter bank to perform low-frequency, medium-frequency and high-frequency decomposition on the input feature to construct a frequency prior; fusing spatial context information, channel interaction information and the frequency prior to generate a dynamic convolution kernel; performing frequency space joint guided dynamic feature extraction on multiple receptive field scales; interacting and fusing channel, height and width dimension information through a multi-dimensional adaptive fusion mechanism; finally, completing detail recovery in the decoding stage and outputting a denoised image; and the application alleviates the confusion between high-frequency texture and noise residual, and improves the detail preservation capability and denoising robustness in complex texture, strong noise and scale change scenes.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Image processing and rock identification method based on high-frequency filtering and mineral composition attention

The application provides an image processing and rock mineral identification method based on high-frequency filtering and mineral component attention, and belongs to the technical field of rock mineral identification. The method is started from the perspective of the frequency domain, the discrete cosine transform is introduced into the CNN architecture, high-frequency features are explicitly enhanced, high-frequency information is directly extracted and enhanced, and the influence of the smoothing effect of the CNN in layer-by-layer transmission on the detail information can be effectively relieved; the multiple attention mechanisms in the channel dimension and the spatial dimension enable the model to automatically determine which features are more important according to the current input rock mineral image, and realize efficient adaptive optimization and extraction of complex texture structures. Experimental results show that the strategy of explicitly enhancing high frequencies in the frequency domain and the spatial-channel double attention collaborative work makes the model show higher identification accuracy when facing actual rock mineral images with complex composition, variable structure and strong background interference.
Owner:NORTHEASTERN UNIV CHINA +1

A hybrid encoding processing method, system, and device based on a video engine

This application relates to the field of hybrid coding in video engines, providing a hybrid coding processing method, system, and device based on a video engine. This includes intelligently switching between H.265 inter-frame predictive coding and MJPEG intra-frame compression coding modes by dynamically monitoring the proportion of moving areas in the video frame. When the scene is static, discrete cosine transform is used to compress spatial redundancy; when there is rapid motion, motion vector compensation is used to eliminate temporal redundancy. Simultaneously, pixel difference analysis and perceptual hashing algorithms are used to dual-filter keyframes, eliminating visually redundant duplicate frames. Finally, a timestamp indexing system is used to establish spatiotemporal correlation between the optimized dual-stream data, achieving efficient storage and accurate reconstruction of the hybrid bitstream. This solves the problems of low coding efficiency, lack of adaptive coding modes, unoptimized redundant frames, and insufficient bitstream management.
Owner:SICHUAN SILICON MICROELECTRONICS TECHNOLOGY CO LTD

A method of digital audio embedding and detection

The application provides a digital audio embedding and detecting method, which comprises the following steps: generating watermark information according to time and watermark load equipment coding; then performing frame processing on the audio to be embedded with the watermark, and performing two-stage wavelet decomposition; embedding a synchronization code in the first half of the low-frequency wavelet coefficient, and performing discrete cosine transform and singular value decomposition processing on the second half of the low-frequency wavelet coefficient to embed the watermark information; finally, in the watermark detecting stage, firstly, the audio subjected to attack in the current sliding window is subjected to two-stage wavelet decomposition through a sliding window method, whether the current sliding window contains the synchronization code is confirmed, the position of the synchronization code is determined, frame synchronization is realized, and then the second half of the audio frame subjected to attack and frame synchronization is processed to extract the watermark information. The synchronization code and the watermark information are embedded in the audio at the same time, the problem that the embedding position of the watermark in the detecting stage is difficult to position and the problem that the watermark is difficult to resist synchronous attacks such as cutting are solved, and the watermark detecting efficiency and accuracy are improved.
Owner:CHINA RES INST OF FILM SCI & TECH

Video recompression evidence obtaining method with intra-frame and inter-frame prediction information fusion

PendingCN121262386ADigital video signal modificationMotion vectorDiscrete cosine transform
The invention discloses a video recompression evidence obtaining method for intra-frame and inter-frame prediction information fusion. The method comprises the following steps: extracting primary code stream characteristics of a prediction frame, constructing a prediction frame matrix and a motion vector, determining a prediction frame attribute mark, determining a prediction frame compression trace image, determining a compression ratio of an intra-frame compression trace image, determining a Fourier transform coefficient of the intra-frame compression trace image, and determining an autocorrelation coefficient of an inter-frame compression trace image. Determining a two-dimensional discrete cosine transform value of the inter-frame compression trace image, determining a high-level code stream feature matrix, and determining a re-compression evidence obtaining result. According to the method, the problem that only intra-frame prediction information is considered and inter-frame prediction information is not considered in the prior art is solved, and compared with an existing video recompression evidence obtaining method, the evidence obtaining accuracy of the method is higher than 80%. According to the method disclosed by the invention, the complexity of evidence obtaining time is reduced, the evidence obtaining precision is obviously improved, the improvement is relatively obvious under a high code rate, and the method can be applied to the technical field of video recompression evidence obtaining.
Owner:XIAN UNIV OF POSTS & TELECOMM +1

High quality transcoding efficient texture format

Techniques for transcoding textures are described that include computing a plurality of endpoint colors of a macroblock. In a first technique, a forward discrete cosine transform (DCT) is applied (406) to a macroblock or portion thereof and used to represent the corresponding portion. In a second technique, a mean value of the endpoint colors and a difference between the endpoint colors are calculated (206), where the mean value and the difference establish (208) a projection vector in a color space. Means and differences are compressed (210), and per-pixel distances along the projection vector are calculated (212) and used together with corresponding mean and differences to represent macroblocks.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Low frequency non-separable transform and use of non-dct2 primary transform

PendingCN122514950AVideo encodingDiscrete cosine transform
Systems, methods, and means for video encoding / decoding are disclosed for the use of Low Frequency Inseparable Transform (LFNST) and Non-Discrete Cosine Transform 2 (DCT2) master transforms. In the examples, a video decoding device can determine whether to disable the Inseparable Master Transform (NSPT) for the current block. If NSPT is disabled for the current block, the video decoding device can determine to use LFNST for the current block. If LFNST is used, the video decoding device can decode the current block based on Multiple Transform Selection (MTS). In the examples, a video encoding device can disable NSPT for the current block. A video encoding device can enable LFNST for the current block. A video encoding device can encode the current block based on MTS.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

Fixed rate intraframe compression and decompression of video based on visual quality

To provide methods, programs and systems for fixed bitrate intraframe compression of video based on 2D discrete cosine transform (DCT) to control a target code amount, on the basis of the visual quality of final decompressed imagery.SOLUTION: A method involves assigning an initial bit budget per macroblock of a given video frame, resulting in degrees of compression, i.e., quantization scale factors, that vary among the macroblocks according to the complexity of the macroblocks. The scale factors are then adjusted while maintaining the overall frame bit budget to reduce the visibility of artifacts in the decompressed frame. The adjustments may include increasing scale factors for simple macroblocks, and reducing scale factors for complex blocks. This can reduce the visibility of compression-related artifacts both in complex and simple portions of the frame.SELECTED DRAWING: Figure 2
Owner:AVID TECHNOLOGY INC

A method and device for detecting facial images generated by local GAN ​​with adaptive frequency perception

The present invention discloses a method and device for detecting facial images generated by a local GAN ​​using adaptive frequency perception, comprising: obtaining a facial image; detecting the facial image using a trained local GAN-generated facial image detection model to obtain a local GAN-generated facial image detection result; the local GAN-generated facial image detection model includes an adaptive frequency perception module and a detection network; wherein the adaptive frequency perception module includes a two-dimensional discrete cosine transform submodule, a learnable filter bank, and a two-dimensional inverse discrete cosine transform submodule arranged in sequence, the learnable filter bank including a high-pass filter and a learnable filter; and the detection network is obtained by deleting the fourth to eleventh residual blocks in an Xception network and adding an ECA attention module to the first to third and twelfth residual blocks. The present invention can accurately detect facial images generated by local GAN ​​and has good generalization and robustness.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

A multi-resolution based audio transform processing method and processing terminal

PendingCN122337216AHuman auditory systemEngineering
This invention discloses a multi-resolution audio transformation processing method, comprising the following steps: preprocessing and framing the obtained discrete time-series audio signal sequentially to obtain a framed audio signal; the preprocessing is used for filtering the audio signal; determining whether each audio frame in the framed audio signal belongs to a transient frame or a steady-state frame; performing wavelet packet decomposition on the audio frames belonging to transient frames to obtain a first type of audio frame; performing an improved discrete cosine transform on the audio frames belonging to steady-state frames to obtain a second type of audio frame; fusing the first type of audio frame and the second type of audio frame to obtain a fused audio signal; and non-uniformly dividing the time-frequency coefficients of the fused audio signal according to the Bark critical frequency band scale to obtain a divided audio signal. This invention can obtain an audio signal with good multi-scale time-frequency domain representation, which can obtain an audio signal that matches the frequency selectivity characteristics of the human auditory system.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

A lightweight sound detection system and method for dry-wood pests based on improved MobileNetV3

This application discloses a lightweight sound detection system and method for borers based on an improved MobileNetV3, comprising: a data preprocessing module for adaptive noise suppression, dynamic duration adjustment, and sampling rate optimization of the original audio signal to obtain a standardized audio signal; a spectrogram extraction and fusion module for performing short-time Fourier transform based on a bidirectional Mel filter bank to obtain a fused spectrogram; a lightweight feature extraction network for extracting multi-scale basic features and outputting an intermediate feature map; a multispectral channel attention module for extracting multi-band frequency features from the spectrogram through a two-dimensional discrete cosine transform, learning frequency weights through a multilayer perceptron and performing channel weighting to obtain a weighted feature map; and a multi-scale feature fusion module for unifying the channel dimensions of the feature map through a feature pyramid network and obtaining a fused feature map through bilinear interpolation upsampling, thereby predicting the probability and type of pest presence. This application can improve the real-time performance and accuracy of detection.
Owner:NORTHWEST A & F UNIV

Method and apparatus for multi-stage vector quantization for audio coding

PendingCN122290614ADiscrete cosine transformVector quantisation
Various embodiments of this disclosure disclose methods and apparatuses for multi-level vector quantization for audio coding. A method implemented by an encoder is disclosed. The method includes: obtaining a (1201) Discrete Cosine Transform (DCT) target vector; performing a (1203) suboptimal pairwise internal search in each segment of a codebook (104) having multiple segments to determine a set of pairwise initial candidates from each of the multiple segments, thereby forming a plurality of pairwise initial candidates from the suboptimal pairwise internal search, wherein each segment has a truncated vector different from the truncated vectors of the other segments in the plurality of segments, wherein the DCT target vector is the target of the suboptimal pairwise internal search for each of the multiple segments; reconstructing (1209) a plurality of final candidates using Inverse Type II Discrete Cosine Transform (DCT-II) to transform the plurality of final candidates into final candidate data in the original domain; and providing the final candidate data (1209) to a second stage of a multi-level vector quantizer.
Owner:TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Multi-rate audio mixing

ActiveUS20250372107A1Speech analysisAutomatic exchangesDiscrete cosine transformAudio frequency
This disclosure provides methods, components, devices and systems for multi-rate audio mixing. Some aspects more specifically relate to mixing audio streams with different sample rates. In some examples, an audio source device may convert audio streams with different sample rates to the frequency domain using a modified discrete cosine transform (MDCT), and the audio source device may mix the audio streams with different sample rates in the frequency domain. The audio source device may apply a pre-emphasis filter after mixing the audio streams in the frequency domain. An audio stream with a higher sample rate may be down-sampled by dropping frequency bins after converting the audio stream to the frequency domain. Additionally, or alternatively, an audio stream with a lower sample rate may be up-sampled by padding frequency bins of the frequency domain-converted audio stream.
Owner:QUALCOMM INC

Low latency audio packet loss concealment

The invention provides a method for real-time concealing errors in audio data packets. A Long Short-Term Memory (LSTM) neural network with a plurality of nodes is provided and pre-trained with audio data. A sequence of packets is received, each packet comprising a set of modified discrete cosine transform (MDCT) coefficients associated with a frame comprising time-domain samples of the audio signal. These MDCT coefficient data are applied to the LSTM neural network, and in case it is identified that a received packet is an erroneous packet, an output from the LSTM neural network is used to generate estimated MDCT co-efficients to provide a concealment packet to replace the erroneous packet. Preferably, the MDCT coefficients are normalized prior to applying to the LSTM neural network. This method can be performed in real-time. A low latency can be obtained and still with a high audio quality.
Owner:RTX AS CO

Image processing and rock mine identification method based on high-frequency filtering and mineral composition attention

The invention provides an image processing and rock and ore identification method based on high-frequency filtering and mineral composition attention, and belongs to the technical field of rock and ore identification. According to the method, from the perspective of a frequency domain, discrete cosine transform is introduced into a CNN architecture, high-frequency features are explicitly enhanced, high-frequency information is extracted and enhanced more directly, and the influence of the smoothing effect of the CNN in layer-by-layer transmission on detailed information can be effectively relieved; a multi-attention mechanism in a channel dimension and a space dimension is adopted, so that the model can automatically judge which features are more important according to a currently input rock and mine image, and efficient self-adaptive optimization and extraction of a complex texture structure are realized. Experimental results show that the strategy of carrying out explicit high-frequency enhancement in the frequency domain and space-channel double attention cooperative work is adopted, so that the model shows higher recognition accuracy when facing actual rock and mine images with complex components, variable structures and strong background interference.
Owner:NORTHEASTERN UNIV CHINA +1