Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Discrete cosine transform" patented technology

A discrete cosine transform (DCT) expresses a finite sequence of data points in terms of a sum of cosine functions oscillating at different frequencies. This is the standard data compression technique widely used by most digital media standards, for image compression (e.g. JPEG and HEIF, where small high-frequency components can be discarded), video coding (e.g. MPEG and H.26x), digital audio (e.g. MP3 and AAC) and digital television (e.g. SDTV and HDTV). DCTs are also important to numerous applications in science and engineering, such as spectral methods for the numerical solution of partial differential equations.

Hybrid coding processing method, system and equipment based on video engine

The invention relates to the field of hybrid coding of video engines, and provides a hybrid coding processing method, system and equipment based on a video engine, which comprises the step of intelligently switching an H.265 inter-frame predictive coding mode and an MJPEG (Multijoint Joint Photographic Experts Group) intra-frame compression coding mode by dynamically monitoring the proportion of a motion area in a video picture. When a scene is static, discrete cosine transform is adopted to compress space redundancy, motion vector compensation is started to eliminate time redundancy when motion is violent, and meanwhile, key frames are doubly screened by using pixel difference analysis and a perceptual hash algorithm, and repeated frames of visual redundancy are eliminated. And finally, space-time association is established for the optimized double-code-stream data through a timestamp index system, and efficient storage and accurate reconstruction of the mixed code stream are realized. The problems that the coding efficiency is low, no self-adaptive coding mode exists, redundant frames are not optimized, and code stream management is insufficient are solved.
Owner:SICHUAN SILICON MICROELECTRONICS TECHNOLOGY CO LTD

An efficient gaussian splatting reconstruction method and device for dynamic endoscopic scenes

The application discloses a kind of high-efficiency Gaussian splashing reconstruction method and device for dynamic endoscope scene in the technical field of medical image processing, comprising: the endoscope video sequence obtained is preprocessed, and initial Gaussian point cloud is generated;The motion trajectory of each Gaussian point in initial Gaussian point cloud is modeled using Gaussian point cloud morphing model based on discrete cosine transform, and dynamic Gaussian scene is generated;The dynamic and static attributes of Gaussian point cloud morphing model based on discrete cosine transform are compressed and stored using residual-aware hybrid precision quantization strategy;Dynamic Gaussian scene is rendered using hardware-aware dense inference strategy, and three-dimensional reconstruction image of endoscope video is generated in real time.The application can effectively improve the expression efficiency of endoscope scene dynamic modeling, while reducing storage / deployment overhead.
Owner:HUBEI UNIV OF ARTS & SCI

Editing tool tracing method and system for multi-tool editing image

The invention discloses an editing tool traceability method for multi-tool editing images, and the method comprises the steps: extracting frequency domain block-level features through sliding window discrete cosine transform based on a dual-domain block feature enhancement mode of a spatial domain and a frequency domain for a to-be-traced image data set, and combining the spatial domain block-level features to obtain a to-be-traced image data set; realizing double-domain feature fusion by adopting a cross-attention mechanism, and outputting fusion features; based on a training data sampling strategy of error feedback, a training sample is divided into a plurality of subsets according to the number of editing steps, the sampling weight is dynamically adjusted based on error feedback of each subset, and model training is optimized; and inputting the fusion features into the optimized model to realize traceability of each editing tool. According to the method provided by the invention, one-time accurate identification of all the editing participating tools in the multi-tool editing image can be realized.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Self-adaptive robust video watermarking method based on H.264AVC video coding

The invention relates to a self-adaptive robust video watermarking method based on H.264AVC video coding. The self-adaptive robust video watermarking method comprises the steps of preprocessing, watermark embedding and extracting. The method comprises the following steps: preprocessing: graying an image, carrying out integer discrete cosine transform and quantization processing on a gray value, then carrying out Z-shaped coefficient scanning, then carrying out data formula transformation and 4-bit unsigned binary coding conversion, and finally introducing an error correction coding link for improving the watermark data recovery capability. When a watermark is embedded, deterministic mapping from a macro block group to a bit is established in a key frame, and the same information is embedded in a distributed manner at different time points. During watermark extraction, a judgment method based on bit-level majority statistics is introduced so as to perform fusion processing on extraction results of the same watermark line in a plurality of key frames. According to the method, the overall performance of the watermarking system in the aspects of synchronism, robustness and invisibility can be improved.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Counterfeit positioning method based on attention enhancement and adaptive frequency selection

PendingCN121937453AImage enhancementImage analysisEngineeringDiscrete cosine transform
The invention discloses a forgery positioning method based on attention enhancement and adaptive frequency selection, and belongs to the technical field of image forgery detection. The invention aims to solve the problem of poor positioning precision and robustness caused by inaccurate feature capture of a forged area and insufficient utilization of multi-scale and frequency information in the existing forged positioning technology. The scheme comprises the following steps: extracting multi-scale spatial features of an input image through an encoder; channel and space attention weighting is carried out on the multi-scale features by using an attention enhancement module so as to enhance and counterfeit related features; carrying out discrete cosine transform and adaptive weight learning on the enhanced features by adopting an adaptive frequency selection module, screening effective frequency components and carrying out cross-domain fusion; and finally, generating a positioning mask of the forged region based on the fusion features. The method can effectively improve the positioning accuracy of the forged area in the image and the robustness in a complex scene, and is suitable for digital media evidence obtaining and content security auditing.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Audio understanding method and system, electronic equipment and storage medium

PendingCN121963770AImprove understanding and analysis efficiencyEffective filteringSpeech analysisEngineeringDiscrete cosine transform
The invention discloses an audio understanding method and system, electronic equipment and a storage medium, and the method comprises the steps: obtaining each coding unit of each audio clip of a target audio, and carrying out the compression through discrete cosine transform based on the information entropy of each coding unit, and obtaining a compressed coding unit; polling each audio clip, and calculating the acoustic significance membership degree of each compressed coding unit and the time sequence correlation membership degree of each current fusion coding unit of the previous audio clip; calculating the contribution weight of each compression coding unit and the contribution weight of each current fusion coding unit to the compression coding unit by using the membership degree of each compression coding unit; respectively utilizing the contribution weight corresponding to each compression coding unit to carry out weighted summation on the compression coding unit and each current fusion coding unit to obtain each fusion coding unit of the current audio clip, inputting the fusion coding units into a large language model, and outputting an understanding analysis result; and fusing the understanding analysis results into a result of the target audio.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Dynamic image compression method

The invention provides an image compression system and method. The method includes acquiring an image frame and dividing it into a set of image blocks. Then obtaining a preliminary quantitative configuration file; each image block is subjected to a discrete cosine transform (DCT) to create a frequency dependent image profile, which is then quantized using a preliminary quantization profile and compressed using at least one compression encoding algorithm. And comparing the size of the compressed image frame with a target size range, and updating the preliminary quantitative configuration file according to a comparison result. In addition, the method further comprises the step of dynamically adjusting the target size range according to the environment or network condition of the system, so that the real-time adaptability and efficiency of processing the dynamic environment and the complex scene are improved.
Owner:BEKEN CORP

An image denoising method based on frequency space joint guided dynamic kernel generation network

The application provides an image denoising method based on frequency space joint guided dynamic kernel generation network, which comprises the following steps: inputting a noisy image into an initial convolution layer to obtain a shallow feature representation; adopting an encoder to perform hierarchical down-sampling and context extraction, and recovering spatial details through a jump connection in a decoder; adopting a fixed two-dimensional discrete cosine transform filter bank to perform low-frequency, medium-frequency and high-frequency decomposition on the input feature to construct a frequency prior; fusing spatial context information, channel interaction information and the frequency prior to generate a dynamic convolution kernel; performing frequency space joint guided dynamic feature extraction on multiple receptive field scales; interacting and fusing channel, height and width dimension information through a multi-dimensional adaptive fusion mechanism; finally, completing detail recovery in the decoding stage and outputting a denoised image; and the application alleviates the confusion between high-frequency texture and noise residual, and improves the detail preservation capability and denoising robustness in complex texture, strong noise and scale change scenes.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Image processing and rock identification method based on high-frequency filtering and mineral composition attention

The application provides an image processing and rock mineral identification method based on high-frequency filtering and mineral component attention, and belongs to the technical field of rock mineral identification. The method is started from the perspective of the frequency domain, the discrete cosine transform is introduced into the CNN architecture, high-frequency features are explicitly enhanced, high-frequency information is directly extracted and enhanced, and the influence of the smoothing effect of the CNN in layer-by-layer transmission on the detail information can be effectively relieved; the multiple attention mechanisms in the channel dimension and the spatial dimension enable the model to automatically determine which features are more important according to the current input rock mineral image, and realize efficient adaptive optimization and extraction of complex texture structures. Experimental results show that the strategy of explicitly enhancing high frequencies in the frequency domain and the spatial-channel double attention collaborative work makes the model show higher identification accuracy when facing actual rock mineral images with complex composition, variable structure and strong background interference.
Owner:NORTHEASTERN UNIV CHINA +1

A hybrid encoding processing method, system, and device based on a video engine

This application relates to the field of hybrid coding in video engines, providing a hybrid coding processing method, system, and device based on a video engine. This includes intelligently switching between H.265 inter-frame predictive coding and MJPEG intra-frame compression coding modes by dynamically monitoring the proportion of moving areas in the video frame. When the scene is static, discrete cosine transform is used to compress spatial redundancy; when there is rapid motion, motion vector compensation is used to eliminate temporal redundancy. Simultaneously, pixel difference analysis and perceptual hashing algorithms are used to dual-filter keyframes, eliminating visually redundant duplicate frames. Finally, a timestamp indexing system is used to establish spatiotemporal correlation between the optimized dual-stream data, achieving efficient storage and accurate reconstruction of the hybrid bitstream. This solves the problems of low coding efficiency, lack of adaptive coding modes, unoptimized redundant frames, and insufficient bitstream management.
Owner:SICHUAN SILICON MICROELECTRONICS TECHNOLOGY CO LTD

Video recompression evidence obtaining method with intra-frame and inter-frame prediction information fusion

PendingCN121262386ADigital video signal modificationMotion vectorDiscrete cosine transform
The invention discloses a video recompression evidence obtaining method for intra-frame and inter-frame prediction information fusion. The method comprises the following steps: extracting primary code stream characteristics of a prediction frame, constructing a prediction frame matrix and a motion vector, determining a prediction frame attribute mark, determining a prediction frame compression trace image, determining a compression ratio of an intra-frame compression trace image, determining a Fourier transform coefficient of the intra-frame compression trace image, and determining an autocorrelation coefficient of an inter-frame compression trace image. Determining a two-dimensional discrete cosine transform value of the inter-frame compression trace image, determining a high-level code stream feature matrix, and determining a re-compression evidence obtaining result. According to the method, the problem that only intra-frame prediction information is considered and inter-frame prediction information is not considered in the prior art is solved, and compared with an existing video recompression evidence obtaining method, the evidence obtaining accuracy of the method is higher than 80%. According to the method disclosed by the invention, the complexity of evidence obtaining time is reduced, the evidence obtaining precision is obviously improved, the improvement is relatively obvious under a high code rate, and the method can be applied to the technical field of video recompression evidence obtaining.
Owner:XIAN UNIV OF POSTS & TELECOMM +1

High quality transcoding efficient texture format

Techniques for transcoding textures are described that include computing a plurality of endpoint colors of a macroblock. In a first technique, a forward discrete cosine transform (DCT) is applied (406) to a macroblock or portion thereof and used to represent the corresponding portion. In a second technique, a mean value of the endpoint colors and a difference between the endpoint colors are calculated (206), where the mean value and the difference establish (208) a projection vector in a color space. Means and differences are compressed (210), and per-pixel distances along the projection vector are calculated (212) and used together with corresponding mean and differences to represent macroblocks.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Low frequency non-separable transform and use of non-dct2 primary transform

PendingCN122514950AVideo encodingDiscrete cosine transform
Systems, methods, and means for video encoding / decoding are disclosed for the use of Low Frequency Inseparable Transform (LFNST) and Non-Discrete Cosine Transform 2 (DCT2) master transforms. In the examples, a video decoding device can determine whether to disable the Inseparable Master Transform (NSPT) for the current block. If NSPT is disabled for the current block, the video decoding device can determine to use LFNST for the current block. If LFNST is used, the video decoding device can decode the current block based on Multiple Transform Selection (MTS). In the examples, a video encoding device can disable NSPT for the current block. A video encoding device can enable LFNST for the current block. A video encoding device can encode the current block based on MTS.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

A multi-resolution based audio transform processing method and processing terminal

PendingCN122337216AHuman auditory systemEngineering
This invention discloses a multi-resolution audio transformation processing method, comprising the following steps: preprocessing and framing the obtained discrete time-series audio signal sequentially to obtain a framed audio signal; the preprocessing is used for filtering the audio signal; determining whether each audio frame in the framed audio signal belongs to a transient frame or a steady-state frame; performing wavelet packet decomposition on the audio frames belonging to transient frames to obtain a first type of audio frame; performing an improved discrete cosine transform on the audio frames belonging to steady-state frames to obtain a second type of audio frame; fusing the first type of audio frame and the second type of audio frame to obtain a fused audio signal; and non-uniformly dividing the time-frequency coefficients of the fused audio signal according to the Bark critical frequency band scale to obtain a divided audio signal. This invention can obtain an audio signal with good multi-scale time-frequency domain representation, which can obtain an audio signal that matches the frequency selectivity characteristics of the human auditory system.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

A lightweight sound detection system and method for dry-wood pests based on improved MobileNetV3

This application discloses a lightweight sound detection system and method for borers based on an improved MobileNetV3, comprising: a data preprocessing module for adaptive noise suppression, dynamic duration adjustment, and sampling rate optimization of the original audio signal to obtain a standardized audio signal; a spectrogram extraction and fusion module for performing short-time Fourier transform based on a bidirectional Mel filter bank to obtain a fused spectrogram; a lightweight feature extraction network for extracting multi-scale basic features and outputting an intermediate feature map; a multispectral channel attention module for extracting multi-band frequency features from the spectrogram through a two-dimensional discrete cosine transform, learning frequency weights through a multilayer perceptron and performing channel weighting to obtain a weighted feature map; and a multi-scale feature fusion module for unifying the channel dimensions of the feature map through a feature pyramid network and obtaining a fused feature map through bilinear interpolation upsampling, thereby predicting the probability and type of pest presence. This application can improve the real-time performance and accuracy of detection.
Owner:NORTHWEST A & F UNIV

Method and apparatus for multi-stage vector quantization for audio coding

PendingCN122290614ADiscrete cosine transformVector quantisation
Various embodiments of this disclosure disclose methods and apparatuses for multi-level vector quantization for audio coding. A method implemented by an encoder is disclosed. The method includes: obtaining a (1201) Discrete Cosine Transform (DCT) target vector; performing a (1203) suboptimal pairwise internal search in each segment of a codebook (104) having multiple segments to determine a set of pairwise initial candidates from each of the multiple segments, thereby forming a plurality of pairwise initial candidates from the suboptimal pairwise internal search, wherein each segment has a truncated vector different from the truncated vectors of the other segments in the plurality of segments, wherein the DCT target vector is the target of the suboptimal pairwise internal search for each of the multiple segments; reconstructing (1209) a plurality of final candidates using Inverse Type II Discrete Cosine Transform (DCT-II) to transform the plurality of final candidates into final candidate data in the original domain; and providing the final candidate data (1209) to a second stage of a multi-level vector quantizer.
Owner:TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Image processing and rock mine identification method based on high-frequency filtering and mineral composition attention

The invention provides an image processing and rock and ore identification method based on high-frequency filtering and mineral composition attention, and belongs to the technical field of rock and ore identification. According to the method, from the perspective of a frequency domain, discrete cosine transform is introduced into a CNN architecture, high-frequency features are explicitly enhanced, high-frequency information is extracted and enhanced more directly, and the influence of the smoothing effect of the CNN in layer-by-layer transmission on detailed information can be effectively relieved; a multi-attention mechanism in a channel dimension and a space dimension is adopted, so that the model can automatically judge which features are more important according to a currently input rock and mine image, and efficient self-adaptive optimization and extraction of a complex texture structure are realized. Experimental results show that the strategy of carrying out explicit high-frequency enhancement in the frequency domain and space-channel double attention cooperative work is adopted, so that the model shows higher recognition accuracy when facing actual rock and mine images with complex components, variable structures and strong background interference.
Owner:NORTHEASTERN UNIV CHINA +1