Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

172 results about "Discrete cosine transforms" patented technology

Intelligent discrimination method for pseudo soldering microcracks based on intelligent visual identification technology

The invention relates to an intelligent visual identification technology-based cold solder joint microcrack intelligent discrimination method, which comprises the steps of collecting an initial RGB image of a to-be-detected welding spot, carrying out two-dimensional discrete cosine transform and inverse two-dimensional discrete cosine transform on the initial RGB image to obtain an enhanced image, and fusing the enhanced image with an R channel of the initial RGB image to obtain a fused image; forming a dual-channel feature map; calculating the phase consistency of the dual-channel feature map, and obtaining a suspected candidate region of the pseudo soldering microcrack through an adaptive threshold segmentation method; acquiring an RGB image sequence of a continuous time sequence of the welding spots, and performing anomaly detection to obtain an abnormal region set; and constructing a welding spot thermal diffusion model, and inputting the geometric parameters and the environmental parameters in the abnormal region set into the welding spot thermal diffusion model to obtain a final judgment result of the pseudo soldering microcracks. According to the method, through multi-dimensional feature fusion and continuous time sequence dynamic tracking, the detection precision of the pseudo soldering microcracks is remarkably improved, the false detection rate is reduced, and the final judgment result is more accurate.
Owner:JUXIN ELECTRONICS TECH MEIZHOU CO LTD

Image tampering detection method and system based on mixed features and RGB features

The invention relates to the technical field of digital image security and authentic identification, and provides an image tampering detection method and system based on mixed features and RGB features, and the method comprises the steps: obtaining a to-be-detected input image, and carrying out the preprocessing of the to-be-detected input image; respectively extracting a Haar wavelet high-frequency component, a discrete cosine transform frequency domain feature and a Bayer convolution noise feature, and carrying out matrix level fusion to obtain a mixed feature; extracting RGB (Red, Green and Blue) features for the preprocessed input image; the mixed features are connected through cross-layer residual errors, and mixed feature learning features are obtained; and integrating the mixed feature learning features and the fused RGB features by using a cross-modal feature interaction architecture to obtain a prediction probability graph. Multi-modal features are fused, high-frequency response is enhanced, and the accuracy of image tampering detection is improved by adopting a dynamic fusion mechanism. The technical problems that an existing tampering detection method is insufficient in feature characterization capacity in a complex scene, low in tampering trace detection sensitivity and the like are solved.
Owner:SHANDONG UNIV

Enhanced systems and methods for synthetic aperture radar image compression with improved phase recovery and unwrapping

A system and method for compressing synthetic aperture radar (SAR) images with enhanced phase recovery and unwrapping capabilities is disclosed. The system performs preprocessing on input SAR images, applies discrete cosine transform (DCT) to create subbands, and utilizes a multi-pass amplitude compression technique. A specialized neural network performs phase unwrapping using compressed amplitude information and interferogram wrapped phase data. The system employs a channel-wise transformer fusion block (CTFB) for feature fusion and a multi-stage context recovery subsystem with optimized loss functions for both amplitude and phase recovery. The method achieves improved compression efficiency and phase recovery accuracy, particularly beneficial for Interferometric SAR (InSAR) applications.
Owner:ATOMBEAM TECH INC

Time-frequency dual-domain isolation time sequence anomaly detection method based on Mamba-self-attention

The invention discloses a time-frequency double-domain isolation time sequence anomaly detection method based on Mamba-self-attention, which can be applied to the fields of industrial manufacturing, medical equipment and the like and can detect anomaly by quantifying time-frequency difference. The method comprises the steps of extracting a multivariable time sequence sample from business data, and obtaining multivariable data through reversible instance normalization; the input time domain representation module is used for independently inputting a Mama network according to a natural time sequence and an inversion time sequence, capturing forward and reverse time features and fusing the forward and reverse time features into time features; a frequency domain representation module is input, seasonal variables are extracted, frequency features are extracted in combination with discrete cosine transform and an attention mechanism, and the frequency features are reconstructed to a time domain through inverse discrete cosine; the time and frequency characteristics are input into a time-frequency difference module, and the inconsistency is quantified through Kullback-Leible (KL) divergence so as to compare and learn a similarity loss function training model; and generating an anomaly score and setting a hyper-parameter to judge anomaly. The method is based on a bidirectional Mama and self-attention time-frequency double-domain isolation architecture, mode specificity discrimination features masked by traditional fusion are reserved, the consistency of normal mode domains is high, the correlation of abnormal performance is collapsed, and the detection accuracy and reliability are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Electronic document tracing method and device based on dynamic encryption and multi-modal watermark

The invention discloses an electronic document tracing method and device based on dynamic encryption and multi-modal watermarking, and relates to the technical field of information security and digital rights management. The method comprises the following steps: acquiring a PDF document uploaded by a user, encrypting the PDF document by adopting a randomly generated main encryption key, and acquiring and storing the encrypted PDF document; the content of the encrypted PDF document is analyzed, secret information is embedded into a picture in the document based on a two-dimensional discrete cosine transform algorithm, and picture watermark embedding is completed; a self-adaptive watermark embedding algorithm is adopted to embed secret information into a text in the document, and text watermark embedding is completed; related information of the checking operation is written into a block chain database; dynamically extracting watermark information from the divulged PDF document by adopting a watermark extraction algorithm; and querying a traceability database according to the extracted watermark information, and outputting a traceability result by comparing the extracted watermark information with information in a database storing various watermark information. According to the invention, the survival rate and the concealment of the watermark can be improved.
Owner:UNIV OF SCI & TECH BEIJING

Implementation method of Raman spectrum multi-component signal unmixing based on multi-modal time-frequency domain transformation and deep learning

According to the invention, the multi-mode time-frequency domain conversion and the deep learning technology are combined, and a multi-component mixed Raman spectrum unmixing method is developed, so that clinical in-vivo and in-situ detection and disease diagnosis of novel Raman probes, instruments and the like are facilitated. The method comprises the following steps: (1) converting a mixed Raman spectrum from a time domain to a frequency domain by using fast Fourier transform (FFT), discrete cosine transform (DCT) and discrete sine transform (DST), and extracting frequency domain features; (2) extracting multi-scale local time-frequency domain characteristics of the mixed spectrum by using short-time Fourier transform (STFT) and discrete wavelet transform (DWT); (3) carrying out spectral unmixing calculation in each mode in combination with a one-dimensional attention mechanism U-shaped neural network model; and (4) fusing various modal unmixing results by using a meta-learning method, and analyzing the weight of each modal to obtain an accurate unmixing spectrum. Compared with a traditional Raman spectrum analysis method, the Raman spectrum multi-component signal unmixing method based on multi-modal time-frequency domain transformation and deep learning can accurately separate independent Raman signals of different tissue structures and biochemical components in a complex environment in a living body, so that the unmixing accuracy of the Raman spectrum multi-component signal is improved. Therefore, convenience is provided for subsequent disease mechanism analysis and diagnosis. The method provides an innovative and potential solution for in-vivo and in-situ detection analysis and disease diagnosis of medical clinical Raman spectroscopy.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

High-fidelity anti-compression image watermarking method and system based on spectrum-airspace decoupling

The invention provides a high-fidelity anti-compression image watermarking method and system based on spectrum-airspace decoupling, and belongs to the field of information security. Firstly, the watermark information is mapped and remodeled; a watermark encoder based on multi-granularity spectrum-spatial domain feature decoupling is constructed, discrete cosine transform is introduced to filter out high-frequency components, Haar wavelet transform is adopted to realize lossless downsampling, a fast Fourier transform dynamic filter is combined to capture global semantic features, and local texture details are combined through multi-scale spatial domain volume accumulation; designing a physical perception and visual self-adaptive dual embedding strategy, and anchoring watermark energy to an anti-compression brightness channel; constructing an anti-attack layer containing differentiable JPEG compression simulation and mixed noise simulation, and participating in network training; and constructing a decoder and designing a loss function to carry out network optimization. According to the method, the robustness of the watermark under strong compression and complex black box attacks is improved, and extremely high visual imperceptibility is realized through physical and visual constraints.
Owner:NANJING UNIV OF INFORMATION SCI & TECH +1

DCT (Discrete Cosine Transform) frequency domain watermark embedding optimization method and system based on potential space

The invention discloses a DCT (Discrete Cosine Transform) frequency domain watermark embedding optimization method and system based on a potential space, and the method comprises the steps: firstly coding an original image into the potential space, segmenting the original image into small blocks, embedding a single-bit watermark into the frequency domain of each small block through weighted voting, then carrying out the inverse transformation to a spatial domain, merging, and forming a potential representation with a watermark; decoding through a stable diffusion model to generate an image with a watermark; iteratively optimizing the potential representation through a multi-objective loss function until the loss is minimum; and finally, decoding the optimized representation again, linearly mixing with the original image, adjusting a mixing coefficient through binary search, and outputting a final image on the premise of meeting preset image quality. According to the method, a multi-dimensional robustness optimization strategy is combined, the stability and the non-erasable property of the watermark under the attack of the generative model are systematically improved, and the core requirements of AI copyright protection and content traceability are met.
Owner:HANGZHOU DIANZI UNIV

Image compression device and image compression method

An image compression device includes a discrete cosine transform (DCT) circuit, a quantization noise shaping (QNS) circuit, and an encoder circuit. The DCT circuit performs a DCT on original image data to generate first data. The QNS circuit performs QNS on block data in first data to determine, based on a first coefficient and a second coefficient of the block data, a QNS score of the first coefficient, and replace the first coefficient with the second coefficient when the QNS score is greater than zero so as to generate second data, wherein the second coefficient is obtained by decreasing an absolute value of the first coefficient. The encoder circuit encodes the second data to generate compressed image data.
Owner:SIGMASTAR TECH LTD

Systems and methods for synthetic aperture radar image compression

For compressing synthetic aperture radar (SAR) images, preprocessing operations are performed on an input SAR image. A discrete cosine transform is performed on the image, and multiple subbands are created, where each subband represents a particular range of frequencies. The subbands are organized into multiple groups, where the multiple groups comprise a first low frequency group, a second low frequency group, and a high frequency group. A latent space representation is generated corresponding to each of the multiple groups of subbands. A first bitstream is created based on the latent space representation, and an alternate representation of the latent space is used for creating a second bitstream, enabling multiple-pass techniques for SAR image data compression, including phase unwrapping for supporting interferometric SAR (InSAR) applications.
Owner:ATOMBEAM TECH INC

Fresh tea leaf sorting method based on frequency domain tree topology network, computer equipment and computer readable medium

The invention discloses a fresh tea leaf sorting method based on a frequency domain tree topology network, computer equipment and a storage medium. The method covers the complete process of image acquisition, preprocessing, frequency domain decomposition, deep modeling, map construction and classification. Firstly, image quality is improved through color normalization and edge enhancement, and frequency domain tree decomposition is carried out through wavelet transform and discrete cosine transform to extract multi-scale features. And then, fusing long and short range dependent modeling and a residual convolution module to realize multi-level feature representation, and constructing a tree topology attention path and a structure map for simulating a bud-leaf-vein relationship to enhance semantic understanding. A tree structure is adopted to perceive a classification function, and fine-grained classification of single bud, one bud and one leaf, one bud and two leaves and one bud and multiple leaves is achieved. In training, the robustness of the model is improved by combining cross entropy loss, data enhancement and a regularization strategy. The method is high in classification accuracy and good in stability on a plurality of tea image data sets, and the practical level of automatic fresh tea leaf sorting is effectively improved.
Owner:JIANGXI ACAD OF AGRI SCI INST OF AGRI ENG

Adaptive real time discrete cosine transform image and video processing with convolutional neural network architecture

A system and method for real time discrete cosine transform image and video processing with convolutional neural network architecture. The system and method transforms degraded inputs into subband images, which are analyzed by a machine learning classification network to identify specific types of blur and compression artifacts. The classification network dynamically adjusts the parameters of separate DC and AC deblurring networks based on its analysis. This adaptive approach optimizes processing for various degradation types, improving the quality of the reconstructed output. The system and method's real-time capability and enhanced adaptability make it suitable for a wide range of imaging and video applications, offering superior performance over traditional methods.
Owner:ATOMBEAM TECH INC

Defense method for defending infrared detection confrontation sample attack

The invention discloses a defense method for resisting an infrared detection confrontation sample attack, and belongs to the technical field of calculation, reckoning or counting. The method comprises the following steps: firstly, partitioning an input infrared image according to a fixed size, executing two-dimensional discrete cosine transform, calculating the mean square difference of the image on each color channel before / after recompression, then converging into a frequency spectrum heat map, and binarizing to obtain a rough mask for positioning an adversarial patch; inputting the original image into the image segmentation model, calculating the intersection-union ratio of the segmentation mask and the rough mask, removing redundancy, and fusing to generate a final patch area mask; then, according to the patch area mask, original image pixels are removed, and a general image completion algorithm is called to recover a removed area; and finally, sending the complemented image into a deep learning target detector to realize robust detection under the infrared confrontation sample attack. According to the method, defense is carried out on adversarial sample attacks under infrared detection for the first time, and adversarial samples are effectively defended by adopting adversarial sample positioning, image segmentation and image restoration.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Method for detecting small defects of high-density integrated circuit packaging substrate

The invention discloses a method for detecting tiny defects of a high-density integrated circuit packaging substrate. The method comprises the following steps: acquiring image data of the surface of the high-density integrated circuit packaging substrate; processing the acquired images in sequence by adopting a mean filtering method, morphological dilation and a texture operator; processing the image by adopting an image binarization processing technology to obtain a binarized image; dividing the binary image into a plurality of energy regions, and calculating an entropy difference through the divided energy regions; performing processing operation on the image data through a discrete cosine transform technology to obtain a DCT coefficient, dividing the DCT coefficient into a high-frequency component and a low-frequency component, and performing calculation to obtain a ratio of a low-frequency coefficient mean value to a high-frequency coefficient mean value; the feature extraction model extracts defect shape features based on the entropy difference and the ratio; comparing a similarity measurement index calculated according to the defect shape characteristics of the packaging substrate with a preset threshold value, and judging whether the packaging substrate has a shape defect or not; and precise detection of small defects of the high-density integrated circuit packaging substrate is realized.
Owner:NANTONG UNIV

Urban water supply prediction method based on non-stationary perception Transform-BiLSTM model

The invention discloses an urban water supply prediction method based on a non-stationary perceptual Transform-BiLSTM model, and the method comprises the steps: independently calculating the mean value and variance of each sequence sample as non-stationary statistical information through introducing a reversible normalization mechanism, enabling a sequence to depend on weight learning reversible mapping on the basis of maintaining the original distribution characteristics of data, and carrying out the calculation of the mean value and variance of each sequence sample as the non-stationary statistical information; and the perception capability of non-stationary components is enhanced, and the training stability is improved. In addition, discrete cosine transform (DCT) is adopted for frequency domain modeling, and the feature extraction capacity of the model for the periodicity and trend of the urban water supply sequence is enhanced. In the aspect of feature modeling, the model fuses the global attention mechanism of Transform and the time sequence modeling capability of BiLSTM, cooperatively captures the long-term dependency relationship and local time sequence dependency features in urban water supply data, and has a remarkable capturing capability effect on the change of a non-stationary structure under the background of multi-scale fluctuation of an urban water supply sequence.
Owner:HENGYANG NORMAL UNIV

Target segmentation system, method and device based on attention mechanism

The invention discloses a target segmentation system, method and device based on an attention mechanism. The system comprises an image acquisition device for acquiring a CT scanning image to be segmented; the image processing equipment is used for segmenting an organ abnormal region by adopting an image segmentation model based on an attention mechanism; the image display equipment is used for displaying the CT scanning image and the mark information; the image segmentation model comprises an image enhancement network, which is used for enhancing boundary high-frequency information and texture details of input by using discrete cosine transform (DCT); the main encoder is used for extracting multi-scale global features and deep global semantic information of the CT scanning image; according to the brain area network, local differentiation features are extracted for the focus through a dynamic adaptation strategy by adopting a partition cooperation mechanism; and the decoder is used for recovering space structure and boundary details by adopting a bilinear interpolation up-sampling and multi-scale feature fusion mechanism on the global features and the local differentiation features. The technical problem that the accuracy of image segmentation of lesions such as lungs is poor in correlation is solved.
Owner:ROCKET FORCE UNIV OF ENG

Efficient transform signaling for small blocks

A size of a transform block is identified. A transform type for the transform block is identified. The transform type includes a horizontal transform type and a vertical transform type. Identifying the transform type includes determining whether the size of the transform block is below a predefined block size; and, in response to determining that the size is below the predefined block size, selecting a default transform type for each of the horizontal transform type and the vertical transform type. The transform type is then applied to the transform block. The default transform type can be the discrete cosine transform (DCT).
Owner:GOOGLE LLC

Cigarette cut tobacco production process quality prediction method based on DCTCN-informer-TSMIX model

The invention discloses a cigarette cut tobacco production process quality prediction method based on a DCTCN-informer-TSMIX model, and the method comprises the steps: collecting the time series data of a loosening and moisture regaining process production line of a cigarette cut tobacco production workshop at a preset time, and forming a sample data set; preprocessing the sample data set to obtain a processed data set; according to a time sequence, dividing the preprocessed data into a training set, a verification set and a test set for model training and performance evaluation; constructing a prediction model fusing a discrete cosine transform enhancement module (DCT), a time sequence convolutional network module (TCN), an Informer model and a time sequence mixer TSMixer; according to the data of the training set and the data of the verification set, the constructed DCTCN-Informer-TSMIX model is trained, and hyper-parameters are adjusted; and predicting the test set data / to-be-tested time sequence data of the loosening and moisture regaining process production line by using the trained model, and outputting a predicted value of the water content of the discharged material at a target time point. The model has the capabilities of frequency domain feature extraction, short-term time sequence modeling, long-term dependence capture and multi-dimensional feature interaction, can effectively improve the accuracy and the stability of prediction of the water content of the discharged material, and provides decision support for intelligent control and quality adjustment in the tobacco shred making process.
Owner:KUNMING UNIV OF SCI & TECH

Human body three-dimensional posture estimation method and system based on Transform and graph convolutional network

The invention belongs to the field of computer vision and artificial intelligence, and particularly discloses a human body three-dimensional posture estimation method and system based on Transform and a graph convolutional network, and the method comprises the steps: firstly receiving a two-dimensional skeleton sequence, sampling a center sub-sequence, and converting the two-dimensional skeleton sequence and the center sub-sequence into an initial spatial-temporal feature containing a spatial position code; through a double-flow parallel space coding module, global dependency features are obtained through a Transform encoder, and local connection features are obtained through a graph convolutional network encoder (based on a human skeleton topological adjacent matrix); discrete cosine transformation is carried out on the complete sequence to obtain a low-frequency coefficient, and frequency domain features are generated; and after the three types of features are spliced, deep fusion is carried out through a time-frequency fusion Transform encoder, and finally, the three-dimensional attitude of the target frame is output through a regression head. The method gives consideration to both efficiency and precision, is high in robustness, and is small in calculation cost increase.
Owner:国网四川省电力公司技能培训中心

Dental fluorosis visual detection image enhancement method

The invention provides a dental fluorosis visual detection image enhancement method, and relates to the technical field of image enhancement. The invention provides a dental fluorosis image data processing flow. The dental fluorosis image data processing flow comprises the steps of obtaining a dental fluorosis image data set, constructing a preliminary enhancement module, constructing a low-high frequency feature interaction module, constructing a deepening enhancement module and constructing a dental fluorosis image enhancement model. Wherein the preliminary enhancement module decomposes a dental fluorosis image into a low-frequency component and a high-frequency component through discrete wavelet transform and preliminarily enhances the dental fluorosis image in combination with discrete cosine transform and an attention mechanism, and the low-frequency and high-frequency feature interaction module is used for interacting features output by a low-frequency branch and a high-frequency branch in the preliminary enhancement module. And the deepening enhancement module utilizes the generative adversarial network to enhance the low-frequency and high-frequency features after interaction of the low-frequency and high-frequency feature interaction module, and fuses the low-frequency and high-frequency features to output a dental fluorosis enhanced image.
Owner:SHENYANG XINWEISHENGKE BIOTECHNOLOGY CO LTD +1

Image forgery detection method and device based on CLIP model

The invention discloses an image forgery detection method and device based on a CLIP model. The method comprises the steps that firstly, a multi-mode tampering data set containing multiple forgery types and corresponding text descriptions is constructed; the CLIP model is finely adjusted through LoRA by using the data set, so that the CLIP model learns pixel-level counterfeit features, and the characterization capability of the difference between a tampered image and a real image is enhanced; the method comprises the following steps of: decomposing low-frequency and high-frequency components of an input image by discrete cosine transform; respectively fusing the fine-tuned characteristic patterns of the CLIP visual encoder with low-frequency characteristics, and outputting a multi-category forged region positioning result through a decoder; meanwhile, the pixel-level splicing trace prediction image is fused with the high-frequency features, and a pixel-level splicing trace prediction image is output through a decoder. According to the method, the generalization ability of the model in a real scene is effectively improved, identification of specific counterfeiting means is realized, edge traces of splicing tampering are accurately revealed, and more comprehensive technical support is provided for image tampering evidence obtaining.
Owner:WUHAN UNIV

Deep learning-based transient thermal field inversion method and system

The invention discloses a transient thermal field inversion method and system based on deep learning, and the method comprises the steps: firstly obtaining thermal field evolution data and space grid coordinates of a measured object in a cooling process, and then designing a thermal field evolution network model composed of two sub-networks, the second sub-network outputs a deduced evolution thermal field, and deduces real material parameters of an object based on a finite difference deep learning model through joint optimization and fusion of internal and boundary thermal evolution equations as physical constraints so as to improve precision and robustness; on the basis, a time-containing inversion operator model is constructed, discrete cosine transform is utilized to map a thermal diffusion coefficient, a time step size and a final-state temperature field evolution sequence to a frequency domain, and accurate reasoning of an initial temperature field is realized through iterative optimization of an operator learning framework. The method has a good application prospect in material science and engineering detection, and can effectively assist in analysis and prediction of a complex heat conduction process.
Owner:ZHEJIANG UNIV

Image compression device and image compression method

An image compression device includes a discrete cosine transform (DCT) circuit, a quantization noise shaping (QNS) circuit, and an encoder circuit. The DCT circuit performs a DCT on original image data to generate first data. The QNS circuit performs QNS on block data in first data to determine, based on a first coefficient and a second coefficient of the block data, a QNS score of the first coefficient, and replace the first coefficient with the second coefficient when the QNS score is greater than zero so as to generate second data, wherein the second coefficient is obtained by decreasing an absolute value of the first coefficient. The encoder circuit encodes the second data to generate compressed image data.
Owner:SIGMASTAR TECH LTD

Mask-based lightweight pedestrian motion prediction method

The invention provides a light-weight pedestrian motion prediction method based on a mask, and is suitable for the technical field of human-computer interaction. According to the method, human body 3D skeleton point data is processed through space and time masks, key features are extracted by using a local perceptron constructed by a lightweight multilayer perceptron module and a cross-frame fusion device, the features are fused by using 1 * 1 convolution and splicing technologies, then a frequency-space domain pedestrian prediction action sequence is generated through a prediction module, and the frequency-space domain pedestrian prediction action sequence is obtained. And converting into a time-space domain sequence through an inverse discrete cosine converter. The method has the advantages of light model, quick response and suitability for real-time interaction. In addition, the strategy of gradually increasing the number of training frames improves the accuracy and stability of prediction.
Owner:SHENZHEN UNIV

Human eye fixation point prediction method based on deep convolutional network and frequency domain feature enhancement

The invention discloses a human eye fixation point prediction method based on a deep convolutional network and frequency domain feature enhancement, and the method comprises the steps: employing the deep convolutional network (namely a residual network ResNet) as a backbone network to form an encoder branch, and employing five layers of coding blocks to extract the spatial features of five layers of an input image; the five layers of spatial features extracted by the encoder are respectively sent to a frequency domain feature enhancement module for frequency domain feature enhancement based on discrete cosine transform (DCT) so as to enhance the features of each layer of human eye fixation area and reduce interference features; sending the frequency-domain-enhanced characteristics of each level into a decoding block of a corresponding layer of a decoder branch, and carrying out decoding and spatial up-sampling operation in sequence from a high layer to a low layer; and the output of the last decoding block of the decoder branch is subjected to 1 * 1 convolution and double up-sampling to obtain a human eye fixation point prediction result. According to the method, frequency domain feature enhancement and sequential decoding from a high layer to a low layer are carried out on the spatial multi-layer convolution features extracted by the backbone network, so that the human eye fixation point prediction precision is improved.
Owner:TONGDA COLLEGE OF NANJING UNIV OF POSTS & TELECOMM

Hardware-efficient disparity estimation using the DCT of interleaved images

A method for disparity estimation between digital images includes generating an interleaved image from two or more digital images with an offset between them, where the interleaved image is subdivided into a plurality of patches, computing discrete cosine transform (DCT) coefficients of each of the plurality of patches, computing, for each of the plurality of patches, a mean DCT descriptor from the DCT coefficients of each patch, and determining a disparity map from the mean DCT descriptor of each of the plurality of patches using a classifier. The disparity map is configured for real-time depth estimation from the two or more digital images.
Owner:SAMSUNG ELECTRONICS CO LTD

Part identification method and device based on industrial large model

The invention relates to the technical field of part identification, and discloses a part identification method and device based on an industrial large model, and the method comprises the steps: carrying out the image block division and embedded transformation of an original image of an industrial part, obtaining a first Token sequence, carrying out the two-dimensional discrete cosine transformation of the first Token sequence, and obtaining a second Token sequence; fusing the second Token sequence with a plurality of prototype feature vectors selected from a category prototype memory library to obtain a memory guide vector; performing gating modulation and residual connection processing on the second Token sequence based on the memory guide vector to obtain a third Token sequence; and performing iterative refinement on the third Token sequence to obtain a fourth Token sequence, and performing global pooling and classification prediction on the fourth Token sequence to obtain a category identification result of the industrial parts, thereby enhancing the discrimination of the part characteristics, and effectively solving the technical problem of difficult identification of small samples of rare parts.
Owner:SHENZHEN ANT FACTORY TECH CO LTD

Method and apparatus for switching between lossy codec and lossless codec

A method and an apparatus for switching between a lossy codec and a lossless codec are provided. The method includes: obtaining a waveform of a previous frame Ti−1, and updating the waveform of the previous frame Ti−1 to an input frame buffer in the lossless codec, where the waveform of the previous frame Ti−1 is encoded by a lossy encoder; performing integer window time domain aliasing cancellation INT winTDAC on a waveform in a lossless encoder buffer to obtain a first transform result, and updating the first transform result to an overlapped buffer in the lossless codec; obtaining a waveform of a current frame Ti, and updating the waveform of the current frame Ti to the input frame buffer; and performing integer modified discrete cosine transform INTMDCT on the waveform in the lossless encoder buffer to obtain a second transform result.
Owner:HUAWEI TECH CO LTD

Multi-modal picture understanding method and device based on discrete cosine transform

The embodiment of the invention discloses a multi-modal picture understanding method and device based on discrete cosine transform, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining an image and a text included in a multi-modal picture; performing discrete cosine transform on the image to obtain a frequency domain feature vector, and converting the frequency domain feature vector into a first visual token; obtaining a text token included in the text by using a text word segmentation device; inputting the text token and the first visual tokens into a Q-former module, and compressing the number of the first visual tokens to obtain a second visual token; and merging the second visual token and the text token, and inputting the merged second visual token and text token into a preset large language model to obtain text description information of the multi-modal picture. According to the embodiment of the invention, the fine-grained sensing capability of the high-resolution picture is improved, the number of visual tokens is reduced, and the computing resources of a large language model are saved.
Owner:ASIAINFO TECH CHINA INC