Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

788 results about "Upsampling" patented technology

In digital signal processing, upsampling, expansion, and interpolation are terms associated with the process of resampling in a multi-rate digital signal processing system. Upsampling can be synonymous with expansion, or it can describe an entire process of expansion and filtering (interpolation). When upsampling is performed on a sequence of samples of a signal or other continuous function, it produces an approximation of the sequence that would have been obtained by sampling the signal at a higher rate (or density, as in the case of a photograph). For example, if compact disc audio at 44,100 samples/second is upsampled by a factor of 5/4, the resulting sample-rate is 55,125.

System and methods for multimodal series transformation for optimal compressibility with neural upsampling

Image series transformation for optimal compressibility is performed with neural upsampling and error resilience. A novel correlation network composed of convolutional layers for feature extraction that extract multi-dimensional features from the image and a channel-wise transformer with attention to capture complex inter-channel dependencies. An angle optimizer enhances compressibility of an image and an error resilience subsystem improves robustness against transmission errors and data loss. The error resilience subsystem applies forward error correction coding, data partitioning based on importance, and embeds error concealment hints. This hybrid approach addresses both local and global features, mitigates compression artifacts, improves image quality, and enhances data integrity during transmission. The correlation network incorporates error correction and concealment techniques during decoding. The model's outputs enable effective image reconstruction, achieving advanced compression while preserving information for accurate analysis.
Owner:ATOMBEAM TECH INC

Remote sensing image change detection system and method based on multi-modal deep learning

The invention relates to a remote sensing image change detection system and method based on multi-modal deep learning. According to the system, the detection precision is improved by constructing a twin network fusing a CNN, a Transform and an attention mechanism; a lightweight backbone network is constructed by adopting multi-size depth separable convolution, and noise is adaptively suppressed and feature expression is enhanced in combination with a soft thresholding channel space attention mechanism; harr wavelet transform downsampling is innovatively introduced to reserve high-frequency details, and traditional pooling operation is replaced to reduce information loss; the twinborn branch semantic deviation is relieved through cross-channel feature exchange, and a Transform codec is used for modeling a long-range dependency relationship; in the decoding stage, edge feature recovery is enhanced by adopting feature splicing and a progressive up-sampling strategy; according to the system, model parameters are reduced, meanwhile, the precision and the anti-interference capability of remote sensing building change detection are remarkably improved, and the system has the advantages of high efficiency and detail keeping.
Owner:CHANGCHUN UNIV

Coal rock fracture intelligent extraction method based on improved U-Net

The invention discloses a coal rock fracture intelligent extraction method based on improved U-Net. The method comprises the following steps: S1, constructing a coal rock fracture CT image data set; s2, constructing an improved U-Net segmentation model, specifically comprising the following steps: S2.1, taking VGG16 as a backbone network, and introducing a depth separable convolution module; s2.2, a PPA attention module is added after each layer of depth separable convolution of the decoder, the PPA attention module is introduced after each up-sampling stage of the decoder, and the output of the PPA attention module is subjected to batch normalization and Dropout layer processing; s2.3, defining a composite loss function; s3, training and optimizing a segmentation model, wherein the specific steps comprise: S3.1, setting hyper-parameters; and S3.2, training the model by using the training set, adjusting hyper-parameters by using the verification set, and evaluating the performance by using the test set, wherein the evaluation indexes comprise MIoU, MAcc and FWIoU. According to the method, the problems of difficult identification of small fractures, large model calculation amount, poor multi-scale information fusion and class imbalance in the coal rock fracture image can be solved, and the robustness, segmentation precision and practicability of the model are improved.
Owner:CHINA UNIV OF MINING & TECH

Multi-modal diffusion-based long video role scene decoupling generation method and system

The invention discloses a long video role scene decoupling generation method and system based on multi-modal diffusion, and relates to the technical field of image processing, and the method comprises the steps: S1, synthesizing the advanced features of a role and a scene through a SigLIP encoder and a DINOv2 encoder; s2, performing cross-modal feature fusion on the advanced features to obtain joint features, and compressing the joint features to obtain compact vectors; s3, generating text features according to the text prompt; s4, potential codes are generated from an input video through a causal 3D convolution encoder, the potential codes pass through a linear projection matrix and then are spliced with a memory state for dimension reduction, and a segmented potential vector sequence is obtained; s5, performing decoupling perception generation on the segmented potential vector sequence through an improved 3D-UNet, and performing deconvolution up-sampling reconstruction after deterministic sampling to obtain an RGB video segmented sequence; according to the method, the key problems of rough dynamic control, limited generation length and over-high resource consumption in long video generation are solved, and the quality and efficiency of the generated video are remarkably improved.
Owner:湖南马栏山视频先进技术研究院有限公司

Unified system for multi-modal data compression with relationship preservation and neural reconstruction

A unified platform for multi-modal data compression and decompression that enables efficient processing of correlated data streams while preserving relationships between different modalities. The platform employs a virtual management layer to analyze and route input streams, implementing correlation analysis to identify temporal and spatial relationships between streams. Multiple compression methods, including neural network-based approaches, are utilized to compress data sets while maintaining cross-modal dependencies. A neural upsampling system leverages learned correlations between streams to enhance reconstruction quality. The platform includes a synchronization manager that maintains temporal alignment and relationship preservation throughout processing. By integrating correlation-aware compression with neural upsampling techniques, the platform provides comprehensive multi-modal compression capabilities while preserving critical relationships between different data types. The system is particularly suited for applications involving synchronized audio-visual data, sensor streams, and other multi-modal content.
Owner:ATOMBEAM TECH INC

System and Methods for Upsampling of Decompressed Audio Data Using a Neural Network

A computer system for upsampling decompressed audio data after lossy compression using specialized neural network techniques. The system processes compressed audio channels through an audio pre-processor that extracts spectral information, detects speech activity, segments audio, and normalizes input levels. A trained deep learning algorithm with multi-channel transformers using channel-wise and self-attention mechanisms recovers information lost during compression. The system further enhances audio quality through a time-frequency domain transformer applying Fourier transforms and Mel-scale frequency processing, while a perceptual quality assessor employing psychoacoustic models evaluates the output. This specialized audio processing approach significantly improves reconstructed audio quality by leveraging correlations between audio channels, addressing both spectral and temporal features, and optimizing for human perception characteristics, resulting in higher fidelity audio reproduction from compressed formats.
Owner:ATOMBEAM TECH INC

Method for constructing remote sensing image defogging network based on wavelet frequency domain heterogeneous enhancement

The invention discloses a method for constructing a remote sensing image defogging network based on wavelet frequency domain heterogeneous enhancement, the network adopts a U-shaped architecture as a basic framework, the network input is a foggy image, and firstly, shallow layer features are extracted through a convolution block; then, a symmetric codec structure is adopted to learn layered representation, a codec comprises up and down sampling and a wavelet frequency domain heterogeneous enhancement module, and the resolution of up and down sampling is controlled through step convolution and deconvolution; the wavelet frequency domain heterogeneous enhancement module separates the high and low frequency features of the image through discrete wavelet transform, and performs heterogeneous enhancement on the separated high and low frequency features by combining the dynamic receptive field advantage of deformable convolution and the global perception capability of Fourier transform; therefore, the recovery of high-frequency local texture details and the removal of low-frequency global haze are effectively promoted. And finally, reconstructing a clear fogless image through the convolution block. According to the research algorithm, the texture features and natural colors of the scene can be precisely reduced.
Owner:CHINA THREE GORGES UNIV

Method and system for controlling rising and falling of PCM (Pulse Code Modulation) audio sampling rate

The invention provides a rising and falling control method and system for a PCM audio sampling rate, and belongs to the technical field of audio processing, and the method comprises the following steps: obtaining audio signal features of pulse code modulation PCM audio data, the audio signal features comprising spectrum energy distribution, instantaneous power, a signal-to-noise ratio and a peak-to-average ratio; monitoring hardware state parameters of the Bluetooth audio equipment in real time, wherein the hardware state parameters comprise a current representative processor load rate, a bandwidth occupancy rate and a current coding and decoding mode parameter; according to the predicted value of the target sampling rate, performing resampling processing on the PCM audio data by adopting a filtering processing method and an interpolation algorithm to obtain PCM audio data matched with the target sampling rate; the resampling processing comprises up-sampling processing or down-sampling processing; and inputting the PCM audio data matched with the target sampling rate to an audio output module of the Bluetooth audio equipment for playing. The audio quality, the computing resources and the transmission efficiency are optimized through self-adaptive resampling processing.
Owner:SHENZHEN HAILINGWEI ELECTRONICS CO LTD

Target detection network for small target and training method and detection method thereof

The invention provides a small target-oriented target detection network and a training method and a detection method thereof. The small target-oriented target detection network comprises an encoder structure and a decoder structure, the encoder structure comprises a double-branch hybrid encoder, wherein the double-branch hybrid encoder comprises a Transform encoder branch and a CNN encoder branch; the Transform encoder branch comprises a plurality of Transform modules, and the CNN encoder branch comprises a plurality of first cascade convolution residual blocks; an interactive fusion module is arranged aiming at each group of Transform modules and the first cascade convolution residual error blocks corresponding to the positions; the interactive fusion module is used for receiving the first processing results output by the corresponding groups and fusing the first processing results; the obtained fusion result and the first processing result are input into a next adjacent group of Transform modules and a first cascade convolution residual block; and the decoder structure is used for performing up-sampling and restoration reconstruction on the features extracted and fused by the encoder structure, and outputting a detection result of the small target in the image. And the detected small target is more accurate.
Owner:NAT SPACE SCI CENT CAS

Super-resolution remote sensing image reconstruction method, system and device, and storage medium

The invention relates to the technical field of remote sensing image data processing, in particular to a super-resolution remote sensing image reconstruction method, system and device and a storage medium, and the method specifically comprises the steps: firstly extracting shallow layer features of a low-resolution remote sensing image, and processing the shallow layer features in two paths: branch 1: channel expansion reactivation, and outputting a first branch feature; and branch 2: after channel expansion, depth separable convolution reactivation processing is carried out, and second branch features are output through snakelike scanning. The two branches are fused into a snakelike feature, are connected with a shallow feature in a jumping manner, then are subjected to channel attention processing, and then are connected in a self-jumping manner to obtain a second connection feature. Grouping is carried out according to channels, multi-scale local features are generated through multiple times of neighborhood attention processing and splicing, and deep features are obtained after iteration. And finally connecting shallow and deep features in a jumping manner, and performing up-sampling to output a reconstructed image. According to the method, the high efficiency of snakelike visual state space processing and the excellent long-distance dependence modeling capability are fully exerted, and efficient and accurate sequence and space feature fusion is realized.
Owner:YANTAI UNIV

Video snapshot compression imaging reconstruction method based on space-time deformable attention

The invention provides a video snapshot compression imaging reconstruction method based on spatio-temporal deformable attention, which improves the reconstruction quality and efficiency, and comprises the following steps: inputting a single frame compression measurement value and a measurement matrix into an initial reconstruction module to obtain an initial reconstruction video frame; inputting the initial reconstructed video frame into a feature extraction encoder, mapping the initial reconstructed video frame to a high-dimensional feature space through multi-layer 3D convolution, and outputting a feature map; the feature map is input into a plurality of stacked DenseRNet Blocks, and the number of the DenseRNet Blocks is one; the DenseRNet Block internally comprises a plurality of DeT Blocks, after the DenseRNet Block dynamically divides an input feature channel, grouping progressive processing and feature fusion are carried out through the plurality of DeT Blocks, and the DeT Blocks comprise a deformable space convolution branch used for modeling local deformation perception, a time self-attention branch used for modeling global time sequence dependence and a feature interaction module used for cross-channel information interaction; and the features processed by the DenseRNet Block are input into a video reconstruction decoder, and a reconstructed video sequence is output through up-sampling of transposition convolution and refining of multilayer 3D convolution.
Owner:DALIAN UNIV

Pipeline defect magnetic flux leakage detection method and device based on multi-scale data driving deep learning

The invention discloses a pipeline defect magnetic flux leakage detection method and device based on multi-scale data driving deep learning, and relates to the technical field of pipeline defect detection. A multi-scale magnetic flux leakage signal data set containing defect global distribution and local details is generated, and multi-level primary features are automatically extracted by using a convolutional neural network; and the defects of poor generalization and easy key information omission of a manual method are overcome. And then dynamically enhancing and performing weighted fusion on multi-scale features by means of a multi-scale convolution branch and an attention mechanism, so that the model can learn global and local features of the defect at the same time, the problem that the global and local features of the defect are difficult to consider in traditional deep learning is solved, and through transverse connection and up-sampling fusion of a feature pyramid, a multi-scale feature is obtained. And the features have high semantic information and high spatial details. And finally, independently carrying out multi-task prediction by virtue of a task decoupling module, and calibrating result consistency by virtue of a task alignment module, so that the detection accuracy in a complex scene is improved, and more accurate and robust pipeline defect magnetic flux leakage detection is realized.
Owner:NORTHEASTERN UNIV CHINA

Mixed CNN-Transform colon polyp image segmentation method combining edge guidance and double attention mechanism

The invention provides a hybrid CNN-Transform colon polyp image segmentation method combining edge guidance and a double attention mechanism, and is applied to the technical field of medical data processing. According to the method, a CNN encoder is adopted to extract multi-scale local features, a Transform encoder is adopted to extract global context, an edge probability graph is generated by introducing an edge guide branch based on shallow layer features, a gating coefficient is calculated according to the edge probability graph and / or statistics obtained by intermediate prediction, weighted fusion is performed on the local and global features, and the local feature and the global feature are integrated. And the decoder performs up-sampling step by step and outputs a segmentation result. Compared with the prior art, the method has the advantages in boundary integrity and small target detection. Experiments show that on a Kvasair-SEG data set, the optimal Dice of a verification set of the scheme is 0.892, the optimal HD95 of the verification set of the scheme is 12.2, the Dice of a test set of the scheme is 0.89, and the optimal HD95 of the test set of the scheme of the scheme is 12.3.
Owner:CIXI PEOPLES HOSPITAL MEDICAL HEALTH GRP (CIXI PEOPLES HOSPITAL) +1

Voice signal compression method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice signal compression method, device, equipment and medium, which comprises the following steps: executing Fourier transform on an initial voice signal to extract an amplitude spectrum and a phase spectrum, respectively processing the amplitude spectrum and the phase spectrum to generate features, and splicing the features to obtain a compressed voice signal; the method comprises the following steps: carrying out quantization by adopting a residual vectorization mode to generate a compressed feature vector, executing entropy coding to obtain a compressed code stream, recovering the compressed feature vector at a decoding side, reconstructing a spliced feature through residual reverse quantization, enhancing feature expression through up-sampling operation, and restoring a voice signal through inverse Fourier transform. According to the method, a dual-path processing structure of amplitude features and phase features is constructed, feature compression is realized in combination with residual vectorization and entropy coding, inverse quantization and up-sampling enhancement are executed in a decoding stage, a spectrum energy structure and phase continuity are effectively reserved, and reconstruction precision and fidelity of voice signals are improved under the condition of low bit rate.
Owner:PING AN TECH (SHENZHEN) CO LTD

Mass spectrum peak identification method and device, computer equipment and storage medium

The invention provides a mass spectrum peak identification method and device based on deep learning, computer equipment and a storage medium. The method comprises the following steps: inputting a mass spectrum, adjusting the channel number through a convolutional layer, then performing convolution and down-sampling operation through a multi-layer encoder to extract image features, extracting key information from a generated high-order feature map through a bottleneck module, and inputting the key information into a decoder; a decoder gradually recovers a feature map to an original size through up-sampling, cross-layer feature fusion is carried out on the decoder and a corresponding layer of the encoder through a jump connection channel, and a channel space attention module is embedded in a connection path to strengthen attention to important features while detail information is reserved; and finally, outputting a segmentation result through point convolution. According to the method, a feature fusion strategy of multi-level feature multiplexing and channel space attention enhancement of a coding and decoding structure and a lightweight network architecture considering precision and speed are adopted, so that the recognition effect is ensured, and meanwhile, the operation efficiency is optimized.
Owner:SHENZHEN UNIV

Screen-shooting-resistant robust image watermark soft fusion network method and system based on UNet architecture

The invention provides an anti-screen-shooting robust image watermark soft fusion network method and system based on a UNet architecture, and the method comprises the steps: 1, carrying out the preprocessing of a cover image, and obtaining a low-frequency component; 2, establishing a multi-stage watermark extension sub-network, and carrying out feature fusion on a watermark information tensor and a low-frequency component; step 3, constructing a watermark soft fusion network ProwerNet based on a Unet architecture: fusing frequency domain information and watermark information embedded with watermark features with features of a jump link and features after continuous up-sampling to realize joint embedding of a space domain and a frequency domain; and step 4, designing a non-micro noise simulation training process, and training the watermark soft fusion network ProwerNet. According to the method, the ProwerNet is used for processing the watermark information and the carrier image, so that multi-scale cross-domain progressive embedding is realized.
Owner:NANJING UNIV OF INFORMATION SCI & TECH +1

Food image denoising method based on dynamic coding and efficient channel perception hybrid upsampling

The invention discloses a food image denoising method based on dynamic coding and efficient channel perception hybrid upsampling, and the method is characterized in that a dynamic coding module CAMixer and a channel perception hybrid upsampling module E-CAMixUp are cooperatively integrated, and an efficient channel perception hybrid denoising network ECAMixDNet is formed. According to the method, the CAMixer utilizes a learnable attention mechanism to adaptively adjust attention so as to contain more useful textures, the representation capability of convolution is improved, and the E-CAMixUp realizes high-quality image reconstruction under multi-scale noise perception through the collaborative design of dual-path feature reconstruction and adaptive channel interactive filtering, so that the image reconstruction efficiency is improved. The problem that artifacts are easily introduced in the reconstruction stage by an inter-channel noise distribution difference suppression mechanism is solved. The network extracts multi-scale features through a Swin Transform and an RBF attention mechanism, a decoder is combined with a residual module to realize high-quality reconstruction, food texture and color sensitivity are effectively kept, and visual restoration quality is improved.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Transform coding based on matrix-based intra prediction

Devices, systems and methods for digital video coding, which includes matrix-based intra prediction methods for video coding, are described. In a representative aspect, a method for video processing includes performing a conversion between a current video block of a video and a bitstream representation of the current video block according to a rule, where the rule specifies a relationship between applicability of a matrix based intra prediction (MIP) mode or a transform mode during the conversion, where the MIP mode includes determining a prediction block of the current video block by performing, on previously coded samples of the video, a boundary downsampling operation, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation, and where the transform mode specifies use of a transform operation for the determining the prediction block for the current video block.
Owner:BYTEDANCE INC +1

Offshore wind power GIS equipment partial discharge mode identification method based on improved SAE network

The invention provides an offshore wind power GIS equipment partial discharge mode identification method based on an improved SAE network, and belongs to the technical field of electrical digital data processing.The method includes the steps that a partial discharge experiment platform meeting the IEC60270 standard is constructed, four typical defect models are designed, and PRPD spectrograms are collected; hSV color space threshold segmentation is utilized to remove interference elements, and image features are enhanced in combination with a partial discharge physical mechanism equation. The improved SAE network integrates a three-layer convolution encoder and an up-sampling decoder, introduces an ECA attention mechanism to optimize a feature channel weight, and combines a pre-trained # imgabs0 # PDCN model to enhance the feature extraction capability. Physical parameter constraints are introduced in the training process, and after global pooling dimensionality reduction and full connection layer processing, high-precision identification of four defect types of point discharge, creeping discharge, air gap discharge and suspended metal discharge is realized, and the technical problems of low accuracy and poor generalization ability of offshore wind power GIS equipment partial discharge mode identification are solved.
Owner:XI AN JIAOTONG UNIV

Video snapshot compression imaging reconstruction method and system

The invention relates to a video snapshot compression imaging reconstruction method and system. The method comprises the following steps: inputting a video frame sequence and a time-varying mask set thereof into a measurement model to obtain initial estimation; constructing a reconstruction network which comprises a feature extraction module, a gating residual network module and a video reconstruction module; the feature extraction module comprises two three-dimensional convolution layers, each three-dimensional convolution layer is connected with an activation function, and the feature extraction module extracts initial features from the initial estimation; inputting the initial features into a gating residual network module, and outputting reconstruction information features; and the video reconstruction module fuses the reconstruction information features, and performs up-sampling and detail refining to reconstruct a video sequence. According to the method, on the premise that parameters and computing power are hardly increased, ghosting and flickering are effectively restrained, the stability of long-time reconstruction is improved, and an effective scheme is provided for SCI reconstruction with the high compression ratio, the super-definition resolution ratio and the long sequence.
Owner:GUANGDONG UNIV OF TECH

Full-parallel arbitrary multiplying power variable sampling system

The invention provides a full-parallel arbitrary multiplying power variable sampling system, and belongs to the technical field of signal processing. The system comprises an up-sampling system and a down-sampling system, the up-sampling system is used for a transmitter and comprises an up-sampling filter and an up-sampling interpolator, and the down-sampling system is used for a receiver and comprises a down-sampling filter, an extractor and a down-sampling interpolator; the up-sampling system can process four paths of input signals in parallel, and four paths of output signals are obtained after the four paths of input signals are processed in parallel; the down-sampling system is formed by cascading a plurality of down-sampling filters and an extractor, an extraction module is responsible for down-sampling of integer 2n times, and a decimal interpolator is connected behind the extractor to realize down-sampling processing of any decimal times. According to the invention, the limitation of traditional integer or simple fractional multiple sampling is broken through, arbitrary-multiplying-power precise resampling of the GHz-level signal is realized, the processing speed and the interpolation precision are obviously improved, the error is low, the anti-aliasing capability is strong, and the method is suitable for the high-requirement fields such as private network communication.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Voice synthesis method and system based on VITS improvement

The invention provides a voice synthesis method and system based on VITS improvement, and the method comprises the steps: optimizing a text encoder of a VITS model, introducing a large language model, and enabling the emotion, intention and speaking style of an input text to be captured when the text is encoded; a random disturbance item is introduced when the Q value is dynamically planned and solved, the alignment flexibility in the initial training stage is improved, meanwhile, monotonicity constraint is strictly kept, and it is avoided that a suboptimal solution is obtained through convergence too early; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is utilized to efficiently reconstruct a time domain signal, so that waveform up-sampling is realized, redundant calculation of traditional transpose convolution is avoided, and reasoning is accelerated. According to the method, the reasoning speed, the emotion expression ability and the style control flexibility of speech synthesis can be effectively improved, a new solution is provided for cross-language diversified speech synthesis, and a reference is provided for the more efficient and more intelligent development of the speech synthesis technology.
Owner:豫章师范学院

Feature hierarchical attention fusion and dynamic optimization method for single-stage target detection

The invention discloses a feature hierarchical attention fusion and dynamic optimization method for single-stage target detection, and relates to the technical field of computer vision. The method comprises the steps of collecting image data, preprocessing an image and dividing the image into a training set, a test set and a verification set; the HAF-DRFPN neck network uses dynamic up-sampling single pixel point sampling to recover the feature resolution; feature extraction is optimized by using a gating residual mechanism and depth separable convolution; a feature refining feed-forward network is introduced through multi-layer perception generation weight, and feature expression is enhanced; respectively capturing and fusing superficial details and deep semantic information by utilizing channel and space attention and cross attention; a multi-branch dynamic sampling stacking framework is adopted, and cross-scale features are adaptively fused and optimized; multi-module cooperative processing is carried out on the whole architecture, and multi-scale features with higher discrimination are provided for the detection head; according to the method, the YOLO12 trunk and the detection head are connected with the HAF-DRFPN to form the model HAF-DRNet suitable for target detection, so that the flexibility and accuracy of single-stage target detection are improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Industrial part abnormal target detection method and system

The invention provides an industrial part abnormal target detection method and system. According to the model, a YOLOv10 framework is used as a basic framework to construct a lightweight model, and the detection performance and efficiency are improved through three-stage optimization. Firstly, a multi-head self-attention mechanism is introduced to reconstruct a feature space, and cross-modal feature interaction is promoted by using CSP structure balance calculation efficiency and feature representation capability and combining a channel grouping shuffling strategy. Secondly, designing a context guide perception module in a feature fusion stage, enhancing multi-scale feature expression through a parallel multi-branch architecture and a spatial self-calibration mechanism, and enhancing up-sampling information density in cooperation with a dynamic interpolation fusion module; the detection head adopts a parameter sharing group convolution structure, and the calculation amount is reduced through a convolution kernel parameter sharing and feature decoupling mechanism. The method effectively solves key problems in industrial part detection, and is of great significance to industrial part automatic anomaly detection scenes on an industrial production line.
Owner:HANGZHOU DIANZI UNIV

Image sonar small target detection method and system based on improved attention mechanism

The invention provides an image sonar small target detection method and system based on an improved attention mechanism, and the method is realized based on YOLOv8, and the method comprises the steps: inserting the following process in the tenth layer of a backbone network of YOLOv8: carrying out the up-sampling of an original image, sequentially carrying out the 5 * 5 point convolution, vertical convolution and horizontal convolution, and finally obtaining data F'through an activation function; performing down-sampling on an original image, then performing 1 * 1 point convolution, horizontal convolution and vertical convolution in sequence, and finally obtaining data F ''through an activation function; and carrying out dot product on the data F ', the data F' 'and the original image data, and splicing with the original image data to obtain a final output result. Compared with the prior art, the method has the advantages that the model learns image features of different dimensions through up-down sampling, convergence is accelerated through aggregation in the horizontal direction and the vertical direction and finally residual connection, a more refined result can be output, background interference is reduced, and calculation overhead is reduced.
Owner:INST OF ACOUSTICS CHINESE ACAD OF SCI

Method for detecting low-confidence small target in radar echo based on hybrid architecture

The invention belongs to the technical field of radar signal processing, and particularly relates to a low-confidence small target detection method in radar echoes based on a hybrid architecture, and the method comprises the steps: firstly carrying out the spectrum symmetric movement and dimension recombination of radar echo data, and generating five-dimensional tensors [B, T, C, H, W] containing time sequence features; then the tensor is input into a detection model formed by cascading a Hurglass 3D module and a YOLOv8 network, the Hurglass 3D module extracts multi-scale spatial-temporal features through a structure of three-dimensional convolution down-sampling, bottleneck layer and three-dimensional transposition convolution up-sampling, and feature fusion is achieved through jump connection; and finally, target detection is completed through a backbone network, a neck network and a decoupling detection head of the YOLOv8 network. According to the invention, through spatio-temporal feature combined extraction and small target feature enhancement, the detection accuracy and the positioning precision of the low signal-to-noise ratio small target in radar echoes are effectively improved.
Owner:ANHUI UNIV

Data transmission method, data modulation method, and electronic device and storage medium

A data transmission method includes transmitting to-be-transmitted data in N frequency domain resource blocks, where each of the N frequency domain resource blocks includes K(n) subcarriers, where n=1, 2, . . . , N, N is greater than or equal to 1, and K(n) is greater than or equal to 1; performing inverse Fourier transform and an upsampling operation on the to-be-transmitted data in each of the N frequency domain resource blocks to form N groups of data sequences; and transmitting the N groups of data sequences.
Owner:ZTE CORP

Super-resolution remote sensing image reconstruction method, system and equipment based on frequency domain enhancement

The invention belongs to the technical field of image data processing, and particularly relates to a super-resolution remote sensing image reconstruction method, system and device based on frequency domain enhancement, and the method comprises the steps: S1, extracting the shallow features of a low-resolution remote sensing image; s2, inputting the shallow features into a plurality of cascaded frequencies for interactive processing, performing double-branch processing on the input features, performing inverse transformation after radial weighting on different frequency components in a frequency domain to obtain first features, and obtaining second features through depth separable convolution, an activation function, a selection scanning module and layer normalization; fusing the two branch features according to the weight, performing jump connection with the original features, performing enhancement through a feedforward network, performing repeated execution for a set number of times, and performing convolution and residual connection to obtain deep features; and S3, fusing the deep features and the shallow features, and outputting a super-resolution image through convolution and pixel rearrangement up-sampling. According to the method, texture details and edge contours of the remote sensing image can be more accurately reconstructed while the structural consistency is kept.
Owner:YANTAI UNIV

Method of video encoding and system for video encoding

A method of video encoding is provided. The method may include receiving, by a head portion of a Lightweight Multi-level mixed Scale and Depth information with Attention mechanism (LMSDA) network, an input image. The method may include extracting, by the head portion of the LMSDA network, a first set of features from the input image. The method may include inputting, by a backbone portion of the LMSDA network, the first set of features through a plurality of LMSDA blocks (LMSDABs). The method may include generating, by the backbone portion of the LMSDA network, a second set of features based on an output of the LMSDABs. The method may include upsampling, by a reconstruction portion of the LMSDA network, the second set of features to generate an enhanced output image.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

A deep learning method for low-frequency SKA broadband effect and synthetic beam effect elimination

ActiveCN119295329BImage enhancementImage analysisAstronomical image processingImaging processing
The application discloses a kind of deep learning methods for low-frequency SKA broadband effect and synthetic beam effect elimination, belong to radio astronomy image processing field, including steps: S1, establish FSAS, FERM and FGFN;S2, based on the FSAS, FERM and FGFN of S1 establishment establishes IFS-Transformer network model;S3, two FERM modules are applied to obtain dirty image I ∈ R H×W×C Low-level features F0 ∈ R H×W×C ;S4, F0 is input FGFN module, and coding part is completed after twice downsampling operation;S5, the feature obtained by coding part is passed through 3 FSAS and twice upsampling, realizes feature learning and decoding operation;S6, after being processed by two FERM, final recovered image is obtained.The effect elimination method based on deep learning can more effectively eliminate the joint of broadband effect and synthetic beam effect, and can greatly reduce the time of effect elimination to a greater extent to restore and reconstruct original sky brightness;The manual operation process of the method of the patent can be greatly simplified by using deep learning technology.
Owner:GUIZHOU UNIV