Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Receptive field" patented technology

A sensory space can be the space surrounding an animal, such as an area of auditory space that is fixed in a reference system based on the ears but that moves with the animal as it moves (the space inside the ears), or in a fixed location in space that is largely independent of the animal's location (place cells). Receptive fields have been identified for neurons of the auditory system, the somatosensory system, and the visual system.

Mixed audio track separation method, system and device and storage medium

The invention relates to the technical field of audio signal processing and artificial intelligence crossing, in particular to a mixed audio track separation method, system and device and a storage medium. The mixed audio track separation method comprises the following steps: extracting time domain features of a mixed audio signal through adaptive receptive field convolution, and extracting frequency domain features of the mixed audio signal through an adaptive spectrum window to obtain output features; generating enhanced features based on the output features, a pre-constructed frequency domain mask of a to-be-separated target musical instrument and feature channel attention; and based on the enhanced features, adaptive waveform reconstruction is carried out through a joint decoder, and separated independent audio tracks are obtained. According to the invention, the efficiency and accuracy of mixed audio track separation can be improved.
Owner:THINKING CHAIN (TIANJIN) INTELLIGENT TECH CO LTD

An artificial visual nervous system based on dual-mode neural devices

ActiveCN120471114BPhysical realisationSynapseOptic nerve
The application relates to an artificial visual nervous system based on a bimodal neural device, comprising a bimodal visual information calculation array composed of a physical convolution kernel array and a synapse calculation array, the physical convolution kernel array is used for receiving an ambient light signal and converting into an electrical signal output, the synapse calculation array is used for receiving the electrical signal and converting into a visual signal, and the physical convolution kernel array and the synapse calculation array are both composed of a plurality of bimodal transistors. The application has excellent bimodal collaborative calculation capability, the receptive field mechanism (excitatory / inhibitory response) of a retinal bipolar cell is simulated through the physical convolution kernel array, image feature extraction and preprocessing are realized, the synaptic plasticity (LTP / LTD mechanism) of a central nervous system is simulated in combination with the synapse calculation array, high-level calculation of the visual signal is completed, and a complete visual nerve bionic system is formed by virtue of an array composed of single devices.
Owner:XIAN JIAOTONG LIVERPOOL UNIV

An acoustic event detection method based on multi-scale spatial feature and coordinate attention fusion

This invention discloses a sound event detection method based on the fusion of multi-scale spatial features and coordinate attention, belonging to the field of sound event detection technology. It aims to address the problems in existing sound event detection methods, such as the limited receptive field of the network, which makes it difficult to capture multi-scale temporal span features, and the inability to accurately focus on target sounds and suppress noise in complex time-frequency spaces. This method innovatively introduces a multi-scale coordinate attention module into the sound event feature extraction network. This module first uses parallel dilated convolutions with an increasing time dilation rate to obtain multi-scale temporal features and performs fusion and dimensionality reduction. Then, it employs a dual-axis time-frequency decoupling mechanism, performing adaptive pooling aggregation along both the time and frequency dimensions to generate directional time attention weights and frequency attention weights. Finally, the dual-axis weights are used to recalibrate the multi-scale features. This method significantly improves the accuracy and robustness of sound event detection.
Owner:GUILIN UNIV OF ELECTRONIC TECH

H-beam surface defect detection method and system

The application provides a H-shaped steel surface defect detection method and system, and belongs to the technical field of profile steel detection. The detection method comprises the following steps: constructing a target area recommendation network based on a visual receptive field, and integrating a YOLOv3 target detection model; performing lightweight processing on the YOLOv3 target detection model by using a depth separable convolution; constructing a multi-scale spatial attention model and a multi-scale channel attention model based on a visual attention mechanism, cascading the multi-scale spatial attention model and the multi-scale channel attention model, and obtaining a double attention model based on the YOLOv3 target detection model; obtaining a surface image of the H-shaped steel, inputting the surface image into the YOLOv3 target detection model, and obtaining a defect detection result through lightweight processing and double attention model processing; and displaying the defect detection result, adjusting a running mode, and displaying a running state. In this way, the detection speed and the detection efficiency are improved, and multiple types of different scale defects can be quickly and accurately detected.
Owner:TIANJIN UNIV

A brain intracranial cavity segmentation method and device based on a U_Net model

The application discloses a brain intracranial cavity segmentation method and device based on a U_Net model, and relates to the technical field of picture semantic segmentation.The application constructs a multistage down-sampling encoder and a corresponding up-sampling decoder in a U-shaped structure, adds an efficient attention mechanism to improve segmentation accuracy, and utilizes a multistage feature fusion module to fuse features; and the optimal segmentation model is trained on a training set.The down-sampling encoder fused with a residual block avoids the gradient vanishing phenomenon caused by excessively deep network, and the attention mechanism is introduced in the middle layer between the down-sampling and the up-sampling to further improve the sensitivity to the target region.Three sizes of convolution kernels are used to segment the image to improve the receptive field of the target region.The application has a deeper network and thus has stronger feature extraction capability, and can mine the target region in a more complex image.Multiple calculation modes make the model structure more flexible and variable, better adapt to business scenarios, and improve the generalization capability of the model.
Owner:HUAQIAO UNIVERSITY

A mixed audio track separation method, system, device, and storage medium

The application relates to the technical field of audio signal processing and artificial intelligence, in particular to a mixed audio track separation method, system, device and storage medium. The mixed audio track separation method comprises the following steps: extracting time domain features of a mixed audio signal through adaptive receptive field convolution, extracting frequency domain features of the mixed audio signal through an adaptive frequency spectrum window, and obtaining output features; generating reinforced features based on the output features, a pre-constructed frequency domain mask of a target instrument to be separated and feature channel attention; and performing adaptive waveform reconstruction through a joint decoder based on the reinforced features, so as to obtain separated independent audio tracks. The application can improve the efficiency and accuracy of mixed audio track separation.
Owner:THINKING CHAIN (TIANJIN) INTELLIGENT TECH CO LTD

System simulating a decisional process in a mammal brain about motions of a visually observed body

A system simulating a decisional process in a mammal brain about characteristics of motions related to body gestures of a visually observed body through a simulated visual path is provided. The system includes an interface toward simulated neuronal structures, the interface at least converting luminous information of the observed body to an optic flow data stream conveying information related to the visually observed body and that can be processed in the simulated neuronal structures, the system being a feed-forward system and comprising hierarchically from the visual observation to the decision: the simulated visual path and its interface, a simulated local motion direction detection neuronal structure for the detection of motion directions with receptive fields, a simulated opponent motions detection neuronal structure, a simulated complex patterns detection neuronal structure, and a simulated motion pattern detection neuronal structure.
Owner:UNIV DE MONTREAL +1

Lightweight field wheat ear detection method and device based on improved RTDETR

The invention relates to an improved RTDETR-based lightweight field wheat ear detection method. The method comprises the steps of obtaining wheat ear image data in a field environment, and performing preprocessing and data expansion; the RTDETR model is improved to obtain an improved RTDETR model, and the improved RTDETR model is used as a wheat ear detection model; training a wheat ear detection model by using the expanded training set; and obtaining a wheat ear detection result. According to the method, the lightweight FasterNet model is introduced as a basic backbone network, so that the efficient feature extraction capability of the wheat ear target is kept while the model parameter quantity and the calculation quantity are reduced; the AIFI-DyT module can inhibit the influence of illumination variation on wheat ear feature distribution, effectively enhance the expression stability of key wheat ear features, and can improve the wheat ear detection capability of the model under different illumination conditions; the DRBNCSPELAN4 module effectively expands the receptive field under the condition of no extra calculation overhead; a matching perception loss function is introduced, gradient feedback of shielding samples is enhanced, and the detection performance and convergence speed of the model in a dense shielding scene are improved.
Owner:ANHUI UNIV +1

Real-time dysarthria speech restoration method based on progressive model distillation

The invention discloses a real-time dysarthria speech restoration method based on progressive model distillation, and belongs to the technical field of speech signal processing. The invention provides a combined solution for the problem of double interference of voice degradation and environmental noise faced by disabled old people in a complex nursing environment. Firstly, a pseudo-parallel corpus is constructed through a self-supervised repair strategy, an ideal target speech is generated by using a high-precision speech conversion model, and the problem of truth value missing in a pathological speech enhancement task is solved. Secondly, constructing a complex convolutional loop network based on voiceprint embedding, and introducing a speaker embedding vector to suppress non-target human voice interference; meanwhile, a gradual receptive field perception distillation strategy is adopted, a non-causal wide convolution teacher model is compressed into a causal narrow convolution student model through micro sparsification, and real-time processing of an embedded terminal is realized. And finally, phoneme perception loss is introduced, and semantic level supervision is performed by using the pre-trained ASR model, so that the intelligibility of the repaired voice is remarkably improved.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Image retrieval method based on visual attention mechanism and symbiotic feature integration

The invention discloses an image retrieval method based on a visual attention mechanism and symbiotic feature integration, and relates to the technical field of image retrieval. The method comprises the following steps: firstly, combining Tamura visual statistic features with depth features, and mining a potential symbiosis key mode learned in a depth model from a data set end side; then, on the basis that a two-dimensional Gabor primary function can approximate receptive field characteristics of visual neurons of mammals, global texture information of the image is extracted by using a multi-direction Gabor filter; besides, in order to relieve the texture bias phenomenon of the depth feature, the geometric moment feature and the depth feature are fused, the depth moment feature is innovatively proposed, and the feature can be used for representing the regional shape attribute of the image. Finally, the multi-direction global features, the depth distance features and the depth semantic features are effectively integrated, and a compact, efficient and universal image representation is constructed and used for an image retrieval task.
Owner:HUNAN INST OF TECH

Improved pathological voice production model and construction method thereof

The application relates to an improved pathological voice generation model and a construction method thereof, which comprises the following steps: S1, using all available pathological / a / vowel speech records under normal pitch in a sound data set as a training data set; S2, a training process is divided into training of a generator and training of a discriminator, both of which are alternately performed, the generator comprises a latent mapping network module and a transposed convolution and multi-receptive field fusion module; the discriminator comprises four discriminators with different cycle settings, which are used for evaluating and training intermediate waveform outputs after each upsampling of the generator; and S3, inputting a random noise vector into the trained generator, converting the random noise vector into new audio data by the generator, performing standardization processing on the generated audio data, and repeating the above steps to generate multiple new data.
Owner:XIAN UNIV OF POSTS & TELECOMM

Diver behavior anomaly detection method and system based on space-time attention mechanism

This invention provides a method and system for detecting abnormal diver behavior based on a spatiotemporal attention mechanism. The method includes: constructing an underwater local optical distortion vector field by tracking the displacement of suspended particles in a video. This vector field is then used in a feature extraction network to perform regularization correction on the image and guide the receptive field shift to extract limb spatial features. When a limb is obscured by a bubble, cross-frame cross-correlation calculation is performed by analyzing the deformation of the bubble boundary and the limb motion velocity vector field of the last frame before obscuration to infer the implicit motion trajectory of the obscured limb, and a compensating feature vector is constructed to generate a spatiotemporal feature matrix. Finally, the divergence value of the gradient field of this matrix is ​​calculated to quantify the degree of behavioral abruptness; when the divergence value breaks through the stable range, it is determined to be abnormal. This invention solves the technical problems of inaccurate behavioral feature extraction and interruption of temporal information caused by underwater environmental interference and bubble obscuration.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Deep learning-based chinese language audio phoneme segmentation method, device and medium

The application discloses a deep learning-based Chinese language audio phoneme segmentation method and device and a medium, belongs to the technical field of speech analysis, and solves the problem of how to improve the precision and robustness of phoneme boundary detection; the application provides an audio phoneme segmentation model combining an audio signal and a mel spectrum feature, simultaneously receives an original audio signal and a mel spectrogram as input; the original audio signal is input into a HuBERT encoder, and the mel spectrogram is input into a Conformer encoder; two output features are firstly subjected to a residual attention mechanism to realize preliminary cross-modal fusion, and then subjected to multi-scale dilated convolution to extract context information in different receptive field ranges; the features subjected to the multi-scale dilated convolution are further input into a gating mechanism for feature selection; the selected features are subjected to a phoneme classification linear layer to obtain classification features, and the application has significant advantages in reducing manual labeling cost, improving segmentation precision and robustness.
Owner:ANHUI UNIV