Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

7 results about "Receptive field" patented technology

A sensory space can be the space surrounding an animal, such as an area of auditory space that is fixed in a reference system based on the ears but that moves with the animal as it moves (the space inside the ears), or in a fixed location in space that is largely independent of the animal's location (place cells). Receptive fields have been identified for neurons of the auditory system, the somatosensory system, and the visual system.

An artificial visual nervous system based on dual-mode neural devices

ActiveCN120471114BPhysical realisationSynapseOptic nerve
The application relates to an artificial visual nervous system based on a bimodal neural device, comprising a bimodal visual information calculation array composed of a physical convolution kernel array and a synapse calculation array, the physical convolution kernel array is used for receiving an ambient light signal and converting into an electrical signal output, the synapse calculation array is used for receiving the electrical signal and converting into a visual signal, and the physical convolution kernel array and the synapse calculation array are both composed of a plurality of bimodal transistors. The application has excellent bimodal collaborative calculation capability, the receptive field mechanism (excitatory / inhibitory response) of a retinal bipolar cell is simulated through the physical convolution kernel array, image feature extraction and preprocessing are realized, the synaptic plasticity (LTP / LTD mechanism) of a central nervous system is simulated in combination with the synapse calculation array, high-level calculation of the visual signal is completed, and a complete visual nerve bionic system is formed by virtue of an array composed of single devices.
Owner:XIAN JIAOTONG LIVERPOOL UNIV

An acoustic event detection method based on multi-scale spatial feature and coordinate attention fusion

This invention discloses a sound event detection method based on the fusion of multi-scale spatial features and coordinate attention, belonging to the field of sound event detection technology. It aims to address the problems in existing sound event detection methods, such as the limited receptive field of the network, which makes it difficult to capture multi-scale temporal span features, and the inability to accurately focus on target sounds and suppress noise in complex time-frequency spaces. This method innovatively introduces a multi-scale coordinate attention module into the sound event feature extraction network. This module first uses parallel dilated convolutions with an increasing time dilation rate to obtain multi-scale temporal features and performs fusion and dimensionality reduction. Then, it employs a dual-axis time-frequency decoupling mechanism, performing adaptive pooling aggregation along both the time and frequency dimensions to generate directional time attention weights and frequency attention weights. Finally, the dual-axis weights are used to recalibrate the multi-scale features. This method significantly improves the accuracy and robustness of sound event detection.
Owner:GUILIN UNIV OF ELECTRONIC TECH

A mixed audio track separation method, system, device, and storage medium

The application relates to the technical field of audio signal processing and artificial intelligence, in particular to a mixed audio track separation method, system, device and storage medium. The mixed audio track separation method comprises the following steps: extracting time domain features of a mixed audio signal through adaptive receptive field convolution, extracting frequency domain features of the mixed audio signal through an adaptive frequency spectrum window, and obtaining output features; generating reinforced features based on the output features, a pre-constructed frequency domain mask of a target instrument to be separated and feature channel attention; and performing adaptive waveform reconstruction through a joint decoder based on the reinforced features, so as to obtain separated independent audio tracks. The application can improve the efficiency and accuracy of mixed audio track separation.
Owner:THINKING CHAIN (TIANJIN) INTELLIGENT TECH CO LTD

Improved pathological voice production model and construction method thereof

The application relates to an improved pathological voice generation model and a construction method thereof, which comprises the following steps: S1, using all available pathological / a / vowel speech records under normal pitch in a sound data set as a training data set; S2, a training process is divided into training of a generator and training of a discriminator, both of which are alternately performed, the generator comprises a latent mapping network module and a transposed convolution and multi-receptive field fusion module; the discriminator comprises four discriminators with different cycle settings, which are used for evaluating and training intermediate waveform outputs after each upsampling of the generator; and S3, inputting a random noise vector into the trained generator, converting the random noise vector into new audio data by the generator, performing standardization processing on the generated audio data, and repeating the above steps to generate multiple new data.
Owner:XIAN UNIV OF POSTS & TELECOMM

Diver behavior anomaly detection method and system based on space-time attention mechanism

This invention provides a method and system for detecting abnormal diver behavior based on a spatiotemporal attention mechanism. The method includes: constructing an underwater local optical distortion vector field by tracking the displacement of suspended particles in a video. This vector field is then used in a feature extraction network to perform regularization correction on the image and guide the receptive field shift to extract limb spatial features. When a limb is obscured by a bubble, cross-frame cross-correlation calculation is performed by analyzing the deformation of the bubble boundary and the limb motion velocity vector field of the last frame before obscuration to infer the implicit motion trajectory of the obscured limb, and a compensating feature vector is constructed to generate a spatiotemporal feature matrix. Finally, the divergence value of the gradient field of this matrix is ​​calculated to quantify the degree of behavioral abruptness; when the divergence value breaks through the stable range, it is determined to be abnormal. This invention solves the technical problems of inaccurate behavioral feature extraction and interruption of temporal information caused by underwater environmental interference and bubble obscuration.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Deep learning-based chinese language audio phoneme segmentation method, device and medium

The application discloses a deep learning-based Chinese language audio phoneme segmentation method and device and a medium, belongs to the technical field of speech analysis, and solves the problem of how to improve the precision and robustness of phoneme boundary detection; the application provides an audio phoneme segmentation model combining an audio signal and a mel spectrum feature, simultaneously receives an original audio signal and a mel spectrogram as input; the original audio signal is input into a HuBERT encoder, and the mel spectrogram is input into a Conformer encoder; two output features are firstly subjected to a residual attention mechanism to realize preliminary cross-modal fusion, and then subjected to multi-scale dilated convolution to extract context information in different receptive field ranges; the features subjected to the multi-scale dilated convolution are further input into a gating mechanism for feature selection; the selected features are subjected to a phoneme classification linear layer to obtain classification features, and the application has significant advantages in reducing manual labeling cost, improving segmentation precision and robustness.
Owner:ANHUI UNIV