Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

16 results about "Time delay neural network" patented technology

Time delay neural network (TDNN) is a multilayer artificial neural network architecture whose purpose is to 1) classify patterns with shift-invariance, and 2) model context at each layer of the network.

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Smart classroom interaction analysis method based on double-layer architecture voice segmentation

The invention provides a smart classroom interaction analysis method based on double-layer architecture voice segmentation, and relates to the technical field of voice segmentation, and the method specifically comprises the following steps: extracting the voice features of a voice signal through employing a Mel-frequency cepstrum coefficient MFCC; designing a text-enhanced multi-scale time sequence-based perception time delay neural network, performing coarse screening on voice features, and dividing an audio clip into a single-speaker clip and a multi-speaker clip; and inputting the coarsely screened multi-speaker segment into a sliding window segmentation model SW-NIF fused with adjacent window information, and positioning speaker conversion points in the multi-speaker segment. And training the constructed model on the data set and verifying the model. According to the technical scheme, the problems that in the prior art, the segmentation problem of the classroom audio is neglected, and only the classroom audio is simply segmented for subsequent tasks, so that speakers in audio clips are mixed, and the analysis effect is affected are solved.
Owner:SHANDONG UNIV OF SCI & TECH

Cross-domain depression detection system based on time-frequency calibration transfer learning

The invention discloses a cross-domain depression detection system based on time-frequency calibration transfer learning. According to the method, firstly, a frequency delay neural network and a time delay neural network are used for respectively extracting frequency domain and time domain features of voice signals, and local and global time-frequency embedding representations are fused through a multi-layer information aggregation module, so that a comprehensive vector capable of representing depression features is obtained. And then, designing a distribution calibration transfer learning module, and through construction of a high-confidence sample pair, reducing the distribution distance between samples of the same kind and enlarging the distribution difference between samples of different kinds, and maintaining the consistency and discrimination of depression features in a domain transfer process. According to the method, the depression detection precision is remarkably improved in cross-corpus and cross-speaker scenes, the problem that a traditional method is insufficient in generalization ability under the situation of domain mismatching is solved, and the method has good robustness and clinical application prospects.
Owner:EAST CHINA UNIV OF SCI & TECH

Underwater acoustic target identification method based on multi-scale feature learning

The invention discloses an underwater acoustic target recognition method based on multi-scale feature learning, and relates to the technical field of underwater acoustic target recognition. The method comprises the following steps: firstly, performing Fbank feature extraction on an obtained audio signal of a target vessel to obtain a spectrogram of the audio signal; then, a SpecAugment data enhancement method is used for carrying out random time mask and frequency mask processing on the extracted spectrogram; extracting multi-scale features through grouping convolution of an enhanced front-end convolution module (EFCM), and re-weighting the multi-scale features and SE Block adaptive features; meanwhile, large kernel convolution, channel halving and global and segment pooling addition are adopted in the enhanced dense connection time delay neural network, and an improved context sensing shielding module is combined to classify and recognize the multi-scale features of the target ship. According to the method, the interference of underwater complex noise can be effectively suppressed, and the recognition precision of the target ship audio signal is improved.
Owner:BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY

TDNN and LLM fused lithium battery electrolyte ultrasonic quantitative detection method

The invention discloses a TDNN and LLM fused lithium battery electrolyte ultrasonic quantitative detection method, and relates to the field of lithium ion battery health management, and the method comprises the steps: obtaining original ultrasonic waveform data, and carrying out the preprocessing; based on a pre-trained time delay neural network TDNN, extracting a high-dimensional depth feature vector; mapping to a text embedding space of a large language model LLM through a linear projection layer, generating a text prototype and constructing a complete prompt sequence; reasoning to obtain an electrolyte content predicted value based on a pre-trained large language model LLM; calculating a global estimated value, and carrying out physical correction on the global estimated value by utilizing the infiltration area ratio to obtain the corrected electrolyte content; and outputting the corrected electrolyte content and the two-dimensional distribution diagram of the electrolyte in the battery. According to the method, deep features of ultrasonic signals are learned through a cascade architecture of TDNN and LLM, reasoning is carried out, physical correction is supplemented, and high-sensitivity and reliable quantitative detection of the content of the lithium battery electrolyte is achieved.
Owner:BEIJING UNIV OF TECH +1

A robust speaker recognition method based on spectrogram denoising and adversarial learning

ActiveCN116469394BSpeech analysisNeural architecturesData setSpeaker recognition system
The present invention provides a robust speaker recognition method based on spectrogram denoising and adversarial learning. First, a spectrogram dataset of clean speech and a noisy spectrogram dataset after the clean speech is noisy are collected; a U-net with a multi-level encoding and decoding structure is trained using a mean square error loss function to remove noise interference from the mel-spectrogram of the noisy speech signal to obtain an enhanced mel-spectrogram; a conditional generative adversarial network based on a time-delay neural network (TDNN-CGAN) is trained using a least squares loss function, a time-delay neural network (TDNN) is used as a generator in the TDNN-CGAN to extract deep features of the enhanced mel-spectrogram, and a multi-layer perceptron (MLP) is used as a discriminator in the TDNN-CGAN; finally, a speaker classifier is trained using cross-entropy loss to identify the speaker's identity, thereby realizing speaker recognition in a noisy environment. The deep features extracted from the noisy speech by the present invention are close to the deep features extracted from the clean speech, thereby improving the performance of the speaker recognition system in a noisy environment.
Owner:NANCHANG UNIV

A voiceprint feature extraction method for specific content voice segments

ActiveCN117649842BEffective modelingEffective identificationNeural architecturesSpeech recognitionFeature extractionMachine learning
This application provides a method for extracting voiceprint features from speech segments with specific content. The method includes: obtaining an acoustic spectrum feature segment through preprocessing; constructing a time-delay neural network module; constructing a residual time-delay neural network module based on the time-delay neural network module, a weighted activation mechanism, and a residual structure; constructing a residual attention time-delay neural network module based on the time-delay neural network module, the residual time-delay neural network module, and an attention pooling mechanism; and inputting the acoustic spectrum feature segment into the residual attention time-delay neural network module to obtain the voiceprint features of the speech segment with specific content. The voiceprint feature extraction method provided here extracts deep-level information from features at multiple scales and, combined with residual networks, weighted activation, and attention pooling mechanisms, can effectively extract voiceprint features from speech segments with specific content.
Owner:CHINA SOUTHERN POWER GRID BIG DATA SERVICE CO LTD

A novel dynamic temporal recurrent neural network for dialect recognition

The application discloses a novel dynamic time delay neural network for dialect recognition, and belongs to the field of acoustic model modeling; specifically, an existing time delay neural network is introduced into k parallel convolution kernels, attention weight parameters of the respective convolution kernels are calculated, and an optimal weight parameter convolution kernel is obtained through weighted average fusion; the optimal weight parameter convolution kernel is used to replace a conventional convolution kernel of the existing time delay neural network, so that an improved dynamic time delay neural network is obtained; according to different dialect inputs, each weight parameter in the optimal weight parameter convolution kernel is dynamically adjusted, deeper feature information of audio is extracted, comparison and analysis are conducted on the extracted deep features and a dialect acoustic template, and finally, a classifier is combined to determine a dialect category; and the application improves the accuracy of determining the dialect category to which the audio belongs.
Owner:BEIJING FANGWEI ZHILIAN TECHNOLOGY CO LTD

Target identification method based on steady-state visual evoked potential brain-computer interface and related device

The invention discloses a target identification method based on a steady-state visual evoked potential brain-computer interface and a related device, and relates to the technical field of brain-computer interfaces, and the method comprises the steps: designing a trained identification model which comprises a segment coding module, a plurality of feature extraction modules and a classification module which are connected in sequence, the feature extraction module comprises a first normalization layer, a self-adaptive frequency spectrum module, an enhanced time delay neural network module, an inverse Fourier transform layer, a second normalization layer, a splicing layer, an interactive convolution module and a first addition layer, the trained recognition model is used for determining the recognition frequency corresponding to the electroencephalogram signals, the recognition frequency serves as the target frequency, and the recognition frequency is used as the target frequency. According to the method and the device, the three core application requirements of high-precision identification, cross-subject strong generalization ability, light weight and low delay of the SSVEP-BCI system under an ultra-short time window can be met at the same time.
Owner:INNER MONGOLIA UNIV OF TECH

A monitoring method, device, storage medium and computer device

The application provides a monitoring method and device, a storage medium and a computer device. When a person to be monitored is monitored, a monitoring picture and a human physiological feature sequence of the person to be monitored when the person to be monitored is active in a monitoring area are acquired, the human physiological feature sequence is input into a time-delay neural network configured in advance, the time-delay neural network is more suitable for processing sequence information due to a certain memory function, the body state of the person to be monitored is output after the sequence information is processed by the time-delay neural network, and when the body state contains a falling state, the falling state and the monitoring picture of the corresponding period of the falling state are sent to a monitoring person for alarm prompt. In this way, the monitoring person can quickly and accurately understand the body state of the person to be monitored from multiple aspects, and can take relevant measures in time through the alarm prompt, so that the monitoring difficulty is effectively reduced.
Owner:DONGGUAN ZKTECO ELECTRONICS TECH

Multi-speaker chinese speech synthesis method based on dense connection delay neural network

The application discloses a multi-speaker Chinese speech synthesis method based on a dense connection time delay neural network, wherein a speaker encoder module in a multi-speaker Chinese speech synthesis network based on the dense connection time delay neural network extracts speaker embedding from a reference speech spectrum; the speaker encoder module is simple in structure and small in parameter quantity; the extracted speaker embedding fuses multi-level information; therefore, the speaker embedding can be optimized together with other modules in the multi-speaker Chinese speech synthesis network, the training process is simplified, and the speaker embedding more suitable for a speech synthesis task can be extracted; secondly, the output of a text encoder module of the multi-speaker Chinese speech synthesis network is taken as a key and a value, the output of the speaker encoder module is taken as a query, and the key, the value and the query are input into an encoder's scaling dot product attention mechanism to generate a conditional text representation as an input of a decoder, so that the speaker embedding can effectively control the style in the synthesized speech and improve the naturalness and similarity of the synthesized speech.
Owner:NANJING UNIV

Radar based object classification

A method for radar based object classification, the method may include obtaining multiple radar samples of an object; the multiple radar samples were acquired at different acquisition times; wherein the multiple radar samples comprise a plurality of first radar sample parameters; calculating second radar sample parameters for the multiple radar samples, by applying one or more non-linear functions on at least some of the plurality of first radar sample parameters of at least some of the multiple radar samples; generating an object signature that comprises temporal information and inter-parameter correlation information; wherein the generating comprises feeding, to each one of a deep neural network (DNN) and a time delay neural network (TDNN), (a) at least some of the plurality of first radar sample parameters, and (b) at least some of the second radar sample parameters; and classifying, by a classifier, the signature to a signature class.
Owner:AUTOBRAINS TECH LTD

Data-driven trajectory tracking control system and method for unmanned mine truck

The present application belongs to the technical field of automatic driving mine car trajectory tracking, and particularly relates to a data-driven unmanned mine car trajectory tracking control system and method. The unmanned mine car trajectory tracking control system comprises a mine area multi-source data set construction module, a data-driven dynamics modeling module and a trajectory tracking cooperative control module. The mine area multi-source data set construction module collects and pre-processes mine car dynamics data. The data-driven dynamics modeling module trains a time delay neural network model and a deep Gaussian process regression model based on the data set, and realizes accurate prediction of mine car dynamics characteristics. The trajectory tracking cooperative control module combines feedforward feedback longitudinal control and model predictive lateral control, introduces model uncertainty constraints, and outputs optimal control instructions. The unmanned mine car trajectory tracking control method can solve the problems of inaccurate representation of traditional physical models, complex parameter calibration and insufficient robustness, and realize high-precision and high-stability trajectory tracking control under complex mine area working conditions.
Owner:SHENHUA BEIDIAN SHENGLI ENERGY +1

A Smart Classroom Interaction Analysis Method Based on Two-Layer Architecture Speech Segmentation

ActiveCN120783757BSpeech recognitionSpeech segmentationNerve network
This invention provides a smart classroom interaction analysis method based on a two-layer architecture speech segmentation, belonging to the field of speech segmentation technology. Specifically, it includes the following steps: extracting speech features from the speech signal using Mel-frequency cepstral coefficients (MFCC); designing a text-enhanced, multi-scale time-aware delay-based neural network to coarsely screen the speech features, dividing audio segments into single-speaker segments and multi-speaker segments; inputting the coarsely screened multi-speaker segments into a sliding window segmentation model (SW-NIF) that integrates neighbor window information to locate speaker transition points within the multi-speaker segments; training and validating the constructed model on a dataset. The technical solution of this invention overcomes the problem of neglecting classroom audio segmentation in existing technologies, which only perform simple segmentation of classroom audio for subsequent tasks, resulting in mixed speakers in the audio segments and affecting the analysis effect.
Owner:SHANDONG UNIV OF SCI & TECH

An ant colony optimization time delay neural network pipelined ADC background calibration method

The application discloses a pipeline ADC background calibration method of an ant colony optimization time delay neural network, relates to the technical field of ADC calibration, and comprises the following steps: connecting a to-be-calibrated pipeline ADC and a high-precision reference ADC to a same signal source, combining a time delay unit and a three-layer BP neural network to construct a time delay neural network, connecting the output of the to-be-calibrated pipeline ADC to the time delay unit of the time delay neural network, and connecting the output of the time delay unit and the output of the to-be-calibrated pipeline ADC to input data of neural network training; performing dimension optimization filtering on training data through an ant colony algorithm; performing global optimization on initial configuration of time delay neural network weights and biases through the ant colony algorithm; and outputting a calibrated result by the time delay neural network.
Owner:HEFEI UNIV OF TECH

Delay neural network parameter ant colony optimization method and device based on pipeline analog-to-digital converter calibration

The invention provides a time delay neural network parameter ant colony optimization method and device based on pipeline analog-to-digital converter calibration, and relates to the technical field of analog-to-digital converter calibration, and the method comprises the steps: firstly configuring the basic parameters of an ant colony algorithm and the initial parameters of a time delay neural network; a joint solution space containing initial weight configuration and initial offset configuration is constructed based on the initial parameters, and ants are controlled to traverse nodes according to greedy and random strategies to form paths (each path corresponds to a complete solution); calculating a fitness function value according to the minimum mean square error of multiple times of training of each frequency under the multi-frequency verification set; and finally, iterative optimization is carried out through a pheromone updating mechanism until a preset number of iterations is reached, and an optimal solution is output, so that the problems of unstable frequency domain performance and limited application bandwidth are solved.
Owner:CHERY AUTOMOBILE CO LTD