Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

130 results about "Cepstrum coefficients" patented technology

The cepstral coefficients are the coefficients of the Fourier transform representation of the logarithm magnitude spectrum.

Wind turbine generator voiceprint fault recognition method

The invention provides a wind turbine generator voiceprint fault recognition method, and relates to the technical field of wind turbine generator state monitoring and fault diagnosis, and the method comprises the steps: carrying out the noise reduction of an original audio signal through variational mode decomposition, screening a target mode of which the frequency, energy and kurtosis accord with features, and reconstructing the signal; extracting a Mel frequency cepstrum coefficient and a sensing noise robust coefficient, and generating multi-dimensional voiceprint data in combination with statistical characteristics such as a frequency spectrum gravity center, a spectrum entropy, energy, kurtosis and a zero-crossing rate; constructing a support set based on the prototype network, realizing small sample fault classification by calculating the Euclidean distance between the feature vector and the prototype vector, and outputting a preliminary result; judging whether the voiceprint is abnormal according to a preset threshold value, if so, storing the voiceprint into a dynamic abnormal voiceprint knowledge base; frequently occurring abnormal samples are manually labeled and added into a support set, the prototype network is retrained to update the model, and continuous optimization of the fault recognition capability is achieved.
Owner:CGN (SHANXI) NEW ENERGY INVESTMENT CO LTD

Steel structure engineering welding quality defect analysis method based on voiceprint monitoring

The invention relates to a steel structure engineering welding quality defect analysis method based on voiceprint monitoring, and the method comprises the steps: carrying out the multi-channel voiceprint synchronous collection, time-frequency feature fusion, wavelet packet analysis and Mel-frequency cepstral coefficient extraction for a plurality of defect features fused in voiceprint data in a welding process; a hierarchical semantic concept space and a dynamic causal relationship generation model are established in combination with a welding physical knowledge base, a causal knowledge graph is constructed, causal association between semantic concepts is deduced through a gating circulation unit and a graph neural network, anti-fact disturbance and path aggregation analysis is carried out on a causal graph structure, and a result is obtained. And finally, defect category probability output and causal traceability graph visual display are realized. According to the scheme, the accuracy, traceability and result interpretability of welding defect recognition are effectively improved, and data support is provided for intelligent diagnosis and continuous model optimization in the welding process.
Owner:GUANGDONG YUECHAO CONSTRUCTION CO LTD

Voice emotion recognition method based on multiple scales and multiple features

The invention discloses a voice emotion recognition method based on multiple scales and multiple features, and belongs to the technical field of artificial intelligence. The method comprises the following steps: firstly, preprocessing an audio signal and extracting a spectrogram and a Mel-frequency cepstral coefficient; then, a residual network, a bidirectional long-short-term memory network and a HuBERT pre-training model are respectively utilized to extract spectrogram high-order spatial features, time sequence context features and voice semantic embedding features; secondly, inputting the first two features into a multi-dimensional multi-scale feature extraction module to extract richer time-frequency features, performing deep fusion by using a multi-layer cross attention mechanism, and performing weighted fusion with speech semantic embedded features; and finally, all the advanced features are spliced, and a final emotion category is recognized through a full-connection classifier. According to the invention, through combination of multi-scale feature extraction and an advanced fusion mechanism, the problem of insufficient complex emotion modeling ability in the prior art is effectively overcome, and the accuracy and robustness of voice emotion recognition are significantly improved.
Owner:NANJING INST OF TECH

Transformer abnormity identification method based on voiceprint feature analysis

The invention discloses a transformer abnormity identification method based on voiceprint feature analysis, and belongs to the field of power equipment state monitoring and intelligent diagnosis. The method comprises the following steps: firstly, analyzing an iron core acoustic mechanism based on a magnetostrictive effect, and establishing a three-dimensional model through finite element simulation to obtain vibration and sound field characteristics; in a complex substation environment, a hybrid noise reduction method combining density peak clustering and a CEEMDAN-wavelet threshold is provided, and the signal-to-noise ratio is effectively improved. Then extracting Mel-frequency cepstrum coefficients (MFCC) and spectrum features, and performing local linear embedding (LLE) dimension reduction to form a compact feature set; in the recognition stage, a convolutional neural network framework is designed, specifically, a spectrogram and an energy spectrum are modeled through a two-dimensional CNN, an MFCC tensor obtained after dimensionality reduction is modeled through a three-dimensional CNN, and accurate diagnosis of mechanical faults such as core looseness is achieved. The method has the advantages of being non-contact, anti-noise and high in recognition precision, and real-time diagnosis and early warning of mechanical abnormity of the transformer can be achieved under complex working conditions.
Owner:YANCHENG POWER SUPPLY CO STATE GRID JIANGSU ELECTRIC POWER CO

Smart home central control system and method based on multi-mode perception

The invention relates to the technical field of smart home control, and particularly discloses a smart home central control system and method based on multi-mode perception, and the system comprises the steps: synchronously collecting a voice audio signal, a gesture image signal, an infrared thermal imaging signal and a millimeter wave radar signal through a plurality of groups of sensors; performing blind source separation processing on the voice and gesture signals, extracting a voice command component and a gesture action component which are independent in statistics, and performing space-time alignment and Kalman filtering fusion on the infrared and radar signals to generate a dynamic environment sensing graph; voice intention features, gesture track features and environment anomaly features are extracted through Mel frequency cepstrum coefficient analysis, skeleton key point tracking and multi-level convolution processing; constructing a three-dimensional decision matrix based on the features, performing weighted evaluation through a fuzzy logic rule base to generate a control instruction priority sequence, and dynamically adjusting an equipment operation mode according to the priority; according to the invention, the problems of control conflict and response delay caused by multi-mode signal coupling are solved.
Owner:XIAN QINGYAO HEZHI INTELLIGENT TECHNOLOGY CO LTD

Spoken language evaluation method and device based on deep learning and medium

The invention discloses a spoken language evaluation method and device based on deep learning and a medium, and relates to the technical field of spoken language evaluation, and the method comprises the steps: collecting spoken language audio signals, carrying out the acoustic feature extraction of the spoken language audio signals through Mel-frequency cepstrum coefficient transformation, and generating an acoustic feature vector sequence; constructing a deep learning pronunciation diagnosis model, inputting the acoustic feature vector sequence into the deep learning pronunciation diagnosis model, calculating a multi-dimensional distance between each voice segment in the acoustic feature vector sequence and the phoneme prototype in a measurement space, and generating a pronunciation diagnosis result; converting the acoustic feature vector sequence into a text sequence through a speech recognition conversion method; and constructing a deep learning role analysis model, and inputting the text sequence into the deep learning role analysis model to generate a semantic role graph. According to the method, the acoustic deviation between the quantized speech segment and the standard phoneme in the measurement space is calculated through the phoneme prototype distance, and accurate space-time positioning and quantitative guidance of the pronunciation defect are realized.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Dialect intelligent customer service and culture knowledge base system

The invention relates to the technical field of agricultural travel services, and provides a dialect intelligent customer service and culture knowledge base system, the system comprises a perception layer, an analysis layer, a knowledge layer and an interaction layer four-dimensional architecture, the perception layer collects and preprocesses dialect voice, noise reduction is performed through spectral subtraction, and Mel frequency cepstrum coefficient features are extracted; the analysis layer identifies dialects based on a fine tuning Wav2Vec2.0 model, converts the dialects into mandarin through a Transform architecture, and completes intention identification and slot filling by using a BERT related model; the knowledge layer constructs a local culture knowledge graph containing non-abandoned, folk and other entities, and supports dynamic updating; and the interaction layer is combined with multiple rounds of dialogue management to generate multiform responses, and can be connected with an external service system. The system also optimizes a feedback module iterative model and knowledge. The system can cover more than ten dialects, realizes dialect interaction, culture interpretation and service conversion closed loop, and assists rural culture revitalizing and rural cultural travel service upgrading.
Owner:SHENZHEN BEIDOU DIGITAL TECH CO LTD

High-pressure pipeline leakage detection method and device based on voiceprint map

The invention discloses a high-pressure pipeline leakage detection method and device based on a voiceprint map, and relates to the technical field of pipeline detection. According to the method, sound signals are collected through a distributed microphone array, multi-dimensional features such as wavelet packet frequency band energy and Mel-frequency cepstral coefficients are extracted after variational mode decomposition is combined with wavelet threshold denoising preprocessing, and dimensionality reduction is performed through principal component analysis; the time delay between the sensors is calculated based on a generalized cross-correlation-phase transformation algorithm, a leakage point is positioned through a particle swarm optimization algorithm, the leakage degree can be evaluated by combining a bidirectional long-short-term memory network of an attention mechanism, and the D-S evidence theory is utilized to fuse multi-sensor result early warning. The system realizes high-sensitivity detection and accurate positioning, is high in anti-interference capability, and is suitable for safety monitoring of high-pressure pipelines in the industries of petroleum, natural gas and the like.
Owner:XINJIANG XINYE ENERGY & CHEM CO LTD

Video content description method and device, electronic equipment and storage medium

The embodiment of the invention provides a video content description method and device, electronic equipment and a storage medium, belongs to the technical field of video processing, and is suitable for the fields of financial science and technology and medical treatment. The method comprises the following steps: acquiring target video data; performing multi-modal data extraction on the target video data to obtain a multi-modal data stream; performing spiking neural coding on the multi-modal data stream to obtain a spatial-temporal feature map; performing hierarchical time sequence transformation on the spatial-temporal feature map to obtain a semantic unit and a Mel-frequency cepstral coefficient; performing cross-modal feature integration on the semantic unit and the Mel-frequency cepstrum coefficient to obtain a cross-modal context representation; and performing text decoding on the cross-modal context representation to obtain target natural language information. According to the embodiment of the invention, the accuracy of video content description can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent grinding process data analysis system for pharmaceutical laboratory

The invention relates to the technical field of data analysis, in particular to a pharmaceutical laboratory intelligent grinding process data analysis system which comprises the steps that a vibration audio signal of a grinding tank is collected in real time through a microphone; the method comprises the following steps: converting a time domain audio signal into a frequency domain frequency spectrum by adopting fast Fourier transform, identifying a first characteristic frequency and amplitude fluctuation thereof, monitoring abnormal harmonic components, extracting harmonic characteristics through a Mel frequency cepstrum coefficient, and judging an abnormal condition of a grinding state; and when the amplitude of the first characteristic frequency in the audio signal exceeds a preset threshold, judging that the particle size of the material reaches a critical value, triggering a rotating speed switching instruction, and finely adjusting the rotating speed of the motor through a PID control algorithm. Audio signals are collected through the microphone array, low-frequency tank vibration can be captured through audio collection, high-frequency abnormal friction noise can be monitored, time-domain signals are converted into frequency-domain frequency spectrums, and energy distribution is visually displayed.
Owner:SHENZHEN ZHUJUNHAO MEDICAL TECHNOLOGY DEVELOPMENT CO LTD

Multi-modal behavior data processing system

The invention discloses a multi-modal behavior data processing system, which comprises a data acquisition module, a feature extraction layer, a dynamic attention weight layer, a multi-modal fusion layer and a downstream task decision-making layer, the data acquisition module guides human-computer interaction through an international neurological and mental interview tool matched with a DSM-5 standard and acquires audio and video stream data; the feature extraction layer extracts a video feature vector (including facial action unit activation intensity and the like), an audio feature vector (including Mel frequency cepstrum coefficient and the like) and a text feature vector (generated by a deep language model after automatic speech recognition transcription) in parallel; the dynamic attention weight layer is combined with data quality, symptomatic priori knowledge and cross-modal correlation to generate a dynamic fusion weight; weighting, splicing and dimensionality reduction are carried out on the multi-modal fusion layer to obtain a fusion feature vector; and the downstream task decision-making layer completes evaluation and generates a multi-modal behavioral index evaluation report. The system is deployed in a non-intrusive manner, the risk is controllable, and the evaluation robustness and accuracy can be improved.
Owner:NEW MAYO HEALTH MANAGEMENT RESEARCH INSTITUTE (CHONGQING) CO LTD +1

Cable extrusion process state online monitoring method and system based on acoustic sensing

The invention belongs to the technical field of extrusion process monitoring, and particularly relates to a cable extrusion process state online monitoring method and system based on acoustic sensing, so as to solve the technical problem of inaccurate cable extrusion process state monitoring in the prior art. The monitoring method comprises the following steps: S1, when a transient impact factor is greater than a preset threshold value, windowing a signal frame by adopting an asymmetric window function with a main lobe in front and a side lobe in back; s2, performing fast Fourier transform on each windowed signal frame to obtain a frequency spectrum of each windowed signal frame; filtering the frequency spectrum of each signal frame by using a filter bank and calculating logarithmic energy of each filtering channel; s3, selecting cepstrum coefficients of which the ratio is greater than a selection threshold to form a process state feature vector; and inputting the process state feature vector into a state classifier to obtain a current cable extrusion process state. According to the method, redundancy and noise interference are eliminated, so that the accuracy of process state monitoring is improved.
Owner:JIANGSU HONGFENG CABLE GROUP

English pronunciation error correction training method based on speech recognition

The invention relates to the technical field of speech recognition and processing, in particular to an English pronunciation error correction training method based on speech recognition, and the method comprises the following steps: S1, collecting a speech signal generated by a learner in a pronunciation training process, digitalizing the speech signal, associating the digitalized speech signal with a target standard text, and generating an original audio data record with a timestamp; according to the invention, phoneme-level decoding is carried out on the voice signal by using the recurrent neural network acoustic model, and accurate alignment of the pronunciation of the learner and the standard phoneme sequence is realized in combination with the dynamic time warping algorithm, so that pronunciation errors such as misreading, missed reading and increased reading can be accurately identified; meanwhile, acoustic features such as Mel frequency cepstrum coefficient, pitch and fundamental frequency are extracted to be quantitatively compared with a standard native language pronunciation database, multi-dimensional evaluation covering accuracy, integrity, fluency and rhythm is generated, and the accuracy and systematicness of oral English pronunciation error correction are remarkably improved.
Owner:吕丽沙

Double-path voice stream real-time identification method, system and application

The invention discloses a double-channel voice stream real-time identification method, system and application, and the method comprises the steps: carrying out the preprocessing of collected VOIP call double-channel audio, and maintaining the time sequence synchronization; extracting Mel-frequency cepstral coefficient features and speech spectrogram features of the preprocessed audio, and inputting the spliced features into a Transform deep neural network model for stream speech recognition to obtain a two-way character sequence; generating a unique identifier based on channel identification and voice energy difference, and establishing a corresponding relation with the character sequence; carrying out punctuation prediction and text standardization by utilizing an LSTM-based model, sorting and aligning character sequences according to timestamp fields, and generating a time sequence dialogue stream; and performing anomaly detection and / or storage management on the time sequence dialogue stream to realize real-time quality inspection and agent assistance. According to the invention, synchronous recognition and role distinguishing of double-channel voice are realized, the recognition delay is low, and the recognition accuracy, the detection precision and the real-time performance are high; and high-efficiency management and safety compliance of data are realized by combining distributed encryption storage.
Owner:XUNMENG COMMUNICATION TECHNOLOGY CO LTD

A music enjoyment degree recognition method based on music-electroencephalogram feature fusion

This invention relates to the field of music signal recognition technology, and more particularly to a method for recognizing the degree of music enjoyment based on music-EEG feature fusion. The method includes: extracting the Mel-frequency cepstral coefficients of the music signal; preprocessing and performing Fast Fourier Transform on the Mel-frequency cepstral coefficients to obtain the spectral feature values ​​of the music signal; and sequentially filtering, reducing dimensionality, and decorrelating the spectral feature values ​​of the music signal; extracting the Mel-frequency cepstral coefficients of the EEG signal; optimizing and updating the order of the Mel-frequency cepstral coefficients of the EEG signal so that the comprehensive error value between the EEG feature reconstruction signal and the original EEG signal meets a preset condition; aligning the music feature coefficients and the EEG feature reconstruction signal in the feature dimension and time axis using the FastDTW algorithm; and recognizing the degree of music enjoyment of the subject based on the alignment result. This invention effectively improves the robustness and accuracy of music recognition.
Owner:LANZHOU UNIV

System and method for detecting stress in audio data

A computerized system and method may process and predict stress levels for audio data using a machine learning based framework. A computerized system including a processor and a memory may calculate a buffer length based on a plurality of audio attributes (e.g., of a given audio input or data item), extract an audio buffer from an audio data item based on the calculated length, and predict, using a machine learning model, a stress level for the audio buffer or data item. Some embodiments of the invention may include extracting a buffer of a length determined dynamically for different audio inputs, e.g., to ensure coherency between audio attributes or features extracted from different audio inputs having different audio characteristics. In some embodiments, audio features which may be considered by the model may include, e.g., a plurality of gradients between mel-frequency cepstrum coefficients computed for relevant audio buffers or inputs.
Owner:NICE LTD

High-voltage circuit breaker operating mechanism voiceprint state detection method and device based on hesitant fuzzy number and medium

The invention relates to a hesitant fuzzy number-based high-voltage circuit breaker operating mechanism voiceprint state detection method and device, and a medium. Extracting features of an original voiceprint data sequence of the high-voltage circuit breaker operating mechanism, wherein the features comprise a Mel-frequency cepstrum coefficient, short-time energy and a zero-crossing rate; forming a symptom feature vector by the symptom parameters, inputting the symptom feature vector into a fuzzy depth residual shrinkage network, constructing a membership function, and outputting a membership value of each state; constructing hesitant fuzzy numbers according to the membership values to form a collective hesitant fuzzy evaluation matrix; determining the evidence weight of each input evaluation model through an optimal-worst method, deriving a state risk weight through a TOPSIS method in combination with a language Z number, and performing weighted fusion on the collective hesitation fuzzy evaluation matrix to obtain a health index; and dividing the health indexes into normal, attention and dangerous health levels by adopting a K-means clustering algorithm. Compared with the prior art, the method has the advantages of high accuracy, high robustness, high flexibility and the like.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Method and device for identifying void defect of steel plate concrete structure

The invention relates to a method and a device for identifying a void defect of a steel plate concrete structure. The method comprises the following steps: acquiring a to-be-identified audio signal of the steel plate concrete structure acquired through excitation conditions; extracting feature information on each frame of the preprocessed audio signal, wherein the feature information at least comprises a Mel-frequency cepstrum coefficient, root mean square energy and two items of a spectral centroid statistical mean value and a standard deviation; the extracted frame feature information is input into a void defect recognition model, void confidence of each frame is obtained, and the void defect recognition model is obtained through training by utilizing a marked audio sample training set based on a support vector machine; in response to the fact that the void confidence coefficient corresponding to each frame of audio is higher than a confidence threshold value, judging that the corresponding frame is a void frame; and taking a judgment result of the majority frames as a final void defect identification result of the audio signal to be identified. The problems that a steel plate concrete structure is low in small-size void defect sensitivity and is greatly interfered by the thickness of a steel plate and an internal complex structure are solved.
Owner:SHANGHAI RESEARCH INSTITUTE OF BUILDING SCIENCES CO LTD +1

Method and system for processing ground surface settlement monitoring data

The invention discloses a method and system for processing ground surface settlement monitoring data, and relates to the technical field of tunnel monitoring data processing, and the method comprises the steps: collecting the ground surface settlement monitoring data of each measuring point, and carrying out the outlier recognition and processing; extracting a feature image of the monitoring data of each measuring point by adopting a Mel-frequency cepstrum coefficient feature extraction method, and converting the ground surface settlement monitoring data into an MFCC feature map; the MFCC feature map is input into a pre-trained MFCC-Transform model for data classification, and measuring point data affected by tunnel excavation construction are reserved; an improved ELM time sequence prediction algorithm is adopted to fill data missing values; and decomposing the monitoring data into a plurality of intrinsic mode functions (IMF) by adopting a VMD time sequence decomposition algorithm, and selecting the IMF with the highest similarity as the monitoring data after noise reduction in combination with a similarity criterion. According to the invention, the full-automatic processing of the ground surface settlement monitoring data is realized, and the reliability of data processing is ensured while the data processing speed is improved.
Owner:KUNMING UNIV OF SCI & TECH

A networked audio product collaborative testing method based on a distributed network

The present application relates to the technical field of acoustic device resource scheduling and conflict determination, first, a standardized multi-type excitation signal is applied to the measured device, audio input and output and near-field sound pressure signals are collected and fused, mel-frequency cepstral coefficients, harmonic-to-noise ratios and formant frequencies are extracted and combined to generate acoustic fingerprints, a multi-dimensional semantic association graph of acoustic fingerprints and physical resource nodes is constructed, and resource real-time mapping and occupation relationship tracking are realized. When scheduling a new task, multi-dimensional features and time sequence overlap queries are performed based on the fingerprint graph, and multi-level conflict state determination is combined with parameters such as phase response and formant shift. Through adaptive adjustment of the graph weight, dynamic optimization of resource conflict prediction is realized. The method improves the accuracy of task scheduling and the resource utilization efficiency in a complex acoustic test environment, and effectively reduces the test risk caused by resource competition.
Owner:SHENZHEN FENDA TECH CO LTD

A ring main unit fault diagnosis method, device and storage medium

PendingCN122362215AMulti source dataAutoencoder
The application discloses a kind of ring network cabinet fault diagnosis method, equipment and storage medium, method includes: through the synchronous acquisition operation data of multiple source sensor group deployed on ring network cabinet;After the pre-processing of the collected multi-source data, respectively extract improved mel frequency cepstrum coefficient, transient voltage feature, ultrasonic feature and electrical feature based on high-order cumulant;The extracted features are spliced to form high-dimensional fusion feature vector, and the compact features are obtained by dimension reduction using sparse autoencoder;The improved aurora optimization algorithm is used to optimize the support vector machine, and the IPLO-SVM model is constructed to identify faults.The application can fully integrate multi-source information, efficiently extract fault features, achieve high-precision classification, and is suitable for intelligent operation and maintenance of distribution network ring network cabinet.
Owner:KEDA INTELLIGENT ELECTRICAL TECH +1

A method for monitoring the wear and fatigue of sluice gate opening and closing equipment

This invention relates to the field of equipment monitoring technology, and more particularly to a method for monitoring the loss and fatigue of a sluice gate opening and closing device. The method includes the following steps: simultaneously acquiring multimodal time-series data using vibration sensors, temperature sensors, and acoustic emission sensors, wherein the multimodal time-series data includes vibration signals, temperature signals, and acoustic emission signals; performing wavelet packet decomposition on the vibration signals to extract the energy proportion of a preset frequency band as vibration energy features; calculating the temperature rise rate and local range of the sliding window of the temperature signal as temperature features; extracting the Mel frequency cepstral coefficients from the acoustic emission signals as acoustic features; and concatenating the vibration energy features, temperature features, and acoustic features according to time steps into a multidimensional time-series feature matrix. This invention achieves accurate identification and prediction of equipment loss and fatigue state by utilizing a hidden Markov model to solve for the maximum probability state transition path.
Owner:GUANGDONG RES INST OF WATER RESOURCES & HYDROPOWER

Early failure diagnosis method for surge absorber based on acoustic characteristics

The invention discloses a surge absorber early failure diagnosis method based on acoustic characteristics. The surge absorber early failure diagnosis method comprises the steps of collecting acoustic signals generated by a surge absorber in an operation state; performing mixed feature extraction by fusing improved variational mode decomposition and frequency cepstrum coefficients, and constructing a high-dimensional initial feature set; pre-screening the high-dimensional feature set by using a minimum redundancy and maximum correlation algorithm, and then performing feature fine optimization by using a support vector machine package method optimized by a quantum behavior particle swarm algorithm to obtain an optimal feature subset; and finally, constructing a plurality of classifier groups based on self-service sampling, and performing decision-making layer fusion on the output of each classifier by using a D-S evidence theory to realize accurate diagnosis of the early failure state of the surge absorber. Through deep fusion of multiple levels and multiple algorithms, the sensitivity, robustness and accuracy of diagnosis are remarkably improved, and non-intrusive online monitoring and early warning of the surge absorber can be realized.
Owner:武汉京品电子科技有限公司

Speech emotion recognition method of multi-dimensional convolutional fusion network

The invention relates to a speech emotion recognition method of a multi-dimensional convolution fusion network, which comprises the following steps of: 1, extracting Mel-Frequency Cepstrum Coefficient (Mel-Frequency Cepstrum Coefficient) features from a speech signal, and respectively inputting the extracted features into a one-dimensional convolution path and a two-dimensional convolution path; 2, processing time sequence features in a one-dimensional path by adopting an adaptive residual extraction module, processing frequency spectrum features in a two-dimensional path by adopting a parallel path module, and fusing feature representations of the two paths through a cross attention conversion mechanism; 3, emotion classification is achieved through a full-connection layer, and the emotion classification is applied to emotion change monitoring and evaluation in the psychological counseling process of the depression patient. According to the method, through multi-dimensional feature extraction and fusion, comprehensive capture of voice emotion information is realized, an objective quantification tool is provided for auxiliary evaluation of depression, the technical problems of high subjectivity, low resource accessibility and the like of a traditional diagnosis method are solved, and the method has a relatively good clinical application value.
Owner:ZHEJIANG UNIV OF TECH

A multi-dimensional fusion feature underwater acoustic target recognition method based on channel attention

The application belongs to the technical field of marine observation, and specifically discloses a multi-dimensional fusion feature underwater acoustic target recognition method based on channel attention. The method discloses a multi-dimensional feature fusion method based on a gamma filter to realize feature fusion of obtained three-dimensional logarithmic gamma spectrum and three-dimensional gamma cepstrum coefficients; the gamma filter has stronger robustness to noise, so that the method for extracting multi-dimensional fusion features can more comprehensively capture useful acoustic information. In addition, in the multi-dimensional feature fusion process, the feature data is denoised through a self-encoding network, noise interference is reduced, feature quality is enhanced, and the underwater acoustic target recognition rate is improved. In addition, the application also discloses a U-Net network combined with a channel attention mechanism to enhance the representation ability of key features, and the underwater acoustic target is classified through a classification head. The method has good robustness in a noisy environment, and has high underwater acoustic target recognition accuracy.
Owner:OCEAN UNIV OF CHINA

Camouflage voice voiceprint recognition method based on Transform model and mixed features

PendingCN121862122Aimprove performanceFitting feature distribution is goodSpeech recognitionFeature extractionGammatone filter
The invention relates to the technical field of speech processing, and particularly provides a disguise speech recognition method based on a Transform model and mixed features, which is carried out from two aspects of feature extraction and model establishment. A resonance peak parameter is calculated by adopting a cepstrum method, a cepstrum coefficient (GFCC) is obtained through a Gammatone filter bank, then the resonance peak, the GFCC and a difference coefficient of the GFCC are combined into a mixed characteristic parameter, and complementary correlation between mixed characteristics is mined. From the perspective of model establishment, the mixed features are used as the input of the model, and the Transform network model is used as the acoustic model of the voiceprint recognition system, so that the feature distribution is better fitted, the classification effect is remarkably improved, and the performance of the camouflage voice voiceprint recognition system is effectively improved. The problem of performance degradation caused by feature redundancy and modal noise in a traditional method is solved.
Owner:CHINA CRIMINAL POLICE UNIV

Ensemble learning intracardiac ultrasonic signal classification method based on multi-feature fusion

The invention discloses an ensemble learning intracardiac ultrasonic signal classification method based on multi-feature fusion. The method comprises the steps that original cardiac blood flow ultrasonic signals are collected to obtain ultrasonic digital signals, and the ultrasonic digital signals are preprocessed; processing the preprocessed ultrasonic digital signal and extracting a Mel-frequency cepstral coefficient as a first feature vector; performing framing processing on the preprocessed ultrasonic digital signal to obtain an envelope self-correlation feature, and taking the envelope self-correlation feature as a second feature vector; extracting an intrinsic mode component of the preprocessed ultrasonic digital signal, and performing Hilbert transform on the intrinsic mode component to obtain a Hilbert marginal spectrum as a third feature vector; extracting a wavelet scattering coefficient as a fourth feature vector; and combining the first feature vector, the second feature vector, the third feature vector and the fourth feature vector to obtain a fusion feature vector, classifying the four features by adopting a K-nearest neighbor algorithm, a support vector machine and a neural network, and performing voting ensemble learning on the obtained three classification results to obtain a final classification result.
Owner:FIRST AFFILIATED HOSPITAL OF DALIAN MEDICAL UNIV

Transformer fault detection method based on double-flow auditory feature fusion and random forest

The invention discloses a transformer fault detection method based on double-flow auditory feature fusion and a random forest, and the method comprises the steps: obtaining a real-time voiceprint signal during the operation of a transformer, carrying out the pre-emphasis, framing and windowing of the real-time voiceprint signal, and obtaining a voiceprint frame; extracting Mel frequency cepstrum coefficient characteristics of the voiceprint frame based on a Mel filter bank; extracting Gammatone frequency cepstrum coefficient characteristics of the voiceprint frame based on a Gammatone filter bank; calculating time domain statistical index characteristics of the voiceprint frames; constructing a joint feature, inputting the joint feature into a pre-trained random forest model, and calculating a Gini importance score so as to obtain an anti-noise robust feature subset; and inputting the anti-noise robust feature subset into a random forest classifier, and outputting a fault type identification result of the transformer by using an integrated voting mechanism of a multi-decision tree. The method can solve the problems that in the prior art, feature extraction is single, noise immunity is poor, attention to transient impact is lacked, and a feature dimension reduction and classification method has defects.
Owner:EAST CHINA JIAOTONG UNIVERSITY

A lightweight identity recognition method combining voiceprint and earprint features

The application provides a lightweight identity recognition method combining voiceprint and earprint features. The method comprises the following steps: obtaining voice and ear canal echo signals of a known registered person and a person to be verified; extracting and fusing 13-dimensional mel-frequency cepstral coefficients (MFCC) of the voice and ear canal echo signals, i.e., obtaining voiceprint and earprint fusion features; inputting the fusion features into a lightweight identity recognition model to extract 128-dimensional embedding features of the known registered person and the person to be verified; calculating the similarity of the embedding features of the two types of persons by using a probabilistic linear discriminant analysis (PLDA) method; and determining whether the person to be verified is the known registered person according to the similarity. In the application, the fusion of earprint and voiceprint features improves the recognition performance of the classification model. The lightweight identity recognition model is obtained through a pre-training process, which reduces the equal error rate (EER) and greatly reduces the parameter quantity of the identity recognition model.
Owner:HOHAI UNIV

Mechanical fault diagnosis method based on mixed attention mechanism

The invention provides a mechanical fault diagnosis method based on a mixed attention mechanism, and the method comprises the steps: collecting a fault audio signal of a rotating machine, carrying out the fault type marking, carrying out the preprocessing of the fault audio signal, and obtaining a training data set; performing Mel frequency cepstrum coefficient feature extraction on the preprocessed signal, and converting the extracted feature into a grayscale image feature; constructing a diagnosis model based on a mixed attention mechanism, and inputting the grayscale image features into the diagnosis model of the mixed attention mechanism based on multi-head self-attention and sequence dimension attention for training to obtain a mechanical fault diagnosis model; and gray level image features are extracted from a to-be-identified mechanical fault audio signal and are input into the mechanical fault diagnosis model to obtain the fault type of the mechanical fault audio signal.
Owner:ZHENGZHOU XINDA ADVANCED TECH RES INST +1