Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

217 results about "Cepstrum coefficients" patented technology

The cepstral coefficients are the coefficients of the Fourier transform representation of the logarithm magnitude spectrum.

Sound anomaly detection method and device based on Transform model, equipment and medium

PendingCN120340527ASpeech analysisAbnormal voiceData acquisition
The invention relates to the technical field of sound anomaly detection, in particular to a sound anomaly detection method and device based on a Transform model, equipment and a medium, and the method comprises the steps: collecting a sound signal during the operation of the equipment through a data collection interface, and obtaining an original sound signal; resampling is carried out on the collected sound signals, and normalization processing is carried out on the resampled data; mel-frequency cepstrum coefficient features are extracted from the sound signals after normalization processing, and the sound signals after normalization processing are input into a pre-training module to output high-dimensional features including time sequence and semantic information; splicing the Mel-frequency cepstrum coefficient features with the high-dimensional features to form a comprehensive feature vector; inputting the comprehensive feature vector into a support vector machine model, and performing abnormal sound recognition through a trained classification hyperplane; and detected abnormal information is fed back to the user in real time. Multi-feature fusion enables the model to identify abnormal sound more accurately.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Voice interaction method and system of AI intelligent robot

The invention relates to the technical field of voice interaction, particularly discloses an AI intelligent robot voice interaction method and system, and aims to solve the problems of low voice interaction accuracy, insufficient reliability and lack of authority control in a complex noise environment. A dynamic noise feature library containing steady-state noise, impact noise and human voice interference features and a pre-stored gesture instruction library are constructed, audio signals are collected in real time, low-frequency-band, middle-frequency-band and high-frequency-band differential noise reduction is executed, Mel-frequency cepstral coefficient features are extracted, noise scenes are matched, corresponding voice recognition models are switched, and voice recognition is achieved. And calculating a confidence value of the voice instruction, outputting multi-modal verification data in combination with a dynamic confidence threshold, and outputting an authority control signal through voiceprint matching, authority verification and instruction consistency judgment. Through multi-modal fusion, dynamic adaptation and authority control, the voice recognition accuracy and interaction safety in a complex noise environment are remarkably improved, and the method is suitable for scenes such as factory intelligent inspection.
Owner:HANGZHOU SOHA TECH CO LTD

Audio feature extraction method and device based on neural network

The invention relates to an audio feature extraction method and device based on a neural network. The method comprises the following steps: performing framing processing and fast Fourier transform on an original audio signal acquired by a vehicle-mounted microphone to obtain a frequency domain feature sequence; extracting short-time energy, a zero-crossing rate, a frequency spectrum centroid and a Mel-frequency cepstrum coefficient from the frequency domain feature sequence, and combining first-order difference and second-order difference features of each feature to obtain a multi-dimensional acoustic feature point sequence; performing time sequence arrangement on the multi-dimensional acoustic feature point sequence, constructing an original feature matrix, and performing principal component analysis on the original feature matrix to obtain a feature description matrix; inputting the feature description matrix into a dynamic time feature extraction network for deep time-frequency feature analysis to obtain vehicle-mounted voice deep time-frequency features; and performing voice detection classification based on the vehicle-mounted voice deep time-frequency features, and outputting a vehicle-mounted voice detection result. According to the invention, various noise interferences in a vehicle-mounted environment are effectively suppressed, and the anti-noise capability of the system is enhanced.
Owner:DONGGUAN HUAZE ELECTRONIC TECH CO LTD

Wind turbine generator voiceprint fault recognition method

The invention provides a wind turbine generator voiceprint fault recognition method, and relates to the technical field of wind turbine generator state monitoring and fault diagnosis, and the method comprises the steps: carrying out the noise reduction of an original audio signal through variational mode decomposition, screening a target mode of which the frequency, energy and kurtosis accord with features, and reconstructing the signal; extracting a Mel frequency cepstrum coefficient and a sensing noise robust coefficient, and generating multi-dimensional voiceprint data in combination with statistical characteristics such as a frequency spectrum gravity center, a spectrum entropy, energy, kurtosis and a zero-crossing rate; constructing a support set based on the prototype network, realizing small sample fault classification by calculating the Euclidean distance between the feature vector and the prototype vector, and outputting a preliminary result; judging whether the voiceprint is abnormal according to a preset threshold value, if so, storing the voiceprint into a dynamic abnormal voiceprint knowledge base; frequently occurring abnormal samples are manually labeled and added into a support set, the prototype network is retrained to update the model, and continuous optimization of the fault recognition capability is achieved.
Owner:CGN (SHANXI) NEW ENERGY INVESTMENT CO LTD

Steel structure engineering welding quality defect analysis method based on voiceprint monitoring

The invention relates to a steel structure engineering welding quality defect analysis method based on voiceprint monitoring, and the method comprises the steps: carrying out the multi-channel voiceprint synchronous collection, time-frequency feature fusion, wavelet packet analysis and Mel-frequency cepstral coefficient extraction for a plurality of defect features fused in voiceprint data in a welding process; a hierarchical semantic concept space and a dynamic causal relationship generation model are established in combination with a welding physical knowledge base, a causal knowledge graph is constructed, causal association between semantic concepts is deduced through a gating circulation unit and a graph neural network, anti-fact disturbance and path aggregation analysis is carried out on a causal graph structure, and a result is obtained. And finally, defect category probability output and causal traceability graph visual display are realized. According to the scheme, the accuracy, traceability and result interpretability of welding defect recognition are effectively improved, and data support is provided for intelligent diagnosis and continuous model optimization in the welding process.
Owner:GUANGDONG YUECHAO CONSTRUCTION CO LTD

Speech recognition method and system based on artificial intelligence

The invention provides a speech recognition method and system based on artificial intelligence, and relates to the technical field of speech recognized.The speech recognition method comprises the steps that speech signals are collected in real time, and speech signal features are extracted through a Mel-frequency cepstrum coefficient after the speech signals are subjected to noise reduction; and combining the Mel-frequency cepstral coefficients and the first-order difference and the second-order difference of the Mel-frequency cepstral coefficients to form a speech feature vector. Meanwhile, a lip moving image is collected to serve as a visual signal, after graying processing is conducted on the image, an image feature vector is generated by calculating LBP values of pixel points in the image, the weight of the voice feature vector and the weight of the image feature vector are dynamically adjusted through a cross-modal attention mechanism, a fusion weight matrix is generated, and a fusion image is obtained. Different fusion weight matrixes correspond to different voice instructions, the original voice signals and the original visual images serve as a training set, the voice instructions corresponding to the fusion weight matrixes serve as labels to train a deep learning network model, and finally real-time voice recognition is carried out by inputting data collected in real time into the trained model.
Owner:DEEPANO

Biological feature recognition method driven by PPG big data

The invention provides a PPG big data driven biological feature recognition method, which comprises the following steps: acquiring PPG signal data of trainees, and performing high-pass filtering and low-pass filtering preprocessing on the PPG signal data to obtain preprocessed PPG signals; framing is carried out on the preprocessed PPG signal, a Mel frequency cepstral coefficient feature and a Gammatone frequency cepstral coefficient feature are extracted respectively, the Mel frequency cepstral coefficient feature and the Gammatone frequency cepstral coefficient feature are fused, then principal component analysis dimension reduction is carried out, and a fusion feature is obtained; a universal background model UBM of a Gaussian mixture model is constructed based on the fusion features, zero-order, first-order and second-order Baum-Welch statistics are calculated, a global difference space matrix is estimated from the Baum-Welch statistics, and i-vector identity authentication vectors are extracted; and inputting the i-vector identity authentication vector into a long short-term memory network for training and classification to obtain an identity recognition result of the trainees. According to the invention, the influence of motion artifacts can be eliminated, the overall variability of PPG signals is captured, and an identity authentication scheme is provided for wearable equipment.
Owner:HUBEI UNIV OF ECONOMICS

Casting industry abnormal sound detection and grading response method based on voiceprint recognition

The invention provides a casting industry abnormal sound detection and grading response method based on voiceprint recognition, and belongs to the technical field of casting industry detection. A high-temperature-resistant microphone array is arranged at an easy-to-leak part of cast aluminum equipment, and three-stage filtering noise reduction and amplitude normalization preprocessing are adopted, so that the problem of poor signal quality caused by noise interference in a complex environment is effectively solved; the characteristics of the molten aluminum leakage sound in different frequency bands and different stages can be captured through variable window long-short time Fourier transform and extended Mel frequency cepstrum coefficient in combination with extraction of an energy change rate and a frequency spectrum gravity center; a Transform-CNN hybrid deep learning model based on an attention mechanism is constructed, and feature screening is optimized through principal component analysis and recursive feature elimination, so that the recognition and generalization ability of the model to the abnormal sound in the casting industry is significantly improved; and meanwhile, graded response measures are made based on the detection result, so that the abnormal conditions of molten aluminum leakage with different severity degrees are processed.
Owner:SHENZHEN POLYTECHNIC

High-voltage circuit breaker voiceprint denoising method based on data enhancement and storage medium

The invention provides a high-voltage circuit breaker voiceprint denoising method based on data enhancement and a storage medium, and the method comprises the steps: processing a collected original voiceprint data sequence, extracting stable and effective Mel-frequency cepstrum coefficient features, and constructing a two-dimensional feature matrix; a parallel mixed data enhancement strategy is adopted to generate a positive sample pair, and an encoder is trained in combination with a contrast learning mechanism, so that the representation robustness of the model under different voiceprint change conditions is improved. The method comprises the following steps: decomposing an original signal containing noise fringes into a plurality of modal components by using variational modal decomposition, extracting low-frequency effective components, introducing Gaussian white noise, constructing a corrosion target signal as a decoder training target, learning through a denoising automatic encoder, and finally outputting a denoised voiceprint feature signal. The method has the advantages of high robustness, high noise suppression capability, excellent feature expression capability and the like, is suitable for the field of online monitoring and intelligent diagnosis of the state of high-voltage circuit breaker equipment, and has good application prospect and engineering value.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Photovoltaic equipment fault detection method and system based on voiceprint recognition

The invention relates to the technical field of voiceprint recognition, in particular to a photovoltaic equipment fault detection method and system based on voiceprint recognition, and the method comprises the following steps: collecting the sound of the surrounding environment of a current photovoltaic power station, and building scene associated equipment spectrum peak features; according to the method, classification of environment sound is introduced in the sound acquisition stage, and scene association is performed on the operation sound of the photovoltaic equipment and the surrounding background sound in combination with Mel-frequency cepstrum coefficient extraction and acoustic scene recognition, so that a more targeted reference framework is provided for subsequent analysis. Modeling of multi-source sound features is achieved by synchronously collecting operation sound of an inverter and a cooling fan in equipment and extracting spectrum peak frequency points and amplitudes. In the signal processing process, multiple sub-bands are divided on the basis of scene correlation spectrum peak frequency points, so that each sub-band is aligned with the acoustic characteristics of equipment, the structural precision of short-time energy envelope extraction is improved, and it is ensured that the cross-correlation analysis result is more reliable.
Owner:WUWEI SHENNENG NORTH ENERGY DEV CO LTD

Speech emotion recognition method and system based on multi-scale adaptive feature fusion

The invention relates to a speech emotion recognition method and system based on multi-scale adaptive feature fusion. A speech signal is acquired and preprocessed to obtain a Mel-frequency cepstrum coefficient; time features and frequency features of Mel-frequency cepstrum coefficients are extracted and fused to obtain multi-scale features, the multi-scale features are divided into a global information estimation branch and an efficient self-attention branch through channel expansion, the efficient self-attention branch extracts fine-grained local features through self-attention, the global information estimation branch extracts low-frequency content through downsampling, and the high-efficiency self-attention branch extracts low-frequency content through down-sampling. Non-local information is captured in combination with global variance modulation; by fusing the global features and the local features obtained by the two branches, the relevance between different features is mined, deep feature representation is obtained, and fused deep time-frequency features are obtained; and classifying the fused deep-layer time-frequency features by using a full-connection network, and determining an emotion category corresponding to the voice signal.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Laying hen voice recognition method and system fusing acoustic features and deep learning features

The invention provides a laying hen voice recognition method and system fusing acoustic features and deep learning features. The method comprises the steps of obtaining a to-be-recognized original audio signal and a voice recognition model; wherein the voice recognition model comprises a feature extraction network, a feature fusion network and a classification recognition network; performing feature extraction on the original audio signal by using the feature extraction network to obtain a spectrogram feature, a Mel-frequency cepstrum coefficient feature and a deep speech feature; the feature fusion network performs feature fusion on the spectrogram features, the Mel-frequency cepstrum coefficient features and the deep speech features by using a collaborative attention mechanism or a multi-head attention mechanism to obtain fused features; and inputting the fused features into a classification recognition network to obtain a voice recognition result. According to the method, the advantages of various characteristics can be fully utilized, and the sound signals are described and analyzed from multiple angles, so that the voiceprint of the laying hen is more accurately recognized, and the voiceprint recognition accuracy of the laying hen is remarkably improved.
Owner:BEIJING RES CENT FOR INFORMATION TECH & AGRI

Transformer fault detection method based on feature fusion

The invention belongs to the technical field of power equipment state monitoring, and discloses a transformer fault detection method based on feature fusion. The method comprises the following steps: acquiring and preprocessing a transformer sound signal; feature extraction of a Mel frequency cepstrum coefficient (MFCC) and a power regularization cepstrum coefficient (PNCC) is carried out on the preprocessed sound signals; using a Fisher criterion-based feature fusion algorithm to perform weighted fusion on the MFCC feature parameters and the PNCC feature parameters to generate fused feature parameters; inputting the fusion feature parameters into a deep neural network model for training; inputting the obtained fusion feature parameters of the sound signals of the transformer to be detected into the trained deep neural network model to obtain a detection result of the fault category of the transformer; according to the method, MFCC and PNCC feature parameters are fused, MPFC features are obtained through dynamic weighting, frequency domain information, anti-noise information and sensitive frequency band information are integrated, and the detection accuracy and the anti-interference capacity are improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Bluetooth earphone AI voice control method and system

The invention relates to the technical field of voice recognition, in particular to a Bluetooth headset AI voice control method and system, and the method comprises the following steps: collecting a voice sample, carrying out the feature extraction of a voice signal through employing a Mel-frequency cepstrum coefficient, and generating voice feature data; and inputting the voice feature data into an acoustic model, and improving the recognition rate of the key instruction words by the acoustic model through learning features to obtain an optimized acoustic model. According to the method, acoustic feature capture is realized through voice signal feature extraction, the recognition accuracy of the key instruction is improved in cooperation with deep learning of acoustic features, the influence of different user speech speeds on the recognition effect is overcome by applying the time alignment technology of voice input, and the stability of the control instruction in different use scenes and speech speed changes is ensured. In addition, the recognition capability of a specific command is further mined and optimized by means of statistical characteristic analysis of voice, and efficient and automatic recognition and extraction of key control commands in continuous speech streams are completed.
Owner:JIANGXI CHANGRONG TECHNOLOGY CO LTD

Artificial intelligence autonomous identification model

The invention relates to the field of mine intelligent management, and discloses an artificial intelligence autonomous identification model, which comprises a data acquisition module used for acquiring visual data, infrared data, audio data and environmental data from a coal mine site; the data preprocessing module is used for carrying out scale normalization, data enhancement, short-time Fourier transform, Mel frequency cepstrum coefficient extraction and normalization processing on the collected data; and the feature extraction module is used for extracting features from the visual data, the infrared data, the audio data and the environment data through a deep neural network. And through a multi-modal information fusion technology, the adaptability to a complex environment in mine operation is improved, the complementarity of different modal data is fully mined, and the recognition precision and robustness are remarkably improved. And an introduced feedback and self-learning optimization mechanism enables the model to be automatically adjusted according to real-time data, so that the stability and adaptability in a mine environment are enhanced.
Owner:NINGXIA ANZHENG SCI & TRADE CO LTD

Water supply network leakage identification method and device based on data enhancement and hybrid neural network architecture, equipment and medium

The invention relates to a water supply network leakage identification method and device based on data enhancement and hybrid neural network architecture, equipment and a medium, and the method comprises the steps: employing a generative adversarial network to generate a simulation pipeline vibration audio signal, and combining the simulation pipeline vibration audio signal with a real pipeline vibration audio to construct a pipeline vibration audio signal data set; performing feature fusion on the Mel-frequency cepstral coefficient feature matrix and a first-order difference matrix of Mel-frequency cepstral coefficients to determine a feature fusion matrix of each pipeline vibration audio signal frame; by taking the feature fusion matrix as a training sample and the leakage state of the pipeline vibration audio signal as a sample label, training a CNN-Bi LSTM hybrid neural network model to determine a water supply network leakage detection model; and inputting the to-be-identified pipeline vibration audio signal into the water supply network leakage detection model to determine whether the to-be-identified pipeline vibration audio signal is in a leakage state or a non-leakage state. According to the invention, the precision, robustness and generalization ability of leakage detection are significantly improved.
Owner:GUANGDONG UNIV OF TECH

Smart classroom interaction analysis method based on double-layer architecture voice segmentation

The invention provides a smart classroom interaction analysis method based on double-layer architecture voice segmentation, and relates to the technical field of voice segmentation, and the method specifically comprises the following steps: extracting the voice features of a voice signal through employing a Mel-frequency cepstrum coefficient MFCC; designing a text-enhanced multi-scale time sequence-based perception time delay neural network, performing coarse screening on voice features, and dividing an audio clip into a single-speaker clip and a multi-speaker clip; and inputting the coarsely screened multi-speaker segment into a sliding window segmentation model SW-NIF fused with adjacent window information, and positioning speaker conversion points in the multi-speaker segment. And training the constructed model on the data set and verifying the model. According to the technical scheme, the problems that in the prior art, the segmentation problem of the classroom audio is neglected, and only the classroom audio is simply segmented for subsequent tasks, so that speakers in audio clips are mixed, and the analysis effect is affected are solved.
Owner:SHANDONG UNIV OF SCI & TECH

Transformer fault voiceprint detection method and system based on multi-spectrum feature fusion

The invention relates to the technical field of power equipment monitoring, and particularly discloses a transformer fault voiceprint detection method and system based on multi-spectrum feature fusion, and the method comprises the steps: obtaining voiceprint sample data of a transformer, carrying out the processing of the de-noised voiceprint data, obtaining a Mel-frequency cepstrum coefficient, and carrying out the detection of the Mel-frequency cepstrum coefficient; calculating a first-order difference coefficient and a second-order difference coefficient based on the Mel-frequency cepstral coefficient to obtain Mel-frequency cepstral coefficient characteristics, first-order difference coefficient characteristics and second-order difference coefficient characteristics; the first-order difference coefficient features and the second-order difference coefficient features are endowed with weights, and in combination with the Mel-frequency cepstral coefficient features, weighted Mel-frequency cepstral coefficient feature vectors are obtained; performing feature extraction on the de-noised voiceprint data to obtain voiceprint slice feature vectors; according to the method, effective features can be accurately extracted in a complex noise environment, and the reliability and adaptability of transformer fault detection are improved.
Owner:ZHANGJIAKOU POWER SUPPLY COMPANY OF STATE GRID JINBEI ELECTRIC POWER COMPANY

Voice emotion recognition method based on multiple scales and multiple features

The invention discloses a voice emotion recognition method based on multiple scales and multiple features, and belongs to the technical field of artificial intelligence. The method comprises the following steps: firstly, preprocessing an audio signal and extracting a spectrogram and a Mel-frequency cepstral coefficient; then, a residual network, a bidirectional long-short-term memory network and a HuBERT pre-training model are respectively utilized to extract spectrogram high-order spatial features, time sequence context features and voice semantic embedding features; secondly, inputting the first two features into a multi-dimensional multi-scale feature extraction module to extract richer time-frequency features, performing deep fusion by using a multi-layer cross attention mechanism, and performing weighted fusion with speech semantic embedded features; and finally, all the advanced features are spliced, and a final emotion category is recognized through a full-connection classifier. According to the invention, through combination of multi-scale feature extraction and an advanced fusion mechanism, the problem of insufficient complex emotion modeling ability in the prior art is effectively overcome, and the accuracy and robustness of voice emotion recognition are significantly improved.
Owner:NANJING INST OF TECH

Loss and fatigue monitoring method for water gate opening and closing equipment

The invention relates to the technical field of equipment monitoring, in particular to a loss and fatigue monitoring method for water gate opening and closing equipment. The method comprises the following steps that multi-mode time sequence data are synchronously collected through a vibration sensor, a temperature sensor and an acoustic emission sensor, and the multi-mode time sequence data comprise a vibration signal, a temperature signal and an acoustic emission signal; performing wavelet packet decomposition on the vibration signal, and extracting an energy ratio of a preset frequency band as a vibration energy feature; calculating a temperature rise rate and a local range of a sliding window of the temperature signal as temperature characteristics; extracting a Mel-frequency cepstrum coefficient from the acoustic emission signal as an acoustic feature; the vibration energy characteristics, the temperature characteristics and the acoustic characteristics are spliced into a multi-dimensional time sequence characteristic matrix according to time steps. According to the method, the maximum probability state transition path is solved by using the hidden Markov model, and accurate identification and prediction of the equipment loss and the fatigue state are realized.
Owner:GUANGDONG RES INST OF WATER RESOURCES & HYDROPOWER

Method and system for remotely diagnosing power system fault of tractor

The invention relates to the technical field of tractor power system fault diagnosis, and discloses a method and system for remotely diagnosing a power system fault of a tractor, and the method comprises the steps: monitoring a vibration signal, a sound signal and a temperature signal of a transmission gear of the tractor; extracting a frequency domain feature, a Mel frequency cepstrum coefficient feature and a temperature change trend feature of the vibration signal; carrying out weighted fusion on the signal features by adopting a dynamic weighting mechanism; constructing a fault prediction model based on an XGBoost algorithm, and outputting the gear wear degree and the fault occurrence time; and generating a fault early warning and transmitting the fault early warning to a mobile terminal APP and a cloud platform through an LTE or 5G network. Compared with a traditional fault diagnosis method in the prior art, particularly under the condition of dynamic working conditions, the technical problem that efficient and accurate fault diagnosis of the tractor power system is difficult to realize is solved, and the accuracy and response efficiency of fault diagnosis are improved by combining the convolutional neural network and the XGBoost algorithm.
Owner:XUZHOU XINGHAO NEW ENERGY TECH CO LTD

Transformer abnormity identification method based on voiceprint feature analysis

The invention discloses a transformer abnormity identification method based on voiceprint feature analysis, and belongs to the field of power equipment state monitoring and intelligent diagnosis. The method comprises the following steps: firstly, analyzing an iron core acoustic mechanism based on a magnetostrictive effect, and establishing a three-dimensional model through finite element simulation to obtain vibration and sound field characteristics; in a complex substation environment, a hybrid noise reduction method combining density peak clustering and a CEEMDAN-wavelet threshold is provided, and the signal-to-noise ratio is effectively improved. Then extracting Mel-frequency cepstrum coefficients (MFCC) and spectrum features, and performing local linear embedding (LLE) dimension reduction to form a compact feature set; in the recognition stage, a convolutional neural network framework is designed, specifically, a spectrogram and an energy spectrum are modeled through a two-dimensional CNN, an MFCC tensor obtained after dimensionality reduction is modeled through a three-dimensional CNN, and accurate diagnosis of mechanical faults such as core looseness is achieved. The method has the advantages of being non-contact, anti-noise and high in recognition precision, and real-time diagnosis and early warning of mechanical abnormity of the transformer can be achieved under complex working conditions.
Owner:YANCHENG POWER SUPPLY CO STATE GRID JIANGSU ELECTRIC POWER CO

Smart home central control system and method based on multi-mode perception

The invention relates to the technical field of smart home control, and particularly discloses a smart home central control system and method based on multi-mode perception, and the system comprises the steps: synchronously collecting a voice audio signal, a gesture image signal, an infrared thermal imaging signal and a millimeter wave radar signal through a plurality of groups of sensors; performing blind source separation processing on the voice and gesture signals, extracting a voice command component and a gesture action component which are independent in statistics, and performing space-time alignment and Kalman filtering fusion on the infrared and radar signals to generate a dynamic environment sensing graph; voice intention features, gesture track features and environment anomaly features are extracted through Mel frequency cepstrum coefficient analysis, skeleton key point tracking and multi-level convolution processing; constructing a three-dimensional decision matrix based on the features, performing weighted evaluation through a fuzzy logic rule base to generate a control instruction priority sequence, and dynamically adjusting an equipment operation mode according to the priority; according to the invention, the problems of control conflict and response delay caused by multi-mode signal coupling are solved.
Owner:XIAN QINGYAO HEZHI INTELLIGENT TECHNOLOGY CO LTD

Underwater target identification method and system based on software and hardware cooperation

The invention provides an underwater target identification method and system based on software and hardware cooperation, and the method comprises the steps: obtaining underwater sound audio data, carrying out the preprocessing of the underwater sound audio data based on a hardware circuit module, and extracting the audio feature data in the underwater sound audio data, the audio feature data comprises a LOFAR spectrum image, a Mel-frequency cepstrum coefficient spectrum image, a continuous wavelet transform spectrum image and a DEMON spectrum image; inputting each kind of extracted audio feature data into a neural network model, and outputting a corresponding underwater target recognition result; and carrying out weighted fusion on the underwater target recognition results corresponding to the various audio feature data to obtain a final underwater target recognition result. According to the underwater target recognition method and system, a hardware circuit is used for preprocessing the underwater sound audio data and extracting the audio feature data, only software is used for processing the target recognition part, and compared with an existing method that software is used for processing in all processing stages, the real-time performance of underwater target recognition is high, and the power consumption of a software module is low.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

Heart sound anomaly detection method and device, electronic equipment and storage medium

PendingCN120340543AStethoscopeSpeech analysisAbnormal heart soundsCardiac cycle
The invention provides a heart sound anomaly detection method and device, electronic equipment and a storage medium. An original heart sound signal is acquired and preprocessed into a standard heart sound signal, a plurality of cardiac cycles are divided to form a periodic heart sound signal, and corresponding Mel filter bank coefficient features and Mel frequency cepstral coefficient features are extracted; fusing the Mel filter bank coefficient features and the Mel frequency cepstrum coefficient features, inputting the fused Mel filter bank coefficient features and Mel frequency cepstrum coefficient features into an encoder in a pre-trained MobileNetV2 network, capturing a long-distance dependency relationship in the periodic heart sound signals by using a self-attention mechanism, and outputting a potential space representation vector; and inputting the potential space representation vector into a classifier of the MobileNetV2 network, and outputting a prediction result through a Softmax function. According to the method, the model complexity can be greatly reduced while the model precision is guaranteed, effective features can be extracted while low calculation overhead is kept, and the recognition capability and generalization performance of the model on abnormal heart sounds can be enhanced.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

Spoken language evaluation method and device based on deep learning and medium

The invention discloses a spoken language evaluation method and device based on deep learning and a medium, and relates to the technical field of spoken language evaluation, and the method comprises the steps: collecting spoken language audio signals, carrying out the acoustic feature extraction of the spoken language audio signals through Mel-frequency cepstrum coefficient transformation, and generating an acoustic feature vector sequence; constructing a deep learning pronunciation diagnosis model, inputting the acoustic feature vector sequence into the deep learning pronunciation diagnosis model, calculating a multi-dimensional distance between each voice segment in the acoustic feature vector sequence and the phoneme prototype in a measurement space, and generating a pronunciation diagnosis result; converting the acoustic feature vector sequence into a text sequence through a speech recognition conversion method; and constructing a deep learning role analysis model, and inputting the text sequence into the deep learning role analysis model to generate a semantic role graph. According to the method, the acoustic deviation between the quantized speech segment and the standard phoneme in the measurement space is calculated through the phoneme prototype distance, and accurate space-time positioning and quantitative guidance of the pronunciation defect are realized.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Dialect intelligent customer service and culture knowledge base system

The invention relates to the technical field of agricultural travel services, and provides a dialect intelligent customer service and culture knowledge base system, the system comprises a perception layer, an analysis layer, a knowledge layer and an interaction layer four-dimensional architecture, the perception layer collects and preprocesses dialect voice, noise reduction is performed through spectral subtraction, and Mel frequency cepstrum coefficient features are extracted; the analysis layer identifies dialects based on a fine tuning Wav2Vec2.0 model, converts the dialects into mandarin through a Transform architecture, and completes intention identification and slot filling by using a BERT related model; the knowledge layer constructs a local culture knowledge graph containing non-abandoned, folk and other entities, and supports dynamic updating; and the interaction layer is combined with multiple rounds of dialogue management to generate multiform responses, and can be connected with an external service system. The system also optimizes a feedback module iterative model and knowledge. The system can cover more than ten dialects, realizes dialect interaction, culture interpretation and service conversion closed loop, and assists rural culture revitalizing and rural cultural travel service upgrading.
Owner:SHENZHEN BEIDOU DIGITAL TECH CO LTD

Rotor unmanned aerial vehicle detection method and device, electronic equipment and storage medium

The invention provides a rotor unmanned aerial vehicle detection method and device, electronic equipment and a storage medium, and belongs to the technical field of unmanned aerial vehicles, and the method comprises the steps: carrying out the Mel frequency cepstrum coefficient MFCC feature extraction of a to-be-processed sound signal, and obtaining an MFCC feature matrix; the MFCC feature matrix is input into an SVM classifier based on a cubic kernel, and a sound source classification result output by the SVM classifier is obtained; determining an unmanned aerial vehicle sound source based on the sound source classification result; and taking the unmanned aerial vehicle sound source as input data, and determining positioning information of the rotor unmanned aerial vehicle based on a three-dimensional SRP-PHAT algorithm. According to the rotor unmanned aerial vehicle detection method and device, the electronic equipment and the storage medium provided by the invention, the unmanned aerial vehicle can be accurately positioned in a complex environment in real time.
Owner:BEIJING GUOYAN RONGXING TECH CO LTD

High-pressure pipeline leakage detection method and device based on voiceprint map

The invention discloses a high-pressure pipeline leakage detection method and device based on a voiceprint map, and relates to the technical field of pipeline detection. According to the method, sound signals are collected through a distributed microphone array, multi-dimensional features such as wavelet packet frequency band energy and Mel-frequency cepstral coefficients are extracted after variational mode decomposition is combined with wavelet threshold denoising preprocessing, and dimensionality reduction is performed through principal component analysis; the time delay between the sensors is calculated based on a generalized cross-correlation-phase transformation algorithm, a leakage point is positioned through a particle swarm optimization algorithm, the leakage degree can be evaluated by combining a bidirectional long-short-term memory network of an attention mechanism, and the D-S evidence theory is utilized to fuse multi-sensor result early warning. The system realizes high-sensitivity detection and accurate positioning, is high in anti-interference capability, and is suitable for safety monitoring of high-pressure pipelines in the industries of petroleum, natural gas and the like.
Owner:XINJIANG XINYE ENERGY & CHEM CO LTD

Video content description method and device, electronic equipment and storage medium

The embodiment of the invention provides a video content description method and device, electronic equipment and a storage medium, belongs to the technical field of video processing, and is suitable for the fields of financial science and technology and medical treatment. The method comprises the following steps: acquiring target video data; performing multi-modal data extraction on the target video data to obtain a multi-modal data stream; performing spiking neural coding on the multi-modal data stream to obtain a spatial-temporal feature map; performing hierarchical time sequence transformation on the spatial-temporal feature map to obtain a semantic unit and a Mel-frequency cepstral coefficient; performing cross-modal feature integration on the semantic unit and the Mel-frequency cepstrum coefficient to obtain a cross-modal context representation; and performing text decoding on the cross-modal context representation to obtain target natural language information. According to the embodiment of the invention, the accuracy of video content description can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD