Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

114 results about "Spectrogram" patented technology

A spectrogram is a visual representation of the spectrum of frequencies of a signal as it varies with time. When applied to an audio signal, spectrograms are sometimes called sonographs, voiceprints, or voicegrams. When the data is represented in a 3D plot they may be called waterfalls.

Earthquake early detection system based on the analysis of spectrograms obtained by continuous wavelet transform using the YOLO classifier

UndeterminedKZ38143BEarthquake detectionAlgorithm
The invention relates to the field of seismology, signal processing and artificial intelligence, namely to automated methods and systems for early detection of earthquakes based on the analysis of seismic data, and can be used to recognize and classify longitudinal waves (P-waves) preceding the main seismic shocks, using deep learning and computer vision methods. The aim of the present invention is to create an automated early earthquake detection system using the classification of seismic signal spectrograms generated by the complex Morlet wave CWT using the YOLO deep neural network architecture. The technical result is an increase in the accuracy and speed of early earthquake detection by analyzing the time-frequency characteristics of seismic signals and automatically localizing P-wave signatures using a neural network model. The device includes a seismic sensor, an analog-to-digital converter, and a microprocessor implementing time-frequency analysis and neural network detection algorithms. The seismic signal is recorded in real time, digitized, segmented into time intervals, and subjected to preliminary digital processing, including noise filtering and amplitude normalization. Each time interval is converted into a time-frequency representation using a continuous wavelet transform, generating a two-dimensional distribution of signal energy over time and frequency. Based on the obtained data, a spectrogram is generated and fed to the YOLO neural network detection model, which is capable of automatically detecting and localizing longitudinal P-wave signatures. Based on the neural network analysis, a determination is made regarding the presence of a P-wave and its arrival time is determined. If a predetermined threshold is exceeded, an early warning signal is generated. The system provides for data accumulation and the possibility of subsequent retraining of the neural network model.
Owner:NON COMMERCIAL JOINT CO KAZAKH NAT UNIV NAMED AFTER AL FARABI

Quantum-derived newton-raphson optimal fractional order spectrogram generation method and system

PendingCN122290569ANonlinear scalingGlobal optimization
This invention provides a quantum-derived Newton-Raphson optimal fractional-order spectrogram generation method, comprising the following steps: Step 1: Acquire the original audio signal and construct a fractional-order spectrogram based on fractional Fourier transform (FRFT); Step 2: Perform nonlinear scaling compression on the fractional-order spectrogram using a Mel filter bank to generate a fractional-order Mel spectrogram; Step 3: Construct an adaptive optimization framework with information entropy minimization as the objective function to measure the information fidelity between the spectrogram and the original signal; Step 4: Use the quantum-derived Newton-Raphson optimization algorithm (QNRBO) to globally optimize the fractional-order order, frame length, and frame shift hyperparameters to generate the optimal fractional-order spectrogram; Step 5: Input the optimal fractional-order spectrogram into a downstream speech recognition model. This technical solution aims to systematically solve core problems such as insufficient traditional time-frequency representation capabilities, rigid hyperparameter configuration, limited optimization algorithm performance, and feature-task disconnect.
Owner:FUZHOU UNIV

A method and system for extracting line spectrum of time-frequency spectrum of underwater acoustic signal

PendingCN122451569ASolve the scarcityAccurately depict blurred boundariesTime domainFrequency spectrum
The application discloses a water acoustic signal time-frequency spectrum line spectrum extraction method and system, and belongs to the technical field of signal processing. A noisy time domain signal is generated through simulation, and a mask label of a time-frequency spectrum of the noisy time domain signal belonging to a line spectrum is generated through a soft threshold function; a denoising model is trained according to the time-frequency spectrum and the mask label; a target water acoustic signal is acquired, the time-frequency spectrum of the target water acoustic signal is input into the denoising model, and a mask label corresponding to the time-frequency spectrum input is output through inference; the time-frequency spectrum of the denoised target water acoustic signal is acquired according to the mask label and the time-frequency spectrum; an initial candidate point set of the time-frequency spectrum of the denoised target water acoustic signal is acquired, and an initial candidate point of a current frame time-frequency spectrum in the initial candidate point set is acquired; a correlation cost matrix is constructed, the initial candidate point and a trajectory are correlated and matched with the minimum difference as a target, and a line spectrum of the trajectory and the candidate point dynamic correlation is acquired. The method can balance denoising fidelity, detection accuracy and real-time performance.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA +1

System and method for identifying sentiment (emotions) in a speech audio input

In a system and method for enabling a user to identify the emotions of speakers during a telephone or online conversation, spoken audio input is pre-processed using a one-dimensional Mel Spectrogram and / or a two-dimensional Mel-Frequency Cepstral Coefficient (MFCC) matrix, reducing the two-dimensional matrix to a single dimension output, and identifying at least one emotion in the audio input using a convolutional or recurrent neural network.
Owner:VALENCE VIBRATIONS INC

A method and system for radio frequency fingerprint recognition of narrowband IoT transceivers

This invention discloses a method and system for radio frequency fingerprinting of narrowband IoT transceivers. The method involves a receiver acquiring wireless messages sent by the transceiver to be identified, obtaining in-phase / orthogonal discrete sample sequences, locating and extracting preamble sample segments, performing a short-time Fourier transform on the preamble sample segments to generate a spectrogram energy matrix, and extracting compressed time-frequency features through singular value decomposition. The main component of the carrier frequency offset is obtained based on the cross-correlation peak position shift between rising and falling chirps, and the carrier frequency offset correction component is obtained based on the phase drift of adjacent rising chirps. The compressed time-frequency features and the carrier frequency offset estimation results are concatenated into a fused feature vector, which is then input into a convolutional neural network classification model to output the device category and category confidence. This invention reduces the dependence on channel location features and decreases model input overhead, making it suitable for online identification of narrowband IoT devices.
Owner:SHANGHAI MARITIME UNIVERSITY

Pet and experimental animal music conditioning system and conditioning method

PendingCN122290627AData connectionHome use
This invention relates to the fields of animal behavior, animal welfare, and mental health conditioning, specifically a music conditioning system and method for pets and laboratory animals. It is a pure PC-based software system, including an audio editing and playback control module, a music library management module, and a music playback module. The output of the audio editing and playback control module is connected to the input of the music playback module, and the music library management module and the audio editing and playback control module have a bidirectional data connection. This invention precisely matches different animals through three parameters: pitch shifting, speed adjustment, and volume adjustment, significantly improving audio compatibility. The effects of emotional stabilization, anxiety relief, and sleep promotion are significantly better than general music. One-click templates are suitable for quick home use. Customizable visual editing is suitable for pet hospitals, laboratory animal centers, and research institutions to establish standardized programs. Waveform graphs, spectrograms, and acoustic spectrograms are displayed synchronously in real time, allowing users to intuitively judge the audio structure. Adjustment accuracy is improved by more than 80%, requiring no professional audio knowledge.
Owner:王子欣

Detecting RF emissions and classifying RF sources using spectrum transformer models

PendingUS20260187193A1Frequency spectrumEdge computing
Edge computing units that are outfitted with signal receivers and provided at local sites or edge locations are configured to capture RF signals and process the RF signals using transformer models. A spectrum encoder of a transformer model performs attention on frequency-encoded features extracted from spectrograms of the RF signals and features extracted from representations of RF signals associated with discrete RF events, and embeddings generated by the spectrum encoder are decoded to classify the RF signals or a source of the RF signals. Additionally, a spectrum decoder of a transformer model is configured to receive a query for information and to generate a response to the query using a spectrum decoder that performs attention on embeddings received from a spectrum encoder, as well as frequency-encoded features representative of the query, and features representative of semantic information regarding RF signals and sources of the RF signals.
Owner:ARMADA SYST INC

A voiceprint recognition method and device

ActiveCN114974256BFeature extractionMedicine
The application relates to a voiceprint recognition method and device. The method comprises the following steps: obtaining a spectrogram of a speech signal, and dividing the spectrogram into a plurality of sub-spectrograms of different frequency bands; using feature extraction networks with different time resolutions to extract feature information of the plurality of sub-spectrograms, wherein the time resolution of a first feature extraction network used for extracting feature information of a high-frequency sub-spectrogram is greater than the time resolution of a second feature extraction network used for extracting feature information of a low-frequency sub-spectrogram; and fusing the feature information extracted by the feature extraction networks with different time resolutions into a voiceprint of the speech signal.
Owner:HUAWEI TECH CO LTD

A method for detecting the fingerprint of chuan she gan oral liquid and a quality control method

The application discloses a kind of chuan shegan oral liquid fingerprint detection method and quality control method, comprising the following steps: using 40-60% ethanol solution to chuan shegan oral liquid is extracted by ultrasound, after centrifugation, filtration, dilution, obtain the liquid to be measured;Using ultra-high performance liquid chromatograph to detect the liquid to be measured, obtain chromatogram and spectrogram, establish the standard atlas of chuan shegan oral liquid;Chromatographic condition is: mobile phase A is acetonitrile, mobile phase B is 0.4% phosphoric acid aqueous solution, column temperature is 39-41 DEG C;The flow rate of mobile phase: 0.3-0.4mL / min;Detection condition: PDA detector, 3D scanning range 190-400nm and extract maximum value diagram, or using the chromatogram of detection channel for wavelength 264nm, can obtain two kinds of fingerprint. By liquid chromatography, shegan glycoside in chuan shegan and other ten or so compounds are determined at a time, simple and fast, good separation effect, high determination accuracy, good result reproducibility, easy to observe, strong specificity, good chromatographic peak peak shape.
Owner:SHANDONG JINZHUJI PHARM CO LTD +1

A method and device for enhancing through-the-wall radar signals of human motion

The application relates to a through-wall radar signal enhancement method and device for human motion, wherein the method comprises the following steps: acquiring a to-be-enhanced through-wall spectrogram; inputting the to-be-enhanced through-wall spectrogram into a generator in a network model based on spectrogram processing to obtain an enhanced through-wall spectrogram, so as to monitor human motion. The training method of the network model comprises the following steps: training a pre-training model through free space spectrograms and corresponding free space flip spectrograms in a data set; inputting a through-wall spectrogram into the generator to obtain a corresponding enhanced through-wall spectrogram; inputting the enhanced through-wall spectrogram and a non-paired free space spectrogram into the pre-training model which has been trained to obtain corresponding channel dimension autocorrelation matrices; inputting the enhanced through-wall spectrogram and the non-paired free space spectrogram into a discriminator to obtain a discrimination result; and constructing a loss function to optimize the generator and the discriminator, so that the generator can perform signal enhancement and improve the accuracy of downstream semantic recognition.
Owner:TIANJIN UNIV

A method and system for intelligent monitoring of soil voiding at the bottom of a caisson

The application discloses a kind of caisson bottom soil body void intelligent monitoring method and system, it is related to caisson construction monitoring technical field, the method includes in caisson bottom preset multiple monitoring positions, obtains the strain time series data of each monitoring position in construction process;Each strain time series data is labeled and length uniform, and the label includes void and non-void;Strain time series data and its label are made into two-class samples after length uniform, and two-class sample set is constructed;The one-dimensional strain time series data in the two-class sample is carried out short-time Fourier transform and generates two-dimensional time-frequency spectrogram;Model training is carried out to the void prediction model pre-set by two-class sample set after processing;The trained void prediction model is deployed and applied, and the intelligent monitoring of caisson bottom soil body void is realized.The application can realize the global, real-time, accurate discrimination of caisson bottom void state.
Owner:中铁桥隧技术有限公司 +2

A light-weight water area human behavior recognition method based on millimeter wave radar multi-spectrum fusion

PendingCN122330841AHuman behaviorFeature extraction
This invention provides a lightweight method for human behavior recognition in water areas based on millimeter-wave radar multispectral fusion, comprising the following steps: S1, processing the raw millimeter-wave radar echo data in the range, Doppler, and rhythm dimensions to construct a multispectral representation; S2, inputting the multispectral representation into a lightweight shared-domain perceptual encoder to extract features from different spectra, obtaining multispectral feature representations; generating temporal and spatial feature representations from the multispectral feature representations using a decoupled spatiotemporal dual-branch structure; S3, performing adaptive fusion processing on the temporal and spatial feature representations to generate spatiotemporal fusion features, and outputting the corresponding human behavior recognition result in water areas based on the spatiotemporal fusion feature representations. This invention, by introducing a range-rhythm spectrogram and combining lightweight shared feature extraction with a low-rank fusion structure, effectively reduces model complexity and computational overhead while ensuring recognition accuracy.
Owner:DALIAN MARITIME UNIVERSITY

Cross-modal silent speech reconstruction method and system based on ear canal air pressure micro-motion perception

The application discloses a cross-modal silent speech reconstruction method and system based on ear canal air pressure micro-motion sensing, and belongs to the technical field of human-computer interaction and wearable computing. The method uses a micro-pressure sensing unit placed in an in-ear earphone to collect a non-acoustic air pressure sequence caused by the movement of a sound-producing organ; through adaptive baseline drift suppression and rhythm perception data enhancement processing, a robust feature space is constructed; further, an end-to-end deep neural network containing domain adversarial adaptation, cross-modal semantic alignment, coarse-grained mel-spectrogram generation and residual detail correction is used to map the TPVS to a high-fidelity acoustic mel spectrum. The application effectively breaks through the technical bottleneck of the lack of high-frequency acoustic features in low-frequency mechanical signals, realizes high-precision silent speech command analysis in a mobile and noisy scene, and introduces a coupled quality evaluation gate and trigger-based start / stop control at the inference end to suppress invalid inference and reduce power consumption when wearing is poor or there is no trigger condition.
Owner:DONGHUA UNIV

Process state real-time correction method based on acoustic features

The application provides a process state real-time correction method based on acoustic characteristics, and relates to the technical field of process control, and comprises the following steps: collecting acoustic signals generated by interaction between a processing medium and a processing unit when the processing unit is executing a processing task; performing real-time spectrum analysis on the acoustic signals to extract characteristic frequencies and energy distribution characteristics thereof which are related to the current physical state of the processing medium; determining whether the current physical state of the processing medium is in an abnormal critical region in real time based on changes of the characteristic frequencies and the energy distribution characteristics thereof relative to a preset reference; the abnormal critical region comprises a first precursor state in which the processing medium tends to be excessively dried or a second precursor state in which the processing medium tends to be excessively accumulated; and if it is determined that the processing medium is in the abnormal critical region, the physical state of the processing medium being currently processed is corrected in real time by adjusting at least one execution parameter which influences the physical state of the processing medium before the processing medium being currently processed reaches a target region.
Owner:TIANJIN MINGJIE INTELLIGENT EQUIP CO LTD

A method, device, equipment, medium and product for labeling a sound sample

The present disclosure relates to a method and device for labeling sound samples, equipment, media and products, the method comprising: obtaining a set of sound samples to be labeled; wherein the set of sound samples to be labeled comprises a plurality of sound samples to be labeled; extracting the mel-spectral spectrogram information and the first sound statistical indicators of each sound sample to be labeled; determining the class of the first sound statistical indicators of each sound sample to be labeled by a target determination rule to obtain an initial label set corresponding to the set of sound samples to be labeled; wherein the target determination rule is a rule generated according to the statistical characteristics corresponding to each sound class in the labeled sound samples; and calibrating the initial label set based on the mel-spectral spectrogram information of each sound sample to be labeled to obtain a target label set. The present disclosure can realize batch labeling of the set of sound samples to be labeled, improve the labeling efficiency of sound samples, and improve the accuracy, consistency and reliability of the labeling results.
Owner:SHENHUA SHENDONG COAL GRP +1

Tibetan speech recognition method fusing tone perception hybrid expert and search correction

PendingCN122347949AElectronic industryVoice transformation
The application discloses a Tibetan speech recognition method fusing tone perception hybrid experts and retrieval error correction, and belongs to the technical field of signal processing in the electronic industry. The specific steps of the recognition method are as follows: I: Tibetan speech signals are acquired, and corresponding log-mel spectrogram features and fundamental frequency contour features are extracted; II: the tone gating weight is calculated based on the fundamental frequency contour features, and is dynamically routed to the corresponding hybrid expert network to acquire dialect-independent acoustic hidden layer features. The application effectively solves the model interference problem caused by the presence or absence of tones among multiple dialects, avoids the parameter conflict between tone dialects and non-tone dialects, significantly improves the recognition performance of multi-dialect hybrid training, corrects the homonym heterograph error commonly existing in Tibetan, improves the performance of multi-dialect Tibetan speech recognition, can be used for converting Tibetan speech into characters, and is helpful for protecting and mining Tibetan culture.
Owner:CHINA UNIVERSITY OF POLITICAL SCIENCE AND LAW

Sound event early detection

Systems and methods for Evidence-based Sound Event Early Detection is provided. The system / method includes parsing collected labeled audio corpus data and real time audio streaming data utilizing mel-spectrogram, encoding features of the parsed mel-spectrograms using a trained neural network, and generating a final predicted result for a sound event based on the belief, disbelief and uncertainty outputs from the encoded mel-spectrograms.
Owner:NEC CORP

A feature extraction method and system for speech signals

The present application relates to the field of data processing, and particularly relates to a feature extraction method and system of a voice signal. The method comprises the steps of: obtaining a voice signal; for any data in the voice signal, obtaining a plurality of data segments with different lengths centered on the data, calculating a stable definition threshold, starting from a reference length, continuously increasing the window size, after any time of window size adjustment, obtaining a window centered on the data based on the adjusted window size, performing Fourier transform on the signal in the window to obtain a frequency spectrum, calculating a stability degree, if the stability degree is less than the stable definition threshold for the first time, taking the window size after this adjustment as a target window size; performing short-time Fourier transform on the voice signal based on the target window size of each data in the voice signal to obtain a spectrogram, and realizing feature extraction of the voice signal based on the spectrogram. The accuracy of the extracted spectrogram is improved by setting a suitable window size.
Owner:GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

A method for detecting fake speech based on frequency band selection

ActiveCN116129913BImprove robustnessReduce feature sizeSpeech synthesisSpeech sound
This invention discloses a method for detecting spoofed speech based on frequency band selection. The method includes: acquiring a target speech signal; transforming the target speech signal to obtain spectrogram features; performing frequency band segmentation on the spectrogram features to obtain low-frequency sub-band features and high-frequency sub-band features; training a speech synthesis spoofed speech detection model using the low-frequency sub-band features; training a recording playback spoofed speech detection model using the high-frequency sub-band features; inputting the low-frequency sub-band features into the speech synthesis spoofed speech detection model; and inputting the cross-matched high- and low-frequency sub-band features into the recording playback spoofed speech detection model to obtain the final speech detection result. In this invention, the robustness of the neural network spoofed speech detection system is improved under conditions such as dataset mismatch, and the feature size is reduced through sub-band selection, thereby reducing the number of parameters and computational load for spoofed speech detection.
Owner:NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT

System and method for identifying sentiment (emotions) in a speech audio input with haptic output

In a system and method for enabling a user to identify the emotions of speakers to a conversation, spoken audio input is pre-processed using a one-dimensional Mel Spectrogram and / or a two-dimensional Mel-Frequency Cepstral Coefficient (MFCC) matrix, reducing the two-dimensional matrix to a single dimension output, identifying at least one emotion in the audio input using a convolutional or recurrent neural network, and providing the user with haptic feedback corresponding to the at least one emotion in the audio input.
Owner:VALENCE VIBRATIONS INC

An audio playing detection method, system, device, storage medium and product

This invention discloses an audio playback detection method, system, device, storage medium, and product. The method includes: collecting audio detection data and processing the audio detection data into a raw spectrogram; inputting the raw spectrogram as input data into a trained audio detection model to obtain a reconstructed spectrogram output by the audio detection model; projecting the reconstructed spectrogram and the raw spectrogram onto the frequency and time axes, respectively; determining the reconstruction error in the time and frequency dimensions based on the projection results; and determining that the audio playback is abnormal if the reconstruction error is greater than an error threshold. The audio playback detection method disclosed in this invention improves the accuracy of audio playback detection and reduces time and labor costs by processing the collected audio detection data into a raw spectrogram and judging whether there is an audio playback abnormality based on the reconstruction error between the reconstructed spectrogram output by the model and the raw spectrogram.
Owner:BEIJING CO WHEELS TECH CO LTD

Method for detecting abnormal activity of driver

PendingUS20260188026A1Frequency spectrumDriver/operator
A method for detecting abnormal activity of a driver implemented by a monitoring device which includes an angle acquisition unit that obtains an angle signal during a preset time period, and a processing unit. The method includes: generating a spectrogram from the angle signal; obtaining a frequency set within the preset time period; obtaining, based on the frequency set, candidate time point(s), each corresponding to a frequency in the spectrogram that is greater than a dynamic frequency threshold; obtaining a base frequency signal based on the spectrogram; converting the base frequency signal to a base angle signal using a signal conversion technique; obtaining a peak time point that corresponds to a maximum angle value of the base angle signal within the preset time period; and obtaining an occurrence period related to an occurrence of the abnormal activity based on the peak time point and the candidate time point(s).
Owner:MITAC DIGITAL TECH CORP

A method and system for partial discharge signal denoising based on time-frequency domain coordination

This application relates to the technical field of partial discharge (PD) detection, and in particular to a PD signal denoising method and system based on time-frequency domain coordination. The method includes: acquiring a prediction model built on an encoder-decoder framework and training it using a composite loss function; acquiring the PD signal and performing a time-frequency transformation to obtain the time spectrum of the noisy signal, the time spectrum of the noisy signal including the clean PD signal and noise components in the time spectrum; inputting the time spectrum of the noisy signal into the trained prediction model and outputting a corresponding prediction signal mask matrix, and multiplying the prediction signal mask matrix element-wise with the time spectrum of the noisy signal to generate the time spectrum of the denoised signal; and performing an inverse time-frequency transformation on the time spectrum of the denoised signal to obtain the denoised time-domain signal. This application aims to improve the noise suppression effect of PD signals in noisy environments.
Owner:ZHEJIANG HONGPU TECH CORP LTD

Intelligent dance choreography method and system based on generative AI and cloud-based motion library

PendingCN122312846AComputer graphics (images)Motion generation
This invention discloses an intelligent dance choreography method and system based on generative AI and a cloud-based motion library, relating to the field of artificial intelligence technology. It constructs a cloud-based motion library and combines it with dance music time frames for candidate motion matching. It utilizes dancer spatial relationships to generate dynamic graph adjacency weights and performs path search to complete motion sequence planning. Based on this, it performs spatial conflict correction and keyframe interpolation to generate continuous motions. Furthermore, it introduces spectrogram theory to enhance and optimize the adjacency structure, achieving secondary motion planning and spatial constraint correction. This generates a target dance motion sequence with good formation structure and temporal continuity, solving the problems of disconnect between motion generation and spatial formation, and insufficient expressive ability of dynamic formation structure in traditional dance choreography.
Owner:SICHUAN NORMAL UNIV

A spectrogram recognition closed loop system and method

The application discloses a spectrum image recognition closed loop system and method. It includes: obtaining spectrum sequence data and view range, generating a segmentation list; rendering a spectrum image for each segment and recording the drawing area boundary; performing target detection inference to obtain a pixel boundary box; based on the drawing area boundary and the segmented frequency, the pixel box is inverted to a frequency interval and aligned to a frequency point grid; performing deduplication, overlap elimination and threshold filtering on the candidate set to output stable candidates; based on the stable candidates, statistical features are calculated on the original sequence and signal parameters are output; the recognition process and results are structured and written to disk; the product is converted into a training data set to generate annotations; the model weight is trained and registered and managed, and the inference service is accessed through the scene binding mechanism to form a closed loop iteration. The application has the beneficial effects of significantly improving the automation level, result reliability and system evolution capability of spectrum monitoring and analysis.
Owner:ZHEJIANG YUANCHU DATA TECH CO LTD +1