Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1190 results about "Spectrogram" patented technology

A spectrogram is a visual representation of the spectrum of frequencies of a signal as it varies with time. When applied to an audio signal, spectrograms are sometimes called sonographs, voiceprints, or voicegrams. When the data is represented in a 3D plot they may be called waterfalls.

Intelligent risk early warning method and system based on multi-dimensional data analysis

The invention relates to the field of enterprise risk early warning analysis, in particular to an intelligent risk early warning method and system based on multi-dimensional data analysis. The method comprises the following steps: acquiring a multi-dimensional enterprise data stream, performing heterogeneous index information analysis and logic hierarchy reconstruction, and constructing an enterprise running state sensing map; performing multi-index local fluctuation amplitude calculation on the enterprise operation state sensing map, and performing dynamic disturbance feature mining to construct a risk disturbance spectrogram; performing deep semantic analysis and node state abrupt change feature analysis on the risk disturbance spectrogram to generate a dynamic transition type of an abrupt change node; and performing full-period time sequence tracing on the risk disturbance spectrogram, performing potential risk node prediction based on the dynamic transition type, and identifying other potential risk propagation ports. Through accurate and efficient enterprise risk perception, the risk intervention decision is made in advance, and the anti-risk capability and the operation stability of the enterprise are improved.
Owner:BEIJING HAOHONGDA XUNJIE TECHNOLOGY DEVELOPMENT CO LTD

Precise interference avoidance method and device in radio system

The invention discloses a precise interference avoidance method and device in a radio system, and relates to the field of signal processing, and the method comprises the steps: constructing a sensing matrix through sensing node data, and capturing a transient interference signal through aperiodic scanning; performing tensor decomposition on the signal data to extract time domain, frequency domain, space domain and modulation domain features, constructing a dual-mode spectrum analysis model, reconstructing an instantaneous spectrogram by using compressed sensing, and predicting an interference mode through LSTM; after the instantaneous spectrogram and the predicted interference graph are fused, threat assessment is carried out through a multi-stage interference classification model; according to the interference category and the threat level, beam forming is optimized, adaptive null is generated, a power density optimization model is constructed, and the transmitting power is dynamically adjusted; an anti-interference frequency hopping sequence is generated based on a chaotic mapping algorithm, and spectrum camouflage and tracking interference resistance are realized. The method has the advantages that accurate identification and dynamic avoidance of interference are realized through multi-dimensional perception, intelligent prediction and adaptive beam forming, and the interference avoidance capability of a wireless system is improved.
Owner:BEIJING BOHONG KEYUAN INFORMATION TECH CO LTD

Identification method of multi-model fusion signal modulation mode based on time-frequency diagram

The invention discloses a method for identifying a multi-model fusion signal modulation mode based on a time-frequency diagram, and belongs to the field of signal processing. Preprocessing the original signal to obtain a signal sample; constructing a data set by using the normalized signal samples; constructing a fusion model; training a modulation identification module in the fusion model; performing modulation mode identification on an input unknown signal by using the fusion model to obtain a preliminary identification result; and performing comprehensive judgment on the preliminary recognition result of the fusion model by using a comprehensive judgment device to obtain a final recognition result. According to the invention, by combining a plurality of time-frequency analysis methods, the time-frequency domain feature information of the signal is fully extracted; the reliability and robustness of a final decision are improved by adopting a multi-model fusion framework, and the modulation recognition performance of the system under the condition of a low signal-to-noise ratio is improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Method for providing video and electronic device supporting the same

An electronic device is provided. The electronic device includes a memory, and at least one processor electrically connected to the memory, wherein the at least one processor is configured to obtain a video including an image and an audio, obtain information on at least one object included in the image from the image, obtain a visual feature of the at least one object, based on the image and the information on the at least one object, obtain a spectrogram of the audio, obtain an audio feature of the at least one object from the spectrogram of the audio, combine the visual feature and the audio feature, obtain, based on the combined visual feature and audio feature, information on a position of the at least one object the information indicating the position of the at least one object in the image, obtain an audio part corresponding to the at least one object in the audio, based on the combined visual feature and audio feature, and store, in the memory, the information on the position of the at least one object and the audio part corresponding to the at least one object.
Owner:SAMSUNG ELECTRONICS CO LTD

Water turbine fault classification diagnosis method based on multi-modal fusion and meta learning

The invention discloses a water turbine fault classification diagnosis method based on multi-modal fusion and meta-learning. The method comprises the following steps: step 1, respectively extracting time domain features and Mel-language spectrogram features of monitoring noise signals of a water turbine through a feature extraction module; step 2, performing cross-modal attention mechanism fusion on the time domain features and the voiceprint features through a multi-modal fusion module, and adjusting a fusion weight based on a dynamic weight distribution mechanism; and step 3, performing small sample training optimization on the fused features through a meta-learning module, and improving the classification capability of the model for new fault types in combination with a Triplet Loss-KNN algorithm and twin network pre-training. The fault diagnosis method based on the time domain-voiceprint fusion network and meta learning has the advantages of being rapid in diagnosis, accurate in classification, high in generalization ability and the like, and the fault diagnosis precision of the water turbine based on noise signals can be effectively improved.
Owner:CHINA THREE GORGES UNIV

Multi-language cross-culture communication auxiliary method and system based on large model

The invention provides a multi-language cross-culture communication assisting method and system based on a large model. The method comprises the following steps: receiving a source language audio stream during a call, calling a multi-language sound frequency harmonic modulation feature library to extract fundamental frequency harmonic intensity distribution and tone turning features, and generating a cultural acoustic fingerprint vector; based on the vector, controlling a microphone array phase difference, directionally enhancing a fundamental frequency harmonic component of a speaker and suppressing noise, and outputting a high signal-to-noise ratio spectrogram; analyzing the pronunciation rhythm and tone turning characteristics of the spectrogram, capturing the pitch jump and duration of the syllable boundary, and generating an acoustic culture label; associating the spectrogram with a target semantic library, matching harmonic distribution and a cultural context rule based on a large model, and outputting a cultural interpretation prompt containing an ambiguity resolution suggestion; and generating a calibration result according to the acoustic tag and the semantic prompt, and overlapping the dynamic floating subtitles to the face area of the speaker in the video conference picture. According to the invention, cultural tone ambiguity in multi-language communication is eliminated.
Owner:LUSTER LIGHTWAVE CO LTD

Systems and methods for multimodal indexing of video using machine learning

Systems, methods, and computer-readable media are disclosed for systems and methods multimodal indexing of video using machine learning. An example method may include deceiving, by a video encoder of an audio-video transformer neural network comprising one or more computer processors coupled to memory, a first frame and a second frame associated with a first segment of a video. The example method may also include receiving, by an audio encoder of the audio-video transformer neural network, an audio spectrogram comprising first audio data associated with the first segment of the video. generating, by the video encoder, a first video embedding. The example method may also include generating, by the audio encoder, a first audio embedding. The example method may also include determining a fusion of the first video embedding and the first audio embedding using a multimodal bottleneck token. The example method may also include determining an output including the first video embedding and the first audio embedding. The example method may also include determining a classification of the first portion of the video based on the output.
Owner:AMAZON TECH INC

Multi-modal forged video detection method based on multi-head addition cross attention mechanism

The invention discloses a multi-mode counterfeit video detection method based on a multi-head addition cross attention mechanism, and belongs to the technical field of video counterfeit detection. The method comprises the following steps: preprocessing a video stream, decomposing a single-frame positioning face, and extracting an audio to generate a Mel spectrogram slice; the 3D convolutional network extracts video spatio-temporal features and motion differences, and the filter bank extracts audio features in combination with the residual network; audio features are mapped to a video alignment space through asymmetric projection, the video features are subjected to bidirectional interaction with an audio input multi-head addition cross attention module after being subjected to time sequence coding, and audio dominant and video dominant features are generated and are cascaded and fused with original features; and constructing cross-modal similarity loss constraint feature distribution, and fusing feature dynamic weighting and time sequence compression to output four classification probabilities of audio-visual double true, audio-visual double pseudo, video pseudo-audio true and video pseudo-audio pseudo. The multi-mode counterfeiting recognition precision is improved, and texture abnormity and audio and video mismatch features are captured.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Hierarchical emotional speech generation method and device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a hierarchical emotional speech generation method, which comprises the following steps: acquiring an input text and an emotional speech sample, extracting a text embedding feature from the input text, extracting a Mel spectrum feature from the emotional speech sample, and obtaining a text embedding feature; the method comprises the following steps: extracting phoneme-level, word-level and statement-level sentiment distribution characteristics through a hierarchical sentiment distribution extraction module, carrying out time dimension alignment, generating a multi-level sentiment guidance matrix, and inputting the multi-level sentiment guidance matrix and text embedding characteristics into a sentiment synthesis module to generate a target Mel spectrum; and finally, converting the target Mel spectrum into target voice through a vocoder. The emotion control is expanded from the statement level to the fine-grained level of phonemes, words and the like, and the multi-level emotion guidance matrix is combined for generation, so that the generated speech is finer and more abundant in emotion expression and conforms to the context, and the naturalness of speech generation and the accuracy of emotion transmission are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Radio frequency fingerprint identification method and system based on convolution-attention mechanism and multi-packet reasoning

The invention relates to a radio frequency fingerprint identification method and system based on a convolution-attention mechanism and multi-packet reasoning, and belongs to the technical field of communication networks and artificial intelligence. Comprising the following steps: step 1, capturing an equipment transmission signal, and preprocessing the transmission signal to obtain a spectrogram; step 2, constructing a radio frequency fingerprint identification model, and inputting the obtained spectrogram into the radio frequency fingerprint identification model for training to obtain a trained radio frequency fingerprint identification model; and step 3, performing prediction by using the trained radio frequency fingerprint identification model to obtain a final prediction result. Aiming at the problems of high calculation complexity and redundant module interaction existing in a traditional model fusion mechanism, a module separation strategy is adopted, so that a convolution module and a Transform module are independently optimized, the calculation complexity is reduced, and the stability of the model is improved.
Owner:SHANDONG NORMAL UNIV

Full-automatic cable vibration high-precision measurement method based on machine vision technology

The invention discloses a full-automatic cable vibration high-precision measurement method based on a machine vision technology, and relates to the technical field of structural vibration measurement. The method comprises the following steps: installing and calibrating equipment, and ensuring that a cable plane coincides with an imaging plane; high-frame-rate video acquisition is carried out, and comparison verification is carried out in combination with an accelerometer; screening effective pixels based on the pixel intensity time sequence variance, and constructing a pixel intensity space-time matrix by using the time sequence of the pixels; sVD decomposition is adopted to extract an accumulated energy ratio gt; 90% of the main modes are reconstructed, and real signals are revealed; a plurality of inhaul cables are automatically positioned and coded through gradient features and Hough transform; sub-pixel-level vertical displacement is calculated by applying an optical flow method, and median filtering is combined to resist noise; fourier transform is carried out to obtain a spectrogram, and a fundamental frequency is identified; and calculating cable force based on a string pulling theorem. According to the method, remote, full-field and multi-target synchronous monitoring is realized, the displacement measurement precision reaches a sub-pixel level, and a high-precision and full-automatic solution is provided for health monitoring of structures such as bridges.
Owner:NINGBO ORIENTAL UNIV OF TECH (TEMPORARY NAME)

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Diffusion-based audio purification for defending against adversarial deepfake attacks

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”). Embodiments implement a machine-learning architecture having a diffusion model that generates purified features that are fed to a deepfake detection model. The machine-learning architecture includes input layers that convert an audio signal into a Gaussian or frequency space representation (e.g., log spectrogram) to extract a set of initial features indicative of spoofing or deepfake attacks. The diffusion model identifies adversarial noise on the audio signal in the initial features and generates purified features or clean version of the input audio signal. A deepfake detector includes a neural network architecture and classifier programmed and trained to generate a deepfake detection score and classify the audio signal as genuine or fraudulent using the purified features.
Owner:PINDROP SECURITY INC

Target detection method and system based on millimeter wave radar

The invention discloses a target detection method and system based on a millimeter wave radar. The target detection method and system are used for realizing accurate detection, positioning and dynamic and static recognition of multiple targets under a complex background. According to the method, a distance-Doppler spectrogram is generated through the technical means of sliding window construction, spectral analysis, clutter suppression and the like, and candidate target points are detected by adopting an SO-CFAR algorithm. Then, determining a target position through high-resolution direction estimation and coordinate transformation, performing spatial clustering in combination with a density-based DBSCAN algorithm, and extracting a target geometric center and a bounding box; in the aspect of target tracking, Kalman filtering is used for predicting and updating the position and speed of the target, and a beam forming technology is used for enhancing a target signal, so that the target recognition stability is improved. And finally, the system performs robust dynamic and static state recognition on the target through a dynamic and static judgment module, so that high precision and robustness of the target detection process are ensured. The method can effectively cope with static background interference and dynamic target changes, and is suitable for target detection and tracking in a complex environment.
Owner:HANGZHOU DIANZI UNIV

Power transmission line fault diagnosis method based on time-frequency multistage fusion and bidirectional time sequence enhancement network

The invention relates to a power transmission line fault diagnosis method based on time-frequency multistage fusion and a bidirectional time sequence enhancement network, and belongs to the field of power system power transmission line fault diagnosis. The method comprises the following steps: firstly, constructing a time-frequency multi-level fusion and bidirectional time sequence enhancement network: converting input power transmission line partial discharge signal data into a time-frequency spectrogram; constructing a multi-scale time-frequency feature joint extraction module, and performing feature extraction on the time-frequency spectrogram by dual offset branches; constructing a dynamic weighted fusion module, and carrying out adaptive fusion on the extracted time-frequency features; a bidirectional time sequence information enhancement module is constructed, modeling is performed on the fusion features through a bidirectional long-short-term memory network, and feature expression of key discharge moments is enhanced by adopting an attention guiding mechanism; and finally, training the time-frequency multi-level fusion and bidirectional time sequence enhancement network by using the training set, and finally realizing fault type identification through a full connection layer and a Softmax classifier. According to the invention, the accuracy of power transmission line fault diagnosis and the model robustness are improved.
Owner:KUNMING UNIV OF SCI & TECH

Sound quality improving method, device and equipment of Bluetooth sound box and storage medium

The invention relates to the technical field of sound quality enhancement, in particular to a sound quality improvement method and device of a Bluetooth loudspeaker box, equipment and a storage medium. The method comprises the following steps: collecting sound box environment acoustic signals, and carrying out peripheral audio time-frequency spectrum analysis and three-dimensional sound field fitting to construct a peripheral three-dimensional sound field model; based on the surrounding three-dimensional sound field model, environment noise source semantic deep analysis is carried out, and user real-time scene sound quality demand prediction is carried out, so that real-time scene sound quality demand features are generated; bluetooth sound box loudspeaker monitoring parameters are collected, and adaptive noise frequency suppression adjustment is carried out based on the surrounding three-dimensional sound field model rate, so that a loudspeaker dynamic noise suppression adjustment strategy is generated; and acquiring a Bluetooth audio output signal, and performing band frequency response characteristic optimization according to the real-time scene tone quality demand characteristics to obtain a frequency response tone quality optimized audio signal. According to changes of different environments and requirements, the output tone quality of the Bluetooth loudspeaker box is optimized in real time, and high-fidelity audio output is kept.
Owner:SHENZHEN HAILINGWEI ELECTRONICS CO LTD

Method for monitoring a rotating machine in order to detect a fault in an aircraft bearing

A method for monitoring a rotating machine in order to detect a fault in a bearing, the method including acquiring, from the rotating machine, a vibration signal measured by a vibration sensor; determining a first-order spectrogram by first-order cyclostationary analysis of the vibration signal using a delta transform and spectral standardisation; determining a second-order spectrogram by second-order cyclostationary analysis of the vibration signal using averaged cyclic coherence, a delta transform and spectral standardisation; and detecting a vibration signature of the fault in the bearing on the basis of the first-order spectrogram and the second-order spectrogram.
Owner:SAFRAN SA +1

Multi-mode emotion recognition method, system, electronic device and storage medium

Disclosed are a multi-mode emotion recognition method, a system, an electronic device, and a storage medium. The method includes obtaining a spectrogram of a voice to be recognized and a corresponding text and inputting the spectrogram and the text into a multi-mode emotion recognition model to obtain an emotion recognition result output by the multi-mode emotion recognition model. The multi-mode emotion recognition model is trained based on a sample spectrogram, and a corresponding sample text, and a sample emotion recognition result, and is configured to extract a feature from the spectrogram and the text by a self-attention mechanism to obtain the voice features and the text feature, fuse the text feature and voice feature to obtain a multi-mode fusion feature, and make an emotion classification decision to obtain an emotion recognition result based on the text feature, the voice feature, and the multi-mode fusion feature.
Owner:HUAZHONG NORMAL UNIV

Vocal music training vowel pronunciation quality evaluation method based on auditory and visual spatio-temporal feature fusion

The invention provides a vocal music training vowel pronunciation quality evaluation method based on auditory and visual spatial-temporal feature fusion, and the method comprises the steps: collecting vowel pronunciation audio signals and corresponding videos of a singer, and constructing a multi-modal data set; generating a fractional order Mel spectrogram for the audio signal through short-time fractional order Fourier transform of an adaptive order; extracting time sequence features and spatial features of the fractional order Mel spectrogram, and fusing the time sequence features and the spatial features through a gating mechanism to generate audio spatio-temporal features; face visual features in the video are extracted and fused with the audio spatio-temporal features through a cross attention mechanism, and the cross attention mechanism is integrated with a periodic modeling network; the fused features are input into a classifier, a dynamic weight multi-mode cosine loss function training model is adopted, the dynamic weight multi-mode cosine loss function dynamically adjusts the sample weight through a confusion matrix, and the weight is increased for the samples with classification errors based on the historical frequency mistaken division times of the samples; and outputting a pronunciation quality evaluation result.
Owner:FUZHOU UNIV

Lightweight YOLOv8n model partial discharge type detection and classification method based on PRPD spectrogram

The invention discloses a light-weight YOLOv8n model partial discharge type detection and classification method based on a PRPD spectrogram, and aims to solve the problems of low partial discharge type detection precision, high model complexity, large calculation amount, poor real-time performance and the like in the prior art. According to the method, PRPD spectrograms of four typical partial discharge types of tip, air gap, suspension and surface are collected, normalization, gray processing, data enhancement and expansion and other preprocessing are carried out, a LabelImg platform is utilized to carry out labeling, then a YOLOv8n model is subjected to lightweight improvement, ShuffleNet-V2 is adopted as a trunk feature extraction network, a CoordAttention mechanism is introduced to enhance the feature extraction capability, and the feature extraction efficiency is improved. An EIOU loss function is applied to optimize target frame regression precision; the partial discharge type detection and classification precision is improved, the model calculation complexity and parameter quantity are reduced, real-time detection is realized, and the method has important practical application value.
Owner:CHINA THREE GORGES UNIV

GIS ultrahigh frequency partial discharge signal separation method based on pattern recognition

The invention relates to the technical field of pattern recognition, in particular to a GIS ultrahigh frequency partial discharge signal separation method based on pattern recognition, which comprises the following steps: synchronously acquiring original data by using various sensors, and preliminarily eliminating interference data to obtain data to be analyzed; generating a first characteristic spectrogram and a second characteristic spectrogram based on the to-be-analyzed data, and extracting spectrogram features; according to the spectrogram characteristics, judging judgment conditions which are met by the to-be-analyzed data, wherein the judgment conditions comprise a first mode, a second mode and a third mode; if the to-be-analyzed data meets the corresponding discrimination condition, classifying the to-be-analyzed data through a corresponding data classification path; and fusing classification results obtained through the data classification paths based on confidence, outputting a classification result of the to-be-analyzed data, and separating identified real data from interference data.
Owner:JIANGSU LIDE INTELLIGENT MONITORING TECH CO LTD

Rapid judgment method for spatial spill electromagnetic signal

The invention, which belongs to the technical field of electromagnetic signal monitoring, relates to a rapid judgment method for a spatial spill electromagnetic signal, and the method specifically comprises the following steps: S1, receiving a spill electromagnetic signal in a space; s2, dividing the monitoring total frequency band into a plurality of sub-frequency bands with the same width, rapidly analyzing the obtained electromagnetic signal frequency spectrum in real time based on multi-scale spectrum entropy analysis, and positioning the frequency band where the characteristic signal is located; and S3, converting the frequency spectrum of the frequency band where the characteristic signal is located into a time domain signal, and simultaneously displaying a spectrogram and a time domain waveform. Through spectrum segmentation, multi-scale spectrum entropy analysis, adaptive threshold judgment and frequency domain-time domain combined display, the problems of poor real-time performance, weak anti-noise capability and single dimension in the prior art are solved.
Owner:ZHONGBEI UNIV

Non-contact pipeline fluid flow velocity measurement method based on frequency wavenumber domain spatial spectrogram

The invention relates to the technical field of pipeline detection, and provides a non-contact pipeline fluid flow velocity measurement method based on a frequency wavenumber domain spatial spectrogram, and the method comprises the steps: collecting a turbulence signal in a pipeline through a piezoelectric film sensor; carrying out snapshot number segmentation on the turbulence signal to obtain a plurality of segmented turbulence time domain signals; fourier transform and narrowband signal processing are carried out on the segmented turbulence time domain signals, and a spatial spectrum corresponding to each narrowband signal component is obtained; constructing a three-dimensional frequency-wave number domain spectrogram according to the spatial spectrum function; and according to the three-dimensional frequency-wavenumber domain spectrogram, calculating the flow velocity of the pipeline fluid. According to the method, a three-dimensional spectrogram characteristic space is constructed through conjoint analysis of the frequency and the wave number, the angle, distance and frequency characteristics of the turbulence signals are decoupled through spatial spectrum estimation methods such as the MUSIC algorithm, multipath interference and noise are effectively restrained, the flow speed of fluid in a pipeline can be accurately estimated, and limitation in a traditional method is overcome.
Owner:NAT ENG RES CENT OF DREDGING TECH & EQUIP

Multi-source fusion spectrum cross-domain high-precision prediction method and system based on transfer learning

The invention provides a multi-source fusion spectrum cross-domain high-precision prediction method and system based on transfer learning. The method comprises the following steps: S1, acquiring spectrograms collected by a spectral imaging device under multiple wavelengths; s2, extracting various types of features from the spectrogram; s3, training a fusion model to obtain a trained fusion model; and S4, inputting a to-be-predicted spectrum into the trained fusion model, and outputting a result. According to the invention, 10 <-4 > nm-level wavelength prediction can be achieved in a common industrial camera system, and the level of a scientific research-level grating spectrometer (0.001-0.002 nm) can be reached / exceeded; on a narrow-band, middle-band and wide-band multi-scene cross-domain data set, the mean absolute error (MAE) can be reduced from the nanoscale to the magnitude of 10 <-4 > nm, and excellent robustness is kept under the conditions of noise, feature deficiency and few samples.
Owner:HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Two-process error correction method and device for real-time speech transcription

The invention provides a two-process error correction method and device for real-time speech transcription, and relates to the technical field of speech processing, and the method comprises the steps: extracting the Mel spectrum features of each segment, inputting each Mel spectrum feature into a lightweight end-to-end model, and obtaining a preliminary transcription text; splicing the segments according to a preset number to obtain a plurality of long segments, and inputting each long segment into a speech recognition model to obtain a high-precision transcription text; performing text comparison on the preliminary transcription text and the high-precision transcription text according to the confidence degree set of the preliminary transcription text to obtain all error vocabularies in the preliminary transcription text; and performing corresponding error correction processing on each error vocabulary in the preliminary transcription text according to the type of the error vocabulary and the high-precision transcription text to obtain a final transcription text. According to the method, through a two-process transcription error correction mechanism of the preliminary transcription text and the high-precision transcription text, transcription error accumulation is reduced on the premise that the real-time performance is not affected, and the transcription accuracy in a complex scene is improved.
Owner:NANJING DOLPHIN INTELLIGENT TECH CO LTD

Laying hen voice recognition method and system fusing acoustic features and deep learning features

The invention provides a laying hen voice recognition method and system fusing acoustic features and deep learning features. The method comprises the steps of obtaining a to-be-recognized original audio signal and a voice recognition model; wherein the voice recognition model comprises a feature extraction network, a feature fusion network and a classification recognition network; performing feature extraction on the original audio signal by using the feature extraction network to obtain a spectrogram feature, a Mel-frequency cepstrum coefficient feature and a deep speech feature; the feature fusion network performs feature fusion on the spectrogram features, the Mel-frequency cepstrum coefficient features and the deep speech features by using a collaborative attention mechanism or a multi-head attention mechanism to obtain fused features; and inputting the fused features into a classification recognition network to obtain a voice recognition result. According to the method, the advantages of various characteristics can be fully utilized, and the sound signals are described and analyzed from multiple angles, so that the voiceprint of the laying hen is more accurately recognized, and the voiceprint recognition accuracy of the laying hen is remarkably improved.
Owner:BEIJING RES CENT FOR INFORMATION TECH & AGRI

AI-based material performance spectrum detection system and method

The invention discloses a material performance spectrum detection system and method based on AI, and relates to the field of artificial intelligence, and the system comprises a spectrogram embedding module, a feature fusion module, a performance prediction module, an uncertainty evaluation module and an output analysis module. According to the method, multi-source information such as spectral data, microstructure images and material composition is fused, material performance characteristics are modeled through a multi-modal attention mechanism, and the integrity of characteristic expression and the accuracy of a prediction model are improved. A heterogeneous neural network structure is constructed, effective representation of high-dimensional spectrogram data is realized, and the feature learning ability is enhanced through a residual fusion mechanism. Through an attention hot area map, an attention distribution map and an LIME causal analysis method, key feature interpretation of each prediction result is provided, and scientificity and reliability of a prediction model are enhanced.
Owner:CHANGSHU INSTITUTE OF TECHNOLOGY

Method and system for converting personalized text into voice, and related equipment

The invention provides a personalized text-to-speech conversion method and system and related equipment, and the method comprises the steps: training a deep learning model through a text / audio corpus of a non-standard speaker, and obtaining a non-standard text-to-speech conversion model; obtaining a single speaker reference audio with customized timbre and a to-be-converted target text; inputting the target text into the non-standard text-to-speech model to obtain a sound spectrum representation of the target text; extracting a timbre embedding vector of a target speaker from the single speaker reference audio by using a voiceprint encoder; and fusing the sound spectrum representation and the timbre embedding vector, and inputting the fused sound spectrum representation and timbre embedding vector into a neural vocoder to obtain a personalized voice waveform. According to the method provided by the invention, the synthesis of the tone migration personalized audio can be realized only through the non-standard language data of the single speaker, the scheme realization difficulty is reduced, and the user demand can be better met.
Owner:SHANGHAI JITU SCI & TECH CO LTD

Method and device for determining obstacle avoidance mechanism of unmanned aerial vehicle and electronic equipment

The invention relates to the technical field of unmanned aerial vehicles, in particular to a method and device for determining an obstacle avoidance mechanism of an unmanned aerial vehicle and electronic equipment. The method comprises the following steps: acquiring a radio signal when the unmanned aerial vehicle flies; the method comprises the following steps: converting a radio signal into a spectrogram through short-time Fourier transform, extracting a time sequence feature representing a time dimension sequence of the spectrogram through a long-short-term memory network, extracting a frequency spectrum feature of the spectrogram through a convolutional neural network, and carrying out weighted fusion on the time sequence feature and the frequency spectrum feature according to a preset weight to form a fusion feature; and inputting the fusion feature to a detection model to obtain an unmanned aerial vehicle model, wherein the detection model is obtained by establishing a mapping relationship between the radio signal and the unmanned aerial vehicle model through machine learning. Based on the model of the unmanned aerial vehicle, determining an obstacle avoidance mechanism of the unmanned aerial vehicle from an obstacle avoidance database, the obstacle avoidance database including obstacle avoidance mechanisms of unmanned aerial vehicles of different models. According to the scheme, the unmanned aerial vehicle obstacle avoidance mechanism can be determined remotely.
Owner:NSFOCUS TECH +1

Deep harmonic finesse: signal separation in wearable systems with limited data

A method for separation of non-stationary quasi-periodic signals when limited data is available is described. The method utilizes prior knowledge of time-frequency patterns in the signals to mask and in-paint spectrograms. In one implementation this is achieved through an application-inspired deep harmonic neural network coupled with an integrated pattern alignment component. The network's structure embeds the implicit harmonic priors within the time-frequency domain, while the pattern-alignment method transforms the sensed signal, ensuring a strong alignment with the network.
Owner:RGT UNIV OF CALIFORNIA