Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

70 results about "Zero-crossing rate" patented technology

The zero-crossing rate is the rate of sign-changes along a signal, i.e., the rate at which the signal changes from positive to zero to negative or from negative to zero to positive. This feature has been used heavily in both speech recognition and music information retrieval, being a key feature to classify percussive sounds. ZCR is defined formally as zcr=1/(T-1)∑ₜ₌₁ᵀ⁻¹𝟙ℝ<₀(sₜsₜ₋₁) where s is a signal of length T and 𝟙ℝ<₀ is an indicator function.

Wind turbine generator voiceprint fault recognition method

The invention provides a wind turbine generator voiceprint fault recognition method, and relates to the technical field of wind turbine generator state monitoring and fault diagnosis, and the method comprises the steps: carrying out the noise reduction of an original audio signal through variational mode decomposition, screening a target mode of which the frequency, energy and kurtosis accord with features, and reconstructing the signal; extracting a Mel frequency cepstrum coefficient and a sensing noise robust coefficient, and generating multi-dimensional voiceprint data in combination with statistical characteristics such as a frequency spectrum gravity center, a spectrum entropy, energy, kurtosis and a zero-crossing rate; constructing a support set based on the prototype network, realizing small sample fault classification by calculating the Euclidean distance between the feature vector and the prototype vector, and outputting a preliminary result; judging whether the voiceprint is abnormal according to a preset threshold value, if so, storing the voiceprint into a dynamic abnormal voiceprint knowledge base; frequently occurring abnormal samples are manually labeled and added into a support set, the prototype network is retrained to update the model, and continuous optimization of the fault recognition capability is achieved.
Owner:CGN (SHANXI) NEW ENERGY INVESTMENT CO LTD

Device working state recognition method based on voiceprint recognition model

The invention discloses an equipment working state recognition method based on a voiceprint recognition model, and relates to the technical field of industrial equipment operation state recognition. The equipment working state recognition method based on the voiceprint recognition model comprises the following steps: collecting operation audio waveform data of target equipment, extracting acoustic representation data containing parameters such as short-time energy, a frequency spectrum centroid, a spectrum flux, MFCC and a zero crossing rate, inputting the acoustic representation data into a pre-trained voiceprint recognition model to extract voiceprint feature representation vectors, and carrying out voiceprint feature representation on the target equipment; according to the method, the audio signal is divided into the frames, the acoustic features such as short-time energy, spectrum centroid, spectrum flux, MFCC and zero crossing rate are extracted, the inter-frame evolution relation is modeled in combination with the bidirectional neural network, and the attention mechanism is introduced to highlight the key frame segment, so that the real-time performance of the audio signal is improved, and the real-time performance of the audio signal is improved. The recognition capability of working conditions such as fuzzy state boundary or unobvious transition is effectively enhanced, and the time sequence analysis and state judgment precision is improved.
Owner:FUJIAN RUIXIN TECH CO LTD

Non-specific person voice recognition intelligent switch control method and system based on deep learning

The invention relates to the technical field of voice recognition intelligent home control, and discloses a non-specific person voice recognition intelligent switch control method and system based on deep learning. The non-specific person voice recognition intelligent switch control method is applied to intelligent switch control equipment, and specifically comprises the following steps of S101, receiving original audio signals continuously collected in a to-be-controlled environment, and preprocessing the collected original audio signals, and then a starting point and an ending point of an effective voice segment are positioned by adopting endpoint detection based on a double-threshold method and combining the characteristic parameters of the short-time energy and the short-time zero-crossing rate. A multi-layer hidden layer structure with Dropout regularization is adopted in a neural network model, the generalization ability of the model is enhanced, a context sensing mechanism is introduced into a semantic understanding module, a composite instruction containing azimuth information can be intelligently analyzed, crossing from recognition to understanding is achieved, and the method has the advantages of being high in practicability and easy to popularize. The system is ensured to maintain a high recognition rate for voice instructions of different users under different environment conditions.
Owner:AIRBEST (SHENZHEN) TECHNOLOGY CO LTD

Voice interactive chart dynamic generation method and system based on large language model

InactiveCN120496560ASpeech recognitionFrequency spectrumNoise power spectrum
The invention discloses a voice interactive chart dynamic generation method and system based on a large language model, and relates to the technical field of chart generation, and the method comprises the steps: collecting a user voice signal, carrying out the preliminary framing, calculating the short-time energy, and analyzing a mute segment set signal; frequency domain signals are extracted from the set signals, gain is calculated through a filter, noise reduction is carried out on the frequency domain signals, the zero-crossing rate and noise reduction short-time energy are calculated, effective voice frames are screened, peak detection is carried out, the average voice speed is calculated, voice speed normalization is carried out, and signal values of sampling points in the frames are extracted. According to the method, the noise power spectrum density is extracted through FFT on the mute section, a frequency domain model of background noise is effectively established, a follow-up filter can accurately act on an actually existing frequency band interference area, weakening of the voice main signal spectrum is avoided, the Mel filter bank is accessed after the power spectrum is calculated through the FFT after frame windowing, and the noise power spectrum density is calculated through the FFT after frame windowing. Energy can be redistributed on the logarithmic Mel scale according to human ear perception characteristics.
Owner:DONGQU INTELLIGENT TRANSPORTATION INFRASTRUCTURE TECH (JIANGSU) CO LTD

Emotion recognition key data marking and extracting method applied to AI artificial customer service and application

The invention relates to the technical field of emotion recognition, and particularly discloses an emotion recognition key data marking and extracting method applied to AI artificial customer service and application, and the method comprises the steps: detecting a voice signal input by a customer in a segmented manner through presetting multi-stage and multi-dimensional energy and a zero-crossing rate threshold value, recognizing a voiced segment boundary, and extracting a voiced segment voice signal; framing the voiced segment voice according to a preset frame length, extracting feature parameters, and generating a dynamic feature vector sequence; performing subtraction on the dynamic feature vectors to obtain real-time voiceprint expression emotional state transition vectors, and determining a real-time emotional state transition adjustment factor in combination with a keyword cloud link state; fusing the two to determine a real-time emotional state transition value, and identifying an emotional turning point; when identification is detected, a pop-up window, a work order red mark and a voice broadcast early warning are synchronously generated, and emotion types, intensity values and timestamps are included; therefore, an artificial customer service staff can timely and intuitively know the emotion change of the customer, the customer service staff is assisted to quickly adjust a communication strategy, and more targeted service is provided.
Owner:GUANGZHOU EAPHONETECH CO LTD

AMC sensor array spatial feature recognition and leakage point detection method based on deep learning

The invention discloses an AMC sensor array spatial feature recognition and leakage point detection method based on deep learning, and relates to the technical field of sensors, and the method comprises the steps: carrying out the spatial correlation analysis of a preprocessing result, and extracting a spatial feature signal reflecting the spatial distribution characteristics of gas; decomposing the denoised spatial features by using an empirical mode decomposition algorithm to obtain a plurality of intrinsic mode components, and generating a reconstruction signal based on each intrinsic mode component; performing framing processing on the reconstructed signal by using a framing windowing technology, calculating an energy value and a zero-crossing rate of a corresponding frame based on a framing processing result, constructing a leakage sensing model in combination with the extracted spatial feature signal, and generating a leakage probability distribution sequence; and inputting the time sequence window sequence into a self-encoder, and performing gas leakage point detection in combination with a knowledge distillation mechanism and a comparative analysis technology. According to the invention, accurate extraction and de-noising processing of the spatial feature signals are realized, the signal quality is improved, and the detection precision is further improved.
Owner:CHINA APPLIED TECH CO LTD

Landslide displacement double-layer fusion prediction method and model

The invention discloses a landslide displacement double-layer fusion prediction method and model, and the method comprises the steps: carrying out the decomposition of an original displacement time sequence through employing an ICEEMDAN algorithm, and obtaining a plurality of IMF components; performing feature engineering on each IMF component, representing a displacement trend by adopting a trend slope and a window mean value, representing mutation early warning by adopting kurtosis and a frequency spectrum entropy, representing a period rule by adopting a main frequency and a zero-crossing rate, representing system stability by adopting a sample entropy and a standard deviation, and constructing a three-dimensional feature space fusing a time domain and a frequency domain; data standardization is carried out on the extracted features, the interference effect of dimensions on the model is eliminated, and it is ensured that all feature dimensions are within a unified calculation scale range; a CNN-BiLSTM model is constructed for each IMF component; a CPO algorithm is used to optimize the CNN-BiLSTM model; according to the method, the data acquisition difficulty during model training and use can be reduced, and the usability of the model in actual deployment is enhanced; and the prediction precision and the accuracy of landslide displacement prediction are improved.
Owner:CHINA COAL TECH & ENG GRP SHENYANG ENG CO

Cloud terminal dynamic audio processing method based on AI

The invention relates to the technical field of audio processing, discloses an AI-based cloud terminal dynamic audio processing method, and aims to solve the problem that an existing scheme cannot dynamically adapt to scene changes. The scheme mainly comprises the steps of collecting environmental noise data, and extracting environmental perception features after preprocessing, framing and frequency domain transformation; performing framing processing on the played audio data, extracting an MFCC feature, a logarithmic short-time energy feature and a zero-crossing rate feature, and obtaining a scene recognition result through a CNN + LSTM scene classification model; splicing the environment perception features and the scene recognition result into a joint feature vector, and inputting the joint feature vector into an AI model to obtain an adaptive gain curve parameter, a scene adaptive noise suppression parameter and an anti-distortion dynamic range control parameter; processing the played audio to obtain optimized audio data; and updating AI model parameters through a PPO algorithm based on the reward value of the user feedback data. According to the invention, intelligent audio processing dynamically adapting to environments and scenes can be realized, and the audio experience of the cloud terminal is effectively improved.
Owner:四川长虹新网科技有限责任公司

Smart home control method and system based on end-side large model

The invention relates to the technical field of smart home control, and discloses a smart home control method and system based on an end-side large model, and the method comprises the steps: deploying a lightweight end-side large model in a terminal device, collecting the voice information of a target user corresponding to a home control scene, extracting a Mel-frequency cepstral coefficient and a zero-crossing rate of the voice information by using a lightweight end-side large model so as to extract key intention words of the voice information; generating a first executable control instruction of the home control scene; identifying linkage intention features of the key intention words, and generating a second executable control instruction of the home control scene; the method comprises the following steps: acquiring a first executable control instruction and a second executable control instruction, analyzing a user behavior of a target user and an abnormal environment of a home control scene, and generating a target execution control instruction of the home control scene under the condition of the first executable control instruction and the second executable control instruction. And privacy protection, localization, individuation and humanization household intelligent control are realized.
Owner:HARBIN SAISI TECH CO LTD

Water supply network leakage detection method and device based on accelerometer and readable storage medium thereof

The invention provides a water supply pipe network leakage detection method and device based on an accelerometer and a readable storage medium thereof, and belongs to the technical field of pipe network detection. The method comprises the following steps: acquiring an original acceleration signal; the acceleration signal is accurately converted into a displacement vibration signal by adopting double-order integral drift suppression processing so as to overcome low-frequency drift; carrying out time-frequency analysis on the displacement signal, and extracting transverse characteristics representing the physical nature of leakage; the transverse characteristic is characterized based on three dimensions of time continuity, frequency broadband and amplitude stability; a self-adaptive threshold value is adopted to judge the transverse characteristics, and the threshold value is dynamically generated according to real-time background noise characteristics; and finally, multi-dimensional comprehensive judgment can be carried out by combining time domain auxiliary conditions such as a peak factor and a zero-crossing rate, and a leakage result is determined. The problems of integral drift, inaccurate feature extraction and poor robustness are solved, and low-cost, high-precision and low-false-alarm leakage detection is realized.
Owner:HANGZHOU ZHIBIN TECH CO LTD

Backflow cable anti-theft method and system based on power line carrier noise suppression

ActiveCN120877441APower distribution line transmissionUnbalanced current interference reductionInterference (communication)Interference elimination
The invention relates to a backflow cable anti-theft method and system based on power line carrier noise suppression, and belongs to the technical field of power system security, and various noise components in a signal are effectively separated by performing sliding window preprocessing and empirical mode decomposition on the carrier communication signal collected by multiple monitoring nodes; then, intelligent suppression and future trend prediction of noise are realized by adopting a mode of combining classification weighting and LSTM network prediction; through the synergistic effect of a dynamic threshold mechanism and an unscented Kalman filtering algorithm, the system can adaptively adjust the noise suppression intensity, and interference components are eliminated to the maximum extent while the signal integrity is ensured; and finally, based on double criteria of short-time energy and zero-crossing rate characteristics and in combination with a multi-node delay inequality positioning algorithm, accurate identification and accurate positioning of theft events are realized.
Owner:HANGZHOU JUQI INFORMATION TECH CO LTD +2

Emotion recognition key data marking and extraction method and application for ai artificial customer service

The application relates to the technical field of emotion recognition, and particularly discloses an emotion recognition key data marking and extracting method applied to an AI artificial customer service and application, which comprises the following steps: detecting a customer input voice signal by preset multi-level and multi-dimensional energy and zero-crossing rate threshold segmentation, recognizing a voiced segment boundary and extracting a voiced segment voice signal; framing the voiced segment voice according to a preset frame length, extracting a characteristic parameter, and generating a dynamic characteristic vector sequence; obtaining a real-time voiceprint expression emotion state transition vector by subtracting the dynamic characteristic vector, combining a keyword cloud link state to determine a real-time emotion state transition adjustment factor; fusing the two to determine a real-time emotion state transition value, and recognizing an emotion turning point; generating a pop-up window, marking a work order red and voice broadcasting early warning in time synchronization, containing an emotion type, an intensity value and a time stamp; and making the artificial customer service know the customer emotion change in time and intuitively, helping the customer service personnel to quickly adjust a communication strategy and provide more targeted service.
Owner:GUANGZHOU EAPHONETECH CO LTD

A backflow cable anti-theft method and system based on power carrier noise suppression

ActiveCN120877441BPower distribution line transmissionUnbalanced current interference reductionInterference (communication)Carrier signal
The application relates to a backflow cable anti-theft method and system based on power carrier noise suppression, and belongs to the technical field of power system security and protection. Through sliding window preprocessing and empirical mode decomposition of carrier communication signals collected by multiple monitoring nodes, various noise components in the signals are effectively separated. Then, a combination of classification weighting and LSTM network prediction is adopted to realize intelligent noise suppression and future trend prediction. Through the synergistic effect of a dynamic threshold mechanism and an unscented Kalman filtering algorithm, the system can adaptively adjust the noise suppression strength, maximize the elimination of interference components while ensuring signal integrity. Finally, based on the double criteria of short-time energy and zero-crossing rate characteristics, combined with a multi-node time delay difference positioning algorithm, accurate identification and precise positioning of the theft event are realized.
Owner:HANGZHOU JUQI INFORMATION TECH CO LTD +2

Audio track switching method and electronic equipment

The embodiment of the invention discloses an audio track switching method and electronic equipment, and the method comprises the steps: collecting multi-modal data through a detection device when a first audio stream corresponding to a first audio track is played; performing normalization processing on the multi-modal data by using a convolutional neural network model to obtain a multi-modal feature vector, inputting the multi-modal feature vector into a context prediction model, and predicting a target audio track matched with the multi-modal data by the context prediction model; detecting a mute section, a zero-crossing rate point, an energy stable section and a harmonic stable section of the first audio stream to obtain a candidate point set; selecting a target candidate point from the candidate point set; and performing phase alignment on the first audio track and the target audio track, and switching to the target audio track at the target candidate point to play a second audio stream corresponding to the target audio track. Therefore, the target audio track conforming to the actual preference of the user is predicted, the switching opportunity of the target audio track is automatically matched, auditory faults are eliminated, and the audio track switching efficiency is improved.
Owner:VIDAA (NETHERLANDS) INT HLDG LTD

High-voltage circuit breaker operating mechanism voiceprint state detection method and device based on hesitant fuzzy number and medium

The invention relates to a hesitant fuzzy number-based high-voltage circuit breaker operating mechanism voiceprint state detection method and device, and a medium. Extracting features of an original voiceprint data sequence of the high-voltage circuit breaker operating mechanism, wherein the features comprise a Mel-frequency cepstrum coefficient, short-time energy and a zero-crossing rate; forming a symptom feature vector by the symptom parameters, inputting the symptom feature vector into a fuzzy depth residual shrinkage network, constructing a membership function, and outputting a membership value of each state; constructing hesitant fuzzy numbers according to the membership values to form a collective hesitant fuzzy evaluation matrix; determining the evidence weight of each input evaluation model through an optimal-worst method, deriving a state risk weight through a TOPSIS method in combination with a language Z number, and performing weighted fusion on the collective hesitation fuzzy evaluation matrix to obtain a health index; and dividing the health indexes into normal, attention and dangerous health levels by adopting a K-means clustering algorithm. Compared with the prior art, the method has the advantages of high accuracy, high robustness, high flexibility and the like.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Bird species identification method, device and storage medium based on bird calls

ActiveCN116259321Bimprove accuracyOptimize kernel parametersSpeech analysisFeature vectorAlgorithm
This invention discloses a method, device, and storage medium for bird species identification based on bird calls. The method includes: (1) acquiring several segments of bird calls and preprocessing them; (2) filtering the power spectrum of the bird calls using two different filter banks, extracting the coefficients of the filtered signals, and then combining the two sets of coefficients, the short-time energy of the bird calls, and the short-time zero-crossing rate of the bird calls to form the feature vector of the current bird calls; (3) constructing a nonlinear classification model and using a prey optimization method to find the optimal kernel function in the nonlinear classification model; (4) inputting the extracted bird call feature vector into the nonlinear classification model for learning; and (5) extracting the feature vector of the bird call to be identified and inputting the feature vector into the learned nonlinear classification model to identify the bird species. This invention has low complexity and high accuracy.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Distributed acoustic sensing road cavity detection method based on active excitation source

The invention discloses a distributed acoustic sensing road cavity detection method based on an active excitation source. The method is based on distributed acoustic sensing (DAS), and comprises the following steps of: firstly, acquiring a DAS signal sampling data sequence to be processed; preprocessing the data of each DAS channel; calculating a short-time zero-crossing rate to remove channels with abnormal data records; then, signal space-time positioning is carried out through sliding window detection, signal envelopes in a signal area are extracted, and space-time distribution information of signals is obtained; and finally, extracting spatial correlation features and energy distribution features of the signals, and realizing automatic detection of the road cavity based on feature judgment. According to the method, all-weather sensing and high-resolution precision characteristics of the DAS array are fully utilized, accurate estimation of road cavity detection is achieved with small calculation amount, large-range sensing can be achieved, and the method is suitable for real-time engineering application occasions.
Owner:SOUTHEAST UNIV

Radar signal pulse sequence extraction method based on cross-correlation and estimation signal-to-noise ratio

The invention discloses a radar signal pulse sequence extraction method based on cross-correlation and an estimated signal-to-noise ratio, and belongs to the field of radar signal processing. The method comprises the following steps: framing a sampling signal, calculating a product sequence of short-time energy and a zero-crossing rate of each frame, constructing a rectangular wave according to a single pulse length and a frame shift, segmenting a signal sequence according to a cross-correlation result of the rectangular wave and the product sequence, each segment comprising an effective pulse and noise before and after the pulse; then, effective pulses in each segment are accurately extracted, a segment of noise is taken in front of the segment, the estimated signal-to-noise ratio of the pulses is calculated, noise normalization processing is conducted on the whole sequence, and finally the starting point and the end point of the pulses are detected according to the value of the estimated signal-to-noise ratio. The method does not depend on threshold setting, extraction errors caused by improper threshold setting are avoided, the method is small in calculation amount and high in practicability, and radar pulses in multiple scenes can be rapidly and accurately extracted.
Owner:SOUTHEAST UNIV

A load prediction method and system based on big data

The application discloses a load prediction method and system based on big data, and relates to the technical field of power grid load prediction.The method comprises the following steps: based on a least squares support vector machine, pre-processing load historical data; through decomposing the load historical data with reduced complexity, obtaining a zero-crossing rate and sample entropy, and determining the multi-frequency components of the load data; through a preset hybrid algorithm and in combination with the multi-frequency components of the load data, training a multi-factor weighted combination analysis model; based on the multi-factor weighted combination analysis model and according to an improved grey wolf algorithm, determining the weight of the load prediction result of each prediction factor module; and weighting and combining the load prediction results of each prediction factor module to obtain a final load prediction result.The application improves the robustness of the model, dynamically updates the model output result, optimizes the weight proportion among the factor modules, and improves the adaptability of load prediction in a complex environment.
Owner:STATE GRID JIANGSU INTEGRATED ENERGY SERVICE CO LTD

An ai-based cloud terminal dynamic audio processing method

The application relates to the technical field of audio processing, and discloses an AI-based cloud terminal dynamic audio processing method, which aims to solve the problem that an existing scheme cannot dynamically adapt to scene changes, and mainly comprises the following steps: collecting environmental noise data, extracting environmental perception features after preprocessing, framing and frequency domain transformation; performing framing processing on playing audio data, extracting MFCC features, logarithmic short-time energy features and zero-crossing rate features, and obtaining a scene recognition result through a CNN+LSTM scene classification model; splicing the environmental perception features and the scene recognition result into a joint feature vector, inputting the joint feature vector into an AI model to obtain adaptive gain curve parameters, scene adaptive noise suppression parameters and anti-distortion dynamic range control parameters; processing the playing audio to obtain optimized audio data; and updating AI model parameters based on a reward value of user feedback data and through a PPO algorithm. The application can realize intelligent audio processing which dynamically adapts to the environment and the scene, and effectively improves the audio experience of the cloud terminal.
Owner:四川长虹新网科技有限责任公司

Transformer defect detection method and device based on width fuzzy neural network, and medium

The invention relates to a transformer defect detection method and device based on a width fuzzy neural network, and a medium, and the method comprises the steps: calculating a time domain feature short-time average amplitude, a short-time zero-crossing rate and a frequency domain feature maximum frequency spectrum of each frame of vibration signal of a transformer, and forming a vibration data feature vector; calculating the Euclidean distance between the feature vector of the current frame and the feature vector of the previous frame, calculating the feature vector deviation between the feature vector of the current frame and a preset normal state reference feature vector, performing fuzzification processing, and inputting the processed feature vector deviation into a preset fuzzy rule set for reasoning; dynamically determining the number of enhanced nodes currently required by the width fuzzy neural network; inputting the normalized voiceprint map, the vibration map, the vibration data feature vector and the number of enhanced nodes into a width fuzzy neural network; the width fuzzy neural network dynamically adjusts the network structure according to the number of enhanced nodes. Compared with the prior art, the method has the advantages of high robustness, high reliability, good real-time effect and the like.
Owner:SHANGHAI ELECTRIC POWER HIGH PRESSURE IND CO LTD +1

Webpage element control method, device, equipment, storage medium and program product

The application provides a webpage element control method and device, equipment, storage medium and program product, which can be applied to the technical field of artificial intelligence and financial technology. The method comprises the following steps: processing a plurality of candidate audio clips through a pre-trained voice activity detection model to obtain a voice existence probability corresponding to each candidate audio clip, the plurality of candidate audio clips being obtained by screening from a plurality of audio clips based on at least one of short-time energy and zero-crossing rate of each audio clip; screening the plurality of candidate audio clips based on the plurality of voice existence probabilities to obtain a target audio clip; performing voice recognition on the target audio clip to obtain voice recognition text; guiding a pre-trained large model to process the voice recognition text and an element index based on a plurality of first prompt words to obtain an element control instruction, and transmitting the element control instruction to a webpage to control the webpage to perform a webpage operation indicated by the element control instruction, the element index indicating a corresponding relationship between a user intention and a webpage element.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

A future community management system and method based on the Internet of Things

The application discloses a future community management system and method based on Internet of Things, and relates to the field of intelligent management.The system is composed of a plurality of function modules, including: an identity recognition module, which acquires facial images and audio data; facial recognition and age interval judgment are performed according to the facial images; an audio data extraction module, which utilizes an attention mechanism to fuse time-frequency dual-channel features, constructs anti-interference footstep sound features, and obtains footstep data; audio energy values, zero-crossing rates and moving speeds are extracted from the footstep data; footstep intervals are calculated according to the zero-crossing rates; feature vectors are synthesized based on the audio energy values and the age intervals, and are used as inputs of a model, are divided into three groups of sets, and are randomly sampled; a plurality of points in a hyperparameter space are selected for evaluation; a supervised learning model is used to judge the physical state of the personnel entering and exiting; the physical state, the moving speed and the footstep interval are acquired.
Owner:ZHEJIANG YANGFAN SMART CITY TECHNOLOGY CO LTD

Intelligent outbound robot dialogue regulation and control system fused with multi-mode emotion recognition

The invention discloses an intelligent outbound robot dialogue regulation and control system fused with multi-modal emotion recognition, and relates to the technical field of artificial intelligence, the intelligent outbound robot dialogue regulation and control system is composed of an audio sensing and structuring module, a multi-modal emotion feature deconstruction module and a dynamic emotion interaction regulation and control module; the method comprises the following steps: firstly, recording target human voice in real time by using recording equipment, judging a voice activity frame through short-time energy and a zero-crossing rate, and accurately segmenting the voice activity frame into a voice segment sequence; secondly, extracting speech speed and loudness features from each segment of audio, and constructing an emotion processing feature sequence in combination with normalization processing; and finally, based on comparison of the emotional feature sequence and a preset threshold table, the emotional state of the target person is judged in real time, dialogue emotional regulation and control parameters of the outbound robot are automatically adjusted according to a dynamic regulation and control strategy, and intelligent and humanized outbound dialogue regulation and control are achieved.
Owner:BEIJING HAOFENG CHUANGYUAN TECH CO LTD

Power supply business hall voice processing and compliance verification method

The invention discloses a voice processing and compliance verification method for a power supply business hall, and the method comprises the steps: collecting voice signals of a user and a teller through an independent microphone array, and integrating a beam forming unit and an adaptive echo cancellation unit to improve the signal-to-noise ratio of a target voice; based on time delay estimation and an array geometric model, the time difference between each microphone pair is calculated, the azimuth angle of a target sound source is determined, a beam forming algorithm is applied to audio data to suppress noise interference, de-noised audio data is generated, echo processing is carried out, and a voice recognition module calculates short-time energy and a zero-crossing rate for the processed audio data. A dual-channel pure audio signal with the signal-to-noise ratio larger than or equal to 15 dB and the sampling rate of 16 kHz is output; the two-way interaction and compliance control module transmits a structured service instruction to a background system, and the system achieves accurate matching of service requirements, common explanation and multilingual output of technical terms and compliance verification of key operations, and remarkably improves communication accuracy, customer understanding degree and service handling efficiency.
Owner:WUXI POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD

Water area rainfall measurement method based on deep learning and rainfall measurement ball

The invention belongs to the field of environmental monitoring and meteorological observation, and discloses a water area rainfall measurement method and rainfall measurement balls based on deep learning, and the measurement method comprises the following steps: firstly, collecting rainfall audio through a plurality of rainfall measurement balls, and carrying out the preprocessing to obtain a pre-preprocessed rainfall audio signal; secondly, extracting an amplitude envelope diagram, a root-mean-square energy diagram, a short-time zero-crossing rate diagram and a Mel-frequency cepstral coefficient diagram, and performing feature fusion on the four diagrams to obtain fusion features; and inputting the fusion features into an audio rainfall model based on deep learning to obtain the rainfall intensity of each rainfall measurement ball in the region, and finally obtaining the regional rainfall intensity by using a spatial interpolation technology according to the rainfall intensity of each rainfall measurement ball to generate a regional rainfall level distribution diagram. The invention provides a new means for water area rainfall observation, and provides important technical support for flood prevention early warning, water resource management and climate change research.
Owner:NANJING INST OF TECH

Robust audio watermarking method based on adaptive quantization strategy and feature classification

The invention discloses a robust audio watermarking method based on an adaptive quantization strategy and feature classification, and belongs to the technical field of digital audio copyright protection. The method comprises the following steps: framing an audio signal, extracting a logarithmic mean feature (DWT-CLM) of discrete wavelet transform, and combining a zero-crossing rate, a variance and energy to form a frame feature vector; a Sigmoid classifier is used to discriminate frame characteristics, and a fixed or variable quantization step size is adaptively selected to embed watermark bits into approximate components; and during extraction, the watermark is accurately extracted through the same feature analysis and classifier discrimination recovery quantization mode. Experiments show that the algorithm has better inaudible property and robustness, can effectively resist attacks such as MP3 compression, resampling, low-pass filtering and re-recording, and is suitable for digital audio copyright protection scenes.
Owner:XINYANG NORMAL UNIVERSITY

A high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction

ActiveCN122050413BStationary noiseNoise
The application provides a high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction. In response to a release event of a push-to-talk button, a two-state noise dictionary is constructed based on impulsive noise components and non-stationary noise components in background noise in a current high-dynamic environment. In response to a press event of the push-to-talk button, when it is detected that there is a transient region matching the impulsive noise components in the noisy intercom audio signal, online updating of the two-state noise dictionary is triggered to obtain an online updating dictionary. The noisy intercom audio signal is reconstructed based on the online updating dictionary to obtain an initial enhanced voice signal. Envelope reconstruction is performed on the voice segment with abnormal zero-crossing rate in the initial enhanced voice signal to generate a final enhanced voice signal. The technical scheme provided by the application can enhance conversation voice in a high-dynamic environment with non-stationary noise and impulsive noise.
Owner:SHENZHEN AIQISHI INTELLIGENT TECHNOLOGY CO LTD

A detection method for frequency shift and abnormal sound of an electric vehicle pedestrian warning device

PendingCN122245347ASpeech analysisNoiseHarmonic
This invention discloses a method for detecting frequency shift and abnormal noise in pedestrian warning devices for electric vehicles. The method triggers audio by sending a simulated operating condition message and eliminates time delay errors by synchronizing with the underlying hardware clock. The acquired audio is sequentially filtered and scaled, and effective signal segments are extracted using short-time energy and zero-crossing rate. After phase-aligned splitting, the psychoacoustic features, harmonic distortion, and frequency shift response parameters of each subsequence are extracted. Finally, a feature local distance matrix is ​​constructed, and a dynamic programming algorithm is used to calculate the minimum cumulative cost path to output the sound defect type, while simultaneously verifying the linear mapping relationship between frequency shift response and vehicle speed. This invention effectively eliminates time delay and noise interference, achieving high-precision automatic classification of abnormal noise and accurate verification of frequency shift characteristics.
Owner:FANGBO TECH (SHENZHEN) CO LTD