Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

45 results about "Zero-crossing rate" patented technology

The zero-crossing rate is the rate of sign-changes along a signal, i.e., the rate at which the signal changes from positive to zero to negative or from negative to zero to positive. This feature has been used heavily in both speech recognition and music information retrieval, being a key feature to classify percussive sounds. ZCR is defined formally as zcr=1/(T-1)∑ₜ₌₁ᵀ⁻¹𝟙ℝ<₀(sₜsₜ₋₁) where s is a signal of length T and 𝟙ℝ<₀ is an indicator function.

Wind turbine generator voiceprint fault recognition method

The invention provides a wind turbine generator voiceprint fault recognition method, and relates to the technical field of wind turbine generator state monitoring and fault diagnosis, and the method comprises the steps: carrying out the noise reduction of an original audio signal through variational mode decomposition, screening a target mode of which the frequency, energy and kurtosis accord with features, and reconstructing the signal; extracting a Mel frequency cepstrum coefficient and a sensing noise robust coefficient, and generating multi-dimensional voiceprint data in combination with statistical characteristics such as a frequency spectrum gravity center, a spectrum entropy, energy, kurtosis and a zero-crossing rate; constructing a support set based on the prototype network, realizing small sample fault classification by calculating the Euclidean distance between the feature vector and the prototype vector, and outputting a preliminary result; judging whether the voiceprint is abnormal according to a preset threshold value, if so, storing the voiceprint into a dynamic abnormal voiceprint knowledge base; frequently occurring abnormal samples are manually labeled and added into a support set, the prototype network is retrained to update the model, and continuous optimization of the fault recognition capability is achieved.
Owner:CGN (SHANXI) NEW ENERGY INVESTMENT CO LTD

Non-specific person voice recognition intelligent switch control method and system based on deep learning

The invention relates to the technical field of voice recognition intelligent home control, and discloses a non-specific person voice recognition intelligent switch control method and system based on deep learning. The non-specific person voice recognition intelligent switch control method is applied to intelligent switch control equipment, and specifically comprises the following steps of S101, receiving original audio signals continuously collected in a to-be-controlled environment, and preprocessing the collected original audio signals, and then a starting point and an ending point of an effective voice segment are positioned by adopting endpoint detection based on a double-threshold method and combining the characteristic parameters of the short-time energy and the short-time zero-crossing rate. A multi-layer hidden layer structure with Dropout regularization is adopted in a neural network model, the generalization ability of the model is enhanced, a context sensing mechanism is introduced into a semantic understanding module, a composite instruction containing azimuth information can be intelligently analyzed, crossing from recognition to understanding is achieved, and the method has the advantages of being high in practicability and easy to popularize. The system is ensured to maintain a high recognition rate for voice instructions of different users under different environment conditions.
Owner:AIRBEST (SHENZHEN) TECHNOLOGY CO LTD

Landslide displacement double-layer fusion prediction method and model

The invention discloses a landslide displacement double-layer fusion prediction method and model, and the method comprises the steps: carrying out the decomposition of an original displacement time sequence through employing an ICEEMDAN algorithm, and obtaining a plurality of IMF components; performing feature engineering on each IMF component, representing a displacement trend by adopting a trend slope and a window mean value, representing mutation early warning by adopting kurtosis and a frequency spectrum entropy, representing a period rule by adopting a main frequency and a zero-crossing rate, representing system stability by adopting a sample entropy and a standard deviation, and constructing a three-dimensional feature space fusing a time domain and a frequency domain; data standardization is carried out on the extracted features, the interference effect of dimensions on the model is eliminated, and it is ensured that all feature dimensions are within a unified calculation scale range; a CNN-BiLSTM model is constructed for each IMF component; a CPO algorithm is used to optimize the CNN-BiLSTM model; according to the method, the data acquisition difficulty during model training and use can be reduced, and the usability of the model in actual deployment is enhanced; and the prediction precision and the accuracy of landslide displacement prediction are improved.
Owner:CHINA COAL TECH & ENG GRP SHENYANG ENG CO

Cloud terminal dynamic audio processing method based on AI

The invention relates to the technical field of audio processing, discloses an AI-based cloud terminal dynamic audio processing method, and aims to solve the problem that an existing scheme cannot dynamically adapt to scene changes. The scheme mainly comprises the steps of collecting environmental noise data, and extracting environmental perception features after preprocessing, framing and frequency domain transformation; performing framing processing on the played audio data, extracting an MFCC feature, a logarithmic short-time energy feature and a zero-crossing rate feature, and obtaining a scene recognition result through a CNN + LSTM scene classification model; splicing the environment perception features and the scene recognition result into a joint feature vector, and inputting the joint feature vector into an AI model to obtain an adaptive gain curve parameter, a scene adaptive noise suppression parameter and an anti-distortion dynamic range control parameter; processing the played audio to obtain optimized audio data; and updating AI model parameters through a PPO algorithm based on the reward value of the user feedback data. According to the invention, intelligent audio processing dynamically adapting to environments and scenes can be realized, and the audio experience of the cloud terminal is effectively improved.
Owner:四川长虹新网科技有限责任公司

Water supply network leakage detection method and device based on accelerometer and readable storage medium thereof

The invention provides a water supply pipe network leakage detection method and device based on an accelerometer and a readable storage medium thereof, and belongs to the technical field of pipe network detection. The method comprises the following steps: acquiring an original acceleration signal; the acceleration signal is accurately converted into a displacement vibration signal by adopting double-order integral drift suppression processing so as to overcome low-frequency drift; carrying out time-frequency analysis on the displacement signal, and extracting transverse characteristics representing the physical nature of leakage; the transverse characteristic is characterized based on three dimensions of time continuity, frequency broadband and amplitude stability; a self-adaptive threshold value is adopted to judge the transverse characteristics, and the threshold value is dynamically generated according to real-time background noise characteristics; and finally, multi-dimensional comprehensive judgment can be carried out by combining time domain auxiliary conditions such as a peak factor and a zero-crossing rate, and a leakage result is determined. The problems of integral drift, inaccurate feature extraction and poor robustness are solved, and low-cost, high-precision and low-false-alarm leakage detection is realized.
Owner:HANGZHOU ZHIBIN TECH CO LTD

Emotion recognition key data marking and extraction method and application for ai artificial customer service

The application relates to the technical field of emotion recognition, and particularly discloses an emotion recognition key data marking and extracting method applied to an AI artificial customer service and application, which comprises the following steps: detecting a customer input voice signal by preset multi-level and multi-dimensional energy and zero-crossing rate threshold segmentation, recognizing a voiced segment boundary and extracting a voiced segment voice signal; framing the voiced segment voice according to a preset frame length, extracting a characteristic parameter, and generating a dynamic characteristic vector sequence; obtaining a real-time voiceprint expression emotion state transition vector by subtracting the dynamic characteristic vector, combining a keyword cloud link state to determine a real-time emotion state transition adjustment factor; fusing the two to determine a real-time emotion state transition value, and recognizing an emotion turning point; generating a pop-up window, marking a work order red and voice broadcasting early warning in time synchronization, containing an emotion type, an intensity value and a time stamp; and making the artificial customer service know the customer emotion change in time and intuitively, helping the customer service personnel to quickly adjust a communication strategy and provide more targeted service.
Owner:GUANGZHOU EAPHONETECH CO LTD

A backflow cable anti-theft method and system based on power carrier noise suppression

ActiveCN120877441BPower distribution line transmissionUnbalanced current interference reductionInterference (communication)Carrier signal
The application relates to a backflow cable anti-theft method and system based on power carrier noise suppression, and belongs to the technical field of power system security and protection. Through sliding window preprocessing and empirical mode decomposition of carrier communication signals collected by multiple monitoring nodes, various noise components in the signals are effectively separated. Then, a combination of classification weighting and LSTM network prediction is adopted to realize intelligent noise suppression and future trend prediction. Through the synergistic effect of a dynamic threshold mechanism and an unscented Kalman filtering algorithm, the system can adaptively adjust the noise suppression strength, maximize the elimination of interference components while ensuring signal integrity. Finally, based on the double criteria of short-time energy and zero-crossing rate characteristics, combined with a multi-node time delay difference positioning algorithm, accurate identification and precise positioning of the theft event are realized.
Owner:HANGZHOU JUQI INFORMATION TECH CO LTD +2

Bird species identification method, device and storage medium based on bird calls

ActiveCN116259321Bimprove accuracyOptimize kernel parametersSpeech analysisFeature vectorAlgorithm
This invention discloses a method, device, and storage medium for bird species identification based on bird calls. The method includes: (1) acquiring several segments of bird calls and preprocessing them; (2) filtering the power spectrum of the bird calls using two different filter banks, extracting the coefficients of the filtered signals, and then combining the two sets of coefficients, the short-time energy of the bird calls, and the short-time zero-crossing rate of the bird calls to form the feature vector of the current bird calls; (3) constructing a nonlinear classification model and using a prey optimization method to find the optimal kernel function in the nonlinear classification model; (4) inputting the extracted bird call feature vector into the nonlinear classification model for learning; and (5) extracting the feature vector of the bird call to be identified and inputting the feature vector into the learned nonlinear classification model to identify the bird species. This invention has low complexity and high accuracy.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

A load prediction method and system based on big data

The application discloses a load prediction method and system based on big data, and relates to the technical field of power grid load prediction.The method comprises the following steps: based on a least squares support vector machine, pre-processing load historical data; through decomposing the load historical data with reduced complexity, obtaining a zero-crossing rate and sample entropy, and determining the multi-frequency components of the load data; through a preset hybrid algorithm and in combination with the multi-frequency components of the load data, training a multi-factor weighted combination analysis model; based on the multi-factor weighted combination analysis model and according to an improved grey wolf algorithm, determining the weight of the load prediction result of each prediction factor module; and weighting and combining the load prediction results of each prediction factor module to obtain a final load prediction result.The application improves the robustness of the model, dynamically updates the model output result, optimizes the weight proportion among the factor modules, and improves the adaptability of load prediction in a complex environment.
Owner:STATE GRID JIANGSU INTEGRATED ENERGY SERVICE CO LTD

An ai-based cloud terminal dynamic audio processing method

The application relates to the technical field of audio processing, and discloses an AI-based cloud terminal dynamic audio processing method, which aims to solve the problem that an existing scheme cannot dynamically adapt to scene changes, and mainly comprises the following steps: collecting environmental noise data, extracting environmental perception features after preprocessing, framing and frequency domain transformation; performing framing processing on playing audio data, extracting MFCC features, logarithmic short-time energy features and zero-crossing rate features, and obtaining a scene recognition result through a CNN+LSTM scene classification model; splicing the environmental perception features and the scene recognition result into a joint feature vector, inputting the joint feature vector into an AI model to obtain adaptive gain curve parameters, scene adaptive noise suppression parameters and anti-distortion dynamic range control parameters; processing the playing audio to obtain optimized audio data; and updating AI model parameters based on a reward value of user feedback data and through a PPO algorithm. The application can realize intelligent audio processing which dynamically adapts to the environment and the scene, and effectively improves the audio experience of the cloud terminal.
Owner:四川长虹新网科技有限责任公司

Transformer defect detection method and device based on width fuzzy neural network, and medium

The invention relates to a transformer defect detection method and device based on a width fuzzy neural network, and a medium, and the method comprises the steps: calculating a time domain feature short-time average amplitude, a short-time zero-crossing rate and a frequency domain feature maximum frequency spectrum of each frame of vibration signal of a transformer, and forming a vibration data feature vector; calculating the Euclidean distance between the feature vector of the current frame and the feature vector of the previous frame, calculating the feature vector deviation between the feature vector of the current frame and a preset normal state reference feature vector, performing fuzzification processing, and inputting the processed feature vector deviation into a preset fuzzy rule set for reasoning; dynamically determining the number of enhanced nodes currently required by the width fuzzy neural network; inputting the normalized voiceprint map, the vibration map, the vibration data feature vector and the number of enhanced nodes into a width fuzzy neural network; the width fuzzy neural network dynamically adjusts the network structure according to the number of enhanced nodes. Compared with the prior art, the method has the advantages of high robustness, high reliability, good real-time effect and the like.
Owner:SHANGHAI ELECTRIC POWER HIGH PRESSURE IND CO LTD +1

Webpage element control method, device, equipment, storage medium and program product

PendingCN122392511AEngineeringZero-crossing rate
The application provides a webpage element control method and device, equipment, storage medium and program product, which can be applied to the technical field of artificial intelligence and financial technology. The method comprises the following steps: processing a plurality of candidate audio clips through a pre-trained voice activity detection model to obtain a voice existence probability corresponding to each candidate audio clip, the plurality of candidate audio clips being obtained by screening from a plurality of audio clips based on at least one of short-time energy and zero-crossing rate of each audio clip; screening the plurality of candidate audio clips based on the plurality of voice existence probabilities to obtain a target audio clip; performing voice recognition on the target audio clip to obtain voice recognition text; guiding a pre-trained large model to process the voice recognition text and an element index based on a plurality of first prompt words to obtain an element control instruction, and transmitting the element control instruction to a webpage to control the webpage to perform a webpage operation indicated by the element control instruction, the element index indicating a corresponding relationship between a user intention and a webpage element.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

A future community management system and method based on the Internet of Things

The application discloses a future community management system and method based on Internet of Things, and relates to the field of intelligent management.The system is composed of a plurality of function modules, including: an identity recognition module, which acquires facial images and audio data; facial recognition and age interval judgment are performed according to the facial images; an audio data extraction module, which utilizes an attention mechanism to fuse time-frequency dual-channel features, constructs anti-interference footstep sound features, and obtains footstep data; audio energy values, zero-crossing rates and moving speeds are extracted from the footstep data; footstep intervals are calculated according to the zero-crossing rates; feature vectors are synthesized based on the audio energy values and the age intervals, and are used as inputs of a model, are divided into three groups of sets, and are randomly sampled; a plurality of points in a hyperparameter space are selected for evaluation; a supervised learning model is used to judge the physical state of the personnel entering and exiting; the physical state, the moving speed and the footstep interval are acquired.
Owner:ZHEJIANG YANGFAN SMART CITY TECHNOLOGY CO LTD

Intelligent outbound robot dialogue regulation and control system fused with multi-mode emotion recognition

The invention discloses an intelligent outbound robot dialogue regulation and control system fused with multi-modal emotion recognition, and relates to the technical field of artificial intelligence, the intelligent outbound robot dialogue regulation and control system is composed of an audio sensing and structuring module, a multi-modal emotion feature deconstruction module and a dynamic emotion interaction regulation and control module; the method comprises the following steps: firstly, recording target human voice in real time by using recording equipment, judging a voice activity frame through short-time energy and a zero-crossing rate, and accurately segmenting the voice activity frame into a voice segment sequence; secondly, extracting speech speed and loudness features from each segment of audio, and constructing an emotion processing feature sequence in combination with normalization processing; and finally, based on comparison of the emotional feature sequence and a preset threshold table, the emotional state of the target person is judged in real time, dialogue emotional regulation and control parameters of the outbound robot are automatically adjusted according to a dynamic regulation and control strategy, and intelligent and humanized outbound dialogue regulation and control are achieved.
Owner:BEIJING HAOFENG CHUANGYUAN TECH CO LTD

Power supply business hall voice processing and compliance verification method

The invention discloses a voice processing and compliance verification method for a power supply business hall, and the method comprises the steps: collecting voice signals of a user and a teller through an independent microphone array, and integrating a beam forming unit and an adaptive echo cancellation unit to improve the signal-to-noise ratio of a target voice; based on time delay estimation and an array geometric model, the time difference between each microphone pair is calculated, the azimuth angle of a target sound source is determined, a beam forming algorithm is applied to audio data to suppress noise interference, de-noised audio data is generated, echo processing is carried out, and a voice recognition module calculates short-time energy and a zero-crossing rate for the processed audio data. A dual-channel pure audio signal with the signal-to-noise ratio larger than or equal to 15 dB and the sampling rate of 16 kHz is output; the two-way interaction and compliance control module transmits a structured service instruction to a background system, and the system achieves accurate matching of service requirements, common explanation and multilingual output of technical terms and compliance verification of key operations, and remarkably improves communication accuracy, customer understanding degree and service handling efficiency.
Owner:WUXI POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD

Robust audio watermarking method based on adaptive quantization strategy and feature classification

PendingCN121506155ASpeech analysisFeature vectorAudio watermark
The invention discloses a robust audio watermarking method based on an adaptive quantization strategy and feature classification, and belongs to the technical field of digital audio copyright protection. The method comprises the following steps: framing an audio signal, extracting a logarithmic mean feature (DWT-CLM) of discrete wavelet transform, and combining a zero-crossing rate, a variance and energy to form a frame feature vector; a Sigmoid classifier is used to discriminate frame characteristics, and a fixed or variable quantization step size is adaptively selected to embed watermark bits into approximate components; and during extraction, the watermark is accurately extracted through the same feature analysis and classifier discrimination recovery quantization mode. Experiments show that the algorithm has better inaudible property and robustness, can effectively resist attacks such as MP3 compression, resampling, low-pass filtering and re-recording, and is suitable for digital audio copyright protection scenes.
Owner:XINYANG NORMAL UNIVERSITY

A high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction

ActiveCN122050413BStationary noiseNoise
The application provides a high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction. In response to a release event of a push-to-talk button, a two-state noise dictionary is constructed based on impulsive noise components and non-stationary noise components in background noise in a current high-dynamic environment. In response to a press event of the push-to-talk button, when it is detected that there is a transient region matching the impulsive noise components in the noisy intercom audio signal, online updating of the two-state noise dictionary is triggered to obtain an online updating dictionary. The noisy intercom audio signal is reconstructed based on the online updating dictionary to obtain an initial enhanced voice signal. Envelope reconstruction is performed on the voice segment with abnormal zero-crossing rate in the initial enhanced voice signal to generate a final enhanced voice signal. The technical scheme provided by the application can enhance conversation voice in a high-dynamic environment with non-stationary noise and impulsive noise.
Owner:SHENZHEN AIQISHI INTELLIGENT TECHNOLOGY CO LTD

A detection method for frequency shift and abnormal sound of an electric vehicle pedestrian warning device

PendingCN122245347ASpeech analysisNoiseHarmonic
This invention discloses a method for detecting frequency shift and abnormal noise in pedestrian warning devices for electric vehicles. The method triggers audio by sending a simulated operating condition message and eliminates time delay errors by synchronizing with the underlying hardware clock. The acquired audio is sequentially filtered and scaled, and effective signal segments are extracted using short-time energy and zero-crossing rate. After phase-aligned splitting, the psychoacoustic features, harmonic distortion, and frequency shift response parameters of each subsequence are extracted. Finally, a feature local distance matrix is ​​constructed, and a dynamic programming algorithm is used to calculate the minimum cumulative cost path to output the sound defect type, while simultaneously verifying the linear mapping relationship between frequency shift response and vehicle speed. This invention effectively eliminates time delay and noise interference, achieving high-precision automatic classification of abnormal noise and accurate verification of frequency shift characteristics.
Owner:FANGBO TECH (SHENZHEN) CO LTD

Airport noise intelligent monitoring and traceability analysis method and system

The invention provides an airport noise intelligent monitoring and traceability analysis method and system, and relates to the technical field of airport noise monitoring, and the method comprises the steps: obtaining airport noise data; extracting a sound pressure level change curve, an energy envelope and a zero-crossing rate to generate time domain features, extracting spectrum distribution, a harmonic structure and a frequency band energy ratio to generate frequency domain features, extracting dynamic spectrum evolution information to generate time-frequency joint features, and summarizing to generate noise features; constructing a noise source intelligent recognition model, extracting local mode features through a convolutional neural network, capturing global dynamic features through a long-short term memory network, and outputting a recognition type; an improved time difference of arrival algorithm is adopted, an identification position is determined according to the time difference of noise arriving at a monitoring point, a propagation physical model is constructed to determine propagation sub-paths and summarize the propagation sub-paths to generate a propagation path set, and finally a spatial-temporal distribution database is output and a noise distribution thermodynamic diagram is generated. Therefore, the coverage range, the recognition precision and the analysis depth of airport noise monitoring are remarkably improved.
Owner:NANJING LUKOU INT AIRPORT AIRPORT TECH CO LTD

Continuous speech recognition method, system and terminal based on large language model

ActiveCN121483230ASpeech recognitionFeature extractionZero-crossing rate
The invention discloses a continuous speech recognition method, system and terminal based on a large language model, and the method comprises the steps: obtaining a speech signal, carrying out the framing preprocessing and feature extraction of the speech signal, and obtaining the short-time energy and zero-crossing rate; performing dynamic mute detection according to the short-time energy and the zero-crossing rate to obtain a trigger audio stream; determining an end-to-end speech recognition model, and performing text transcription on the trigger audio stream through the end-to-end speech recognition model to obtain an original text sequence; determining a large language model, and performing semantic error correction and context optimization on the original text sequence through the large language model to obtain an identified text sequence; and reprocessing the recognition text sequence by adopting a cache and paragraph fusion mechanism to obtain a final speech recognition result. According to the invention, the end-to-end speech recognition model and the large language model are deeply fused, so that the recognition precision and semantic coherence of the speech recognition result are effectively improved.
Owner:GALAXY WENJIE (CHANGCHUN) DIGITAL TECHNOLOGY CO LTD

Method for identifying spatial features of amc sensor array and detecting leakage point based on deep learning

The application discloses an AMC sensor array space feature recognition and leakage point detection method based on deep learning, and relates to the technical field of sensors.The method comprises the following steps: performing space correlation analysis on a pretreatment result to extract a space feature signal reflecting the space distribution characteristics of a gas; using an empirical mode decomposition algorithm to decompose the denoised space feature to obtain a plurality of intrinsic mode components, and generating a reconstruction signal based on the intrinsic mode components; using a frame windowing technology to perform frame processing on the reconstruction signal, calculating the energy value and zero-crossing rate of the corresponding frame based on the frame processing result, combining the extracted space feature signal, constructing a leakage perception model, and generating a leakage probability distribution sequence; and inputting a time sequence window sequence into a self-encoder, combining a knowledge distillation mechanism and a comparative analysis technology, and performing gas leakage point detection.The application realizes accurate extraction and denoising processing of the space feature signal, improves the signal quality, and further improves the detection accuracy.
Owner:CHINA APPLIED TECH CO LTD

Water meter operation state online monitoring method and system based on internet of things

PendingCN122282069AMicrocontrollerAccelerometer
This invention relates to the field of water meter status monitoring technology, specifically an online monitoring method and system for water meter operation status based on the Internet of Things (IoT). The method includes: when the water meter microcontroller unit is in sleep mode, an accelerometer collects vibration signals using a dual-threshold hysteresis comparison method. When the ratio of short-time energy to long-time energy exceeds an energy ratio threshold and the duration exceeds a set duration, a first-level wake-up signal is output to wake up the coprocessor. The coprocessor performs first-order differential processing on the vibration signal, calculates the zero-crossing rate and peak-to-average power ratio (PAPR) of the differential sequence, and uses the weighted sum of the zero-crossing rate and PAPR as a water flow characteristic factor. This invention addresses false alarm anomalies by using an exponentially weighted moving average mechanism to smoothly correct the wake-up threshold and judgment boundary. This allows the system to continuously improve its identification benchmark based on historical operating data during long-term service, enhancing the long-term reliability of water meter operation status monitoring and the overall endurance of the equipment.
Owner:HENAN XIDAO INSTR R & D CO LTD

A landslide displacement double-layer fusion prediction method and model

The application discloses a landslide displacement double-layer fusion prediction method and model, uses an ICEEMDAN algorithm to decompose an original displacement time sequence, obtains a plurality of IMF components, carries out feature engineering on the IMF components, adopts a trend slope and a window mean value to represent a displacement trend, adopts kurtosis and spectral entropy to represent a mutation early warning, adopts a main frequency and a zero-crossing rate to represent a periodical law, adopts sample entropy and a standard deviation to represent system stability, constructs a three-dimensional feature space fusing time domain and frequency domain, carries out data standardization on the extracted features, eliminates the interference effect of dimensions on the model, and ensures that all feature dimensions are in a unified calculation scale range, constructs a CNN-BiLSTM model for each IMF component, and uses a CPO algorithm to optimize the CNN-BiLSTM model, so that the data acquisition difficulty during model training and use can be reduced, the usability of the model in actual deployment can be enhanced, the prediction precision is improved, and the accuracy of landslide displacement prediction is improved.
Owner:CHINA COAL TECH & ENG GRP SHENYANG ENG CO

Long-distance Internet of Things intercom system and method based on low-power wide area network

The invention relates to the technical field of internet of things talkback, in particular to a long-distance internet of things talkback system and method based on a low-power wide area network. The system comprises a node creation module, a detection distribution module and a conflict adjustment module. The node weight is calculated through the node creating module and updated in real time, reasonable election and dynamic adaptation of the main node and the standby main node are completed, the detection distribution module improves the accuracy of voice state and silence state judgment based on short-time energy and zero-crossing rate double-time-domain features, and the accuracy of voice state and silence state judgment is improved. Meanwhile, through on-demand time slot scheduling and dynamic super-frame capacity expansion based on the time slot resource pressure ratio, the time slot resource utilization rate is maximized, and a conflict adjustment module adopts weight priority arbitration and hash randomization to solve time slot request conflicts between nodes; the communication conflict rate is reduced and the communication reliability is improved by combining the double judgment of continuous heartbeat missing and time stamp timeout and the quick take-over and whole group state synchronization of the standby nodes.
Owner:NANJING WANGTAI COMMUNICATION TECHNOLOGY CO LTD

Real-time mute detection method based on zero-crossing rate and energy value optimization

PendingCN121528234ASpeech analysisNoiseZero-crossing rate
The invention relates to the technical field of audio signal processing, and discloses a real-time mute detection method based on zero-crossing rate and energy value optimization, which comprises the following steps: calculating an energy value and a noise stability value of each frame of audio signal, classifying noise, and dividing noise intensity grades; performing framing processing on the input signal, and calculating a current frame signal peak value; setting a zero-crossing amplitude threshold value to obtain an optimized zero-crossing rate; generating a self-adaptive energy value according to the dynamic threshold value; constructing a feature pair sequence, calculating a correlation coefficient of the feature pair sequence and judging a feature relationship; dynamically calculating a zero-crossing rate threshold value; judging a mute candidate state and a non-mute candidate state according to the dynamic threshold value and the zero-crossing rate threshold value; and outputting a judgment result according to the states of two continuous frames. According to the method, the zero-crossing rate calculation mode is dynamically adjusted, the self-adaptive energy threshold value is designed, the double-feature collaborative decision logic is constructed, and high-precision and low-delay silence detection in the complex noise environment is achieved.
Owner:AEROSPACE XINTONG TECH CO LTD

Voice instruction dynamic recognition and separation method based on multi-person voice scene

ActiveCN121565172ASpeech recognitionFundamental frequencyZero-crossing rate
The invention discloses a voice instruction dynamic recognition and separation method based on a multi-person voice scene, relates to the technical field of voice instruction recognition, and is used for solving the problem of low multi-task voice control efficiency. The method comprises the following steps: extracting frame energy and a zero-crossing rate to identify a human voice segment, collecting a fundamental frequency from the human voice segment and calculating acoustic deviation characteristics, creating an instruction sending target according to acoustic deviation, detecting the sounding duration of each target to obtain an effective instruction, calling the sending time of the effective instruction, and calculating overlapping time. And selecting to trigger a parallel execution or overlapping sorting mechanism based on the overlapping time, in the overlapping sorting mechanism, obtaining the operation state and the audio amplitude of the corresponding cooking equipment, setting instruction priorities in combination with historical registration information, and performing one-by-one execution on an execution sequence generated by effective instruction sorting. And the response speed and accuracy of kitchen multi-task voice instruction processing are improved.
Owner:STELLA IND CO LTD

Voice response data processing method and device, equipment, storage medium and product

The invention provides a voice response data processing method and device, equipment, a storage medium and a product. Comprising the steps that in the voice interaction process of a user, voice response data generated for the user is acquired; acquiring a voice signal of the voice response data; performing framing processing on the voice signal to obtain multiple frames of voice signals; performing feature extraction on each frame of voice signal to obtain an energy feature and a zero-crossing rate feature; recognizing a continuous mute frame sequence from the multi-frame voice signal according to the energy feature and the zero-crossing rate feature; performing mute truncation on the voice signal according to the continuous mute frame sequence to obtain optimized voice response data; and outputting the optimized voice response data. The resource waste is reduced; and the user experience is improved.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Speech instruction dynamic recognition and separation method based on multi-person speech scene

ActiveCN121565172BFundamental frequencyZero-crossing rate
The application discloses a voice instruction dynamic recognition and separation method based on a multi-person voice scene, relates to the technical field of voice instruction recognition, and is used for solving the problem of low multi-task voice control efficiency. After a kitchen voice control device is started, ambient audio streams are listened to and divided into audio frames, frame energy and zero-crossing rate are extracted to recognize human voice segments, the fundamental frequency of the human voice segments is collected, and acoustic deviation features are calculated, instruction issuing targets are created according to the acoustic deviation, the sound duration of each target is detected to obtain effective instructions, the issuing time of the effective instructions is called, overlapping time is calculated, and parallel execution or overlapping sorting mechanism is selected based on the overlapping time. In the overlapping sorting mechanism, the operation state of the corresponding cooking device and the audio amplitude are obtained, the instruction priority level is set in combination with historical registration information, the effective instructions are sorted to generate an execution sequence, and the execution sequence is executed in sequence one by one, so that the response speed and accuracy of kitchen multi-task voice instruction processing are improved.
Owner:STELLA IND CO LTD

Electric energy quality transaction monitoring method and system based on window function optimization multi-dimensional characteristics

The invention discloses an electric energy quality transaction monitoring method and system based on window function optimization multi-dimensional characteristics, and the method comprises the steps: segmenting obtained historical electric energy quality signal data, carrying out the weighting processing through employing a Bartlett triangular window function, calculating the five statistical characteristics of a data segment, i.e., the variance, skewness, kurtosis, a normalized energy characteristic, and a zero-crossing rate characteristic, and carrying out the calculation of the variance, skewness, kurtosis, normalized energy characteristic, and zero-crossing rate characteristic of the data segment; combining the five statistics to construct five pairs of two-dimensional feature planes, and stacking the five pairs of two-dimensional features into a feature matrix; a power quality transaction depth monitoring model is constructed and trained, and power quality transaction monitoring is performed by using the trained model; according to the method, high-precision automatic identification of seven typical disturbances such as normal disturbance, voltage sag disturbance, voltage rise disturbance, harmonic disturbance, flicker disturbance, voltage interruption disturbance and transient oscillation disturbance is realized, and timely and accurate disturbance alarm and decision support are provided for operation and maintenance personnel of a power system.
Owner:TIANJIN UNIV +2