Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4856 results about "Audio signal" patented technology

An audio signal is a representation of sound, typically using a level of electrical voltage for analog signals, and a series of binary numbers for digital signals. Audio signals have frequencies in the audio frequency range of roughly 20 to 20,000 Hz, which corresponds to the lower and upper limits of human hearing. Audio signals may be synthesized directly, or may originate at a transducer such as a microphone, musical instrument pickup, phonograph cartridge, or tape head. Loudspeakers or headphones convert an electrical audio signal back into sound.

Headset antenna and connector for the same

ActiveUS20090033574A1Increase the equivalent impedanceImprove rendering capabilitiesAntenna supports/mountingsAntenna adaptation in movable bodiesHeadphonesAudio frequency
A headset antenna and a connector for the same are provided. The headset antenna includes an audio signal line, an antenna and a high impedance element in specified application frequency ranges. The audio signal line is adapted for transmitting an audio signal and the antenna is adapted for receiving an RF signal. The high impedance element is disposed on a transmission path of the audio signal and generates a high impedance at a specified frequency band of the RF signal, so that the audio signal line is equivalent to an open circuit and the antenna obtains a better receiving capability.
Owner:HTC CORP

Compensation for nonuniform delayed group communications

A method for synchronizing audio reproduction in collocated end devices is presented. Each of the devices auto-correlates using noise to determine a threshold prior to the antenna receiving an audio signal. When the devices receive a common audio signal, they provide audio outputs. Each device cross-correlates its audio output with the audio outputs of the other devices. The timing of the audio output of each device is then adjusted such that the audio outputs of all of the devices align temporally with the lagging or leading device.
Owner:MOTOROLA SOLUTIONS INC

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Wind turbine generator voiceprint fault recognition method

The invention provides a wind turbine generator voiceprint fault recognition method, and relates to the technical field of wind turbine generator state monitoring and fault diagnosis, and the method comprises the steps: carrying out the noise reduction of an original audio signal through variational mode decomposition, screening a target mode of which the frequency, energy and kurtosis accord with features, and reconstructing the signal; extracting a Mel frequency cepstrum coefficient and a sensing noise robust coefficient, and generating multi-dimensional voiceprint data in combination with statistical characteristics such as a frequency spectrum gravity center, a spectrum entropy, energy, kurtosis and a zero-crossing rate; constructing a support set based on the prototype network, realizing small sample fault classification by calculating the Euclidean distance between the feature vector and the prototype vector, and outputting a preliminary result; judging whether the voiceprint is abnormal according to a preset threshold value, if so, storing the voiceprint into a dynamic abnormal voiceprint knowledge base; frequently occurring abnormal samples are manually labeled and added into a support set, the prototype network is retrained to update the model, and continuous optimization of the fault recognition capability is achieved.
Owner:CGN (SHANXI) NEW ENERGY INVESTMENT CO LTD

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Physiological monitoring soundbar

A soundbar for medical monitoring which may comprise a speaker, a sensor, and a hardware processor. The speaker can be configured to emit audio signals. The sensor can be configured to obtain sensor data relating to a physiology of a subject. The sensor can include a camera and the sensor data can include image data. The hardware processor can be configured to access the sensor data and determine a health status of the subject based on at least the sensor data.
Owner:MASIMO CORP

Bird identification method and device based on sound-image multi-modal fusion

The invention discloses a bird identification method based on sound-image multi-modal fusion. The bird identification method comprises the following steps: S1, carrying out standardized frame-level preprocessing on bird audio signals; s2, acoustic features are extracted and enhanced, and an acoustic high-level feature vector which highlights birdsong discrimination information and suppresses environmental noise is obtained; s3, visual image standardization preprocessing; s4, performing visual feature extraction and multi-scale fusion to obtain a visual high-level feature vector which enhances correspondence to the bird key form area and inhibits background interference; s5, performing dynamic weighted fusion on the decision-making layer to obtain a bird existence probability; and S6, comparing the bird existence probability with a preset threshold value of the corresponding bird, and judging whether the bird exists or not and the type of the existing bird. Through cross-modal feature enhancement and adaptive fusion, the precision, robustness and real-time performance of bird recognition in a complex orchard environment are significantly improved, and a core technical support is provided for green intelligent bird repelling.
Owner:NANJING FORESTRY UNIV

Automation for inserting a reference accessing a screen-shared file in meeting summaries and transcripts

The disclosed techniques provide a system for automatically inserting reference to a file in meeting transcripts or meeting summaries. In general, the disclosed techniques manage and enrich meeting transcripts, meeting summaries, meeting recordings, and screen-shared files during an online meeting. During an online meeting, when a presenter screenshares an application file, such as Word doc, PowerPoint, Excel, etc., a system creates meeting transcripts or a summary for the real-time discussion based on audio signals and / or chat messages relating to the shared contents of the application file. The system also determines the location of the application file and inserts a reference to the application file in the transcript or summary.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Grid telephone traffic quality inspection intelligent analysis system and method based on large language model

The invention relates to the technical field of intelligent telephone traffic quality inspection, and discloses a grid telephone traffic quality inspection intelligent analysis system and method based on a large language model, and the system comprises the steps: collecting and obtaining a multi-role call audio signal in real time, and preliminarily carrying out the speaker separation and role marking through a voice recognition module and a voiceprint recognition module; forming a preliminary role recognition result; detecting a suspected role identity error region in combination with multi-dimensional features, and when a detection result meets a preset condition, triggering a dynamic correction mechanism, generating auxiliary judgment information in combination with identity declaration keywords, business term matching and dialogue context logic inference, adopting a multi-dimensional weight decision strategy, and re-correcting a role identity tag, so as to obtain a role identity error region. Updating a role recognition result; and setting an accurate evaluation module of a dynamic correction result, feeding back and adjusting a multi-feature weight and a trigger threshold in real time, and forming an iterative optimization mechanism of dynamic correction. The method has the advantage of improving the dynamic correction capability.
Owner:XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER

Processing parametrically coded audio

A method comprising receiving a first input bit stream for a first parametrically coded input audio signal, the first input bit stream including data representing a first input core audio signal and a first set including at least one spatial parameter relating to the first parametrically coded input audio signal. A first covariance matrix of the first parametrically coded audio signal is determined based on the spatial parameter(s) of the first set. A modified set including at least one spatial parameter is determined based on the determined first covariance matrix, wherein the modified set is different from the first set. An output core audio signal is determined, which is based on, or constituted by, the first input core audio signal. An output bit stream for a parametrically coded output audio signal is generated, the output bit stream including data representing the output core audio signal and the modified set.
Owner:DOLBY LABORATORIES LICENSING CORP +1

Earthquake early warning equipment and method based on campus broadcast

The invention discloses an earthquake early warning device and method based on campus broadcast, and the method comprises the steps: carrying out the building response coupling analysis through receiving an early warning signal of an earthquake monitoring center, generating an earthquake magnitude-building coupling feature, and mapping the earthquake magnitude-building coupling feature into an enhanced early warning source adaptive to an acoustic environment; acoustic cavity detection is carried out based on threatened area identification, and differential playing topology is constructed; extracting a high-risk gathering area through personnel density scanning, generating personalized evacuation guiding content and determining a playing priority; high-precision time sequence control is realized by adopting a synchronous reference anchor point technology, and a multi-mode playing domain is constructed through carrier modulation and frequency division multiplexing; and finally, an acoustic navigation field is utilized to guide accurate scheduling and playing of audio signals, differential playing and accurate coverage of early warning information can be realized according to the anti-seismic characteristics and acoustic environment characteristics of different buildings, and an intelligent acoustic solution is provided for campus earthquake early warning.
Owner:FUZHOU BENYANG INFORMATION TECH CO LTD

Atmosphere lamp control method, electronic device and program product

The invention discloses an atmosphere lamp control method, electronic equipment and a program product. The method comprises the following steps: acquiring PCM data of an audio signal in real time in the music playing process of a vehicle; performing frequency domain analysis on the PCM data corresponding to each audio frame, and extracting frequency information and loudness information; the method comprises the following steps: calculating frequency information and loudness information of a preset number of continuous audio frames, and respectively calculating a frequency change rate and a loudness change rate within a preset time; respectively comparing the frequency change rate and the loudness change rate with corresponding change rate thresholds; and if at least one of the frequency change rate and the loudness change rate exceeds the corresponding change rate threshold value, a corresponding light updating instruction is sent to the atmosphere lamp control module. The music rhythm function of the automotive interior atmosphere lamp can be enhanced, and the cooperative interaction ability of music and the atmosphere lamp is improved.
Owner:ANHUI KAIYANG TECHNOLOGY CO LTD +1

Human body gesture generation method and related equipment

The invention discloses a human body gesture generation method and related equipment, and relates to the technical field of computer vision, and the method comprises the steps: obtaining multi-modal input information, the multi-modal input information comprises a voice audio signal, text transcription information and a reference gesture sequence, the text transcription information comprises a semantic annotation, and the reference gesture sequence comprises a reference gesture sequence; the reference gesture sequence is a basic gesture template in the target scene; generating a multi-modal feature based on the multi-modal input information; performing time alignment processing operation on the multi-modal features to obtain alignment condition features; performing space-time decoupling modeling operation on the alignment condition features to obtain optimized gesture potential features; inputting the optimized gesture potential features into a diffusion model for iterative denoising to obtain denoised gesture potential features; and generating a target human body gesture sequence based on the de-noised gesture potential features.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Vehicle NVH comfort assessment method and device based on sound and vibration fusion and electronic equipment

The invention provides a vehicle NVH comfort assessment method and device based on sound and vibration fusion and electronic equipment. The method comprises the following steps: performing in-vehicle interference noise identification processing on window-level data in a candidate window set to obtain a candidate effective window set; calculating a sound and vibration consistency index between the in-vehicle audio signal and the in-vehicle vertical acceleration signal for the window-level data in the candidate effective window set, and screening an external noise dominant window to obtain an effective window set; respectively calculating an acoustic index and a vibration index based on the effective window set, and performing calibration compensation on the acoustic index and the vibration index based on the sensor calibration information; and performing fusion calculation based on the data volume information of the effective window set, the working condition coverage information, the uncertainty information corresponding to the sensor calibration information and the proportion information of the effective window set in the candidate window set to obtain a comprehensive confidence index. According to the method, the evaluation accuracy and stability can be improved, the cross-equipment comparability is enhanced, and the result credibility is improved.
Owner:CAR CONTROL (BEIJING) TECH CO LTD

Voice data recognition method and system based on AI voice algorithm

The invention discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem that the voice data recognition capability is low. The method comprises the following steps: S1, multi-mode cooperative triggering collection: synchronously collecting lip electromyographic signals and voiceprint features through a multi-mode sensor, an activation instruction is generated through feature fusion, and voice acquisition starting is triggered; s2, AI adaptive noise reduction processing: carrying out noise separation on the original audio signal by adopting a generative adversarial network, separating environmental noise features to generate a dynamic noise reduction mask, and keeping the integrity of human voice features; s3, beam dynamic optimization adjustment: analyzing real-time audio quality based on a reinforcement learning algorithm, dynamically adjusting beam pointing and gain parameters of a microphone array, and focusing a target sound source; and S4, semantic association cache enhancement: carrying out real-time semantic analysis on the collected voice data. According to the invention, the voice data recognition capability of an AI voice algorithm is greatly improved.
Owner:HUAQIAO UNIVERSITY

Real-time noise reduction method and system supporting Bluetooth audio interaction

The invention discloses a real-time noise reduction method and system supporting Bluetooth audio interaction, and the method comprises the steps: collecting an environment audio signal and a Bluetooth interaction audio signal, and carrying out the timestamp alignment and spatial calibration of the environment audio signal and the Bluetooth interaction audio signal; an improved multivariate variational mode decomposition algorithm is adopted to decompose and extract noise features, the noise features are combined with user historical noise data, a dynamic noise model is constructed through a long-short-term memory network, parameters are updated in real time, Bluetooth interaction audio content features and a user behavior data recognition scene are analyzed, and a noise reduction strategy is determined. And according to the dynamic noise model and the scene result, reinforcing learning to optimize filter parameters, generating a directional anti-noise signal to carry out noise reduction on the Bluetooth interaction audio signal, collecting the noise reduction intensity manually adjusted by a user, calculating voice definition and total harmonic distortion, and carrying out noise reduction on the Bluetooth interaction audio signal. And inputting a feedback and evaluation result into the modeling and noise reduction steps, carrying out dynamic range adjustment, sound channel balance and power amplification on the audio after noise reduction, and outputting the audio through a loudspeaker.
Owner:SHENZHEN GAOWEI COMM TECH CO LTD

Audio signal generation model and training method using generative adversarial network

A generative adversarial network-based audio signal generation model for generating a high quality audio signal may comprise: a generator generating an audio signal with an external input; a harmonic-percussive separation model separating the generated audio signal into a harmonic component signal and a percussive component signal; and at least one discriminator evaluating whether each of the harmonic component signal and the percussive component signal is real or fake.
Owner:ELECTRONICS & TELECOMM RES INST +1

Multi-channel audio calibration method, system and device and storage medium

The invention discloses a multi-channel audio calibration method, system and device and a storage medium, and relates to the technical field of audio, and the method comprises the steps: collecting and separating a main sound channel signal and a bass sound channel signal in a multi-channel audio signal; detecting time delay between a main sound channel and a bass sound channel based on a phase correlation algorithm, and performing adaptive delay compensation; the state change of the audio signal is monitored in real time, and before or when a signal switching event is detected, gradient gain envelope is applied to the bass sound channel signal, so that the signal amplitude gradually changes to a target value according to a preset nonlinear curve; a sonic boom event is analyzed and identified based on multi-dimensional characteristics, a self-adaptive amplitude limiter with a prospective processing capability is started for suppression, and harmonic distortion compensation and phase correction are carried out on a processed signal, so that the problems of signal asynchronization caused by low-frequency output delay of an audio system in the prior art, and the system reliability is improved are solved. Therefore, the low-frequency sound quality and the system stability are improved.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Box body and signal transmission method

The invention provides a box body and a signal transmission method. The box body comprises an upper shell, a lower shell, a first microphone, keys, a rotating shaft and an antenna assembly, the lower shell and the upper shell are installed in a matched mode to form a containing cavity, the containing cavity comprises a first surface, opposite to the upper shell, in the lower shell, the lower shell comprises an annular side face connected with the first surface, and the annular side face comprises a first side face; the second side face and the third side face are connected with the first side face, the second side face and the third side face are oppositely arranged, a first microphone pickup hole is formed in the first side face, the first microphone is arranged in the lower shell, communicated with the first microphone pickup hole and used for collecting a first audio signal outside the box body, and the key is arranged on the second side face and used for collecting a second audio signal outside the box body. The rotating shaft is arranged corresponding to the third side surface and is connected between the lower shell and the upper shell, the upper shell and the lower shell are rotatably connected through the rotating shaft, and the antenna assembly is arranged on the first surface and is used for at least transmitting at least part of the first audio signal to the mobile terminal.
Owner:SHENZHEN NAXIN TECHNOLOGY R&D CO LTD

Emotion recognition method and system based on multi-modal feature retrieval, terminal and storage medium

The invention relates to the technical field of data processing, and discloses an emotion recognition method and system based on multi-modal feature retrieval, a terminal and a storage medium, and the method comprises the steps: collecting a video signal and an audio signal of a subject, converting the audio signal into a text signal, and extracting a video feature, an audio feature and a text feature; retrieving in a single-mode feature retrieval library according to the video features, the audio features and the text features to obtain enhanced features; aligning the enhanced features to a unified feature space through a mapping module, inputting a double-branch structure, dynamically adjusting the weight of the enhanced features through a modal weight distributor to obtain weighted enhanced features, and querying in a multi-modal feature retrieval library according to the enhanced features to obtain multi-modal retrieval features; and fusing the weighted enhanced features and the multi-modal retrieval features, inputting the fused features into a multi-modal large model for processing, and outputting an emotion recognition result of the subject. According to the invention, accurate perception of the individual emotional state is realized.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

Lightweight robust unsupervised feature selection audio denoising method based on algorithm

The invention discloses a lightweight robust unsupervised feature selection audio denoising method based on an algorithm, and the method comprises the steps: collecting a noise-containing audio signal in a real-time scene, and taking the noise-containing audio signal as to-be-denoised audio data; based on the noise frequency data to be denoised, constructing a robust unsupervised feature selection model based on joint subspace learning, and constructing a target optimization function through multi-task complementary learning; adopting an alternating optimization method to iteratively solve a target optimization function step by step; based on an algorithm expansion technology, converting an iterative solution process into a modular neural network structure, and constructing a lightweight unsupervised feature selection network; on the basis of the constructed lightweight unsupervised feature selection network, outputting an optimal feature weight matrix after the to-be-denoised audio data is subjected to iterative computation; evaluating feature importance based on the optimal feature weight matrix, and selecting key audio features corresponding to the target feature subset; and outputting final audio data after feature selection. By adopting the method, efficient denoising and quality improvement of the audio data are realized.
Owner:SHENZHEN UNIV

Emotion recognition method and system based on behavior and physiological modality of orthogonal fusion

The invention discloses a behavior and physiological modal emotion recognition method and system based on orthogonal fusion, and the method comprises the steps: obtaining multi-modal data comprising a video frame sequence, an audio signal and an electroencephalogram signal, and carrying out the preprocessing of the multi-modal data; constructing a preliminary emotion recognition network; training the constructed preliminary emotion recognition network by using the preprocessed multi-modal data to obtain a trained emotion recognition model; and obtaining a to-be-recognized video frame sequence, an audio signal and an electroencephalogram signal, and inputting the to-be-recognized video frame sequence, the audio signal and the electroencephalogram signal into the emotion recognition model to obtain a corresponding emotion recognition result. According to the method, a modularized emotion recognition model is constructed, all functional modules cooperate with one another, redundant information is reduced, complementary information of multiple modes is fused, emotion feature characterization accuracy is improved, and therefore the emotion recognition effect is improved.
Owner:NANJING MEDICAL UNIV

Emotion recognition method and device based on multi-modal consensus and diversity decoupling

The invention relates to an emotion recognition method and device based on multi-modal consensus and diversity decoupling. The method comprises the following steps: firstly, collecting multi-modal input data including language, vision and audio signals and carrying out corresponding preprocessing; then, constructing a multi-modal consensus and diversity decoupling emotion recognition model which comprises a multi-modal decoupling coding module, a prototype-Gram unification module, a feature enhancement module, a diversity classification module and an emotion prediction head; then, inputting the preprocessed multi-modal input data into the multi-modal consensus and diversity decoupling emotion recognition model, and performing model training optimization based on a total loss function formed by emotion prediction task loss, decoupling loss, unified target loss and diversity loss; and finally, inputting the multi-modal data to be recognized into the trained multi-modal consensus and diversity decoupling emotion recognition model, and outputting an emotion recognition result. And the accuracy, robustness and interpretability of the multi-modal emotion recognition system are improved.
Owner:SICHUAN UNIV

Children mouth breathing monitoring method, device and equipment and medium

The invention discloses a child mouth breathing monitoring method, device and equipment and a medium, and the method comprises the steps that a non-contact sensor group is used for synchronously collecting multi-mode physiological data of a target child in the sleep period, and the multi-mode physiological data at least comprises a face thermal imaging video stream, a thoracic and abdominal micro-motion signal and an environment audio signal; the multi-modal physiological data is processed, multi-modal features related to respiration are extracted respectively, and the multi-modal features comprise thermodynamic change features of mouth and nose areas extracted based on a face thermal imaging video stream, respiratory effort waveforms extracted based on thoracic and abdominal micro-motion signals, and the respiratory effort waveforms extracted based on the thoracic and abdominal micro-motion signals; extracting a sound source spatial position and a spectrum feature of breathing sound based on the environment audio signal; inputting the multi-modal features into a pre-trained mouth breathing recognition model, and outputting a judgment result of the functional mouth breathing event of the target child within a specific time period; and generating a mouth breathing monitoring report of the target child based on the judgment result.
Owner:AFFILIATED CHILDRENS HOSPITAL OF CAPITAL INST OF PEDIATRICS

Garden underground pest monitoring method, electronic equipment and storage medium

The invention provides a garden underground pest monitoring method, electronic equipment and a storage medium, and belongs to the technical field of pest detection. The method comprises the following steps: collecting background noise signals and audio signals of an underground pest activity period; denoising the audio signal based on the background noise signal; outputting an edge recognition result and confidence through a lightweight recognition model deployed at an edge computing node; when the confidence coefficient is greater than or equal to a preset threshold value, determining that the edge recognition result is a pest recognition result; when the confidence coefficient is smaller than a preset threshold value, outputting a cloud recognition result as a pest recognition result through a deep learning model deployed in the cloud; and according to the pest identification result, the soil environment data, the pest type weight and a preset diffusion coefficient, obtaining a risk index corresponding to each pest type to generate early warning information. And a garden underground pest monitoring solution which is more real in data, resistant to environmental interference, more economical and efficient in system deployment and has a prospective early warning capability is achieved.
Owner:GUANGDONG HONGJING INTELLIGENT TECH CO LTD

Permanent magnet synchronous host bearing state detection method and system

The invention relates to a permanent magnet synchronous host bearing state detection method and system, and the method comprises the steps: 1, enabling a sharp end of a special measurement rod to directly contact with the surface of a host bearing end cover, and collecting an audio signal and a sound wave signal in the operation of a bearing in real time; secondly, the elevator is made to operate under the three different load working conditions of full load, 50% load and no load, and bearing operation signals under all the working conditions are collected; and step 3, transmitting the collected signals to an intelligent AI analysis platform for preprocessing and feature extraction. According to the invention, the intelligent decision-making system can generate a gradient maintenance scheme considering economy and reliability based on an optimization algorithm of reinforcement learning, effectively prolongs the service life of the bearing, reduces the maintenance cost, guarantees the data safety and traceability through the application of the block chain technology, achieves the collaborative diagnosis and knowledge sharing of cross-brand equipment, and improves the reliability of the bearing. The method has remarkable advantages in the aspects of improving the equipment reliability, optimizing the maintenance strategy, reducing the operation and maintenance cost and the like.
Owner:SHENZHEN FULING BUILDING TECH CO LTD

Non-specific person voice recognition intelligent switch control method and system based on deep learning

The invention relates to the technical field of voice recognition intelligent home control, and discloses a non-specific person voice recognition intelligent switch control method and system based on deep learning. The non-specific person voice recognition intelligent switch control method is applied to intelligent switch control equipment, and specifically comprises the following steps of S101, receiving original audio signals continuously collected in a to-be-controlled environment, and preprocessing the collected original audio signals, and then a starting point and an ending point of an effective voice segment are positioned by adopting endpoint detection based on a double-threshold method and combining the characteristic parameters of the short-time energy and the short-time zero-crossing rate. A multi-layer hidden layer structure with Dropout regularization is adopted in a neural network model, the generalization ability of the model is enhanced, a context sensing mechanism is introduced into a semantic understanding module, a composite instruction containing azimuth information can be intelligently analyzed, crossing from recognition to understanding is achieved, and the method has the advantages of being high in practicability and easy to popularize. The system is ensured to maintain a high recognition rate for voice instructions of different users under different environment conditions.
Owner:AIRBEST (SHENZHEN) TECHNOLOGY CO LTD

Voice emotion recognition method based on multiple scales and multiple features

The invention discloses a voice emotion recognition method based on multiple scales and multiple features, and belongs to the technical field of artificial intelligence. The method comprises the following steps: firstly, preprocessing an audio signal and extracting a spectrogram and a Mel-frequency cepstral coefficient; then, a residual network, a bidirectional long-short-term memory network and a HuBERT pre-training model are respectively utilized to extract spectrogram high-order spatial features, time sequence context features and voice semantic embedding features; secondly, inputting the first two features into a multi-dimensional multi-scale feature extraction module to extract richer time-frequency features, performing deep fusion by using a multi-layer cross attention mechanism, and performing weighted fusion with speech semantic embedded features; and finally, all the advanced features are spliced, and a final emotion category is recognized through a full-connection classifier. According to the invention, through combination of multi-scale feature extraction and an advanced fusion mechanism, the problem of insufficient complex emotion modeling ability in the prior art is effectively overcome, and the accuracy and robustness of voice emotion recognition are significantly improved.
Owner:NANJING INST OF TECH