Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

39961 results about "Audio frequency" patented technology

An audio frequency (abbreviation: AF) or audible frequency is a periodic vibration whose frequency is in the band audible to the average human. The SI unit of audio frequency is the hertz (Hz). It is the property of sound that most determines pitch.

Headset antenna and connector for the same

ActiveUS20090033574A1Increase the equivalent impedanceImprove rendering capabilitiesAntenna supports/mountingsAntenna adaptation in movable bodiesHeadphonesAudio frequency
A headset antenna and a connector for the same are provided. The headset antenna includes an audio signal line, an antenna and a high impedance element in specified application frequency ranges. The audio signal line is adapted for transmitting an audio signal and the antenna is adapted for receiving an RF signal. The high impedance element is disposed on a transmission path of the audio signal and generates a high impedance at a specified frequency band of the RF signal, so that the audio signal line is equivalent to an open circuit and the antenna obtains a better receiving capability.
Owner:HTC CORP

Conference summary processing method and system using AI

The invention relates to the technical field of intelligent conference processing, and relates to a conference summary processing method and system using AI, and the method comprises the steps: carrying out the real-time noise suppression of a collected conference audio stream and associated text data through a noise suppression algorithm, and carrying out the cross-modal alignment of the denoised data through a cross-modal alignment algorithm; a domain-specific attention head is inserted into an attention layer of the pre-trained Transform model, a domain-enhanced speech recognition model is constructed, and audio is converted into a text sequence with a speaker tag; adopting a heterogeneous graph neural network to construct a structured topic evolution graph; key decision nodes in the structured topic evolution graph are extracted based on a reinforcement learning strategy, and a final conference summary document is generated. In the decoding stage, the fusion proportion of the acoustic model and the language model is dynamically adjusted based on the real-time acoustic confidence coefficient, the recognition rate of the vocabularies in the professional field is increased, and the problems of frequent term transcription errors and poor semantic coherence in the professional conference are effectively solved.
Owner:GUANGZHOU DAZZLE VIEW INTELLIGENT TECH CO LTD

Intelligent sound box voice processing method and system based on artificial intelligence

The invention provides an intelligent sound box voice processing method and system based on artificial intelligence, and the method comprises the steps: obtaining audio data and mouth shape video data, carrying out the processing of the audio data and the mouth shape video data, and carrying out the multi-modal feature fusion, and obtaining a fusion feature; performing bimodal voice activity detection on the fusion features to obtain effective voice data; performing context sensing recognition of audio and video fusion on the effective voice data to obtain a first text; constructing a user feature model, and performing semantic understanding on the text based on the model to obtain an understanding result; performing intention recognition and slot filling based on the understanding result to obtain user intention and key information; generating a response strategy in combination with the user intention, the key information and the environment perception data; generating response voice according to the response strategy; and monitoring feedback information of the user to the response voice in real time, and updating the user feature model and the response strategy evaluation model based on feedback. According to the scheme, the voice can be recognized more accurately, and the safety and robustness of the system are enhanced.
Owner:SHENZHEN ZHANDIAN SMART TECH CO LTD

Method and system for early diagnosis of parkinson's disease based on multimodal deep learning

A method for early diagnosis of Parkinson's disease based on multimodal deep learning is provided. Audio-visual data of a to-be-diagnosed subject while performing a speech task is acquired. The audio-visual data are preprocessed to extract a plurality of audio segments and a plurality of video segments. A face image sequence is extracted from each of the plurality of video segments. A Mel-spectrogram of each of the plurality of audio segments is calculated. The face image sequence and the Mel-spectrogram are input into a multimodal deep learning model to output a classification result for Parkinson's disease early diagnosis of the to-be-diagnosed subject. A system for early diagnosis of Parkinson's disease based on multimodal deep learning is also provided.
Owner:SHANDONG UNIV

Real-time interactive digital human system supporting high concurrency and implementation method thereof

The invention discloses a real-time interactive digital human system supporting high concurrency and an implementation method thereof, and relates to the technical field of digital human interaction.The system comprises a model instance pool module used for loading the weight of a deep learning model to a shared memory area through the memory mapping technology; the multi-thread scheduling module is used for scheduling user requests by adopting a lock-free queue and a dynamic priority algorithm; the asynchronous pipeline processing module is composed of decoupled micro-services, and the modules are connected in series through asynchronous message queues; the client SDK is used for dynamically switching a rendering mode according to terminal hardware performance and network conditions; the elastic capacity expansion and contraction module is used for monitoring resource loads in real time and automatically adjusting the number of service instances; and the audio and video synchronization calibration module is used for ensuring that the synchronization error of the audio and video frames is lower than a preset threshold value through a timestamp alignment algorithm. According to the scheme, core challenges in a high-concurrency digital human interaction scene can be systematically solved, and innovative support is provided for large-scale real-time application.
Owner:LIANGSHENG DIGITAL ARTIFICIAL INTELLIGENCE (SHENZHEN) CO LTD

Multimodal intelligent agent system for dynamic environmental monitoring and human-centered support

A multimodal intelligent agent system for dynamic environmental monitoring and user-centered support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Multimedia equipment operation and maintenance management system based on AI

The invention relates to the technical field of audio and video analysis and data processing, in particular to an AI-based multimedia equipment operation and maintenance management system, which comprises a data acquisition module, an analysis module, an equipment state modeling module, a fault prediction module, a maintenance strategy optimization module, a resource allocation module and a visual platform. The data acquisition module acquires temperature, vibration, audio and video feature data of the equipment through the sensor and the video equipment, and the data are preprocessed and then transmitted to the cloud; the equipment state modeling module constructs a dynamic baseline model based on historical data, abnormal signals are generated through real-time comparison, and the visual platform realizes multi-dimensional state monitoring through 3D topology rendering, AR auxiliary diagnosis and early warning linkage. The system solves the problems of response lag and low resource utilization rate caused by dependence on manual operation and maintenance in the prior art through a closed-loop management mechanism driven by a data stream, and improves the fault diagnosis precision and the operation and maintenance efficiency.
Owner:BEIJING AIWEIKANG TECHNOLOGY CO LTD

Intelligent security data analysis decision method and system based on artificial intelligence

The invention provides an intelligent security and protection data analysis and decision method and system based on artificial intelligence. The method comprises the following steps: firstly, collecting heterogeneous data streams such as video frame sequences, infrared sensing signals, audio waveforms and equipment state logs uploaded by a plurality of security and protection monitoring equipment in a target area; multi-modal feature extraction is carried out through the spatial-temporal feature fusion model, and spatial-temporal feature vectors containing equipment deployment coordinates and the like are generated; dynamically dividing equipment clusters based on a self-adaptive clustering algorithm, so that the spatial-temporal feature vector similarity of the equipment in the same cluster is high; secondly, for each cluster, analyzing abnormal signal relevance by using an abnormal propagation network model to obtain a cluster-level risk prediction result and an equipment collaborative response instruction; and finally, according to the risk prediction result, dynamically adjusting the node weight of the abnormal propagation network model and the clustering center of the adaptive clustering algorithm, optimizing subsequent cluster division, and improving the accuracy and intelligence of intelligent security data analysis decision.
Owner:CHENGDU YUNJUXIANG TECH CO LTD

Method and system for adaptively adjusting ambient noise of Bluetooth headset

The invention relates to the technical field of audio noise reduction, and discloses a Bluetooth headset environment noise adaptive adjustment method and system, and the method comprises the steps: collecting multi-band sound data of headset environment noise through a preset acoustic sensor array, generating a noise spectrogram, extracting a noise masking threshold through time-frequency joint analysis, and collecting a headset audio signal; next, noise suppression is carried out on the audio signal by using a multi-order digital filter bank to obtain a primary noise reduction signal, the primary noise reduction signal is optimized into an optimized noise reduction signal based on a signal-to-noise ratio, and the frequency response matching degree of the optimized noise reduction signal and the original audio signal is calculated; then, performing phase correction and amplitude-frequency equalization processing according to the frequency response matching degree to obtain an audio equalization parameter; and finally, monitoring residual noise energy based on the parameter, generating a noise regulation and control instruction, driving adaptive sampling, and generating an adaptive noise reduction scheme. According to the invention, the audio noise reduction effect of the Bluetooth earphone can be improved.
Owner:SHENZHEN SHENYU ELECTRONICS TECH CO LTD

Techniques for determining conversational intent

The present disclosure relates to systems and methods for enhancing the interaction between users and automated agents, such as digital assistants, by employing Large Language Models (LLMs) to infer the intent of spoken language. The invention involves continuously monitoring ambient audio, converting speech to text, and utilizing LLMs to determine whether spoken language is intended for the automated agent. A structured prompt, including the converted text and specific instructions, is sent to the LLM, which is fine-tuned to process domain-specific prompts. The LLM provides a structured output in a standardized format, indicating the user's intent. The system may involve multiple prompts to perform separate tasks, such as identifying intent and generating additional context-specific data. This approach facilitates a more natural and intuitive user experience by eliminating the need for wake words and allowing seamless conversational interaction with virtual assistants across various platforms and devices.
Owner:SNAP INC

Intelligent control method and system for tunnel loudspeaker

The invention discloses an intelligent control method and system for tunnel loudspeakers, and relates to the technical field of tunnel audio control, environmental parameters in a tunnel are collected by adopting a mode of deploying sampling equipment in a distributed manner, and data preprocessing is performed in a targeted manner for different environmental parameters; a sound propagation model is established, and attenuation and delay of sound in different environments are simulated. According to the intelligent control method and system for the tunnel loudspeakers, various temperature and humidity sensors are arranged in the tunnel, and the absolute humidity is calculated in combination with the air pressure data, so that the sound velocity is accurately corrected, and the phase difference of the multiple loudspeakers is reduced; an adaptive Kalman filtering algorithm is adopted to process wind speed data, and reliable input is provided for a sound propagation model; a deep reinforcement learning algorithm is used to carry out collaborative optimization on parameters such as amplitudes and directional angles of multiple loudspeakers, a Bayesian network is used to detect loudspeaker faults, and Delaunay triangulation and a distributed consistency algorithm are combined to realize rapid fault reconstruction.
Owner:陕西省西咸新区秦汉新城城市管理中心

Audio signal authenticity verification method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an audio signal authenticity verification method, device, equipment and medium, and the method comprises the steps: constructing an original audio text data set, and generating an adversarial sample set, inputting the original audio text data set and the adversarial sample set into an audio detection model for joint training to obtain an audio detection model subjected to adversarial training; the method comprises the steps of obtaining a to-be-detected audio signal and extracting an acoustic feature of the to-be-detected audio signal, obtaining a non-acoustic feature associated with the to-be-detected audio signal, constructing a multi-dimensional feature vector according to the acoustic feature and the non-acoustic feature, inputting the multi-dimensional feature vector into an audio detection model to generate an abnormal index, and executing a hierarchical response operation based on the abnormal index. According to the method, the robustness of the model is enhanced by introducing adversarial sample training, and the multi-dimensional feature vector is constructed by fusing the multi-modal features, so that accurate recognition and hierarchical response to the voice cloning attack are realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Large model Agent intelligent decision-making method and system fusing multi-modal data

The invention discloses a multi-modal data fused large model Agent intelligent decision-making method and system, belongs to the technical field of artificial intelligence, multi-modal data processing, deep learning, reinforcement learning and intelligent decision-making, and aims to solve the technical problem of how to improve the performance and adaptability of intelligent decision-making in processing complex tasks and dynamic environments. According to the technical scheme, the method comprises the steps of multi-modal data fusion, wherein text, image and audio data from different modals are integrated, and unified feature representation is generated through feature extraction and feature fusion technologies; intelligent decision-making: decision-making reasoning is carried out based on the fused feature representation, and a final decision-making result is generated by adopting a deep learning model and a reinforcement learning algorithm; adaptive learning: monitoring data changes and decision-making effects in real time, and dynamically adjusting deep learning model parameters and strategies; and feedback optimization: further optimizing the performance of the deep learning model by collecting the feedback information of the decision result.
Owner:浪潮智慧城市科技有限公司

Ear worn device and case

A system, including: an ear worn device configured to provide audio to a user's ear; a switch configured to receive physical contact from the user and, in response, alter an activation state of the switch; and a case configured to receive the ear worn device when not in use, wherein the case includes an electromagnet configured to magnetically attract the ear worn device, wherein changing the activation state of the switch causes the electromagnet to reduce attraction to the ear worn device to allow the user to more easily remove the ear worn device from the case.
Owner:MASIMO CORP

Audio noise reduction method, device and system based on deep learning

The invention relates to an audio noise reduction method, device and system based on deep learning, and the method comprises the steps: obtaining an input audio signal with noise, and carrying out the multi-scale time-frequency decomposition, and obtaining a mixed time-frequency feature and a noise fingerprint spectrum; performing parameter parallel processing on the noise fingerprint spectrum through a preset dynamic kernel generation network, and performing preliminary noise reduction processing on the mixed time-frequency characteristics to obtain noise-reduced mixed data; performing dual-path processing structure construction on the noise reduction mixed data to obtain amplitude optimization data and phase optimization data; performing dynamic time-frequency domain cross fusion on the amplitude optimization data and the phase optimization data to obtain fused audio data; and carrying out differentiable acoustic equation constraint adversarial training on the fused audio data, and carrying out inverse time-frequency transformation processing to obtain a target noise-reduced audio signal. According to the invention, the overall efficiency and effect of audio signal processing can be effectively improved.
Owner:DONGGUAN HUAZE ELECTRONIC TECH CO LTD

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Voice interaction method and device based on lip language enhancement, equipment and storage medium

The invention discloses a voice interaction method and device based on lip language enhancement, equipment and a storage medium, and the method comprises the steps: extracting lip language features based on an image sequence of a lip region, and carrying out the feature extraction of a voice signal, and obtaining an audio feature; performing cross-modal fusion coding on the lip language features and the audio features to generate mixed features containing audio-visual information; inputting the mixed features into a large language model, understanding the intention of the interaction object and generating a corresponding semantic reply; and finally, synthesizing into voice and / or converting into characters. According to the invention, by introducing the lip features, additional visual clues are provided for speech recognition, and the robustness and accuracy of speech recognition can be significantly improved; effective fusion coding is carried out on the lip language features and the sound features, and semantic information splitting caused by simple and independent recognition is avoided; and the capability of the large model is fully utilized, so that more natural and more intelligent interaction experience is realized.
Owner:SHENZHEN WANRUI INTELLIGENT TECH CO LTD

Intelligent microphone pickup and speech enhancement method, system and device

The invention relates to an intelligent microphone pickup and speech enhancement method, system and device, and the method comprises the steps: obtaining an audio signal collected by a dual-silicon microphone array, and carrying out the cross-correlation analysis of the audio signal, and obtaining a corresponding human voice signal correlation feature; performing sound source positioning analysis on the audio signal based on the human voice signal correlation feature to obtain a corresponding target human voice signal; performing frequency characteristic analysis on the target human voice signal, and performing segmented dynamic gain processing on the signal according to a preset frequency band range to obtain a corresponding voice enhancement signal; performing noise component adaptive filtering processing on the target human voice signal to obtain a corresponding noise suppression parameter; and inputting the speech enhancement signal and the noise suppression parameter into a preset echo cancellation model for joint optimization to obtain a corresponding output human voice signal. According to the invention, the coupling problem of multiple acoustic interferences can be effectively solved.
Owner:SHENZHEN SHIDU DIGITAL TECH CO LTD

Insurance claim settlement-oriented multi-modal image video evidence analysis method and system

The invention discloses an insurance claim settlement-oriented multi-modal image video evidence analysis method and system. The method comprises the following steps of: acquiring video / image and multi-source data such as metadata, audio, IMU (Inertial Measurement Unit), GPS (Global Positioning System), OBD (On-Board Diagnostic) and the like; calculating content Hash of the video and the audio according to frames, connecting the content Hash with time information in series to form chained Hash, and adding a verification digital signature and a credible timestamp; realizing cross-modal time sequence alignment based on self-adaptive time anchor-attitude coupling; tampering detection is carried out in combination with PRNU fingerprints, noise field consistency, dual compression, copy-movement and the like; multi-view geometry and monocular depth are fused, IMU scale constraint and micro rendering are introduced, three-dimensional reconstruction and re-projection optimization are completed, and collision dynamics verification is carried out; and constructing an event cause and effect graph, judging responsibility in combination with traffic rules, outputting a confidence coefficient vector and a structured report, and generating a verifiable evidence packet. The scheme has the advantages of high efficiency and traceability in the aspects of space-time restoration and interpretable responsibility judgment.
Owner:国任财产保险股份有限公司

Speech feature processing method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice feature processing method, device, equipment and medium. Performing time resolution analysis based on the fused Mel band energy to generate a multi-scale Mel spectrum amplitude value, and performing nonlinear transformation on the multi-scale Mel spectrum amplitude value according to the noise intensity parameter to generate a noise suppression Mel component; and generating a perception weighting coefficient according to an auditory perception model, and executing frequency domain energy adjustment on the noise suppression Mel component to generate Mel spectrum representation. On the basis of frequency resolution self-adaption, time resolution dynamic adjustment and auditory perception modeling, nonlinear transformation and perception weighting processing are applied to the multi-scale Mel spectrum amplitude value, the influence of noise interference on voice features can be effectively reduced, and the key information retention capacity of voice signals is enhanced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Audio analysis method and device based on feature fusion, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology, culture and art and the like, and discloses a feature fusion-based audio analysis method, which comprises the following steps of: obtaining a target audio in a target field and a target text associated with the target audio, extracting a music feature vector of the target audio, extracting a text semantic feature vector of the target text, fusing the music feature vector and the text semantic feature vector to generate fusion features, constructing a knowledge graph containing knowledge nodes of the target domain, inputting the fusion features into the knowledge graph for semantic analysis, and generating an analysis result. According to the method, deep semantic analysis is performed by fusing audio and text features and combining a knowledge graph, so that multi-modal understanding of audio data is realized, the accuracy and interpretability of an analysis result are improved, and the relevance between the analysis result and industry knowledge is enhanced; therefore, the applicability of the audio analysis technology in the fields of culture and art, medical health, finance and the like is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Dynamic calibration method and system of vehicle-mounted emotion recognition system

The invention provides a dynamic calibration method and system for a vehicle-mounted emotion recognition system, and the method comprises the steps: S1, obtaining multi-source data which comprises a facial image, a voice signal and a physiological signal; the obtained multi-source data are preprocessed, and preprocessed multi-source data are obtained; s2, performing feature extraction based on the preprocessed facial image, the voice signal and the physiological signal to obtain a facial expression feature vector, an audio feature vector and a physiological state feature vector; s3, evaluating the current environment credibility based on an environment credibility evaluation function; s4, dynamically distributing the weight of the multi-source data according to the credibility of the current environment and the real-time scene; and S5, constructing a multi-modal fusion vector based on the dynamically distributed weight of the multi-source data, the facial expression feature vector, the audio feature vector and the physiological state feature vector, and performing emotion recognition by using the constructed emotion recognition model based on the multi-modal fusion vector.
Owner:SHANGHAI PUFAFEN ELECTRONIC TECH CO LTD

Audio coding and decoding method, device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an audio coding and decoding method, device, equipment and medium, and the method comprises the steps: carrying out the sliding window segmentation processing of an input audio signal, and generating signal segments; processing the signal segments through an encoder containing a multi-layer self-attention mechanism to generate continuous potential representations; performing decomposition vector quantization processing on the continuous potential representation to generate a discrete code; the discrete codes are processed through a decoder comprising a multi-layer self-attention mechanism, and reconstructed signal segments are generated; and splicing the reconstructed signal segments to generate a complete audio signal. According to the method, a traditional convolution structure is replaced by a multi-layer self-attention mechanism, global time sequence dependence modeling is carried out on the audio signals after sliding window segmentation, effective compression of potential representation is realized in combination with decomposition vector quantization, and continuity and fidelity of audio reconstruction are improved on the premise that calculation complexity is not increased.
Owner:PING AN TECH (SHENZHEN) CO LTD

Large-scene monitoring video abnormal event early warning method based on multi-modal large model

The invention relates to the technical field of abnormal event early warning, and provides a large-scene monitoring video abnormal event early warning method based on a multi-mode large model. According to the invention, the problems of delay, low accuracy and limited coverage range of abnormal event early warning of large-scene monitoring videos in the prior art are solved. According to the main scheme, multiple paths of high-resolution monitoring videos are spliced and preprocessed to generate a panoramic video; synchronously acquiring and preprocessing audio and sensor data to construct a multi-modal data set; video key frames are extracted by adopting a traditional small model, and the video key frames and multi-modal data are jointly input into a multi-modal large model based on a Transform architecture for deep feature fusion; abnormal events such as tumble, congestion and fight are identified based on the fusion features; triggering an early warning mechanism to send event type and position information in real time; and storing the full-dimensional data of the abnormal event for tracing analysis. The real-time processing performance is optimized through edge calculation, the complex scene understanding ability is enhanced in combination with a multi-modal large model, and the detection precision and the response speed are remarkably improved.
Owner:PEKING UNIV (TIANJIN BINHAI) NEW GENERATION INFORMATION TECH RES INST +1

Wind turbine generator data analysis and fault diagnosis method and system based on big data and artificial intelligence

The invention discloses a wind turbine generator data analysis and fault diagnosis method and system based on big data and artificial intelligence. According to the method, a blade image, a vibration signal, audio data and operation parameters are synchronously acquired through an unmanned aerial vehicle multi-mode sensor and a ground monitoring system, and a multi-source heterogeneous data set is constructed; after the data is classified and preprocessed, image features, vibration time-frequency domain features and operation parameter key value pairs are extracted respectively; dimensionality reduction is carried out by using an auto-encoder, feature-level space-time alignment is realized through an improved DTW algorithm, and a multi-dimensional fault feature matrix is generated; a hierarchical diagnosis model including a GRU auto-encoder, an MLP network and an attention mechanism CNN is constructed, and training is carried out by taking minimization of sub-model deviation as an optimization target; and finally, fusing multi-source features to realize fault classification, and generating a visual diagnosis report. According to the method, efficient fusion and accurate diagnosis of multi-source heterogeneous data are realized, and the accuracy and the real-time performance of fault detection of the wind turbine generator are remarkably improved.
Owner:NAT ENERGY GRP DONGTAI OFFSHORE WIND POWER CO LTD

Apparatus and method for end-to-end text-to-speech synthesis

An apparatus for end-to-end text-to-speech synthesis is provided. The apparatus comprises input interface circuitry configured to receive first input data indicative of a phoneme and second input data indicative of a first target duration for the phoneme. The apparatus further comprises processing circuitry configured to, using a trained machine-learning model, map the phoneme to a state using an encoder sub-model of the trained machine-learning model, estimate a second target duration for the phoneme based on the state and determine an attention weight based on the first target duration and the second target duration. The processing circuitry is further configured to map the state to audio data based on the attention weight using a decoder sub-model of the trained machine-learning model, wherein the audio data are indicative of an audio waveform representing speech.
Owner:SONY GROUP CORP

Multi-modal intention recognition method and system

The invention relates to a multi-mode intention recognition method and system, and the method comprises the steps: carrying out the time domain and frequency domain enhancement of the features of text, video and audio modes, carrying out the splicing to obtain non-language mode fusion features, combining the features of an original text, modeling the time synchronization relation of audio-text and video-text, and carrying out the recognition of a multi-mode intention. Standardized audio features, video features and text features are obtained through context alignment processing; fusing the standardized features of the three modalities to obtain a fused feature vector, and mapping the fused feature vector back to the text modal space to be connected with the weighted residual error of the original text feature to obtain a fused semantic vector; extracting global semantic anchor points and mask positions from the fused semantic vector, and splicing the global semantic anchor points and the mask positions with the original text features and the fused semantic vector to obtain input features; and obtaining probability distribution of multiple intention categories by using the input features. Three types of heterogeneous modal input can be supported, and the accuracy and robustness of intention recognition are improved through fine-grained semantic supervision and enhancement strategies.
Owner:XINJIANG UNIVERSITY

Audio and video dual-mode emotion recognition method and system based on adapter fusion

The invention relates to the technical field of artificial intelligence and emotion calculation, in particular to an audio and video dual-mode emotion recognition method and system based on adapter fusion. The method comprises the following steps: acquiring a video frame sequence and an audio signal, and preprocessing the video frame sequence and the audio signal; constructing an emotion recognition model; based on a bimodal feature extraction module, a space adapter and a global adapter are embedded in sequence, and corresponding modal enhanced space features and global features are obtained in sequence; generating intermediate representations of the corresponding modes based on the global features, and performing feature fusion according to the intermediate representations to obtain fusion features of the corresponding modes; the fusion features are spliced, time sequence features are extracted, and final features are obtained; inputting the final features into a classifier to obtain a predicted emotion category, training an emotion recognition model by adopting a loss function, and determining an optimal emotion recognition model; and inputting a to-be-recognized video frame sequence and an audio signal into the emotion recognition model, and outputting a recognition result.
Owner:NANJING MEDICAL UNIV

Voice enhancement method and device based on noise perception, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice enhancement method, device and equipment based on noise perception and a medium. Environment feature information is extracted and input into an audio enhancement model to generate an enhanced audio signal; obtaining a reference audio sample, extracting a personalized feature vector, and carrying out personalized processing on the enhanced audio signal; and collecting playing feedback data, determining a playing time domain adjustment parameter and a playing frequency domain adjustment parameter, adjusting the personalized enhanced audio signal, and generating an optimized audio signal. According to the method, dynamic adjustment is realized in combination with the feedback parameters in the playing process by fusing the environmental perception information and the personalized speaker characteristics, clear and natural optimized audio output with personalized styles can be generated in a complex environment, and the voice interaction quality and adaptability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice interaction method and system of AI intelligent robot

The invention relates to the technical field of voice interaction, particularly discloses an AI intelligent robot voice interaction method and system, and aims to solve the problems of low voice interaction accuracy, insufficient reliability and lack of authority control in a complex noise environment. A dynamic noise feature library containing steady-state noise, impact noise and human voice interference features and a pre-stored gesture instruction library are constructed, audio signals are collected in real time, low-frequency-band, middle-frequency-band and high-frequency-band differential noise reduction is executed, Mel-frequency cepstral coefficient features are extracted, noise scenes are matched, corresponding voice recognition models are switched, and voice recognition is achieved. And calculating a confidence value of the voice instruction, outputting multi-modal verification data in combination with a dynamic confidence threshold, and outputting an authority control signal through voiceprint matching, authority verification and instruction consistency judgment. Through multi-modal fusion, dynamic adaptation and authority control, the voice recognition accuracy and interaction safety in a complex noise environment are remarkably improved, and the method is suitable for scenes such as factory intelligent inspection.
Owner:HANGZHOU SOHA TECH CO LTD