Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6542 results about "Audio signal" patented technology

An audio signal is a representation of sound, typically using a level of electrical voltage for analog signals, and a series of binary numbers for digital signals. Audio signals have frequencies in the audio frequency range of roughly 20 to 20,000 Hz, which corresponds to the lower and upper limits of human hearing. Audio signals may be synthesized directly, or may originate at a transducer such as a microphone, musical instrument pickup, phonograph cartridge, or tape head. Loudspeakers or headphones convert an electrical audio signal back into sound.

Headset antenna and connector for the same

ActiveUS20090033574A1Increase the equivalent impedanceImprove rendering capabilitiesAntenna supports/mountingsAntenna adaptation in movable bodiesHeadphonesAudio frequency
A headset antenna and a connector for the same are provided. The headset antenna includes an audio signal line, an antenna and a high impedance element in specified application frequency ranges. The audio signal line is adapted for transmitting an audio signal and the antenna is adapted for receiving an RF signal. The high impedance element is disposed on a transmission path of the audio signal and generates a high impedance at a specified frequency band of the RF signal, so that the audio signal line is equivalent to an open circuit and the antenna obtains a better receiving capability.
Owner:HTC CORP

Audio signal authenticity verification method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an audio signal authenticity verification method, device, equipment and medium, and the method comprises the steps: constructing an original audio text data set, and generating an adversarial sample set, inputting the original audio text data set and the adversarial sample set into an audio detection model for joint training to obtain an audio detection model subjected to adversarial training; the method comprises the steps of obtaining a to-be-detected audio signal and extracting an acoustic feature of the to-be-detected audio signal, obtaining a non-acoustic feature associated with the to-be-detected audio signal, constructing a multi-dimensional feature vector according to the acoustic feature and the non-acoustic feature, inputting the multi-dimensional feature vector into an audio detection model to generate an abnormal index, and executing a hierarchical response operation based on the abnormal index. According to the method, the robustness of the model is enhanced by introducing adversarial sample training, and the multi-dimensional feature vector is constructed by fusing the multi-modal features, so that accurate recognition and hierarchical response to the voice cloning attack are realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice enhancement method and device based on noise perception, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice enhancement method, device and equipment based on noise perception and a medium. Environment feature information is extracted and input into an audio enhancement model to generate an enhanced audio signal; obtaining a reference audio sample, extracting a personalized feature vector, and carrying out personalized processing on the enhanced audio signal; and collecting playing feedback data, determining a playing time domain adjustment parameter and a playing frequency domain adjustment parameter, adjusting the personalized enhanced audio signal, and generating an optimized audio signal. According to the method, dynamic adjustment is realized in combination with the feedback parameters in the playing process by fusing the environmental perception information and the personalized speaker characteristics, clear and natural optimized audio output with personalized styles can be generated in a complex environment, and the voice interaction quality and adaptability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice interaction method and system of AI intelligent robot

The invention relates to the technical field of voice interaction, particularly discloses an AI intelligent robot voice interaction method and system, and aims to solve the problems of low voice interaction accuracy, insufficient reliability and lack of authority control in a complex noise environment. A dynamic noise feature library containing steady-state noise, impact noise and human voice interference features and a pre-stored gesture instruction library are constructed, audio signals are collected in real time, low-frequency-band, middle-frequency-band and high-frequency-band differential noise reduction is executed, Mel-frequency cepstral coefficient features are extracted, noise scenes are matched, corresponding voice recognition models are switched, and voice recognition is achieved. And calculating a confidence value of the voice instruction, outputting multi-modal verification data in combination with a dynamic confidence threshold, and outputting an authority control signal through voiceprint matching, authority verification and instruction consistency judgment. Through multi-modal fusion, dynamic adaptation and authority control, the voice recognition accuracy and interaction safety in a complex noise environment are remarkably improved, and the method is suitable for scenes such as factory intelligent inspection.
Owner:HANGZHOU SOHA TECH CO LTD

Compensation for nonuniform delayed group communications

A method for synchronizing audio reproduction in collocated end devices is presented. Each of the devices auto-correlates using noise to determine a threshold prior to the antenna receiving an audio signal. When the devices receive a common audio signal, they provide audio outputs. Each device cross-correlates its audio output with the audio outputs of the other devices. The timing of the audio output of each device is then adjusted such that the audio outputs of all of the devices align temporally with the lagging or leading device.
Owner:MOTOROLA SOLUTIONS INC

Methods and systems of text-conditioned audio-visual speech generation with multi-modal latent diffusion models

Methods, systems, and computer programs are presented for audio-visual speech generation with multi-modal latent diffusion models. One method includes encoding raw audio signals and video frames into respective latent spaces using audio and visual autoencoders. A text transcript is processed into phoneme sequences using a text transcript processor. The audio and visual latent spaces are conditioned using the text transcript and a conditioning variable. Joint distributions of the visual and audio latent spaces, text transcript, and conditioning variable are learned using a multi-modal latent diffusion model. The model adds noise to the latent audio-visual representations and predicts the noise through denoising neural networks. An inverted diffusion process is utilized to generate diverse speech content and speaker characteristics, resulting in realistic audio-visual speech. The technology presented provides a novel approach to conditional speech generation with potential applications in speech synthesis, voice conversion, and speech recognition.
Owner:TENSORTYPE INC

Auxiliary expectoration control method and device

The invention provides an expectoration auxiliary control method and device. The method comprises the steps that audio signals of breathing sounds and cough sounds at multiple positions of the chest of a patient are collected through an audio sensor array, and a standardized audio feature data set is obtained through noise reduction, segmentation and feature extraction; a dynamic weighted graph model is constructed, and a random edge sampling algorithm and a parallel batch processing dynamic algorithm are combined to analyze and obtain a sputum viscosity index and a sputum distribution position map; based on the result, mapping the vibration parameter space into an unweighted disk diagram, and obtaining a personalized vibration treatment scheme by using a shortest path algorithm; in combination with the real-time breathing cycle of the patient, working parameters of the sound wave vibrator and the negative pressure suction device are synchronously controlled, and a coordinated and consistent multi-mode treatment execution instruction is obtained; and collecting real-time feedback data in treatment, and dynamically adjusting working parameters through reinforcement learning to obtain an optimized expectoration adjuvant therapy scheme. Intelligent adjustment can be achieved based on the real-time breathing state and sputum characteristics of the patient, the expectoration efficiency is improved, and discomfort of the patient is reduced.
Owner:THE FIRST AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Wind turbine generator voiceprint fault recognition method

The invention provides a wind turbine generator voiceprint fault recognition method, and relates to the technical field of wind turbine generator state monitoring and fault diagnosis, and the method comprises the steps: carrying out the noise reduction of an original audio signal through variational mode decomposition, screening a target mode of which the frequency, energy and kurtosis accord with features, and reconstructing the signal; extracting a Mel frequency cepstrum coefficient and a sensing noise robust coefficient, and generating multi-dimensional voiceprint data in combination with statistical characteristics such as a frequency spectrum gravity center, a spectrum entropy, energy, kurtosis and a zero-crossing rate; constructing a support set based on the prototype network, realizing small sample fault classification by calculating the Euclidean distance between the feature vector and the prototype vector, and outputting a preliminary result; judging whether the voiceprint is abnormal according to a preset threshold value, if so, storing the voiceprint into a dynamic abnormal voiceprint knowledge base; frequently occurring abnormal samples are manually labeled and added into a support set, the prototype network is retrained to update the model, and continuous optimization of the fault recognition capability is achieved.
Owner:CGN (SHANXI) NEW ENERGY INVESTMENT CO LTD

System and method for enhancing speech of target speaker from audio signal in an ear-worn device using voice signatures

An ear-worn device is provided that operates to isolate and individually treat the received speech of a target speaker or multiple target speakers from an audio input signal detected in a multi-speaker environment. The ear-worn device uses a machine learning model that receives a voice signature of each of one or more target speakers as input signals, to identify and isolate the component of the audio input signal attributable to the target speaker(s). Once isolated, the target speaker's speech may be enhanced, de-emphasized, or otherwise processed in a manner desired by the wearer of the ear-worn device. The wearer may use an external electronic device, e.g., a phone, to select one or more target speakers in a conversation and / or configure various settings associated with processing the speech on the ear-worn device.
Owner:FORTELL RESEARCH INC

Unsupervised wind power equipment blade fault detection method based on phase perception parallel attention mechanism

The invention relates to a wind power equipment blade fault detection technology, discloses an unsupervised wind power equipment blade fault detection method based on a phase perception parallel attention mechanism, and solves the problems that an existing wind power equipment blade fault detection method is high in dependence on labeled data, insufficient in generalization ability under strong noise and variable working conditions and high in fault detection efficiency. And a weak transient fault signal and a dynamic change characteristic are difficult to capture robustly. According to the scheme of the invention, the method comprises the steps: collecting a blade operation audio signal, and extracting a dual-channel time-frequency feature containing an amplitude spectrum and a phase spectrum through improved short-time Fourier transform; a deep adversarial auto-encoder is constructed by using an encoder containing a phase perception parallel attention module, a decoder and an auxiliary encoder, and normal working condition feature distribution is learned by reconstructing an error loss, potential representation consistency loss, adversarial loss and phase consistency loss optimization model during off-line training; in the reasoning stage, the fault is judged based on the feature distance score and the reconstruction error score.
Owner:CHINA HYDROELECTRIC ENGINEERING CONSULTING GROUP CHENGDU RESEARCH HYDROELECTRIC INVESTIGATION DESIGN AND INSTITUTE

Room sound correction method and device of audio equipment, equipment and storage medium

The invention relates to the technical field of audio signal processing, and discloses a room sound correction method, device and equipment for audio equipment and a storage medium, which are used for improving the accuracy and correction effect of correction parameters and effectively improving the acoustic effect of the audio equipment in a room. The room sound correction method of the audio equipment comprises the following steps: acquiring response data of a room to a multi-band preset test signal, and extracting room acoustic data by adopting a time-frequency conjoint analysis method, the room acoustic data comprising frequency attenuation, reflection paths, phase shifts and reverberation time information of different bands; based on the room acoustic data, performing compensation processing by using preset microphone calibration data to obtain a calibrated frequency response curve; performing fitting calculation on a preset ideal frequency curve and the calibrated frequency response curve according to an optimization algorithm to obtain specific correction parameters of the room; and correspondingly adjusting the output of the audio equipment according to the correction parameter.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

Security camera abnormal behavior identification method and system based on multi-modal fusion

The invention relates to a security camera abnormal behavior identification method and system based on multi-modal fusion. The method comprises the steps of obtaining multi-modal data such as visible light image data, infrared image data and audio signal data of a security camera in a target monitoring period; performing feature extraction on the data of different modes to obtain respective feature vector sets; the features of all the modes are fused, and in the multi-mode feature fusion process, a weighted fusion algorithm for dynamically adjusting the fusion weight according to the confidence coefficient weight of feature vectors of all the modes is adopted; inputting the fusion feature vector set into an abnormal behavior classification model for classification and early warning; according to the scheme, the accuracy and effectiveness of abnormal behavior recognition in a complex environment can be improved, and then the overall efficiency of security monitoring is improved.
Owner:SHENZHEN KEAN DIGITAL CO LTD

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Intelligent predictive maintenance system for audio equipment fault

The invention discloses an intelligent predictive maintenance system for an audio equipment fault, and the system comprises a data collection module which collects the internal sensor data during the operation of audio equipment, and outputs an audio signal feature parameter, an external environment parameter, and a historical operation log; the feature preprocessing module is used for carrying out standardization and noise reduction processing on the multi-source data based on a multi-modal feature fusion algorithm; the health state evaluation module outputs a health state evaluation result of the audio equipment in real time through a hybrid analysis model combining a convolutional neural network, a long and short-term memory network and an attention mechanism; the fault risk prediction module is used for performing real-time fault risk prediction according to the evaluation result and generating predictive maintenance decision parameters; and the maintenance decision and early warning module is used for outputting fault early warning information according to the prediction parameters and automatically generating maintenance operation suggestions when the early warning level reaches a preset condition. According to the invention, the operation reliability of the audio equipment can be effectively improved, and intelligent prediction and advanced maintenance of faults are realized.
Owner:SHENZHEN JIEYU INFORMATION TECH CO LTD

Physiological monitoring soundbar

A soundbar for medical monitoring which may comprise a speaker, a sensor, and a hardware processor. The speaker can be configured to emit audio signals. The sensor can be configured to obtain sensor data relating to a physiology of a subject. The sensor can include a camera and the sensor data can include image data. The hardware processor can be configured to access the sensor data and determine a health status of the subject based on at least the sensor data.
Owner:MASIMO CORP

Diffusion-based audio purification for defending against adversarial deepfake attacks

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”). Embodiments implement a machine-learning architecture having a diffusion model that generates purified features that are fed to a deepfake detection model. The machine-learning architecture includes input layers that convert an audio signal into a Gaussian or frequency space representation (e.g., log spectrogram) to extract a set of initial features indicative of spoofing or deepfake attacks. The diffusion model identifies adversarial noise on the audio signal in the initial features and generates purified features or clean version of the input audio signal. A deepfake detector includes a neural network architecture and classifier programmed and trained to generate a deepfake detection score and classify the audio signal as genuine or fraudulent using the purified features.
Owner:PINDROP SECURITY INC

Lamp strip module atmosphere creating method based on scene induction control

The invention discloses a lamp strip module atmosphere creating method based on scene induction control, and relates to the technical field of intelligent illumination control, and the method comprises the following steps: S1, collecting an original audio signal of a target space, extracting amplitude fluctuation, a spectrum structure and energy mutation features, recording the starting time and duration of a sound event, and constructing a sound source time sequence; according to the invention, through sound source time sequence modeling, abnormal behavior identification and environment trend analysis, in combination with a multi-dimensional consistency judgment mechanism, accurate response to an abnormal sound source is realized, safe transition is carried out through neutral color temperature lighting effect, and then light effect smooth regression is realized through dynamic recovery evaluation and a slow changing algorithm, so that the accuracy of the abnormal sound source is improved. An intelligent light control system with real-time perception, context understanding and closed-loop regulation and control capabilities is constructed, immersion experience and safety are improved, and the intelligent light control system has remarkable practical value and technical innovation.
Owner:GUANGZHOU MAGIC MASTER TECHNOLOGY CO LTD

Noise reduction method combining different noise reduction algorithms of motorcycle riding earphone

The invention relates to the technical field of earphone noise reduction processing, and discloses a motorcycle riding earphone noise reduction method combining different noise reduction algorithms, comprising the following steps: acquiring an original audio signal, and distributing the original audio signal to each processing module; a preprocessing signal is generated, voice activity information is determined, and environmental noise energy information is calculated; according to the environmental noise energy or the auxiliary information, adaptively adjusting a high-low noise energy threshold value; the original signals are input into the traditional and AI noise reduction module in parallel to generate two paths of noise reduction signals; based on the environment noise, the voice information and the threshold value, performing weighted fusion to generate an output signal; and carrying out equalization and compression processing on the output signal, and driving the loudspeaker to output. According to the invention, comprehensive judgment on environmental noise energy information, voice activity information and riding speed is introduced, an output signal of a traditional noise reduction module is set to be zero in a high-speed voice-free environment, and AI noise reduction is combined, so that tone quality distortion of a traditional noise reduction algorithm in a noise environment is effectively avoided.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD +1

Aluminum electrolytic capacitor state evaluation method and device based on multi-dimensional data analysis

The invention relates to the field of capacitor state evaluation, in particular to an aluminum electrolytic capacitor state evaluation method and device based on multi-dimensional data analysis. The method comprises the following steps: performing multi-band excitation on an aluminum electrolytic capacitor, and performing three-dimensional fault distribution reconstruction and structure degradation trend prediction so as to construct an impedance tomography degradation trend prediction map; collecting audio signals in the charging and discharging process of the capacitor, carrying out acoustic texture analysis, carrying out voiceprint-structure degradation correlation mining according to the impedance tomography degradation trend prediction map, and constructing an acoustic texture damage evolution evaluation model; obtaining a thermal imaging image of the capacitor, carrying out local overheating diffusion mining, and constructing a three-dimensional thermal anomaly detection model; performing electrolyte chemical component change tracking and electrolyte degradation situation analysis on the electrolyte of the capacitor, and constructing an electrolyte degradation situation curve; according to the invention, efficient and accurate capacitor state evaluation is realized. The operation reliability and stability of capacitor equipment are improved, and the service life of the equipment is effectively prolonged.
Owner:ZHUHAI LEAGUER CAPACITOR

Vocal music training vowel pronunciation quality evaluation method based on auditory and visual spatio-temporal feature fusion

The invention provides a vocal music training vowel pronunciation quality evaluation method based on auditory and visual spatial-temporal feature fusion, and the method comprises the steps: collecting vowel pronunciation audio signals and corresponding videos of a singer, and constructing a multi-modal data set; generating a fractional order Mel spectrogram for the audio signal through short-time fractional order Fourier transform of an adaptive order; extracting time sequence features and spatial features of the fractional order Mel spectrogram, and fusing the time sequence features and the spatial features through a gating mechanism to generate audio spatio-temporal features; face visual features in the video are extracted and fused with the audio spatio-temporal features through a cross attention mechanism, and the cross attention mechanism is integrated with a periodic modeling network; the fused features are input into a classifier, a dynamic weight multi-mode cosine loss function training model is adopted, the dynamic weight multi-mode cosine loss function dynamically adjusts the sample weight through a confusion matrix, and the weight is increased for the samples with classification errors based on the historical frequency mistaken division times of the samples; and outputting a pronunciation quality evaluation result.
Owner:FUZHOU UNIV

Method, apparatus and system for neural network hearing aid

The disclosure generally relates to a method, system and apparatus to improve a user's understanding of speech in real-time conversations by processing the audio through a neural network contained in a hearing device. The hearing device may be a headphone or hearing aid. In one embodiment, the disclosure relates to an apparatus to enhance incoming audio signal. The apparatus includes a controller to receive an incoming signal and provide a controller output signal; a neural network engine (NNE) circuitry in communication with the controller, the NNE circuitry activatable by the controller, the NNE circuitry configured to generate an NNE output signal from the controller output signal; and a digital signal processing (DSP) circuitry to receive one or more of controller output signal or the NNE circuitry output signal to thereby generate a processed signal; wherein the controller determines a processing path of the controller output signal through one of the DSP or the NNE circuitries as a function of one or more of predefined parameters, incoming signal characteristics and NNE circuitry feedback.
Owner:FORTELL RESEARCH INC

Method, apparatus and system for neural network hearing aid

The disclosure generally relates to a method, system and apparatus to improve a user's understanding of speech in real-time conversations by processing the audio through a neural network contained in a hearing device. The hearing device may be a headphone or hearing aid. In one embodiment, the disclosure relates to an apparatus to enhance incoming audio signal. The apparatus includes a controller to receive an incoming signal and provide a controller output signal; a neural network engine (NNE) circuitry in communication with the controller, the NNE circuitry activatable by the controller, the NNE circuitry configured to generate an NNE output signal from the controller output signal; and a digital signal processing (DSP) circuitry to receive one or more of controller output signal or the NNE circuitry output signal to thereby generate a processed signal; wherein the controller determines a processing path of the controller output signal through one of the DSP or the NNE circuitries as a function of one or more of predefined parameters, incoming signal characteristics and NNE circuitry feedback.
Owner:FORTELL RESEARCH INC

Bird identification method and device based on sound-image multi-modal fusion

The invention discloses a bird identification method based on sound-image multi-modal fusion. The bird identification method comprises the following steps: S1, carrying out standardized frame-level preprocessing on bird audio signals; s2, acoustic features are extracted and enhanced, and an acoustic high-level feature vector which highlights birdsong discrimination information and suppresses environmental noise is obtained; s3, visual image standardization preprocessing; s4, performing visual feature extraction and multi-scale fusion to obtain a visual high-level feature vector which enhances correspondence to the bird key form area and inhibits background interference; s5, performing dynamic weighted fusion on the decision-making layer to obtain a bird existence probability; and S6, comparing the bird existence probability with a preset threshold value of the corresponding bird, and judging whether the bird exists or not and the type of the existing bird. Through cross-modal feature enhancement and adaptive fusion, the precision, robustness and real-time performance of bird recognition in a complex orchard environment are significantly improved, and a core technical support is provided for green intelligent bird repelling.
Owner:NANJING FORESTRY UNIV

Method, apparatus and system for neural network enabled hearing aid

The disclosure generally relates to a method, system and apparatus for processing audio through a neural network contained in a hearing device. In one embodiment, the disclosure relates to an apparatus to enhance incoming audio signal. The apparatus includes a controller to receive an incoming signal and provide a controller output signal; neural network engine (NNE) circuitry in communication with the controller, the NNE circuitry activatable by the controller, the NNE circuitry configured to generate an NNE output signal from the controller output signal; and digital signal processing (DSP) circuitry to receive one or more of controller output signal or the NNE circuitry output signal to thereby generate a processed signal; wherein the controller determines a processing path of the controller output signal through one of the DSP or the NNE circuitries as a function of one or more of predefined parameters, incoming signal characteristics and NNE circuitry feedback.
Owner:FORTELL RESEARCH INC

Automation for inserting a reference accessing a screen-shared file in meeting summaries and transcripts

The disclosed techniques provide a system for automatically inserting reference to a file in meeting transcripts or meeting summaries. In general, the disclosed techniques manage and enrich meeting transcripts, meeting summaries, meeting recordings, and screen-shared files during an online meeting. During an online meeting, when a presenter screenshares an application file, such as Word doc, PowerPoint, Excel, etc., a system creates meeting transcripts or a summary for the real-time discussion based on audio signals and / or chat messages relating to the shared contents of the application file. The system also determines the location of the application file and inserts a reference to the application file in the transcript or summary.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Grid telephone traffic quality inspection intelligent analysis system and method based on large language model

The invention relates to the technical field of intelligent telephone traffic quality inspection, and discloses a grid telephone traffic quality inspection intelligent analysis system and method based on a large language model, and the system comprises the steps: collecting and obtaining a multi-role call audio signal in real time, and preliminarily carrying out the speaker separation and role marking through a voice recognition module and a voiceprint recognition module; forming a preliminary role recognition result; detecting a suspected role identity error region in combination with multi-dimensional features, and when a detection result meets a preset condition, triggering a dynamic correction mechanism, generating auxiliary judgment information in combination with identity declaration keywords, business term matching and dialogue context logic inference, adopting a multi-dimensional weight decision strategy, and re-correcting a role identity tag, so as to obtain a role identity error region. Updating a role recognition result; and setting an accurate evaluation module of a dynamic correction result, feeding back and adjusting a multi-feature weight and a trigger threshold in real time, and forming an iterative optimization mechanism of dynamic correction. The method has the advantage of improving the dynamic correction capability.
Owner:XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER

Headphones (B137)

1. Name of the product of this design: Headphones (B137). 2. Purpose of this design product: used for receiving audio signals and making calls. 3. The key point of the design of this product lies in its shape. 4. The picture or photo that best illustrates the design points: three-dimensional picture.
Owner:马跃

Sound box sound effect intelligent adjustment method and system based on data analysis, and storage medium

The invention relates to the technical field of audio signal processing, and discloses a sound box and sound effect intelligent adjusting method and system based on data analysis and a storage medium, and the sound box and sound effect intelligent adjusting method based on data analysis comprises the steps: constructing a user feature modeling engine, and extracting user auditory characteristic data; constructing an environment characteristic analysis engine, and collecting environment acoustic characteristic data; constructing a content feature extraction engine, and analyzing audio content semantic data; constructing a three-dimensional fusion optimizer, inputting the three feature vectors into an auditory scene fusion model, calculating an auditory experience score through tensor fusion operation, and solving an optimal sound effect parameter by applying a multi-objective optimization algorithm; constructing a parameter generation controller, and converting the optimal sound effect parameter into a specific audio processing parameter; according to the invention, the problem of mutual interference caused by traditional separated processing is solved, and the accuracy of sound effect adjustment and the user satisfaction are improved.
Owner:SHENZHEN ZUNTE DIGITAL CO LTD