Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

206 results about "Voice activity detection" patented technology

Voice activity detection (VAD), also known as speech activity detection or speech detection, is a technique used in speech processing in which the presence or absence of human speech is detected. The main uses of VAD are in speech coding and speech recognition. It can facilitate speech processing, and can also be used to deactivate some processes during non-speech section of an audio session: it can avoid unnecessary coding/transmission of silence packets in Voice over Internet Protocol applications, saving on computation and on network bandwidth.

Intelligent sound box voice processing method and system based on artificial intelligence

The invention provides an intelligent sound box voice processing method and system based on artificial intelligence, and the method comprises the steps: obtaining audio data and mouth shape video data, carrying out the processing of the audio data and the mouth shape video data, and carrying out the multi-modal feature fusion, and obtaining a fusion feature; performing bimodal voice activity detection on the fusion features to obtain effective voice data; performing context sensing recognition of audio and video fusion on the effective voice data to obtain a first text; constructing a user feature model, and performing semantic understanding on the text based on the model to obtain an understanding result; performing intention recognition and slot filling based on the understanding result to obtain user intention and key information; generating a response strategy in combination with the user intention, the key information and the environment perception data; generating response voice according to the response strategy; and monitoring feedback information of the user to the response voice in real time, and updating the user feature model and the response strategy evaluation model based on feedback. According to the scheme, the voice can be recognized more accurately, and the safety and robustness of the system are enhanced.
Owner:SHENZHEN ZHANDIAN SMART TECH CO LTD

Voice intention recognition method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice intention recognition method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-processed voice signal, carrying out the voice activity detection processing of the voice signal, dividing the voice signal into a plurality of voice segments, analyzing semantic contents of the plurality of voice segments, determining semantic correlation information of each voice segment, analyzing sound source attributes of the plurality of voice segments, determining sound field type information of each voice segment, screening out a target voice segment from the plurality of voice segments according to the semantic correlation information and the sound field type information, and executing intention recognition processing based on the target voice segment to generate an intention recognition result. According to the invention, through a dual analysis mechanism of semantic correlation information and sound field type information, effective screening of voice segments is realized before voice recognition, non-target voice or interference segments are effectively prevented from being sent to an intention recognition model, and the accuracy of a recognition result is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Zero-configuration adaptive speaker recognition method and system

The invention discloses a zero-configuration adaptive speaker recognition method and system, and relates to the technical field of voice signal processing. The method comprises the following steps: receiving an audio stream, carrying out voice activity detection, obtaining a single-person voice segment, extracting a voiceprint embedding vector of the single-person voice segment, and carrying out online clustering to generate a speaker identity pool; and the identity pool is updated by calculating the multi-dimensional fusion similarity between the remaining single-person voice segments and the voice model, and temporary identity tags corresponding to the single-person voice segments are output. According to the speaker recognition method provided by the invention, a voiceprint template does not need to be registered in advance, and the flexibility and real-time performance of the system are improved.
Owner:北京文聿科技有限公司

Automatic delay calibration method for smart speaker system, device, and storage medium

Disclosed in the present invention are an automatic delay calibration method for a smart speaker system, a device, and a storage medium. The method comprises: initializing a smart speaker system and performing environmental testing to preliminarily configure system parameters; exciting a target sound channel and playing a target test signal; locating a first test signal on the basis of a voice activity detection (VAD) algorithm, determining the start and end of the first test signal, and identifying a time difference of the first test signal to obtain a delay of main channels; locating a time of arrival of a second test signal on the basis of a fast peak search algorithm, and identifying a delay of a subwoofer; and performing delay calibration on the subwoofer and the main channels, and dynamically adjusting delay calibration parameters of the channels on the basis of a delay calibration amount, thereby achieving accurate time alignment between the subwoofer and the main channels. The present invention solves the problem of low-frequency signal delays that are difficult to process in conventional technologies, and particularly makes a breakthrough progress in the coordination of a subwoofer and other sound channels.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Real-time translation method and interaction system based on streaming voice segmentation and semantic verification

PendingCN122050365ANatural language translationSpeech recognitionSpeech segmentationFrame based
The invention provides a real-time translation method and interaction system based on streaming voice segmentation and semantic verification. The method comprises the following steps: receiving a conference audio stream, identifying effective audio frames based on dual-channel voice activity detection, and pressing the effective audio frames into a dynamic buffer area; according to a preset dynamic segmentation strategy, outputting an initial voice segment from the dynamic buffer area for voice recognition, and obtaining a corresponding initial text segment; performing multi-stage reliability verification on the initial text fragment, and dynamically correcting or complementing the initial text fragment according to a verification result to obtain a reliable recognition result; and after the reliable identification result is obtained, triggering an asynchronous parallel translation task. According to the method, a semantic correction mechanism cooperating with the adaptive truncation depth is designed, so that the semantic fragmentation problem of the long-sequence audio during streaming truncation is solved, and the translation accuracy of a complex word order language is greatly improved while low delay is ensured.
Owner:WUHAN UNIV

Internal noise source filtering for active acoustic sensing

Techniques and devices for performing internal noise source filtering for active acoustic sensing are described. With internal noise source filtering (202), interference caused by the hearing-worn device (102) performing other operations (e.g., rendering audio content (204), performing active noise cancellation (206), or operating according to a pass-through mode (214)) may be attenuated within the received ultrasound signal to improve the sensitivity and accuracy of audio volumetric tracing. This performance improvement enables audio plethysmography to be performed when the audible wearable device (102) performs these other operations. Further, it improves the ability of audio plethysmography for speech processing, which may include speech activity detection, speech recognition, and / or dialog detection.
Owner:GOOGLE LLC

Voice activity detection device and method

A voice activity detection method includes a voice pickup module, a processing module coupled to the voice pickup module and a prompt module coupled to the processing module. The processing module is used to perform a voice activity detection method, including: receiving voice information by the voice pickup module; acquiring a volume value of the voice information by the processing module; determining by the processing module whether the volume value is less than or equal to a first volume threshold; when it is determined that the volume value is less than or equal to the first volume threshold, generating by the prompt module a first prompt message for increasing the volume value; and when it is determined that the volume value is more than the first volume threshold, generating by the prompt module a second prompt message that indicates a volume value standard has been satisfied.
Owner:GETAC TECH CORP

Air traffic control dialogue type speech recognition method and device based on multi-task learning

The invention provides an air traffic control dialogue type speech recognition method and device based on multi-task learning. Cooperative optimization of voice segmentation and content recognition is realized by constructing a joint learning model of a voice feature representation learning module and a voice activity detection module. The method comprises the following steps: firstly, designing a segmented shielding comparison pre-training strategy, and constructing a supervision signal through a product quantization codebook, so that a speech feature representation learning module learns fine-grained acoustic features in unlabeled data; secondly, a double-branch joint training mode fine tuning model is designed, a main branch optimizes a voice recognition task through connection time sequence classification loss, an auxiliary branch improves voice segment boundary detection precision through binary cross entropy loss, and meanwhile, a dynamic task weight adjustment strategy is combined to balance a multi-target optimization direction; and finally, innovatively introducing a dynamic local window attention mechanism, and focusing context modeling of an effective voice segment according to a real-time segmentation index.
Owner:SICHUAN UNIV

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

Method and apparatus for training audio processing model, storage medium, and electronic device

PCT designated stage expiredWO2025152852A1Speech recognitionEngineeringSpeech sound
The present disclosure relates to the technical field of artificial intelligence, and provides a method for training an audio processing model, an apparatus for training an audio processing model, a computer storage medium, and an electronic device. The method for training an audio processing model comprises: acquiring a training sample set; using a first sample set to pre-train a first branch network of an audio processing model to be trained, so as to obtain a pre-trained first branch network, and using a second sample set to pre-train a second branch network of said audio processing model, so as to obtain a pre-trained second branch network; and using the training sample set to jointly train the pre-trained first branch network and the pre-trained second branch network, so as to obtain a trained audio processing model, wherein the first branch network is used for executing echo cancellation and speech enhancement tasks, and the second branch network is used for executing a voice activity detection task. In the present disclosure, multiple audio processing tasks can be executed by means of one model, thereby reducing system power.
Owner:JINGDONG CITY BEIJING DIGITS TECH CO LTD +1

Conference summary generation method and device

The invention provides a conference summary generation method and device, and the method comprises the steps: obtaining an audio and video file of a target conference, and the audio and video file comprises an audio file and / or a video file; performing voice activity detection processing, text conversion processing and speaker recognition processing on the audio and video file to obtain a target dialogue text; constructing prompt words based on the application scene of the target conference and the target dialogue text; and inputting the cue word into a conference summary generation large model, and generating a conference summary corresponding to the target conference. According to the invention, the problem of low conference summary generation efficiency is solved, and more time is saved; the generation result of the conference summary generation large model is optimized by adopting a prompt word engineering mode, so that the accuracy of generating the conference summary can be further improved; the conference summary is generated by adopting the conference summary generation large model, compared with a traditional neural language programming technology, feature engineering construction is omitted, and the performance is better.
Owner:SANY HEAVY MACHINERY

Real-time multilingual transcription system and method

Disclosed are a method, system, and apparatus of a real-time multilingual transcription system and method. In one embodiment, a method includes continuously capturing an audio data and segment it into short segments; implementing a pre-trained enterprise-grade voice activity detection (“VAD”) system on each of the short segmental and filtering out non-speech segments to reduce computational waste, focusing resources on relevant audio data and minimizing latency. If speech is detected, a particular short segment is added to a processing queue. If speech is not detected, declining to add the particular segment to the processing queue, thereby reducing unnecessary processing.
Owner:GOVERNMENTGPT INC

Cough sound recognition method based on PSO-GBDT-LR model

The invention discloses a cough sound recognition method based on a PSO-GBDT-LR model, and belongs to the technical field of signal recognition. The method includes acquiring audio signals; de-noising the audio signal by using a Berouti spectral subtraction method to obtain a de-noised audio signal; the audio event detection VAD is used for segmenting the part, where the sound appears, of the audio; 7-dimensional time domain features are extracted from each segmented audio sample; performing short-time Fourier transform (STFT) on each segmented audio sample, and extracting two-dimensional frequency domain features from a frequency spectrum; combining the extracted 7-dimensional time domain features and the extracted 2-dimensional frequency domain features to form a 9-dimensional feature vector combination; and marking the feature vector of the cough audio sample as a class 1, and marking the feature vector of the non-cough audio sample as a class 0. According to the method, the problems of excessive noise features and abnormal features of the sound in the cough sound recognition process are solved, the features of the cough and non-cough sound can be accurately distinguished, and the generalization ability is high.
Owner:KUNMING UNIV OF SCI & TECH

System and method for efficiency among devices

A wearable multifunction device or earpiece or a pair of earpieces includes one or more processors, at least one microphone coupled to the one or more processors, a biometric sensor coupled to the one or more processors, and a memory coupled to the one or more processors, the memory having computer instructions causing the one or more processors to perform the operations of sensing a remaining battery life and based on the sensing, prioritizing one or more of the functions of always on recording, biometric measuring, biometric recording, sound pressure level measuring, voice activity detection, key word detection, key word analysis, personal audio assistant functions, transmission of data to a tethered phone, transmission of data to a server, transmission of data to a cloud device.
Owner:THE DIABLO CANYON COLLECTIVE LLC

End-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method

The invention discloses an end-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method, and relates to the technical field of voice interaction control, and the method comprises the steps: an end side collects microphone audio, obtains loudspeaker reference audio, judges a broadcast state, and switches a voice activity detection strategy; end-side framing detection is carried out, voice segment events are output through a three-state state machine, characteristics such as duration, relative volume and broadcast overlapping degree are extracted, and candidate interruption events are generated in combination with an automatic voice recognition intermediate text; the end side coarse screening reports a gray area event to the cloud side semantic research and judgment, and pre-suppression, buffer foresight contraction and block buffer are executed during waiting; and the cloud side issues an instruction containing a period of validity, a deduplication key and an interruption type, and the execution side stops generating, synthesizing and playing and empties buffer so as to play a completion receipt, submit a dialogue history and roll back unplayed content, thereby reducing false triggering and missed triggering, shortening interruption delay, reducing residual broadcast and inhibiting context drift.
Owner:SUZHOU MENGWU INTELLIGENT TECHNOLOGY CO LTD

Voice activity detection method and system, and voice enhancement method and system

A voice activity detection method and system and a voice enhancement method and system are provided. A voice presence probability of a target voice signal present in microphone signals may be determined by calculating a linear correlation between a signal subspace where the microphone signals are located and a target subspace where the target voice signal is located. The voice enhancement method and system may be used to calculate filter coefficients based on the voice presence probability, so as to perform voice enhancement on the microphone signals. The calculation accuracy of the voice presence probability is improved, and the voice enhancement effect is also improved.
Owner:SHENZHEN SHOKZ CO LTD

AI glasses automatic shooting method and system based on voice control

The invention provides an AI glasses automatic shooting method and system based on voice control, and relates to the technical field of intelligent wearable equipment, accurate interaction is realized through a multi-channel directional microphone array and an end-cloud collaborative voice recognition engine, the microphone array optimizes the pickup angle based on the wearing position characteristics of a user, and the user experience is improved. In combination with real-time voice activity detection, environmental noise is filtered out, it is ensured that a clear voice instruction can still be captured in a noisy environment, an end-cloud cooperation mode operates a lightweight model locally to guarantee the off-line response speed, a cloud large model is called when a network is available to improve the complex instruction analysis capability, and a composite statement containing parameter adjustment can be recognized; and an operation intention and parameters are automatically bound through a natural language processing engine, so that one-step execution of the instruction is realized.
Owner:MIODAO CLOUD COMPUTING (HANGZHOU) CO LTD

Method and system for processing remote active speech during a call

A method performed by a first device, which includes performing an audio call with a second device by transmitting a microphone signal as an uplink signal and receiving a downlink signal for driving a first speaker and while performing the audio call, performing a joint media playback session in which both devices independently stream a piece of media content for synchronous playback such that both devices receive an audio signal of the piece of media content for driving respective speakers at the same time, determining that a voice activity detection (VAD) signal indicates that the downlink signal includes speech, in response to determining that the VAD signal indicates that the downlink signal includes speech, processing the audio signal of the piece of media content by applying a scalar gain, and driving the first speaker with a mix of the downlink signal and the audio signal.
Owner:APPLE INC

Voice interaction control method and device, safety helmet, storage medium and program product

The invention discloses a voice interaction control method and device, a safety helmet, a storage medium and a program product, and relates to the technical field of voice interaction. According to the intelligent safety helmet, a throat vibration sensor and a lip movement detection camera are integrated on the intelligent safety helmet, a throat vibration signal collected by the throat vibration sensor and a lip movement signal collected by the lip movement detection camera are obtained respectively, whether a wearer generates throat vibration or not is determined based on the throat vibration signal, and whether the wearer generates throat vibration or not is determined based on the throat vibration signal. And determining whether the wearer generates lip vibration or not based on the lip movement signal, and performing voice activity detection on the wearer when the wearer generates throat vibration and lip vibration. Therefore, according to the added throat vibration and lip vibration detection, the accuracy of the voice activity detection function integrated on the intelligent safety helmet is improved, and the probability of semantic misjudgment is reduced.
Owner:SHENZHEN JURUIYUN TECHNOLOGYCO LTD

Teaching interaction quality evaluation method and system based on large language model

The invention relates to the field of teaching interaction quality evaluation, in particular to a teaching interaction quality evaluation method and system based on a large language model. The method comprises the following steps: audio transcription: converting classroom audio into an original transcription text through voice activity detection, speaker classification, automatic voice recognition and punctuation recovery; transcriptional refining: performing context-based text error correction on the original transcriptional text by using a large language model in combination with a preschool education field knowledge base to generate a refined transcriptional text; a quality evaluation step: based on a preschool education quality evaluation scale, using few sample example guidance and thinking chain reasoning for each scoring point, judging whether a voice segment conforming to the scoring point exists in the refined transcriptional text, performing binary scoring, and determining whether the voice segment conforms to the scoring point; and generating an interactive quality evaluation report containing the standard-reaching rate of each evaluation dimension, teaching bright spot analysis and staged teaching optimization suggestions. The evaluation efficiency is remarkably improved. The method is suitable for teaching interaction quality evaluation.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN)

Desktop robot microphone array sound source positioning system and method

The invention relates to the technical field of robots, and particularly discloses a desktop robot microphone array sound source positioning system and method. Comprising a main control chip, acquisition microphones and a power supply module, an ESP32 main control board is connected with a circular array formed by the four acquisition microphones through double I2S bus pins, the array diameter of the four acquisition microphones is 50 mm, the included angle between every two adjacent acquisition microphones is 90 degrees, the acquisition microphones are grounded through GND lines, and the power supply module is connected with the ESP32 main control board. And the ESP32 main control board provides 3.3 V voltage for the acquisition microphone through a power supply line. A four-microphone circular array with the diameter of 50 mm and ESP32 double I2S synchronous acquisition are adopted, 360-degree full horizontal positioning is achieved, the hardware size is small, the miniaturization requirement of a desktop accompanying robot is met, a VAD voice activity detection module is introduced, algorithm operation is triggered only when human voice is detected, power consumption in a dormant state is small, and the endurance of the robot is prolonged.
Owner:HANGZHOU XINGMENGDAO TECHNOLOGY CO LTD

Multi-person voice processing system and separation method based on voiceprint recognition and clustering

The invention discloses a multi-person voice processing system and separation method based on voiceprint recognition and clustering, and belongs to the technical field of multi-person voice separation processing, and the system comprises a voice activity detection unit, an audio cache unit, a voiceprint processing unit and a data processing unit. According to the method, voice activity detection, voiceprint technology, speaker separation technology and caching technology are fused, voice activity detection is performed on a section of voice, non-human voice such as silence, noise and music in the audio is removed, only human voice is reserved, the caching technology is used for caching the audio data, and when the cached data reach a preset threshold value, the voice data are cached. The method comprises the following steps: caching audio data, segmenting the cached audio data according to a certain frame shift and frame length, then extracting speaker features in sub-audio segments by using a voiceprint recognition technology, finally clustering the features, recognizing the audio segment of each person, and outputting corresponding time information.
Owner:HEFEI KEXIN ZHILIAN TECHNOLOGY CO LTD

Speech recognition method and device and electronic equipment

The invention provides a voice recognition method and device and electronic equipment, and the method comprises the steps: obtaining a to-be-recognized original audio, carrying out the voice activity detection of the original audio, and determining each segmented first audio segment and each second audio segment of the original audio; the first audio segment and the second audio segment are obtained through segmentation based on different detection results of whether voice is included; extracting time-frequency domain features corresponding to each first audio segment and each second audio segment; inputting the time-frequency domain features into a speech recognition model, and outputting a target text corresponding to the original audio; the speech recognition model has the capability of distinguishing speech segments from non-speech segments after being trained.
Owner:SMARTER SILICON (SHANGHAI) TECH CO LTD

Multimedia conference room sound system based on artificial intelligence

The invention relates to the technical field of conference room sound equipment control, in particular to a multimedia conference room sound equipment system based on artificial intelligence, which acquires image positions and head postures of participants in real time through a personnel positioning module, automatically establishes a spatial mapping relation between the personnel and channels in combination with microphone layout, and realizes dynamic binding of microphone channels. A main speaking channel is automatically recognized based on voice activity detection of audio data, differentiated control is performed on channel gains through a sound output module, and background noise of a non-speaking microphone is effectively suppressed. The behavior characteristics and the interaction intention of a main speaker are extracted, joint modeling is carried out on the behavior characteristics and the interaction intention based on a reinforcement learning model, prediction of a next speaker and dynamic updating of a main channel are achieved, and strategy parameters are continuously optimized based on speech feedback. Manual operation is reduced, and voice output definition and natural fluency of conference interaction are improved.
Owner:GUANGZHOU QINSHENG ELECTRIC CO LTD

Real-time communication audio equipment detection method based on multiple threads

The invention discloses a multi-thread-based real-time communication audio equipment detection method, which realizes comprehensive coverage of various system environments and driver versions by parallelly starting a plurality of threads corresponding to bottom audio acquisition interfaces APIs of different operating systems, and introduces a mechanism for automatically adjusting the sampling rate, the sound channel number and the sampling format. The problem of silence or abnormal noise caused by mismatching of sampling parameters is effectively solved, and the effectiveness of audio data acquisition is improved. Meanwhile, a neural network voice activity detection model and a comprehensive scoring strategy of time domain statistical indexes are combined, so that the system can accurately recognize effective voice in a complex noise environment, and misjudgment and missed judgment of traditional time domain threshold detection are avoided. The dynamic incremental detection duration strategy balances the detection speed and accuracy, ensures that a user can quickly obtain a preliminary detection result, performs full analysis for abnormal conditions, and improves the detection robustness.
Owner:BEIJING ANXIN ZHITONG TECH CO LTD

Voice emotion recognition method and device, computer equipment and storage medium

The invention belongs to the technical field of artificial intelligence, and relates to a voice emotion recognition method and device, computer equipment and a storage medium, and the method comprises the steps: collecting the voice data of a customer in a call process between a customer service and the customer; caching the voice data into a buffer area; segmenting the voice data in the buffer area based on a target voice activity detection algorithm to obtain voice segments; performing feature extraction on the voice segments based on the extraction model to obtain voice feature vectors; obtaining customer call data in a call process, and extracting a steady-state emotion vector from the customer call data; performing feature fusion on the voice feature vector and the steady-state emotion vector to obtain a target feature vector; and performing emotion recognition on the target feature vector based on an emotion recognition model to generate an emotion recognition result of the voice segment. In addition, the emotion recognition result can be stored in the block chain. The method can be applied to voice emotion recognition scenes in the financial field, and the accuracy and timeliness of emotion recognition are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Family life event atlas construction method and system based on cooperation of multiple intelligent calculation engines

The invention discloses a family life event atlas construction method and system based on cooperation of multiple intelligent computing engines, and relates to the technical field of knowledge atlas construction.The method comprises the steps that a family environment continuous audio stream is collected through an end side, effective fragments are screened through voice activity detection, and original audio is destroyed after acoustic feature vectors are extracted; and then a plurality of parallel intelligent calculation engines respectively recognize non-language environment sounds, emotion and voiceprint identities, semantics and potential intentions, a standardized family event object is generated through multi-modal fusion, the standardized family event object is stored in an independent research and development data warehouse of a time sequence + graph mixed framework, a family life knowledge graph is constructed, and finally graph changes are monitored through an inference engine, so that the family life knowledge graph is obtained. And actively triggering home equipment adjustment or pushing suggestions. According to the invention, active perception and semantic understanding of family events are realized, the risk of privacy disclosure is avoided, and the problems of passive response, data islands and privacy safety hidden troubles of the existing smart home are solved.
Owner:BEIJING CLOUDWAVE TIMES TECH CO LTD

Streaming video subtitle generation method and system based on localized large model, and storage medium

The invention discloses a streaming video subtitle generation method and system based on a localized large model, and a storage medium, and belongs to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: constructing an asynchronous streaming audio extraction queue, and realizing parallel processing of millisecond-level response and background fragmentation of a first audio segment of a video; a decoupled voice activity detection model is adopted to carry out refined slicing on the audio stream, and noise filtering and muting are carried out; text decoding is carried out through a speech recognition model subjected to fine tuning of synthetic data, and term enhancement in professional fields is supported; carrying out multi-modal fusion and time sequence calibration on an identification result to realize sound and picture synchronization and text standardization; and finally, rendering subtitles through a real-time callback mechanism and a cross-platform interface. The method solves the problems of high delay of long video subtitle generation, time axis drift and low terminology recognition rate, is especially suitable for Linux / Windows cross-platform localization deployment, and has high privacy security and low-cost field adaptive capability.
Owner:WUHAN UNIV OF SCI & TECH

BCM (Body Control Module) cooperative control method based on voice instruction recognition in vehicle-mounted high-noise environment

The invention relates to the technical field of artificial intelligence, and discloses a BCM module cooperative control method based on voice instruction recognition in a vehicle-mounted high-noise environment, and the method comprises the following steps: collecting original voice instruction data and vehicle state parameter data in the vehicle-mounted high-noise environment; performing voice signal preprocessing in a noise environment based on the original voice instruction data to generate de-noised voice instruction data; according to the method, the engine noise and the wind noise steady-state background noise are filtered out through multi-level noise suppression processing combining adaptive filtering and spectral subtraction, residual noise elimination is carried out for sudden impact noise, and the definition of voice signals is improved. Meanwhile, through a voice activity detection algorithm, a detection threshold is dynamically adjusted according to vehicle state parameters, an effective voice segment and a noise segment can be separated, the accuracy of voice feature extraction in a complex time-varying noise environment is ensured, and thus the robustness and reliability of voice instruction recognition are improved.
Owner:XIAMEN FAJOINT-IOT TECH CO LTD

Audio processing method and related device

The invention discloses an audio processing method and a related device. The method comprises the following steps: performing voice activity detection on an obtained first audio clip; if the first audio clip is non-voice data, determining a first continuous non-voice duration according to the duration of the first audio clip; if the first continuous non-voice duration is greater than or equal to the first dynamic duration threshold, acquiring a first text set; generating a first intermediate text according to the first text set, and adding the first intermediate text into the first sentence segmentation result set; generating a first to-be-detected text according to each intermediate text included in the first sentence segmentation result set; inputting the first to-be-detected text and the first interaction text into a trained semantic sentence segmentation model to obtain an output result; and if the output result represents that the semantics of the first to-be-detected text is complete, taking the first to-be-detected text as a first sentence segmentation result, and outputting the first sentence segmentation result. According to the invention, the flexibility and accuracy of voice sentence segmentation of the electronic equipment can be improved.
Owner:ZHAOLIAN CONSUMER FINANCE CO LTD