Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

154 results about "Voice activity detection" patented technology

Voice activity detection (VAD), also known as speech activity detection or speech detection, is a technique used in speech processing in which the presence or absence of human speech is detected. The main uses of VAD are in speech coding and speech recognition. It can facilitate speech processing, and can also be used to deactivate some processes during non-speech section of an audio session: it can avoid unnecessary coding/transmission of silence packets in Voice over Internet Protocol applications, saving on computation and on network bandwidth.

Automatic delay calibration method for smart speaker system, device, and storage medium

PCT designated stageWO2026066023A1Signal processingTransducer circuitsSignal onTime alignment
Disclosed in the present invention are an automatic delay calibration method for a smart speaker system, a device, and a storage medium. The method comprises: initializing a smart speaker system and performing environmental testing to preliminarily configure system parameters; exciting a target sound channel and playing a target test signal; locating a first test signal on the basis of a voice activity detection (VAD) algorithm, determining the start and end of the first test signal, and identifying a time difference of the first test signal to obtain a delay of main channels; locating a time of arrival of a second test signal on the basis of a fast peak search algorithm, and identifying a delay of a subwoofer; and performing delay calibration on the subwoofer and the main channels, and dynamically adjusting delay calibration parameters of the channels on the basis of a delay calibration amount, thereby achieving accurate time alignment between the subwoofer and the main channels. The present invention solves the problem of low-frequency signal delays that are difficult to process in conventional technologies, and particularly makes a breakthrough progress in the coordination of a subwoofer and other sound channels.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Real-time translation method and interaction system based on streaming voice segmentation and semantic verification

PendingCN122050365ANatural language translationSpeech recognitionSpeech segmentationFrame based
The invention provides a real-time translation method and interaction system based on streaming voice segmentation and semantic verification. The method comprises the following steps: receiving a conference audio stream, identifying effective audio frames based on dual-channel voice activity detection, and pressing the effective audio frames into a dynamic buffer area; according to a preset dynamic segmentation strategy, outputting an initial voice segment from the dynamic buffer area for voice recognition, and obtaining a corresponding initial text segment; performing multi-stage reliability verification on the initial text fragment, and dynamically correcting or complementing the initial text fragment according to a verification result to obtain a reliable recognition result; and after the reliable identification result is obtained, triggering an asynchronous parallel translation task. According to the method, a semantic correction mechanism cooperating with the adaptive truncation depth is designed, so that the semantic fragmentation problem of the long-sequence audio during streaming truncation is solved, and the translation accuracy of a complex word order language is greatly improved while low delay is ensured.
Owner:WUHAN UNIV

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

System and method for efficiency among devices

A wearable multifunction device or earpiece or a pair of earpieces includes one or more processors, at least one microphone coupled to the one or more processors, a biometric sensor coupled to the one or more processors, and a memory coupled to the one or more processors, the memory having computer instructions causing the one or more processors to perform the operations of sensing a remaining battery life and based on the sensing, prioritizing one or more of the functions of always on recording, biometric measuring, biometric recording, sound pressure level measuring, voice activity detection, key word detection, key word analysis, personal audio assistant functions, transmission of data to a tethered phone, transmission of data to a server, transmission of data to a cloud device.
Owner:THE DIABLO CANYON COLLECTIVE LLC

End-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method

The invention discloses an end-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method, and relates to the technical field of voice interaction control, and the method comprises the steps: an end side collects microphone audio, obtains loudspeaker reference audio, judges a broadcast state, and switches a voice activity detection strategy; end-side framing detection is carried out, voice segment events are output through a three-state state machine, characteristics such as duration, relative volume and broadcast overlapping degree are extracted, and candidate interruption events are generated in combination with an automatic voice recognition intermediate text; the end side coarse screening reports a gray area event to the cloud side semantic research and judgment, and pre-suppression, buffer foresight contraction and block buffer are executed during waiting; and the cloud side issues an instruction containing a period of validity, a deduplication key and an interruption type, and the execution side stops generating, synthesizing and playing and empties buffer so as to play a completion receipt, submit a dialogue history and roll back unplayed content, thereby reducing false triggering and missed triggering, shortening interruption delay, reducing residual broadcast and inhibiting context drift.
Owner:SUZHOU MENGWU INTELLIGENT TECHNOLOGY CO LTD

Voice activity detection method and system, and voice enhancement method and system

A voice activity detection method and system and a voice enhancement method and system are provided. A voice presence probability of a target voice signal present in microphone signals may be determined by calculating a linear correlation between a signal subspace where the microphone signals are located and a target subspace where the target voice signal is located. The voice enhancement method and system may be used to calculate filter coefficients based on the voice presence probability, so as to perform voice enhancement on the microphone signals. The calculation accuracy of the voice presence probability is improved, and the voice enhancement effect is also improved.
Owner:SHENZHEN SHOKZ CO LTD

AI glasses automatic shooting method and system based on voice control

The invention provides an AI glasses automatic shooting method and system based on voice control, and relates to the technical field of intelligent wearable equipment, accurate interaction is realized through a multi-channel directional microphone array and an end-cloud collaborative voice recognition engine, the microphone array optimizes the pickup angle based on the wearing position characteristics of a user, and the user experience is improved. In combination with real-time voice activity detection, environmental noise is filtered out, it is ensured that a clear voice instruction can still be captured in a noisy environment, an end-cloud cooperation mode operates a lightweight model locally to guarantee the off-line response speed, a cloud large model is called when a network is available to improve the complex instruction analysis capability, and a composite statement containing parameter adjustment can be recognized; and an operation intention and parameters are automatically bound through a natural language processing engine, so that one-step execution of the instruction is realized.
Owner:MIODAO CLOUD COMPUTING (HANGZHOU) CO LTD

Teaching interaction quality evaluation method and system based on large language model

The invention relates to the field of teaching interaction quality evaluation, in particular to a teaching interaction quality evaluation method and system based on a large language model. The method comprises the following steps: audio transcription: converting classroom audio into an original transcription text through voice activity detection, speaker classification, automatic voice recognition and punctuation recovery; transcriptional refining: performing context-based text error correction on the original transcriptional text by using a large language model in combination with a preschool education field knowledge base to generate a refined transcriptional text; a quality evaluation step: based on a preschool education quality evaluation scale, using few sample example guidance and thinking chain reasoning for each scoring point, judging whether a voice segment conforming to the scoring point exists in the refined transcriptional text, performing binary scoring, and determining whether the voice segment conforms to the scoring point; and generating an interactive quality evaluation report containing the standard-reaching rate of each evaluation dimension, teaching bright spot analysis and staged teaching optimization suggestions. The evaluation efficiency is remarkably improved. The method is suitable for teaching interaction quality evaluation.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN)

Desktop robot microphone array sound source positioning system and method

The invention relates to the technical field of robots, and particularly discloses a desktop robot microphone array sound source positioning system and method. Comprising a main control chip, acquisition microphones and a power supply module, an ESP32 main control board is connected with a circular array formed by the four acquisition microphones through double I2S bus pins, the array diameter of the four acquisition microphones is 50 mm, the included angle between every two adjacent acquisition microphones is 90 degrees, the acquisition microphones are grounded through GND lines, and the power supply module is connected with the ESP32 main control board. And the ESP32 main control board provides 3.3 V voltage for the acquisition microphone through a power supply line. A four-microphone circular array with the diameter of 50 mm and ESP32 double I2S synchronous acquisition are adopted, 360-degree full horizontal positioning is achieved, the hardware size is small, the miniaturization requirement of a desktop accompanying robot is met, a VAD voice activity detection module is introduced, algorithm operation is triggered only when human voice is detected, power consumption in a dormant state is small, and the endurance of the robot is prolonged.
Owner:HANGZHOU XINGMENGDAO TECHNOLOGY CO LTD

Family life event atlas construction method and system based on cooperation of multiple intelligent calculation engines

The invention discloses a family life event atlas construction method and system based on cooperation of multiple intelligent computing engines, and relates to the technical field of knowledge atlas construction.The method comprises the steps that a family environment continuous audio stream is collected through an end side, effective fragments are screened through voice activity detection, and original audio is destroyed after acoustic feature vectors are extracted; and then a plurality of parallel intelligent calculation engines respectively recognize non-language environment sounds, emotion and voiceprint identities, semantics and potential intentions, a standardized family event object is generated through multi-modal fusion, the standardized family event object is stored in an independent research and development data warehouse of a time sequence + graph mixed framework, a family life knowledge graph is constructed, and finally graph changes are monitored through an inference engine, so that the family life knowledge graph is obtained. And actively triggering home equipment adjustment or pushing suggestions. According to the invention, active perception and semantic understanding of family events are realized, the risk of privacy disclosure is avoided, and the problems of passive response, data islands and privacy safety hidden troubles of the existing smart home are solved.
Owner:BEIJING CLOUDWAVE TIMES TECH CO LTD

Streaming video subtitle generation method and system based on localized large model, and storage medium

The invention discloses a streaming video subtitle generation method and system based on a localized large model, and a storage medium, and belongs to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: constructing an asynchronous streaming audio extraction queue, and realizing parallel processing of millisecond-level response and background fragmentation of a first audio segment of a video; a decoupled voice activity detection model is adopted to carry out refined slicing on the audio stream, and noise filtering and muting are carried out; text decoding is carried out through a speech recognition model subjected to fine tuning of synthetic data, and term enhancement in professional fields is supported; carrying out multi-modal fusion and time sequence calibration on an identification result to realize sound and picture synchronization and text standardization; and finally, rendering subtitles through a real-time callback mechanism and a cross-platform interface. The method solves the problems of high delay of long video subtitle generation, time axis drift and low terminology recognition rate, is especially suitable for Linux / Windows cross-platform localization deployment, and has high privacy security and low-cost field adaptive capability.
Owner:WUHAN UNIV OF SCI & TECH

BCM (Body Control Module) cooperative control method based on voice instruction recognition in vehicle-mounted high-noise environment

The invention relates to the technical field of artificial intelligence, and discloses a BCM module cooperative control method based on voice instruction recognition in a vehicle-mounted high-noise environment, and the method comprises the following steps: collecting original voice instruction data and vehicle state parameter data in the vehicle-mounted high-noise environment; performing voice signal preprocessing in a noise environment based on the original voice instruction data to generate de-noised voice instruction data; according to the method, the engine noise and the wind noise steady-state background noise are filtered out through multi-level noise suppression processing combining adaptive filtering and spectral subtraction, residual noise elimination is carried out for sudden impact noise, and the definition of voice signals is improved. Meanwhile, through a voice activity detection algorithm, a detection threshold is dynamically adjusted according to vehicle state parameters, an effective voice segment and a noise segment can be separated, the accuracy of voice feature extraction in a complex time-varying noise environment is ensured, and thus the robustness and reliability of voice instruction recognition are improved.
Owner:XIAMEN FAJOINT-IOT TECH CO LTD

Audio processing method and related device

The invention discloses an audio processing method and a related device. The method comprises the following steps: performing voice activity detection on an obtained first audio clip; if the first audio clip is non-voice data, determining a first continuous non-voice duration according to the duration of the first audio clip; if the first continuous non-voice duration is greater than or equal to the first dynamic duration threshold, acquiring a first text set; generating a first intermediate text according to the first text set, and adding the first intermediate text into the first sentence segmentation result set; generating a first to-be-detected text according to each intermediate text included in the first sentence segmentation result set; inputting the first to-be-detected text and the first interaction text into a trained semantic sentence segmentation model to obtain an output result; and if the output result represents that the semantics of the first to-be-detected text is complete, taking the first to-be-detected text as a first sentence segmentation result, and outputting the first sentence segmentation result. According to the invention, the flexibility and accuracy of voice sentence segmentation of the electronic equipment can be improved.
Owner:ZHAOLIAN CONSUMER FINANCE CO LTD

A baby crying detection method, apparatus and electronic device

The application discloses a baby crying detection method and device and electronic equipment. The method comprises the following steps: calculating the spectrum, fundamental frequency and voice activity detection result of a new frame of voice signal each time a new frame of voice signal is acquired; updating a pre-established matrix set according to the spectrum, fundamental frequency and voice activity detection result of the new frame of voice signal according to the first-in first-out principle; calculating the first variance of the effective fundamental frequency elements in the fundamental frequency matrix after updating the matrix set each time and extracting all blocks composed of preset detection values from the detection result matrix after singular point smoothing processing when the spectrum matrix is full of columns; counting the number of elements of each block and calculating the second variance of all element numbers; judging whether there is baby crying in the voice signal according to the second variance and a preset second variance threshold; and continuing to update the matrix set when it is determined that there is no baby crying. The application can improve the accuracy of baby crying detection.
Owner:TP-LINK INT SHENZHEN CO LTD

Voice AI interaction state management method for low-power-consumption terminal

The invention discloses a voice AI interaction state management method for a low-power-consumption terminal, and aims to solve the problems of unbalanced energy consumption and response speed, lack of interaction content and high hardware cost of voice AI interaction of the low-power-consumption terminal in the prior art. According to the method, a'dormancy-standby-wakeup-interaction 'multi-stage state machine closed-loop management system is constructed, and a dynamic power consumption adjustment algorithm and a'local wakeup-cloud reasoning' mixed architecture are combined. Rapid triggering and intention confirmation are realized locally through hardware-level wake-up word detection and voice activity detection, an LTE module is activated only after effective wake-up to establish cloud connection, and a cloud pre-trained language large model supports flexible semantic analysis and interaction content generation. Meanwhile, a module on-demand start-stop strategy is adopted, power supply of a redundant module is cut off in an inactive stage, a communication link maintains connection with a low-volume heartbeat packet, and the communication link rapidly falls back to a low-power-consumption state after interaction. According to the invention, collaborative optimization of low power consumption and AI interactive experience is realized.
Owner:LIERDA SCI & TECH GRP

BCM module cooperative control method based on voice instruction recognition in vehicle-mounted high-noise environment

The application relates to the technical field of artificial intelligence, and discloses a BCM module cooperative control method based on voice instruction recognition in a vehicle high-noise environment, which comprises the following steps: collecting original voice instruction data and vehicle state parameter data in the vehicle high-noise environment; performing voice signal preprocessing in the noise environment based on the original voice instruction data to generate denoised voice instruction data; through multi-level noise suppression processing combining adaptive filtering and spectral subtraction, engine noise, wind noise and steady-state background noise are filtered out, and residual noise is eliminated for sudden impact noise, so that the intelligibility of the voice signal is improved. Meanwhile, through a voice activity detection algorithm, and by dynamically adjusting the detection threshold according to the vehicle state parameters, effective voice segments and noise segments can be separated, the accuracy of voice feature extraction in a complex time-varying noise environment is ensured, and the robustness and reliability of the voice instruction recognition are improved.
Owner:XIAMEN FAJOINT-IOT TECH CO LTD

Voice activity detection method, system, terminal and storage medium

The application relates to the technical field of audio processing, and discloses a voice activity detection method, a system, a terminal and a storage medium. The method comprises the following steps: continuously monitoring an input audio signal through a first detection layer, performing preliminary voice judgment on each frame of audio signal based on basic acoustic characteristics, and generating a wake-up signal to start a second detection layer when the judgment result meets a first preset trigger condition; the second detection layer extracts at least two acoustic characteristics reflecting voice harmonic structure and time-frequency distribution characteristics of the audio signal in response to the wake-up signal; all acoustic characteristics reflecting voice harmonic structure and time-frequency distribution characteristics are subjected to fusion analysis, and when the fusion analysis result meets a second preset trigger condition, it is determined that the current audio signal is a candidate audio signal, and an activation signal is generated to start a third detection layer; the third detection layer inputs the candidate audio signal into a preset deep learning model after processing the candidate audio signal in response to the activation signal, so as to output a voice activity detection result.
Owner:SHENZHEN AI LING TECHNOLOGY CO LTD

A distributed speech enhancement system based on maximum likelihood

ActiveCN116524943Bincrease diversityGood noise cancellation performanceSpeech analysisData compressionNoise
The present application belongs to the technical field of distributed speech enhancement, and particularly relates to a distributed speech enhancement system based on maximum likelihood. In order to expand the diversity of speech enhancement technology in WASN and complete good noise elimination performance, the system comprises a discrete Fourier transform module, a speech activity detection module, a steering vector estimation module, a data compression module, a result output module, a signal construction module, a weighted correlation matrix estimation module, a filter update module, and a discrete inverse Fourier transform module. The present application is a distributed speech enhancement technology which can be applied to a wireless acoustic sensor network without a data processing center. The technology estimates a weighted correlation matrix through a local signal constructed by a node and a variance of an output result, and updates a filter by combining the estimated weighted correlation matrix with a constructed local steering vector, so as to complete distributed speech enhancement.
Owner:ZHONGBEI UNIV

Multimedia conference room sound system based on artificial intelligence

The present application relates to conference room sound control technical field, especially in kind based on artificial intelligence's multimedia conference room sound system, through personnel positioning module real-time acquisition of the image position and head posture of the participant, the automatic establishment of the space mapping relationship of personnel and channel is combined with microphone layout, the dynamic binding of microphone channel is realized. Voice activity detection based on audio data automatically identifies the main speaking channel, and differentiates the control of the channel gain through the sound output module, effectively suppresses the background noise of the non-speaking microphone. By extracting the behavior characteristics and interaction intention of the main speaker, the behavior characteristics and interaction intention are jointly modeled based on the reinforcement learning model, the prediction of the next speaker and the dynamic update of the main channel are realized, and the strategy parameters are continuously optimized based on the speech feedback. Reduce manual operation, improve the clarity of voice output and the natural fluency of conference interaction.
Owner:GUANGDONG RUIZHAO AUDIO EQUIPMENT CO LTD

Sound replication method and related apparatus

The application provides a sound replication method and related device, and relates to the technical field of sound processing. The reference audio is subjected to audio verification to obtain first audio, the first audio is subjected to a speech enhancement operation to achieve the purposes of noise reduction, dereverberation and improvement of the signal-to-noise ratio of the audio, thereby obtaining second audio, the second audio is subjected to a speech activity detection and segment division operation to obtain a candidate speech segment, a target speech segment meeting the sound replication requirement is selected from the candidate speech segment, the optimal segment with clear timbre, high signal-to-noise ratio and stable pronunciation is selected from the candidate speech segment for timbre embedding extraction and speech generation, the probability of timbre deviation of the generated speech is reduced, the accuracy of the speech generation is improved, and the user experience is improved.
Owner:BEIJING SOHU NEW MEDIA INFORMATION TECH

Intelligent audio optimization method

PendingCN121545535ASpeech analysisTransmissionComfort noiseVoice communication
The invention relates to an intelligent audio optimization method, and the method comprises the steps: inputting a voice signal and a network state parameter into a deep learning model for conjoint analysis, and generating a corresponding environment evaluation result; dynamically generating a voice activity detection threshold, and detecting the voice signal by using the voice activity detection threshold to determine whether the user is in a voice state; and if the user is not in the voice state, calling network state information to execute discontinuous transmission processing, and if the user is in the voice state and a voice gap exists in the voice signal, executing comfortable noise supplement processing, and outputting a target communication audio stream. Discontinuous transmission processing reduces bandwidth occupancy and energy consumption. The comfortable noise supplementary processing ensures the natural continuity and immersion of the communication process by inserting simulated background sound from a preset noise library or personalized configuration of a user into a voice gap. The stable voice communication quality is kept in a changeable environment, and the equipment power consumption is reduced.
Owner:SHENZHEN SOUNDFIT TECH CO LTD

Sound signal detection method and device, computer readable storage medium, terminal

A sound signal detection method and device, a computer readable storage medium and a terminal, the method comprising: determining a plurality of sound signals collected by a plurality of microphones from a same input signal within a first preset time period; determining a correlation coefficient, an error energy cumulative value and an energy difference cumulative value between each two sound signals of the plurality of sound signals; and if one or more of the following conditions are met, determining that wind noise exists in the input signal: at least one of the correlation coefficients is less than a preset correlation threshold, at least one of the error energy cumulative values is greater than a first preset difference value, and at least one of the energy difference cumulative values is greater than a second preset difference value. The present application can accurately determine whether wind noise exists in the sound signal collected by the microphone, and reduce unnecessary voice activity detection caused by excessive environmental wind noise.
Owner:SPREADTRUM COMMUNICATION (SHANGHAI) CO LTD

A voice directional transmission method and system

The application provides a voice directional transmission method and system, and relates to the technical field of communication. The method comprises the following steps: collecting audio data of a sound source in a directional manner, generating sound source direction information and calculating the change rate thereof; performing voice activity detection on the audio, counting the voice frame density and the dialogue round change rate, and generating user communication state data; obtaining real-time quality parameters of a communication link and predicting a short-time prediction value of the link quality; constructing a multi-dimensional decision model, taking the short-time prediction value of the link, the real-time quality parameters, the user communication state data and the sound source direction change rate as inputs, and dynamically calculating a mixing coefficient; dynamically generating a target coding mode and its pre-processing parameters from at least two coding modes according to the mixing coefficient; and sending the coding result obtained according to the target parameters to a mobile terminal through a low-power channel. The voice acquisition and transmission are adaptively and cooperatively optimized according to the sound source dynamics, the dialogue rhythm and the link fluctuation, and the communication quality and user experience are improved.
Owner:SHENZHEN ZHILIANMAO TECHNOLOGY CO LTD

Webpage element control method, device, equipment, storage medium and program product

The application provides a webpage element control method and device, equipment, storage medium and program product, which can be applied to the technical field of artificial intelligence and financial technology. The method comprises the following steps: processing a plurality of candidate audio clips through a pre-trained voice activity detection model to obtain a voice existence probability corresponding to each candidate audio clip, the plurality of candidate audio clips being obtained by screening from a plurality of audio clips based on at least one of short-time energy and zero-crossing rate of each audio clip; screening the plurality of candidate audio clips based on the plurality of voice existence probabilities to obtain a target audio clip; performing voice recognition on the target audio clip to obtain voice recognition text; guiding a pre-trained large model to process the voice recognition text and an element index based on a plurality of first prompt words to obtain an element control instruction, and transmitting the element control instruction to a webpage to control the webpage to perform a webpage operation indicated by the element control instruction, the element index indicating a corresponding relationship between a user intention and a webpage element.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Voice quality detection method and device, medium and equipment

The embodiment of the invention discloses a voice quality detection method, and the method comprises the steps: determining a voice classification result of each frame through a preset voice activity detection algorithm, dividing audio data into a voice segment and a non-voice segment, taking each frame of audio of which the classification result is the voice data in the non-voice segment as an interference frame, and carrying out the elimination. And calculating a signal-to-noise ratio of the voice segment to determine a quality detection result. And the interference frame which is misjudged as voice is eliminated from the non-voice segment, and purer noise estimation is obtained. Based on the calculated signal-to-noise ratio, the voice signal quality can be reflected more truly, and the problems of noise power overestimation and signal-to-noise ratio underestimation caused by VAD false detection are effectively solved, so that the accuracy and reliability of voice quality detection are improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Intelligent dialogue system and method supporting real-time voice interruption

The invention discloses an intelligent dialogue system and method supporting real-time voice interruption, and relates to the technical field of man-machine interaction. The system comprises a voice activity detection module, a voice recognition module, a dialogue management and LLM interaction module, a voice synthesis module and an interruption control module. Wherein the dialogue management and LLM interaction module is used for managing dialogue states and contexts and calling LLM to generate reply texts; when the dialogue is interrupted, re-integrating the context and the new instruction text, and then calling the large language model to generate a new reply text; the interruption control module is used for sending an interruption signal to the voice synthesis module when the voice activity detection module detects effective user input, and triggering the dialogue management and LLM interaction module to re-integrate the context to generate a new reply; the voice activity detection module and the voice synthesis module operate at the same time and are not blocked. The method supports interruption at any time, and ensures semantic coherence of the interrupted dialogue through a context management mechanism, and the user experience is smooth.
Owner:FUJIAN STAR NET WISDOM TECH CO LTD

Human-computer voice interaction method and device, storage medium and electronic device

The application discloses a human-computer voice interaction method and device, a storage medium and an electronic device, relates to the technical field of smart homes, and comprises the following steps: automatically performing voice recognition on voice interaction information of a target object, obtaining a first intermediate recognition text and a second intermediate recognition text, a pause duration threshold set in a first voice activity detection strategy for detecting a moment when the target object stops speaking is lower than a pause duration threshold set in a second voice activity detection strategy for detecting the moment when the target object stops speaking; comparing the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text is consistent with the first intermediate recognition text, obtaining a target recognition result of the voice interaction information based on a voice recognition end identifier of the monitored voice interaction information; and performing natural language processing on the target recognition result to obtain a first interaction intention, thereby solving the technical problem that automatic voice recognition is not accurate enough in the human-computer voice interaction process.
Owner:HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD

Integrated voice-video collaborative perception all-in-one machine interaction system

The invention discloses an integrated voice-video collaborative perception all-in-one machine interaction system, which comprises a camera module, a microphone array module, a loudspeaker module, a processor module and a memory module, and is characterized in that the processor module is configured to execute the following steps: S1, processing an audio signal acquired by the microphone array module, obtaining sound source direction information and carrying out voice activity detection; s2, processing a video frame collected by the camera module, detecting and tracking a face or human body area of at least one target, and extracting mouth motion features; s3, based on a unified master clock, stamping unified timestamps on the audio data in the S1 and the video data in the S2 to realize cross-modal time synchronization; the method has the advantages that audio and video bimodal information is deeply integrated, fusion decision making is carried out under the unified time reference, and therefore stable, accurate and self-adaptive spokesman tracking and video composition are achieved.
Owner:深圳百利鸿成科技有限公司

Intelligent collection method and device based on audio analysis

The invention discloses an intelligent collection method and device based on audio analysis, and is applied to the technical field of data processing.The method comprises the steps that multi-mode audio analysis serves as the core, and basic information such as debtor credit data and legal document templates is collected and stored in a structured mode; call audio is transferred into a text stream with a timestamp in real time through an ASR technology, and audio framing, voice activity detection and other features are synchronously extracted. Through semantic, acoustic feature and dialogue rhythm three-dimensional analysis, a communication strategy and a switching threshold are determined by a multi-modal fusion and strategy decision engine, an adaptive verbal skill is generated, and a conversation link is established through a third-party platform. The debtor feedback is analyzed in real time, the verbal skill is dynamically adjusted until an effective result is obtained, finally key information such as committed repayment is extracted, a compliance legal document is automatically generated and sent out through one key after manual auditing, and whole-process intelligent collection is achieved.
Owner:ZHANGZHOU SEETEC OPTOELECTRONICS TECH CO LTD

Speaker log analysis method and system fused with spatial representation, and storage medium

The invention relates to the technical field of audio recognition and analysis, in particular to a speaker log analysis method and system fused with spatial representation and a storage medium, and the method comprises the steps: obtaining a multi-channel audio and a single-channel audio corresponding to the multi-channel audio, and carrying out the voice activity detection of the single-channel audio, and determining an effective voice segment; extracting voiceprint representation vectors from the effective voice segments; the multi-channel audio is processed and then input to the spatial representation extraction model, and the model outputs a spatial representation vector; according to a voice activity detection result, performing time alignment and segmentation on the obtained spatial representation vector to obtain a segmented spatial representation vector; performing feature splicing on the voiceprint representation vector and the segmented space representation vector to form a representation fusion vector; and clustering the representation fusion vector, and generating a speaker log with a timestamp according to a clustering grouping result. According to the method and the device, the original multi-channel audio can be converted into the low-dimensional space representation vector, and then the low-dimensional space representation vector is fused with the voiceprint representation vector to realize a high-precision speaker log task.
Owner:AISPEECH CO LTD