Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

544 results about "Speech characteristics" patented technology

Speech characteristics are features of speech that in varying may affect intelligibility. They include: Articulation. Pronunciation. Speech disfluency. Speech pauses. Speech pitch. Speech rate.

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Man-machine interaction method for electronic screen control

The invention belongs to the field of man-machine interaction, and particularly discloses a man-machine interaction method for electronic screen control, which comprises the following steps: acquiring touch track data of a user according to a touch sensor, and performing feature extraction on the touch track data to obtain a touch feature vector; the method comprises the following steps: acquiring gesture action data of a user according to image acquisition equipment, and performing key point detection on the gesture action data to obtain a gesture feature vector; the method comprises the following steps: acquiring voice instruction data of a user according to an audio acquisition device, and performing acoustic feature extraction on the voice instruction data to obtain a voice feature vector; acquiring current environment parameters in real time through an environment detection module, wherein the environment parameters comprise illumination intensity, environment noise level and distance between a user and a screen; the objective of the invention is to solve the problem of insufficient reliability of a man-machine interaction mode in a complex environment in the prior art.
Owner:SHENZHEN SAIBO YUHUA ELECTRONIC TECH CO LTD

Speech encoder training method and apparatus, device, medium, and program product

Disclosed are a speech encoder training method performed by a computer device. The method includes: masking a first sub-feature representation at a first feature position in a first text feature representation to obtain a first masked feature representation; performing feature prediction on a masked first feature position in the first masked feature representation based on a first speech feature representation to obtain a first predicted feature representation; and training a first speech encoder based on a difference between the first predicted feature representation and the first sub-feature representation to obtain a second speech encoder. The first speech encoder is trained by combining data in a speech modality with data in a text modality, and information included in the data in the text modality is adopted so that the first speech encoder can learn relatively high-level semantic representations of speech, thereby improving the prediction accuracy of representations.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Fan blade voiceprint monitoring method and device based on deep learning

The embodiment of the invention provides a fan blade voiceprint monitoring method and device based on deep learning, and the method comprises the steps: obtaining noise reduction voice signal data, obtaining a multi-scale voice feature matrix, carrying out the fusion of a fan blade rotation phase label and an initial high-dimensional acoustic embedding vector, and obtaining a phase-guided acoustic embedding vector; inputting the locally enhanced acoustic embedding vector into a multi-head self-attention module of an improved Conform network, and combining position coding to obtain a global context acoustic embedding vector; inputting the first feed-forward module into a second feed-forward module of the improved Conform network to obtain a Conform output embedded sequence; inputting the Conformer output embedding sequence into a frame-level attention mechanism to obtain a frame-level weighted embedding sequence, inputting the segment-level weighted acoustic embedding vector into a joint loss optimization network, outputting a voiceprint feature vector through the joint loss optimization network, performing similarity comparison on the voiceprint feature vector and a preset fan blade operation state voiceprint library, and obtaining a voiceprint feature vector; and judging the current blade acoustic state category.
Owner:DATANG ENVIRONMENT IND GRP

Array microphone noise reduction recording method based on cascade noise reduction and blind source separation

PendingCN121237113ASpeech analysisBiological modelsLossless codingNoise
The invention relates to an array microphone noise reduction recording method based on cascade noise reduction and blind source separation, which belongs to the technical field of voice signal processing and recording, and comprises the following steps: configuring a multi-channel array microphone, ensuring that the amplitude and phase of a channel signal are consistent, and collecting an original multi-channel voice signal; a weighted kernel function blind source separation algorithm is adopted, signal-to-noise ratio distribution characteristics of signals are extracted, kernel function weights are given, and target voice and interference signal components are obtained through decoupling of an independent component analysis model; executing target-oriented adaptive cascade noise reduction, locking the voice of a keynote speaker through directional pickup, reducing noise, filtering out reverberation, and enhancing the voice of a far-field target by combining a voice mask neural network with a far-field pickup algorithm in sequence; and processing the target voice through voice feature perception lossless coding and storing the target voice. According to the invention, stable acquisition of multi-channel signals, accurate separation of mixed signals and layered suppression of noise reverberation are realized, the signal-to-noise ratio and definition of far-field voice are significantly improved, and the method is suitable for single-person speaking or multi-person dialogue scenes.
Owner:SHANGHAI RONGDA DIGITAL TECH CO LTD

Intelligent control method and device, electronic equipment and readable storage medium

PendingCN121956605AImprove Fusion Accuracyimprove accuracyComputer controlModal dataEngineering
The invention discloses an intelligent control method and device, electronic equipment and a readable storage medium, and belongs to the technical field of intelligent control. Under the condition that the gesture instruction does not conflict with the voice instruction, performing multi-dimensional normalization on the gesture feature vector to obtain a to-be-fused gesture vector; performing normalization processing on the voice feature vector based on context semantic perception to obtain a to-be-fused voice vector; fusing the to-be-fused gesture vector and the to-be-fused voice vector to obtain a fusion vector, and determining a fusion instruction corresponding to the fusion vector as a target execution instruction; and under the condition that the gesture instruction conflicts with the voice instruction, selecting a target execution instruction according to a first confidence coefficient corresponding to the gesture instruction and a second confidence coefficient corresponding to the voice instruction. The self-adaptive normalization strategy is designed for different modal data, the data fusion accuracy is improved, meanwhile, the gesture and voice instruction conflict can be intelligently and flexibly processed, and the accuracy of determining the real intention of the user is improved.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1

Automatic voice response fault diagnosis method and system based on multi-source information fusion

The invention discloses an automatic voice response fault diagnosis method and system based on multi-source information fusion. The method comprises the following steps: receiving user voice input, synchronously obtaining user side intelligent electric meter data, meteorological environment information and historical service records, and extracting multi-modal features; semantic association is realized through an electric power service knowledge graph, and voice features, equipment data and environmental parameters are fused through space-time alignment; a machine learning dynamic decision tree engine is combined with a power consumption behavior analysis model to form a fault reasoning model, and a diagnosis result is effectively verified; and outputting the grading disposal scheme and triggering intelligent chemical order distribution. According to the invention, various information sources of the calling platform are fully utilized, a fault reasoning model with high accuracy is formed, the technical problems of inaccurate fault positioning and insufficient multi-source data collaboration of traditional voice response in power customer service are effectively solved, the diagnosis accuracy of customer power consumption problems is effectively improved, and the average processing time is shortened.
Owner:国家电网有限公司客户服务中心

Conformer-based mixed self-attention and convolution improved speech recognition method and system

The invention discloses a Conformer-based mixed self-attention and convolution improved speech recognition method and system. The recognition method comprises the steps of obtaining speech feature representation; the speech feature representation is input to a Conformer encoder for encoding processing, high-dimensional feature representation is obtained, and the Conformer encoder comprises a convolution module, an attention linear enhancement module and a feedforward network residual scaling module; according to the high-dimensional feature representation, bidirectional decoding is carried out through a bidirectional Transform decoder, and bidirectional decoding output is obtained; and carrying out merging processing on the bidirectional decoding output to obtain a speech recognition result. According to the method, the modeling capacity for long-time dependence and global context is enhanced through a bidirectional decoder architecture, the processing capacity of the model in a long sequence is improved through ALiBi relative position coding, and the model can be flexibly adjusted in a noise environment through a right decoder weighting adjustment mechanism.
Owner:HEBEI UNIV OF ENG

Estimation method, recording medium, and estimation device

An estimation method includes: obtaining a first voice feature group of a plurality of persons who speak a first language; obtaining a second voice feature group of a plurality of persons who speak a second language; obtaining a voice feature of a subject; correcting the voice feature of the subject according to a relationship between the first voice feature group and the second voice feature group; estimating, from the voice feature of the subject that has been corrected, an oral function or a cognitive function of the subject by using an estimation process for an oral function or a cognitive function based on the second language; and outputting a result of estimation of the oral function or the cognitive function of the subject.
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Multi-mode control method for voice interaction and gesture recognition of atmosphere lamp

The invention discloses a multi-mode control method for voice interaction and gesture recognition of an atmosphere lamp, and belongs to the field of intelligent control. The method comprises the following steps: S1, constructing a multi-mode sensing module, and arranging a bidirectional linear array microphone, a dual-lens infrared depth camera and a space calibration laser transmitter; s2, extracting voice features by adopting an extended spectrogram domain attention network, and generating a first semantic vector; s3, dynamic attitude features are extracted through the time-airspace sparse convolutional network; s4, performing modal alignment and cooperative coding on the two semantic vectors based on a bidirectional gating fusion mechanism, generating a joint interaction instruction vector, and solving modal conflicts by a time sequence synchronization and confidence regulation strategy; s5, inputting the vector into a multi-task intention classifier, analyzing a user intention and outputting a unique atmosphere lamp control code; and S6, transmitting the control code to a driving unit to realize light control. The beneficial effects are that voice and gesture fusion control is realized, and atmosphere lamp interaction intelligence and response precision are improved.
Owner:NINGBO ZHONGJUN SHANGYUAN AUTO PARTS

Multilingual full-speech processing method and device based on speech recognition, and medium

The invention discloses a multilingual full-speech processing method and device based on speech recognition and a medium, and relates to the technical field of speech recognition, and the method comprises the steps: calculating a language explicit trajectory based on a multilingual speech feature set, carrying out the association with a speech segment through language preference information in a historical session, and constructing a language implicit trajectory, integrating and generating a language weight track; dividing the language weight trajectory into voice trajectory nodes, recording multi-language voice trajectory composite features, connecting the multi-language voice trajectory composite features into a voice trajectory chain, calculating inter-node continuity indexes, and generating a voice trajectory node structure; constructing a track node continuity credibility field according to the voice track node structure, and adjusting multilingual voice recognition decision parameters to generate a multilingual transcription candidate set; and carrying out time sequence splicing and language mark arrangement on the language transcription candidate set to generate a multi-language full-voice transcription result set. According to the method, self-adaptive decoding of structure perception is realized, language switching is optimized, and the transcription precision is improved.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Digital human question and answer system

The invention provides a digital human question-answering system, and the system comprises a voice receiving and processing module which is used for receiving the voice input of a user in real time, splitting the input voice, extracting the acoustic features of each frame of voice in real time, and caching the acoustic features of continuous frames to form a streaming data queue; the voice feature reasoning module is used for inputting the streaming data queue into a pre-training question and answer model and outputting a question and answer text result and a voice synthesis instruction in real time; and the digital human synchronous driving module is used for generating a synchronous voice signal based on the question and answer text result and the voice synthesis instruction, mapping the synchronous voice signal to a mouth shape mapping library and an action template library, generating a mouth shape sequence and a limb action sequence matched with the voice time sequence, and completing digital human broadcasting. According to the method, user text or voice input is received, the question and answer result is returned after model reasoning, and the digital person is driven to complete voice broadcasting.
Owner:ZHONGCHUANG (WUHAN) TECH CO LTD

Voice wake-up method and device of camera unit, electronic equipment and storage medium

The embodiment of the invention provides a voice wake-up method and device for a camera unit, electronic equipment and a storage medium, first user voice sent by interaction equipment is received, corresponding voice feature data is obtained according to the first user voice, and the voice feature data represents the voice content of the first user voice; the voice feature data are processed through a large language model, a first execution instruction is generated, and the instruction content of the first execution instruction is matched with the user intention corresponding to the voice content; and sending the first execution instruction to the interaction device to call a camera unit of the interaction device to shoot a target image. The first user voice is processed through the large voice model, the first execution instruction matched with the user intention corresponding to the voice content is generated, then the first execution instruction is sent to the interaction device to call the camera shooting unit to shoot the target image, the camera shooting unit does not need to be triggered to shoot by using a fixed voice wake-up word, the false wake-up rate is reduced, and the user experience is improved. And the man-machine interaction experience is improved.
Owner:SHENZHEN DASHI FUTURE TECH CO LTD

Voice interaction method and device, electronic equipment and storage medium

The invention provides a voice interaction method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a real-time voice stream; under the condition that the wake-up word is detected from the real-time voice stream, voiceprint features of a target speaker are extracted from a voice segment, corresponding to the wake-up word, in the real-time voice stream; based on the correlation between the voiceprint feature and the voice feature of the real-time voice stream, performing voice enhancement on the real-time voice stream to obtain an enhanced voice stream of the target speaker; and performing voice interaction based on the enhanced voice stream of the target speaker. According to the method and device, the electronic equipment and the storage medium provided by the invention, under the condition that the wake-up word is detected, the voiceprint feature of the target speaker is extracted, and targeted voice enhancement is performed based on the voiceprint feature, so that the enhanced voice stream for the target speaker is accurately extracted from the complex real-time voice stream, and the voice enhancement efficiency is improved. The influence of interference voice on voice interaction is greatly reduced, and the accuracy and reliability of voice interaction are ensured.
Owner:XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD

Interaction method, interaction device and interaction system of refrigeration equipment and refrigeration equipment

The invention discloses an interaction method, an interaction device and an interaction system of refrigeration equipment and the refrigeration equipment, and belongs to the technical field of refrigeration equipment. The interaction method of the refrigeration equipment comprises the steps that interaction data of a user and the refrigeration equipment are obtained, the interaction data at least comprise image interaction data and voice interaction data, and food material storing and taking position information of the refrigeration equipment is obtained; extracting image feature information of the image interaction data, and inputting the image feature information and the food material access position information into the visual large model to obtain a first interaction result output by the visual large model; extracting first voice feature information of the voice interaction data, and inputting the first voice feature information into the large voice model to obtain a second interaction result output by the large voice model; and outputting target interaction information based on the first interaction result and the second interaction result. The intention and the demand of the user can be more comprehensively understood, so that the interaction accuracy is improved.
Owner:QINDAO HAIER REFRIGERATOR CO LTD +1

Zero sample speech synthesis method and device, computer equipment and storage medium

The invention relates to a zero-sample speech synthesis method and device, computer equipment, a storage medium and a program product, and the method comprises the steps: obtaining a target coding feature according to a reference speech and a target text; inputting the target coding features into a stream matching model to obtain a conditional velocity field and an unconditional velocity field; inputting the target coding feature into a prior model to obtain a prior speech feature; obtaining a prior generation flow field according to the prior voice features and the standard Gaussian noise; calculating a KL divergence value between the priori generated flow field and a preset real generated flow field, and taking a moment when the KL divergence value is smaller than or equal to a preset KL divergence threshold as an initial moment of the priori generated flow field; fusing the conditional velocity field and the unconditional velocity field to obtain a fused velocity field; and inputting the fusion velocity field from the initial moment to the target moment and the priori generated flow field into an ordinary differential equation solver to obtain the target speech features, thereby improving the speech quality of the synthesized speech.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

AI call content optimization method and device based on voiceprint recognition

The invention relates to the technical field of intelligent voice processing and communication, and discloses an AI call content optimization method and device based on voiceprint recognition. The method comprises the following steps: acquiring an original voice data stream from a user call, and separating voice speed, tone and frequency characteristics to obtain a dynamic voice characteristic set; determining a microphone frequency response deviation and a speech speed change rate according to the microphone frequency response deviation and the speech speed change rate to form a speech feature parameter set; if the parameter exceeds the threshold value, redistributing a speech speed weight to generate an adjusted speech data stream; noise reduction is carried out to obtain a pure data stream, and features are fused to generate a personalized sound effect adjustment curve; adapting the equipment difference to obtain an adaptation curve; compressing the voice data according to the adaptive curve and optimizing the transmission priority to obtain an optimized transmission data stream; and combining the transmission data stream with the adaptive curve to generate final call voice output. According to the method, conversation content dynamic adaptation and whole-process optimization are realized, and conversation quality stability and cross-scene applicability are improved.
Owner:QUANZHOU YUANZHISHI ELECTRONIC COMMERCE CO LTD

Speech translation method and device

The invention provides a speech translation method and device, and the method comprises the steps: translating source language speech data based on a speech translation model, and obtaining a target language text; a training target of the speech translation model comprises minimizing a difference between a first target language prediction text generated based on source language sample speech data and a translation label corresponding to the source language sample speech data, and minimizing a difference between speech features of the source language sample speech data and text features of a first source language sample text. And minimizing the difference between a second target language prediction text generated based on the pseudo source language speech features and a translation label corresponding to the second source language sample text. According to the method and the device, the problem of scarcity of annotated voice data is solved by efficiently utilizing relatively rich text data in a low-resource language scene, so that the performance of a voice translation model is improved.
Owner:IFLYTEK CO LTD

A government affair text auxiliary writing system and method based on a multi-modal large model

The application discloses a government affair text auxiliary writing system and method based on a multi-modal large model, and relates to the technical field of artificial intelligence and informationization; comprising: step 1: converting heterogeneous government affair data into unified semantic representation, realizing cross-modal deep alignment through fine-grained contrast learning; encoding text, image and voice features by using an encoder respectively, and constructing a training set containing 100,000 pairs of government affair image-text samples; weighting and fusing the text, image and voice features, dynamically adjusting the weight through an attention mechanism, training a large model, step 2: performing lightweight deployment of the large model, and step 3: using the large model to understand and identify user intention, and outputting government affair text.
Owner:INSPUR SOFTWARE TECH CO LTD

Depression detection method fused with multi-modal attention mechanism

The invention discloses a depression detection method fusing a multi-modal attention mechanism, and relates to the technical field of auxiliary psychological health diagnosis, and the method comprises the steps: respectively extracting a facial image and a voice signal from an original video, extracting a global feature map from the facial image through a deep convolutional network, and introducing a local fusion module and a global fusion module; combining a Gaussian pyramid attention module with a spatial domain attention module; speech signals are converted into a Mel spectrogram through preprocessing, speech features are extracted through a SincNet network, and the speech features are further sent to a time-frequency attention module to highlight depression-related intonation and energy changes; in the feature fusion stage, a self-adaptive attention mechanism is adopted, self-attention modeling is performed on face and voice modes, cross-mode fusion and context modeling are completed through a Query-Key-Value structure, and a depression score is output through a full connection layer. According to the scheme, the problem that cross-modal nonlinear correlation is difficult to dynamically capture is solved, and effective technical support is provided for mental health assessment and intervention.
Owner:ANHUI NORMAL UNIV

Digital human interaction action generation method and device, electronic equipment and storage medium

The invention provides a digital human interaction action generation method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: inputting multi-modal data of interaction between a target user and a digital human into a feature extraction module of an interaction model, and obtaining multi-modal features; wherein the multi-modal features comprise a visual focus feature, a voice feature and a text emotion feature corresponding to the voice feature; the multi-modal features are input into a fusion module of the interaction model, the fusion module adjusts the initial attention score of the fusion module based on the time sequence change features of the visual focus features, and the adjusted attention score is obtained; and the fusion module performs fusion processing on the multi-modal features according to the adjusted attention score to obtain multi-modal fusion features, and generates control parameters for driving at least one body part of the digital human to move according to the multi-modal fusion features.
Owner:IFLYTEK CO LTD

Model training method and device, electronic equipment, storage medium and program product

The invention provides a model training method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of data processing, and the method comprises the steps: carrying out the dynamic masking processing of a voice feature sample sequence corresponding to a voice signal sample, and obtaining a voice feature sample sequence after the masking processing, under the condition that parameters of a quantizer in the self-supervised learning model are kept fixed, training an encoder in the self-supervised learning model based on the voice feature sample sequence and the quantization tag sequence after masking processing, and obtaining a first self-supervised learning model under the condition that a first preset training condition is met; and training the quantizer in the self-supervised learning model on the basis of the voice feature sample sequence and the quantization tag sequence after masking processing under the condition that the parameter fixed limitation of the quantizer is removed and the parameter of the encoder is kept fixed, and obtaining a second self-supervised learning model under the condition that a second preset training condition is met. And the model stability and the training efficiency are improved.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Feature extraction method based on end-side speech recognition, electronic equipment and readable medium

The invention relates to a feature extraction method based on end-side speech recognition, electronic equipment and a readable medium, and the method comprises the steps: obtaining a to-be-recognized target speech waveform, and extracting an FBank feature matrix of the target speech waveform; identifying audio content characteristics corresponding to the FBank feature matrix through an attention mechanism, and determining frame rate compression parameters matched with the audio content characteristics; generating a first feature sequence corresponding to the FBank feature matrix according to the frame rate compression parameter; converting the first feature sequence into a second feature sequence with a fixed length through semantic driving; and converting the second feature sequence into a normalized feature sequence through incremental learning so as to adapt to the input of an end-side speech recognition model. Through dynamic frame rate compression based on audio content and semantic-driven feature filling, voices with different lengths are uniformly converted into feature sequences with fixed lengths and complete semantics, and the problems of information loss and redundant filling during end-side voice feature extraction by using fixed parameters are solved.
Owner:FIBOCOM WIRELESS

Self-adaptive voice semantic communication method based on hierarchical time sequence importance

The invention relates to a self-adaptive voice semantic communication method based on hierarchy time sequence importance, which belongs to the field of voice analysis, and comprises the following steps: a sending end performs priority division on a discrete feature matrix according to hierarchy and time sequence attributes of voice features, and screens the features according to importance scores in a feature matrix packaging and selecting link; the method comprises the following steps: interacting with wireless communication through a channel adaptive scheduling module, introducing a channel feedback mechanism, and dynamically adjusting transmission power and a resource allocation strategy by sensing current channel state information; and after completing signal demodulation, a receiving end inputs the acquired sparse feature flow into a voice restoration module, and globally reconstructs the received features by using voice priori knowledge of deep pre-training. According to the method, a hierarchical speech feature extraction technology based on a discrete codebook and a generative semantic repair technology are combined, so that the speech word error rate in a severe channel environment is reduced to the maximum extent, and the semantic intelligibility of a receiving end is improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

Voice signal decoding method and apparatus and electronic device

This disclosure provides a voice signal decoding method and apparatus and an electronic device. An encoding apparatus encodes an original voice signal, to obtain an encoded bitstream. The encoded bitstream includes an acoustic feature encoding result. A decoding apparatus obtains the acoustic feature encoding result in the encoded bitstream, obtains a style feature from the acoustic feature encoding result, obtains an excitation feature, performs style fusion processing on the excitation feature and the style feature, to obtain a fused voice feature, and reconstructs a decoded voice signal based on the voice feature. Because the style feature indicates a voice style of an original voice signal, a voice feature in the original voice signal can be restored from the voice feature obtained by performing style fusion on the excitation feature and the style feature.
Owner:HUAWEI TECH CO LTD

Method and related device for decoding natural continuous voice by using neural electrophysiological signal

The invention discloses a method and a related device for decoding natural continuous voice by a neural electrophysiological signal, and relates to the technical field of neural signal processing and voice synthesis, and the method comprises the steps: obtaining a neural electrophysiological signal generated when a user reads a target text or listens to a corresponding audio, breaking the single signal modal limitation, and obtaining a voice signal; various non-intrusive application scenes are considered; slicing according to a preset time window to obtain continuous neural electrophysiological signal fragments; inputting the segments into a pre-trained deep learning physiological signal decoding model, outputting corresponding speech feature segments, directly mapping neural signals and speech features, omitting an intermediate conversion step, and reducing error accumulation; and converting the feature fragments into time-domain voice waveforms through a pre-trained neural vocoder, and splicing the time-domain voice waveforms to obtain natural continuous voices. Practicality and naturalness of brain-computer interface communication are remarkably improved, and an efficient and convenient communication way is provided for people with language dysfunction.
Owner:SHENZHEN UNIV

A speech recognition method, system, device, and medium for elevator entrapment scenarios.

This invention provides a speech recognition method, system, device, and medium for elevator entrapment scenarios. The method includes: acquiring speech data in an elevator scenario and preprocessing the speech data to obtain a first speech feature; using the first speech feature as input to a deep neural network to output recognized text and an entrapment probability value; determining entrapment based on the recognized text and the entrapment probability value, and outputting an entrapment determination result. By using a deep neural network and a specific encoding and decoding process, this method can more accurately process speech data in elevator scenarios, thereby improving the accuracy of entrapment detection. By comparing the entrapment probability value with a threshold and combining it with the recognized text for comprehensive judgment, this method can reduce false detections and false negatives, improving the accuracy of entrapment determination.
Owner:ZHEJIANG NEW ZAILING TECH CO LTD

Synthetic data driven paralinguistic labeling method, apparatus, device and storage medium

The present disclosure relates to the technical field of speech recognition, and specifically provides a synthetic data driven paralinguistic labeling method and device, equipment and a storage medium. The synthetic data driven paralinguistic labeling method comprises: obtaining language information of paralinguistic to be labeled; identifying the language information through a preset paralinguistic labeling model to obtain target text, the target text containing paralinguistic labeling, the paralinguistic labeling model being obtained by adjusting a speech recognition model architecture through output text output by the speech recognition model architecture and label data included in a sample pair, the output text being obtained by fusing a speech feature vector and a semantic feature vector, and converting the obtained fusion vector into text, the speech feature vector and the semantic feature vector being obtained by respectively extracting features of speech and text included in the sample pair, and converting the obtained speech features and semantic features into vectors. Through the present disclosure, paralinguistic in speech can be accurately labeled.
Owner:BEIJING HAITIAN RUISHENG SCI TECH CO LTD

System for voice data processing and method for operation thereof

This system for voice data processing comprises: an analog voice processing unit that receives an input of an analog voice signal and generates a digital voice signal; and a feature extraction unit that, when the digital voice signal is input, operates as an infinite impulse response (IIR) filter to filter the digital voice signal, and operate as an RNN to extract a voice feature from the filtered digital voice signal.
Owner:KOREA ADVANCED INST OF SCI & TECH

Outbound call processing methods, devices, and computer program products

This application discloses a method, apparatus, and computer program product for processing outbound call services. Relating to the field of artificial intelligence, the method includes: executing an outbound call task to a user; processing the target service according to a preset service processing procedure; collecting the user's voice signal during the target service processing; extracting voice features and voice content from the voice signal; determining the current service processing node corresponding to the voice content in the preset service processing procedure; determining the service path map corresponding to the preset service processing procedure; determining a path deviation index value based on the current service processing node and the service path map; inputting the voice features into a target model to obtain the user's behavior recognition result; determining a risk index value for the user interrupting the target service based on the behavior recognition result and the path deviation index value; and adjusting the preset service processing procedure according to the risk index value. This application solves the problem of poor accuracy in recognizing user intent in outbound call services in related technologies.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA