Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

497 results about "Speech processing" patented technology

Speech processing is the study of speech signals and the processing methods of signals. The signals are usually processed in a digital representation, so speech processing can be regarded as a special case of digital signal processing, applied to speech signals. Aspects of speech processing includes the acquisition, manipulation, storage, transfer and output of speech signals. The input is called speech recognition and the output is called speech synthesis.

Embedding-based large language model tuning

Systems and methods for embedding-based LLM tuning include generating a first embedding of received user input data and utilizing a translation model trained to associate user input with device names to generate a second embedding that differs at least in part from the first embedding. Reference embeddings corresponding to devices associated with user account data may be generated and a subset of the reference embeddings that satisfy a threshold similarity to the second embedding may be determined. A large language model (LLM) configured to determine a response to the user input data may utilize data for a subset of devices that correspond to the subset of the reference embeddings for performing speech processing.
Owner:AMAZON TECH INC

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Task processing method and device based on semantic understanding, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a task processing method, device and equipment based on semantic understanding and a medium. The method comprises the following steps: extracting semantic elements to generate structured task information, initiating a parameter supplement request and updating the structured task information when detecting that the semantic elements are missing, decomposing the structured task information into a plurality of pieces of structured sub-task information, generating a task scheme based on the plurality of pieces of structured sub-task information, and sending the task scheme to a server. And monitoring the execution state of the sub-task information, dynamically adjusting the task scheme, and outputting adjustment feedback information. According to the method, the task scheme is automatically generated through semantic element extraction and subtask decomposition, manual intervention is reduced in combination with a state data dynamic adjustment mechanism, and the processing accuracy and the execution flexibility are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

High-performance voice processing method

The invention relates to a high-performance voice processing method, and the method comprises the steps: obtaining voice data transmitted through a network, generating corresponding scene fingerprint information, determining a frame parameter of a lightweight multi-scene adaptation frame, and a corresponding processing parameter, and if the duration of a lost segment is smaller than a preset threshold value, determining that the segment is lost. If yes, generating a corresponding first compensation result according to the acoustic compensation parameter; if the duration is larger than or equal to a preset threshold value, whether the acoustic compensation process and the semantic processing process are performed in parallel or not is judged according to the frame parameters, if not, a corresponding second compensation result is generated, and if yes, a corresponding second compensation result is generated according to the acoustic compensation process and the semantic processing process; and outputting the corresponding target voice transmission data. The definition and the stability of the voice signal in a complex environment can be effectively improved, different compensation modes are flexibly switched or executed in parallel according to the packet loss duration and the resource state, and the effectiveness and the stability of a recovery mechanism are improved.
Owner:SHENZHEN SOUNDFIT TECH CO LTD

Medical record automatic filling method and system based on voice input

The invention provides a medical record automatic filling method and system based on voice input, and relates to the technical field of medical information processing.The method comprises the steps that voice input data are received by stages through voice collection equipment during diagnosis and treatment, real-time diagnosis and treatment stage identifiers are obtained, and a staged voice processing model is constructed based on the real-time diagnosis and treatment stage identifiers; therefore, semantic hierarchical processing is performed on the voice input data to generate a hierarchical semantic result, and a dynamic mapping relationship between the hierarchical semantic result and the medical record field is established and is converted and filled into a template to generate a staged medical record document. Semantic inheritance processing is carried out on the documents in different stages to correct semantic faults, and a complete medical record document is obtained and uploaded to a hospital electronic medical record system to complete automatic filling. The medical record recording efficiency and accuracy are improved.
Owner:四川互慧软件有限公司

Audio-visual feature fusion-based multi-mode voice separation method and audio-visual feature fusion-based multi-mode voice separation system

The invention discloses a multi-mode voice separation method and system based on audio-visual feature fusion. The method comprises the following steps: a voice and video acquisition step; a voice short-time Fourier transform step; adopting short-time Fourier transform to convert the time domain signal of the mixed voice into a complex spectrum; a face video processing step; preprocessing the face video by using a face recognition model, and extracting a video frame sequence of a lip region of a single speaker; a voice separation step; and a voice short-time inverse Fourier transform step. The system comprises a voice preprocessing module, a video preprocessing module, an audio-visual voice separation module and a voice post-processing module. The multi-mode voice separation method and system based on audio-visual feature fusion have the advantages that the parameter quantity and the calculation quantity in the voice processing process can be reduced, meanwhile, the available information quantity in the voice separation process is increased through the time domain and frequency domain features, and the voice separation accuracy is improved.
Owner:ANHUI UNIV

Multilingual full-speech processing method and device based on speech recognition, and medium

The invention discloses a multilingual full-speech processing method and device based on speech recognition and a medium, and relates to the technical field of speech recognition, and the method comprises the steps: calculating a language explicit trajectory based on a multilingual speech feature set, carrying out the association with a speech segment through language preference information in a historical session, and constructing a language implicit trajectory, integrating and generating a language weight track; dividing the language weight trajectory into voice trajectory nodes, recording multi-language voice trajectory composite features, connecting the multi-language voice trajectory composite features into a voice trajectory chain, calculating inter-node continuity indexes, and generating a voice trajectory node structure; constructing a track node continuity credibility field according to the voice track node structure, and adjusting multilingual voice recognition decision parameters to generate a multilingual transcription candidate set; and carrying out time sequence splicing and language mark arrangement on the language transcription candidate set to generate a multi-language full-voice transcription result set. According to the method, self-adaptive decoding of structure perception is realized, language switching is optimized, and the transcription precision is improved.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Voice control in a healthcare facility

Systems for voice control of medical devices in a healthcare facility are disclosed herein. The systems employ continuous speech processing software, voice recognition software, natural language processing software, and other software to permit voice control of the medical devices. Systems are also provided for distinguishing which medical device from among multiple medical devices in a patient room is the particular medical device to be controlled by voice input from a caregiver or a patient.
Owner:HILL ROM SERVICES INC

Adaptive voicemail and IVR detection for AI-driven call automation

The present disclosure provides a system for adaptive voicemail and interactive voice response (IVR) detection in outbound calls. The system includes a call initialization module configured to establish an outbound call connection, a speech processing module configured to convert incoming audio signals into text in real-time, a classification module configured to analyze the text and determine whether the call has reached a live recipient, a voicemail system, or an IVR menu, a decision-making module configured to determine an appropriate course of action based on the classification, and a response generation module configured to generate and deliver appropriate responses based on the determined course of action. The system enables efficient handling of outbound calls by accurately detecting and responding to different call scenarios.
Owner:SALESCLOSER TECHNOLOGIES INC

Method and apparatus for voice processing, device, storage medium, and program product

A method and apparatus for voice processing, a device, a storage medium, and a program product. The method comprises: acquiring a response text for a question voice of a target user (410); at least on the basis of a speech conversion requirement of the response text, selecting at least one of a first speech synthesis model on a local device and a second speech synthesis model on a server for executing a speech synthesis function on the response text (420); and acquiring a response voice corresponding to the response text, the response voice being generated after the speech synthesis function is executed on the response text by using the selected at least one speech synthesis model (430).
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Input processing with profile context

A speech-processing system may be configured to provide certain endpoints (e.g., skills and / or applications) with data reflecting an explicitly selected user profile. When handling requests from a device configured to support a session context (e.g., a personal device), the skill may handle multiple requests in an interaction (e.g., a “session”) using the selected profile as the session context, even if the system recognizes a different user for one or more of the requests. Thus, if the system recognizes more than one user using a personal device during an interaction with the skill (e.g., playing a game or requesting songs), the skill may handle each request of the interaction according to the selected profile. For a device not configured to support a session person, however, (e.g., a communal device) the skill may handle each request according to the recognized profile regardless of an explicitly selected profile.
Owner:AMAZON TECH INC

Law enforcement method with AI voice processing and video structured storage

The invention discloses a law enforcement method with AI voice processing and video structured storage. The law enforcement method comprises the following steps: acquiring building video data and on-site voice data in a hidden engineering scene; processing the on-site voice data through an AI noise reduction module to obtain a clear voice signal; converting the clear voice signal into a structured text through a voice-to-text technology, and associating a video timestamp of a voice time period corresponding to the structured text; according to the video timestamp, the structured text and the building construction stage information, generating a superposed watermark and embedding a picture of the building video data; packaging and storing the building video data embedded with the superimposed watermark and structured information, wherein the structured information comprises a structured text, a video timestamp and a building state tag; and establishing a retrieval index according to a preset dimension, wherein the preset dimension comprises a construction stage and a marking keyword.
Owner:BEIJING SLINGSHOT TECH CO LTD

English vocabulary follow-up pronunciation correction method based on English teaching

The invention discloses an English vocabulary follow-up pronunciation correction method based on English teaching, and relates to the technical field of voice processing, and the method comprises the following steps: S1, constructing a standard vocabulary pronunciation unit; s2, carrying out phoneme difference positioning by using a standard vocabulary pronunciation unit; s3, performing syllable dynamic comparison by using the phoneme deviation distribution of the learner; s4, performing vocabulary level comparison by using the learner phoneme deviation distribution and the syllable deviation mapping; and S5, performing difference mapping output by using the vocabulary pronunciation difference set. By setting the difference comparison between the actual pronunciation of the learner and the standard pronunciation, an accurate alignment relationship can be established at the phoneme level, and the difference between the actual pronunciation and the standard pronunciation in the aspects of acoustic characteristics, pronunciation duration, phoneme boundaries and the like is quantized, so that compared with the prior art, only the overall similarity or fuzziness judgment can be given, and the accuracy of pronunciation is improved. And a difference comparison mechanism can provide a clearer error correction reference, so that the learner can clearly recognize the position and degree of the problem when receiving the result.
Owner:GUANGZHOU COLLEGE OF COMMERCE

Speech processing architecture interfaces

Systems and methods for input processing architecture interfaces include receiving first audio data representing a first voice command within a first input processing architecture. A first action domain associated with the first voice command may be determined by the first input processing architecture. A first domain API may be determined, wherein the first domain API is predefined for interfacing with a second input processing architecture. Utilizing an application of multiple applications, a first input processing result associated with the first voice command may be determined. The first input processing result may be provided to the first domain API. The first input processing architecture may be caused to utilize the first input processing result from the first domain API to determine a first action to be performed responsive to the first voice command.
Owner:AMAZON TECH INC

Noise isolation and target sound enhancement system based on specific voice pre-storage

The invention relates to the technical field of voice processing and enhancement, in particular to a noise isolation and target voice enhancement system based on specific voice pre-storage, and the system extracts pre-stored features from a pre-stored target voice database module through collecting environment audio signals in real time and carrying out frame segmentation and frequency domain transformation; adaptive modulation of amplitude and phase of each sub-band signal is realized by combining energy sensing sub-band modulation, then environmental noise energy is identified through dynamic noise isolation and multi-stage suppression is carried out, and finally continuous and clear target voice output is generated through enhanced fusion. The method can achieve the effective enhancement and environmental noise suppression of the target voice in a complex and changeable environment, improves the voice recognition precision and definition, supports the dynamic switching of multiple teachers and multiple classes, adapts to the target voiceprint change, and achieves the high-robustness voice enhancement and noise reduction in teaching, conference and multi-target scenes.
Owner:UNIV OF JINAN

Digital conference voice processing method, system and device and storage medium

The invention relates to a digital conference voice processing method, system and device and a storage medium, and the method comprises the following steps: carrying out the pickup of a conference voice, obtaining a mixed voice signal, carrying out the framing sampling, and forming a voice sampling sequence; carrying out sound source direction estimation based on the sequence to obtain multi-sound-source position information, and carrying out beam forming and spatial filtering on the voice sampling sequence according to the multi-sound-source position information to obtain a sound source separation signal; performing voice segment segmentation on the signal to obtain a voice segment sequence, extracting voiceprint features of each segment, and generating a speaker feature mark; and performing time sequence recombination on the voice fragment sequence by using the mark, constructing a speaking time sequence table, selectively outputting the voice fragment sequence according to the table, and finally generating clear and ordered conference voice. The technical problems that due to the fact that a traditional voice processing method lacks effective space-voiceprint joint constraint, voice separation is not thorough, identities of speakers are confused, and the speaking time sequence is disordered are solved.
Owner:SHENZHEN YUXUN IOT CO LTD

Natural speech emotion synthesis and recognition method and system fused with deep learning

The invention provides a natural speech emotion synthesis and recognition method and system fused with deep learning, and relates to the technical field of speech processing, and the method comprises the steps: obtaining to-be-processed speech data and corresponding text content, extracting acoustic feature representation and semantic feature representation, building edge connection between a time sequence frame node of an acoustic feature and a semantic unit node of a semantic feature, constructing a bidirectional emotion association graph, calculating an edge weight, and propagating and updating based on graph convolution operation to obtain fusion feature representation; inputting the fusion features into an emotion classifier to obtain an emotion state identifier, calculating an emotion target area mask according to edge connection weight distribution, and generating an emotion regulation and control parameter; and performing speech synthesis based on the emotion regulation and control parameters and performing consistency verification to obtain synthesized speech and emotion deviation feedback information. According to the method, deep fusion of acoustics and semantics is realized through the bidirectional emotion association graph, and the emotion recognition accuracy and the emotion expressive force of speech synthesis are improved.
Owner:SMIC WANYE TECHNOLOGY CO LTD

Language identification method, device and equipment

The invention provides a language recognition method, device and equipment, and is applied to the field of voice processing. The language recognition method comprises the following steps: acquiring voice data; language recognition is carried out on the voice data to obtain a recognition result, and the recognition result comprises the first language; converting the voice data into text data in a first language to obtain first text data; performing quality inspection on the first text data to obtain a first inspection result; and under the condition that the first check result meets the text quality requirement, determining that the target language of the voice data is the first language. Therefore, by converting the voice data into the text data under the recognition language and performing quality inspection on the text data, the accuracy of language recognition on the voice data is improved.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

A role separation-based sales voice dialogue speaker segmentation and labeling method

The present application relates to the technical field of speech processing, in particular to a sales speech dialogue speaker segmentation and marking method based on role separation, comprising the following steps: generating a time-stamped speech segment sequence, synchronously extracting a voiceprint feature vector of each speech segment, starting a silence intention analysis engine, predicting the attribution role of the current silence according to the semantic content of the speech segments before and after the silence, generating a virtual voiceprint feature vector and inserting it into the speech segment sequence; performing role clustering on the voiceprint feature vector, splitting to generate a new speaker cluster; when detecting that there is speech overlap or audio loss in the speech segment, triggering mixed speech separation and generation compensation. The present application improves the abnormal recognition capability in the scene of sales role transformation and sound line disguise, reduces the risk of sales and customer speech confusion, and ensures the coherent and accurate role marking of each speech segment.
Owner:GUANGDONG INFORMATION NETWORK CO LTD

Voice processing method and device and electronic equipment

The invention discloses a voice processing method and device and electronic equipment. Under the condition that the electronic equipment collects voice signals of a user, sound signals collected by a microphone array of the electronic equipment are obtained, the voice signals of the user in the sound signals are blocked through a target blocking matrix in a plurality of preset blocking matrixes, and target to-be-eliminated signals are obtained, the target to-be-eliminated signal comprises an interference voice signal and / or an environmental noise signal, the plurality of preset blocking matrixes correspond to different postures of the electronic equipment for collecting the sound signal, the relative positions of the electronic equipment and the sound source are different under different postures, and the target to-be-eliminated signal in the sound signal is eliminated to obtain a user voice signal. By presetting the blocking matrixes corresponding to different voice input postures, the interference voice signals and the environmental noise signals are determined from the acquired sound signals and are eliminated from the sound signals, so that the purposes of protecting the privacy of a user and enabling the voice to be clear and understandable are achieved.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Data provenance method, apparatus and device based on cooperative training and medium

The present application relates to the technical field of speech processing, which can be applied to business scenarios such as financial technology and medical health, and discloses a data tracing method and device based on collaborative training, equipment and medium, comprising: obtaining an initial intermediate feature representation output by a generation model and a discrimination model, generating speech data and inputting the discrimination model to produce a discrimination loss, updating the generation model based on the discrimination loss to obtain a traceable intermediate feature representation, generating an updated intermediate feature representation using the updated generation model, generating training speech data and inputting the training speech data and real data into the discrimination model to obtain an updated discrimination model, and using the updated discrimination model to discriminate and output a detection result of whether the input speech data is generated by the generation model when receiving the input speech data. The present application trains the generation model and the discrimination model collaboratively, so that the generation model outputs speech embedded with traceable features, while improving the discrimination model's ability to reject unknown sources, thereby reducing waveform disturbance and accurately determining the source.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech recognition model training method, speech recognition method and device

The application provides a speech recognition model training method, a speech recognition method and device, and relates to the technical field of speech processing. The method comprises the following steps: in each iteration process, a training sample set is obtained, the training sample set comprises multiple training samples, each training sample comprises a sample speech signal, a text corresponding to the sample speech signal and a label of the text, the label is used for indicating whether the text is a complete sentence, for each training sample in the training sample set, taking the acoustic feature of the sample speech signal in the training sample as the input of a speech recognition model, outputting the speech recognition text and the predicted label of the sample speech signal, adjusting the parameters of the speech recognition model according to the speech recognition text and the predicted label of the sample speech signal obtained in each iteration process, and the text corresponding to the sample speech signal and the label of the text, until a stop training condition is met, and a trained speech recognition model is obtained. The accuracy of continuous speech recognition can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

System for voice data processing and method for operation thereof

This system for voice data processing comprises: an analog voice processing unit that receives an input of an analog voice signal and generates a digital voice signal; and a feature extraction unit that, when the digital voice signal is input, operates as an infinite impulse response (IIR) filter to filter the digital voice signal, and operate as an RNN to extract a voice feature from the filtered digital voice signal.
Owner:KOREA ADVANCED INST OF SCI & TECH

Voice processing method and device, electronic equipment and medium

The embodiment of the invention discloses a voice processing method and device, electronic equipment and a medium, and relates to the technical field of cloud games, and one specific implementation mode of the method comprises the steps that according to user voice input by a user in a current virtual scene and a virtual character state generated when the user operates a virtual character, the virtual character is processed according to the user voice input by the user and the virtual character state generated when the user operates the virtual character; and determining a user voice emotion and a virtual character emotion, determining an emotion matching result between the user voice emotion and the virtual character emotion in combination with an emotion matching algorithm, and adjusting voice parameters of the user voice according to a target parameter range corresponding to the virtual character emotion when the emotion matching result is that the emotion is not matched. According to the process, the user voice can fit the emotional state of the role in the current scene in real time, the dynamic adaptation of the user voice and the virtual role emotion is realized, the disjoint illegal feeling of the user emotion and the role emotion is avoided, and the immersion of voice interaction and the scene adaptability are remarkably improved.
Owner:MIGU COMIC CO LTD +2

Speech translation method, system, and storage medium

The application discloses a speech translation method and system and a storage medium, relates to the technical field of speech processing, and comprises the following steps: acquiring a real-time audio stream of source audio, and incrementally generating source language text corresponding to the real-time audio stream; detecting the semantic integrity of the generated source language text in real time; in the case where the semantic integrity is greater than or equal to the integrity threshold value corresponding to a minimum semantic unit, taking the generated source language text as a source language text block; determining target translation corresponding to the source language text block, and outputting the target translation. After the source language text is incrementally generated, the translation of the source language text block and the output are driven based on the semantic integrity, online segmentation translation of long sentences is realized, the translation waiting time is shortened, and the target translation and the original sound are synchronized.
Owner:ZHUHAI MOJIE TECH CO LTD

Model training methods, speech processing methods, devices, electronic devices, computer-readable storage media, and computer program products

ActiveCN122067509BTraining phaseEngineering
This application provides a model training method, a speech processing method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The method includes: acquiring a first phoneme sequence sample, a first word sequence sample, and a first alignment relationship between the first phoneme sequence sample and the first word sequence sample; in a first training phase, training an initial prediction model based on the first phoneme sequence sample, the first word sequence sample, the first alignment relationship, the first alignment window, the first phoneme loss weight, and a joint loss function of the initial prediction model; determining a state evaluation index for the initial prediction model; and triggering entry into a second training phase when the state evaluation index meets the phase switching conditions. This application can improve the accuracy of the model's prediction of alignment relationships.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Voice processing method and device, equipment, storage medium and program product

The embodiment of the invention provides a voice processing method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring a voice signal; performing signal processing on the voice signal to obtain a first time frequency signal; performing adaptive normalization processing on the first time-frequency signal to obtain a first feature map; performing two-dimensional modeling processing on the first feature map to obtain a second feature map; and performing synthesis processing on the second feature map to obtain an enhanced voice signal. The method can improve the voice quality.
Owner:RDA CHONGQING MICROELECTRONICS TECH CO LTD

Voice processing method and device, equipment, storage medium and program product

The invention provides a voice processing method and device, equipment, a storage medium and a program product, and relates to the technical field of data processing, and the method comprises the steps: carrying out the feature extraction of to-be-processed voice data, and obtaining the acoustic features of the voice data; determining a first emotion category of the voice data based on the acoustic features; determining a first emotion parameter corresponding to the first emotion category; and generating voice output data based on the first emotion parameter and the acoustic feature, and outputting the voice output data. According to the invention, the output voice output data can reflect emotional changes, so that the user experience is improved.
Owner:CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD +1

Voice keyword recognition method based on CNN-LSTM and online knowledge distillation

The invention discloses a voice keyword recognition method based on CNN-LSTM and online knowledge distillation, and relates to the technical field of voice processing. According to the method, an improved contraction residual attention module is introduced into a CNN-LSTM-based neural network model and is used for discovering and inhibiting redundant information and noise in features, and the expression ability of the features is enhanced. According to the method, a training method based on online knowledge distillation is introduced, the current training iteration is supervised by using information in the previous two training iterations, and an extra regularization effect is provided for the training process by using the difference in prediction probabilities output by a neural network model. The overfitting risk of the neural network model is reduced, and the generalization ability of the neural network model is enhanced.
Owner:GUILIN UNIV OF ELECTRONIC TECH