Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

851 results about "Speech processing" patented technology

Speech processing is the study of speech signals and the processing methods of signals. The signals are usually processed in a digital representation, so speech processing can be regarded as a special case of digital signal processing, applied to speech signals. Aspects of speech processing includes the acquisition, manipulation, storage, transfer and output of speech signals. The input is called speech recognition and the output is called speech synthesis.

Audio signal authenticity verification method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an audio signal authenticity verification method, device, equipment and medium, and the method comprises the steps: constructing an original audio text data set, and generating an adversarial sample set, inputting the original audio text data set and the adversarial sample set into an audio detection model for joint training to obtain an audio detection model subjected to adversarial training; the method comprises the steps of obtaining a to-be-detected audio signal and extracting an acoustic feature of the to-be-detected audio signal, obtaining a non-acoustic feature associated with the to-be-detected audio signal, constructing a multi-dimensional feature vector according to the acoustic feature and the non-acoustic feature, inputting the multi-dimensional feature vector into an audio detection model to generate an abnormal index, and executing a hierarchical response operation based on the abnormal index. According to the method, the robustness of the model is enhanced by introducing adversarial sample training, and the multi-dimensional feature vector is constructed by fusing the multi-modal features, so that accurate recognition and hierarchical response to the voice cloning attack are realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice enhancement method and device based on noise perception, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice enhancement method, device and equipment based on noise perception and a medium. Environment feature information is extracted and input into an audio enhancement model to generate an enhanced audio signal; obtaining a reference audio sample, extracting a personalized feature vector, and carrying out personalized processing on the enhanced audio signal; and collecting playing feedback data, determining a playing time domain adjustment parameter and a playing frequency domain adjustment parameter, adjusting the personalized enhanced audio signal, and generating an optimized audio signal. According to the method, dynamic adjustment is realized in combination with the feedback parameters in the playing process by fusing the environmental perception information and the personalized speaker characteristics, clear and natural optimized audio output with personalized styles can be generated in a complex environment, and the voice interaction quality and adaptability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice emotion recognition method and device based on context information, equipment and medium

PendingCN120636474ASpeech recognitionSingle sentenceSpeech sound
The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a context information-based voice emotion recognition method, device, equipment and medium, which comprises the following steps: receiving an original voice stream and generating an independent voice segment, recognizing a text and determining a speaker role type, and extracting an acoustic feature index; and generating a preliminary emotion label, generating context information in combination with the historical dialogue text, and inputting the context information, the preliminary emotion label, the speaker role type and the acoustic feature index into a multi-modal fusion module to generate an emotion judgment result. According to the method, multi-modal fusion is realized on the basis of context information by combining voice, text and role information, so that the emotion change of each role can be accurately recognized and understood in a complex dialogue scene, the problems of large single sentence emotion judgment error and neglect of the context information in a traditional method are avoided, and the accuracy and stability of emotion recognition are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Beamforming using image data

A device capable of using image data for purposes of determining a location of a user and audio beam selection to isolate audio in the direction of the user. The beamforming / beam-steering may occur after determining the user's location in order to conserve computing resources that would otherwise have been spent determining beams for non-desired directions. The beamformed audio may be used for speech processing, a communication session involving the device, or other purposes.
Owner:AMAZON TECH INC

Environmental acoustic simulation voice generation method, apparatus and device, and medium

ActiveCN120612917ASpeech recognitionSpeech synthesisPhonetic environmentData set
The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses an environmental acoustics simulation voice generation method, device and equipment and a medium, and the method comprises the steps: obtaining and separating mixed voice data, and generating original voice content and original environmental acoustics information; converting the original voice content into first text information; determining a target environment acoustic tag in combination with the original environment acoustic tag, the first text information and the target geographical location information; acquiring target environmental acoustic information from a preset sound data set based on the tag, and adjusting the amplitude characteristic of the target environmental acoustic information to match the original environmental acoustic information; and synthesizing the adjusted target environment acoustic information and the original voice content into simulated voice data. According to the method, the target geographic position information is introduced to participate in acoustic feature determination and amplitude adjustment, so that the generated simulated voice data is more consistent in geographic semantics and acoustic performance, and the authenticity and concealment of voice environment disguise are effectively improved.
Owner:PING AN TECH (BEIJING) CO LTD

System and method for enhancing speech of target speaker from audio signal in an ear-worn device using voice signatures

An ear-worn device is provided that operates to isolate and individually treat the received speech of a target speaker or multiple target speakers from an audio input signal detected in a multi-speaker environment. The ear-worn device uses a machine learning model that receives a voice signature of each of one or more target speakers as input signals, to identify and isolate the component of the audio input signal attributable to the target speaker(s). Once isolated, the target speaker's speech may be enhanced, de-emphasized, or otherwise processed in a manner desired by the wearer of the ear-worn device. The wearer may use an external electronic device, e.g., a phone, to select one or more target speakers in a conversation and / or configure various settings associated with processing the speech on the ear-worn device.
Owner:FORTELL RESEARCH INC

Voice text bidirectional conversion method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice text bidirectional conversion method, device, equipment and medium, and the method comprises the steps: respectively executing voice recognition or voice synthesis operation according to the type of input information; for the voice information, noise suppression parameters are generated in combination with the lip movement video data, noise reduction processing is executed, and the recognition accuracy is improved; for text information, a pre-generated speaker style vector is obtained, the vector is cited in the speech synthesis process to generate natural personalized speech, and lip movement information and tactile feedback which are synchronous with speech output are generated. According to the method, complex noise is suppressed by fusing lip movement data, personalized voice is generated by using the style vector, and lip movement and touch information is output, so that bidirectional real-time conversion of voice and text in a complex environment is realized, and recognition accuracy, voice naturalness and interaction synchronism are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-speaker dialogue voice analysis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a multi-speaker dialogue voice analysis method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-analyzed multi-speaker dialogue voice, determining a naturalness score based on an acoustic feature, and obtaining a multi-speaker dialogue voice analysis result; determining a semantic consistency score based on voice embedding and semantic embedding corresponding to a preset text, determining a speaker consistency score based on embedding of a plurality of speakers of the same speaker, determining an interaction rationality score based on voice alternate overlapping duration, determining a diversity score based on a variance of voice features, and fusing the scores, the comprehensive mass fraction is obtained. According to the invention, through quantitative evaluation of five dimensions of naturalness, semantic consistency, speaker consistency, interaction rationality and diversity, a comprehensive quality scoring system is established, so that the evaluation result simultaneously reflects voice fluency, content matching degree, identity stability, interaction rhythm rationality and feature richness.
Owner:PING AN TECH (SHENZHEN) CO LTD

Conversation-based skill component for assessing a user's state

The present application provides techniques for implementing a skill component, configured to perform an assessment of a user, as part of a speech processing system. The system may receive a natural language user input requesting assistance. The skill component may, using one or more machine learning models, determine at least one characteristic of the natural language input (e.g., lexical embedding, acoustic embedding, topic, tone, etc.). The skill component may determine state data for a present session, where the state data indicates a topic of the natural language user input and / or a user state associated with the natural language user input. The skill component may determine past state data of one or more past sessions, and generate a question to the user based on the state data for the natural language user input and the past state data.
Owner:AMAZON TECH INC

Embedding-based large language model tuning

Systems and methods for embedding-based LLM tuning include generating a first embedding of received user input data and utilizing a translation model trained to associate user input with device names to generate a second embedding that differs at least in part from the first embedding. Reference embeddings corresponding to devices associated with user account data may be generated and a subset of the reference embeddings that satisfy a threshold similarity to the second embedding may be determined. A large language model (LLM) configured to determine a response to the user input data may utilize data for a subset of devices that correspond to the subset of the reference embeddings for performing speech processing.
Owner:AMAZON TECH INC

Voice recognition processing method, system and equipment based on conference scene and medium

The invention relates to a voice recognition processing method, system and device based on a conference scene and a medium, and belongs to the technical field of voice processing. The voice recognition processing method comprises the following steps: acquiring an original conference audio stream collected by a microphone array; performing signal preprocessing on the original conference audio stream collected by the main channel, and outputting a pure voice signal; generating a sound source orientation thermodynamic diagram based on the original conference audio stream; extracting multi-dimensional voiceprint feature vectors from the pure voice signals, performing dynamic grouping, outputting a voice fragment set marked with voiceprint IDs, and generating an initial transcription text; dynamically correcting the initial transliteration text, and outputting a transliteration text stream with an industry term tag; and performing periodic memory enhancement processing on the transliteration text stream, outputting and analyzing a long text, and generating structured conference summary data. According to the invention, the automation level and accuracy of conference voice processing can be improved.
Owner:CHINA TRANSPORT INFORMATION TECH GRP CO LTD

Two-process error correction method and device for real-time speech transcription

The invention provides a two-process error correction method and device for real-time speech transcription, and relates to the technical field of speech processing, and the method comprises the steps: extracting the Mel spectrum features of each segment, inputting each Mel spectrum feature into a lightweight end-to-end model, and obtaining a preliminary transcription text; splicing the segments according to a preset number to obtain a plurality of long segments, and inputting each long segment into a speech recognition model to obtain a high-precision transcription text; performing text comparison on the preliminary transcription text and the high-precision transcription text according to the confidence degree set of the preliminary transcription text to obtain all error vocabularies in the preliminary transcription text; and performing corresponding error correction processing on each error vocabulary in the preliminary transcription text according to the type of the error vocabulary and the high-precision transcription text to obtain a final transcription text. According to the method, through a two-process transcription error correction mechanism of the preliminary transcription text and the high-precision transcription text, transcription error accumulation is reduced on the premise that the real-time performance is not affected, and the transcription accuracy in a complex scene is improved.
Owner:NANJING DOLPHIN INTELLIGENT TECH CO LTD

Rhythm migration method and device, electronic equipment and storage medium

The invention relates to the technical field of voice processing, and provides a rhythm migration method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a decoupled rhythm feature based on a source rhythm voice, and a decoupled tone feature based on the voice of a target speaker, the decoupled rhythm feature represents the rhythm of the source rhythm voice, and the decoupled tone feature represents the tone of the target speaker; the decoupled timbre features represent the timbre of the voice of the target speaker; generating a target voice vector sequence based on the text features of the target text, the voice features of the voice of the target speaker, the decoupled rhythm features and the decoupled timbre features; and synthesizing a target audio based on the target voice vector sequence. According to the method and the device, the decoupled rhythm features and the decoupled timbre features are acquired, and the target voice is generated based on the features, so that the problem of feature mixing is effectively relieved, the timbre purity of the target speaker in cross-person rhythm migration is ensured, the expressive force of rhythm migration is improved, and the synthesized audio is more natural and vivid.
Owner:IFLYTEK CO LTD

Intelligent cabin multi-mode voice interaction system and method

The invention belongs to the technical field of voice processing, and discloses an intelligent cockpit multi-mode voice interaction system and method, and the system comprises a voice triggering unit which collects the environment audio and video information in a cockpit, judges whether to enter a voice interaction mode or not through combining with the environment perception parameters in a vehicle, and sends the voice interaction mode to the vehicle; when the voice interaction triggering condition is satisfied, generating a voice interaction input signal matched with the current environment; the mouth shape analysis unit is used for carrying out acoustic feature extraction on the voice interaction input signal, synchronously analyzing the lip motion trail of the driver in the video information, establishing a corresponding relation between voice phonemes and mouth shape motion, and forming a joint analysis feature; the candidate generation unit is used for carrying out segmented alignment on the joint analysis features and constructing a continuous multi-modal fragment sequence; performing time synchronization on the multi-modal fragment sequence, and projecting the multi-modal fragment sequence to a predefined intention space to obtain a candidate intention set containing different candidate intentions; and the man-machine interaction experience of the intelligent cabin is improved.
Owner:SHENZHEN SHENHANG HUACHUANG AUTOMOBILE TECH CO LTD

Speech processing method and apparatus, device, and medium

PendingUS20250329334A1Speech analysisSpeech segmentationAcoustics
A speech processing method includes: obtaining overlapping speech data; obtaining reference speech data of a specified object; extracting a voiceprint representation vector of the specified object from the reference speech data, the voiceprint representation vector representing a voiceprint characteristic of the specified object, and inputting the overlapping speech data and the voiceprint representation vector into a preset speech segmentation model, and segmenting, by the speech segmentation model based on an attention mechanism, the overlapping speech data to obtain a target speech signal matching the voiceprint characteristic; and generating a speech file of the specified object based on the speech signal.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Pilot earphone hearing protection method and system based on voice recognition compensation

The invention relates to the field of aviation voice signal processing, and discloses a voice recognition compensation pilot earphone hearing protection method and system, and the method comprises the following steps: collecting multi-modal data, and separating a sound source through tensor decomposition; inferring a pilot state by using a dynamic network; performing context recognition and semantic evaluation on the attention target voice; and finally, dynamically modulating the sound field based on deep reinforcement learning, enhancing the voice in a personalized manner, and outputting after noise suppression. The system comprises a multi-mode perception data acquisition module, a sound source decoupling module, a pilot state inference module, a voice processing and semantic evaluation module and a sound field modulation and output module. According to the invention, high-fidelity speech extraction is realized through multi-modal perception and tensor decomposition; evaluating priority key information in combination with attention and semantics; and deep reinforcement learning and model prediction control are adopted to dynamically optimize the sound field, so that the voice recognition accuracy and the pilot information acquisition efficiency are improved.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY

Ambient sound feature-based authenticity analysis method, apparatus and device, and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical treatment and health and the like, and discloses an authenticity analysis method, device and equipment based on environmental sound characteristics and a medium. Performing voice separation processing on the original voice data to generate environment voice data and pure voice data, analyzing the environment voice data and retrieving knowledge information associated with the environment voice data, and identifying the content of the pure voice data to generate dialogue text data, and inputting the environment sound features, the knowledge information, the dialogue text data and the user declaration information into an analysis model, and outputting a authenticity analysis result. According to the method, the environment sound data and the voice content are separated, the available features of the environment sound data and the voice content are extracted respectively, and the background knowledge and the user declaration information are combined to perform fusion reasoning in the unified analysis model, so that the accuracy of authenticity judgment and the adaptability to complex scenes can be improved.
Owner:PING AN TECH (BEIJING) CO LTD

Business data voice processing method and system based on artificial intelligence

The invention provides a service data voice processing method and system based on artificial intelligence, and relates to the field. The service data voice processing method based on artificial intelligence comprises multi-mode voice preprocessing, domain enhancement voice recognition, knowledge graph driven semantic analysis, a dynamic dialogue management mechanism and a closed-loop feedback optimization system. According to the method, high-precision recognition, semantic understanding and structured processing of the service voice data are realized by fusing the domain knowledge graph, the dynamic model optimization and the context sensing technology, and the service data processing efficiency and accuracy are improved; secondly, the system comprises a voice acquisition and preprocessing module, an AI voice processing module, a business data processing module, a model optimization module and a data interface module, the efficiency and accuracy of business data voice processing can be remarkably improved, and the system is suitable for various professional scenes such as finance, customer service and medical treatment.
Owner:HEBEI VOCATIONAL & TECH UNIV OF SCI & TECH

System for intelligent interaction of software as service platform

The invention provides a system for intelligent interaction of a software as a service platform. The system comprises a large language model LLM module, a voice processing module, a model context protocol MCP communication layer and a context management module. According to the system, function operation of the SaaS platform is directly driven through natural language interaction, combination of a menu or button control of a traditional graphical user interface (GUI), an LLM module and historical data of a context management module is not needed, the intention recognition accuracy is larger than or equal to 92%, and repeated correction of a user is reduced.
Owner:SUZHOU SANRUN LANDSCAPE ENG

Task processing method and device based on semantic understanding, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a task processing method, device and equipment based on semantic understanding and a medium. The method comprises the following steps: extracting semantic elements to generate structured task information, initiating a parameter supplement request and updating the structured task information when detecting that the semantic elements are missing, decomposing the structured task information into a plurality of pieces of structured sub-task information, generating a task scheme based on the plurality of pieces of structured sub-task information, and sending the task scheme to a server. And monitoring the execution state of the sub-task information, dynamically adjusting the task scheme, and outputting adjustment feedback information. According to the method, the task scheme is automatically generated through semantic element extraction and subtask decomposition, manual intervention is reduced in combination with a state data dynamic adjustment mechanism, and the processing accuracy and the execution flexibility are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-voice common exchange type anti-interference Bluetooth earphone translation system

The invention discloses a multi-voice common exchange type anti-interference Bluetooth earphone translation system, and relates to the technical field of voice processing. Comprising an audio dynamic state acquisition module, an audio feature preprocessing and data construction module, a dynamic range abnormal feature extraction and quantitative analysis module, an audio system state intelligent evaluation module and a sound source weight adaptive regulation and control module, through a high-precision microphone array, a digital signal processing chip and an audio processing module which are integrated in the earphone, audio dynamic range state characteristic data generated in the operation process of each sound source channel are monitored and captured in real time. According to the method, the audio dynamic range imbalance state is recognized in real time, the sound source weight is adjusted in a self-adaptive mode, interference channel local suppression and main voice channel stable maintenance are achieved, and translation task continuity and recognition accuracy are effectively guaranteed.
Owner:VISION INTELLIGENCE CO LTD

Cross-lingual ai voice cloning method, system, and storage medium thereof

This invention relates to the field of speech recognition technology, specifically disclosing a cross-language AI voiceprint cloning method, system, and storage medium. The method includes: a speech collection end gating the raw microphone speech, with qualified samples entering preprocessing; gating, restricted spectrum, and unified conditional signals are integrated across the upstream and downstream processes to significantly improve robustness under noise / echo conditions; a speech processing end using AI adaptive filtering to denoise and obtaining representational data according to restricted parameters; a feature extraction and recognition end extracting voiceprints from the spectrum and embedding them into parallel language recognition, storing the voiceprint-language-quality association; a voiceprint cloning end performing cosine search in a template library to obtain a similarity queue, which is then weighted and aggregated after being rearranged based on quality and language consistency to obtain a target template; small-sample adaptation improves cross-language generalization and scalability; and finally, obtaining the target language and generating cloned speech by combining it with the target template.
Owner:HUNAN BOJI LIFE TECHNOLOGY CO LTD

A method of mixed voice processing, an electronic device, a computer readable medium

The present application relates to the technical field of speech processing, and particularly relates to a mixed speech processing method, an electronic device and a computer readable medium. The method comprises the following steps: collecting mixed speech and environmental influence parameters; performing bionics frequency domain analysis on the mixed speech to obtain low-frequency attenuation compensation feature data; performing multipath effect propagation analysis on the low-frequency attenuation compensation feature data through the environmental influence parameters to generate channel distortion data; performing time domain-frequency domain joint deconvolution processing on the mixed speech by using the channel distortion data to generate direct sound components and reflected sound components; performing adversarial training based on the direct sound components and the reflected sound components to generate anti-multipath speech enhancement data; and constructing a dynamic frequency compensation filter based on preset environmental acoustic characteristics. The present application improves the output quality of mixed speech through multi-stage signal processing, frequency compensation and real-time optimization technology.
Owner:GUANGZHOU ZHIYU CLOUD NETWORK COMMUNICATIONS CO LTD

Data tracing method and device based on cooperative training, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a data tracing method, device, equipment and medium based on cooperative training, which comprises the following steps: obtaining a generative model and a discrimination model to output initial intermediate feature representation, generating voice data and inputting the voice data into the discrimination model to generate discrimination loss; updating the generation model based on the discrimination loss to obtain a traceable intermediate feature representation, generating an updated intermediate feature representation by using the updated generation model, generating training voice data, and inputting the training voice data and real data into the discrimination model for training to obtain an updated discrimination model; and when input voice data is received, the updated discrimination model is used for discriminating and outputting a detection result whether the input voice data is generated by the generation model or not. According to the invention, through cooperative training of the generative model and the discrimination model, the generative model outputs voice embedded with traceable features, the rejection capability of the discrimination model to unknown sources is improved, and waveform disturbance reduction and accurate source determination are realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

High-performance voice processing method

The invention relates to a high-performance voice processing method, and the method comprises the steps: obtaining voice data transmitted through a network, generating corresponding scene fingerprint information, determining a frame parameter of a lightweight multi-scene adaptation frame, and a corresponding processing parameter, and if the duration of a lost segment is smaller than a preset threshold value, determining that the segment is lost. If yes, generating a corresponding first compensation result according to the acoustic compensation parameter; if the duration is larger than or equal to a preset threshold value, whether the acoustic compensation process and the semantic processing process are performed in parallel or not is judged according to the frame parameters, if not, a corresponding second compensation result is generated, and if yes, a corresponding second compensation result is generated according to the acoustic compensation process and the semantic processing process; and outputting the corresponding target voice transmission data. The definition and the stability of the voice signal in a complex environment can be effectively improved, different compensation modes are flexibly switched or executed in parallel according to the packet loss duration and the resource state, and the effectiveness and the stability of a recovery mechanism are improved.
Owner:SHENZHEN SOUNDFIT TECH CO LTD

Speech processing utilizing customized embeddings

Systems and methods for speech processing utilizing customized embeddings include receiving a first textual representation of first audio data and intent data indicating an intent of a first voice command. A first embedding may be generated from the textual representation and stored on the device. Second audio data representing a second voice command may be received and a second embedding may be generated therefrom. The first embedding may be determined to have at least a threshold similarity to the second embedding, and an intent may be determined that is associated with the second voice command. An action may be performed utilizing the intent.
Owner:AMAZON TECH INC

State determination method and device based on audio information, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a state determination method, device, equipment and medium based on audio information. The method comprises the following steps: uniformly performing voice-to-text processing after a preset fixed time length to obtain text information, extracting acoustic features and linguistic features, generating a multi-dimensional feature vector after fusion, inputting a pre-trained analysis model, generating a state probability value, and determining a target state corresponding to an audio signal based on the state probability value. According to the method, the multi-dimensional feature vectors are fused on the basis of acoustic features and linguistic features, and the pre-trained analysis model is introduced to judge the state probability value, so that the problems of incomplete feature extraction, insufficient feature fusion and poor judgment result accuracy and generalization ability are effectively solved; and the accuracy and the stability of audio signal state judgment are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Medical record automatic filling method and system based on voice input

The invention provides a medical record automatic filling method and system based on voice input, and relates to the technical field of medical information processing.The method comprises the steps that voice input data are received by stages through voice collection equipment during diagnosis and treatment, real-time diagnosis and treatment stage identifiers are obtained, and a staged voice processing model is constructed based on the real-time diagnosis and treatment stage identifiers; therefore, semantic hierarchical processing is performed on the voice input data to generate a hierarchical semantic result, and a dynamic mapping relationship between the hierarchical semantic result and the medical record field is established and is converted and filled into a template to generate a staged medical record document. Semantic inheritance processing is carried out on the documents in different stages to correct semantic faults, and a complete medical record document is obtained and uploaded to a hospital electronic medical record system to complete automatic filling. The medical record recording efficiency and accuracy are improved.
Owner:四川互慧软件有限公司

Speech synthesis method, system and device, storage medium and program product

The invention provides a voice synthesis method, system and device, a storage medium and a program product, and relates to the technical field of artificial intelligence and voice processing, and the method comprises the steps: obtaining Mel spectrum data corresponding to to-be-synthesized voice text data; inputting the Mel spectrum data into a neural vocoder based on a selective state space model; and performing long sequence processing on the Mel spectrum data by using the selective state space model based on the neural vocoder to obtain synthesized audio data corresponding to the to-be-synthesized voice text data. According to the invention, the neural vocoder can be constructed based on the state space model for speech synthesis, the high-frequency reconstruction capability is improved, the loss of high-frequency details is avoided, and better synthetic tone quality is obtained.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Synthetic speech processing related to prosody prediction

A speech-processing system receives input data representing text. A prosody prediction component processes the input data to determine prosody embedding data corresponding to prosody of the text. A decoder processes the prosody embedding data and phoneme encoded data derived from the input data to determine audio output data corresponding to the text and the prosody.
Owner:AMAZON TECH INC