Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

377 results about "Voice transformation" patented technology

Control method and artificial intelligence experiment system

The invention provides a control method, which is applied to an artificial intelligence experiment system, at least comprises a robot hardware platform, a controller, a wheeled robot chassis and an industrial mechanical arm, and the method comprises the following steps: starting an operation system and executing hardware self-inspection; loading a large language model service and a traditional AI model, wherein the traditional AI model comprises a speech recognition model, a target detection model and a semantic analysis model; a natural language instruction of a user is received, voice is converted into a structured text through a voice recognition model, an instruction intention is analyzed through a large language model and a semantic analysis model, robot control parameters are generated, and a multi-modal interaction control signal is obtained; based on the multi-mode interaction control signal, a camera is called to collect image data, the position and category of a target object to be grabbed by the robot are recognized through a target detection model, the moving path of the robot and the grabbing track of the mechanical arm are calculated in combination with laser radar data, and the wheeled robot chassis and the industrial mechanical arm are controlled to execute coordinated actions.
Owner:BEIJING ETERNAL CREATIVE TECH CO LTD

Method and device for converting lip language into voice, computer storage medium and terminal

The invention discloses a lip language-to-voice conversion method and device, a computer storage medium and a terminal, and aims to solve the problems that a lip language-to-voice conversion technology cannot be deployed on terminal equipment and voice quality cannot meet application requirements. The method is combined with a neural vocoder which reduces calculation complexity and resource requirements, system configuration requirements are reduced while system parameters are reduced, a design basis is provided for deploying a lip language-to-speech method on terminal equipment, and context-related visual feature sequences and user audio embedding vectors are fused, so that the user audio-to-speech conversion efficiency is improved, and the user audio-to-speech conversion efficiency is improved. A Mel spectrum acoustic feature sequence used for being converted into a voice waveform is obtained, and a user audio vector fused with the Mel spectrum acoustic feature sequence is obtained, so that the output voice waveform is more consistent with the real voice of a user; according to the embodiment of the invention, technical support is provided for deploying and applying the lip language-to-speech method meeting the speech quality requirement on the terminal equipment.
Owner:BEIJING WATERTEK INFORMATION TECH

Systems and methods for orchestrating interaction with an artificial intelligence application

Systems and methods for orchestrating interaction with an artificial intelligence (AI) application in a contact center environment receive, via an AI agent, a voice message from a user; convert the message from voice to text; generate an initial computational inference process based on the text message; determine whether or not all information required to execute the initial computational inference process is available to the processor; when a determination is made that all information required is available: execute the initial computational inference process; generate a text reply based on the initial computational inference process; convert the text reply to a voice reply; and send the voice reply to the user via the AI agent; when a determination is made that information is unavailable: generate a text query requesting the information; convert the text query to a voice query; and send the voice query to the user via the AI agent.
Owner:THE BANK OF NEW YORK MELLON

Speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering

The invention discloses a speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering, and relates to the technical field of speaker segmentation. The method comprises the following steps: extracting at an extraction point of a target voice to obtain an extraction point voiceprint; selecting a reference point, and collecting a reference speaker and a reference voiceprint; if the voiceprint of the extraction point is not consistent with the reference voiceprint, judging whether multi-person talking overlapping exists or not according to the voiceprint perception density, and if the multi-person talking overlapping does not exist, judging whether the extraction point is used as a voice conversion point or not according to the health information and the environment information of the reference speaker; if yes, judging whether the extraction point is used as a voice conversion point or not according to the voice content of the extraction point and the background noise probability; and classifying the voice segments according to the voice conversion points to realize speaker segmentation. According to the invention, the resource utilization rate of speaker segmentation based on multi-scale feature fusion and voiceprint perception density clustering is improved.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Sign language-voice conversion system

The invention discloses a sign language-voice conversion system, and belongs to the technical field of auxiliary communication and wearable computing. The system comprises a wearable myoelectricity acquisition module used for acquiring double-arm myoelectricity signals when a user executes sign language; the mobile terminal module is wirelessly connected with the acquisition module and is used for receiving and preprocessing the signal and uploading the signal; the cloud processing module is used for receiving the signal, converting the signal into text information through a sign language recognition model, and further calling a voice synthesis service to convert the text into voice data; and the wearable audio output module is used for receiving and playing the voice data. Through an innovative end-to-end hardware system architecture, natural, accurate and real-time translation and voice output of sign language gestures are realized, communication barriers between hearing-impaired people and healthy hearing people are effectively solved, and the system has the advantages of flexible deployment, user friendliness and privacy protection.
Owner:宋飞 +1

AI-based emotional text voice conversion method and device

The invention discloses an AI-based emotional text speech conversion method and device. The method comprises the following steps: acquiring speech segments and text records from historical data of a user; performing noise reduction processing and feature extraction according to the voice segments and the text records to obtain voice features; inputting the voice features into a pre-constructed emotional tendency model, and outputting emotional tendency and emotional intensity; according to the emotional tendency and the emotional intensity, adjusting a tone weight, a speech speed and a volume to obtain a speech parameter; extracting new voice features according to the voice parameters to perform scene emotion label matching, and determining voice adjustment parameters through a linear regression model; according to the emotional tendency and the new voice features, generating an emotional type through a pre-established emotional intention classification model, and calculating a voice parameter weight in combination with a pre-established emotional mapping table; and according to the voice adjustment parameter and the voice parameter weight, performing language synthesis to generate personalized voice. According to the method, personalized expression can be accurately generated according to the scene.
Owner:FUJIAN YUANZHI UNIVERSE CULTURE COMMUNICATION CO LTD

Multi-agent-based old people information acquisition voice interaction system and method

The invention discloses a multi-agent-based old people information acquisition voice interaction system and method. The system comprises a voice acquisition module for forming an original voice signal; the voice recognition module receives an original voice signal, converts the voice into a text through compression and coding processing and back-end voice recognition service, and forms a user text instruction signal; the agent processing module receives a domain specific request signal, executes task processing and information retrieval by calling a knowledge base or a data service interface of a corresponding domain, and generates an integrated result signal containing structured data and a natural language answer; and the bimodal output module receives the integration result signal and synchronously generates a high-readability text display signal and a high-recognizability voice broadcast signal, the text display signal is rendered and presented through an interface suitable for aging, and the voice broadcast signal is output through an audio component. According to the invention, the problems of complex operation, dispersed service and difficult interaction when the old use the digital application can be solved.
Owner:BEIJING FUDI INTELLIGENT TECHNOLOGY CO LTD

Intelligent Technical Protocol Based Approach Leveraging AI-ML to Block Vishing Scammers

Systems and methods detect and prevent vishing attacks through an integrated framework combining SIP header customization, STIR / SHAKEN frameworks, AI / ML analysis, and real-time speech analysis using the Viterbi algorithm. The system begins with call initiation, embedding authentication information in the SIP header. The SIP data is transmitted and verified using STIR / SHAKEN frameworks, ensuring the authenticity of the caller's identity. Verified data is cross-referenced with third-party databases and analyzed by an AI / ML engine to detect anomalies. If potential fraud is detected, the call is blocked, and the customer is notified. Calls that pass initial checks are further analyzed using the Viterbi algorithm, which converts speech to text and identifies suspicious patterns. An anomaly pattern detector processes the converted text to detect vishing indicators, terminating the call if a match is found. This multi-layered approach ensures robust protection against vishing, enhancing the security and reliability of voice communications while safeguarding users from fraud.
Owner:BANK OF AMERICA CORP

Audio communication method, audio conversion method, apparatus, electronic device, computer-readable storage medium, and computer program product

PCT designated stageWO2025237010A1Speech analysisComputer hardwareFeature coding
An audio communication method, an audio conversion method, a bitstream processing method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The audio communication method comprises: in response to a first communication request for an audio signal, acquiring, from a plurality of communication modes, a voice transformation mode for the audio signal (101); performing feature coding on the audio signal, so as to obtain a coded feature of the audio signal (102); acquiring a target timbre corresponding to the voice transformation mode, and determining a timbre feature of the target timbre (103); performing timbre conversion on the coded feature on the basis of the timbre feature, so as to obtain a target coded feature (104); and performing signal coding on the target coded feature, so as to obtain a target audio bitstream conforming to the target timbre, and transmitting the target audio bitstream to a decoding terminal (105).
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and apparatus for voice processing, device, storage medium, and program product

A method and apparatus for voice processing, a device, a storage medium, and a program product. The method comprises: acquiring a response text for a question voice of a target user (410); at least on the basis of a speech conversion requirement of the response text, selecting at least one of a first speech synthesis model on a local device and a second speech synthesis model on a server for executing a speech synthesis function on the response text (420); and acquiring a response voice corresponding to the response text, the response voice being generated after the speech synthesis function is executed on the response text by using the selected at least one speech synthesis model (430).
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Telecommunications switch-type infrastructure for communications with call router and audio record server for computational source-to-target language conversion via application of artificial intelligence agents

A class-4 telecommunications switch hosts artificial-intelligence agents that convert live speech to text, translate the text between languages, and synthesize natural speech in real time. The switch proxies calls among public trunks, PBX / media gateways, and cloud ACD / CRM services, embedding diacritic-rich transcripts and user-specific language-model personalization. Deployable at the customer edge, in the PSTN core, or as SaaS, the system supports one-to-one, one-to-many, many-to-one, and many-to-many call patterns. FPGA, ASIC, or SoC accelerators minimize latency and bandwidth, cutting capital cost while improving global voice interoperability and cybersecurity.
Owner:CUNNINGHAM CHERYL EE LIN

Speech conversion algorithm based on state space model and fusion strategy

The invention discloses a voice conversion algorithm based on a state space model and a fusion strategy, belongs to the field of voice signal processing, and aims to solve the problems of simplification of a feature fusion strategy (such as simple feature addition), high model calculation complexity, insufficient long-distance modeling dependency capability and the like in an existing voice conversion method. The invention provides two core innovations: 1, a novel feature fusion strategy based on cross attention and a gating mechanism, and 2, an efficient encoder architecture combined with a state space model. The voice conversion method mainly comprises two stages: the first stage is a data preprocessing stage, and cutting, voice conversion and labeling are mainly performed on an original voice data set so as to provide subsequent training; the second stage is a training stage of a voice conversion algorithm, namely a core stage, and mainly comprises a data feature extraction stage, a voice synthesis encoder and a decoder. The method has good performance in the traditional voice conversion task, and the current system is expected to have better adaptive capability.
Owner:BEIJING UNIV OF TECH

Cross-modal data retrieval method based on voice intention

The invention discloses a voice intention-based cross-modal data retrieval method, which comprises the following steps of: 1, inputting voice corresponding to a query request, and converting the voice into a text through a voice recognition algorithm; 2, intention analysis is conducted on the text, structured intention representation is output, word segmentation is conducted on keywords of the voice recognition result, and domain labels are extracted in combination with a domain word bank; 3, generating an image feature vector, a visual feature vector, a text feature vector and a structured feature for local data; storing each feature and the domain tag thereof into an ES vector database, wherein each vector is associated with the corresponding domain tag; 4, two-stage matching is carried out according to query information input by a user, and multi-modal retrieval is carried out; 5, filtering the search result based on the intention analysis result, and screening by utilizing a filtering condition to obtain a final query result; and step 6, performing model optimization based on user feedback, and maintaining and updating the vector library in real time through an online learning module.
Owner:THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

Context adaptive speech recognition method integrating entity replication and vector retrieval

The invention provides a context adaptive speech recognition method integrating entity replication and vector retrieval, and belongs to the technical field of artificial intelligence. The method comprises the following steps: converting input voice into acoustic feature representation; converting the external entity dictionary into vector representation; screening out candidate entities through an index technology, and determining to generate output or copy entities from a standard word list according to a matching degree calculated by an attention mechanism; jointly optimizing the model by comprehensively optimizing the objective function, wherein acoustic modeling, language modeling, entity copying accuracy and entity distinguishing capability are included; continuously updating the index during training; in reasoning, in combination with beam search and confidence threshold mechanisms, selection is made between generation of outputs from a standard word list and copying of entities. In the aspect of recognition performance, a dynamic copy-generation fusion mechanism and a contrast learning strategy are adopted in the method, the accuracy of named entity recognition is effectively improved, particularly, the method is excellent in performance when proper nouns such as personal names and place names are processed, and meanwhile the distinguishing capacity for entities similar in pronunciation is enhanced.
Owner:DALIAN MARITIME UNIVERSITY

Artificial intelligence-based speech synthesis method and device, computer equipment and medium

The application is suitable for the technical field of speech synthesis, and particularly relates to a speech synthesis method and device based on artificial intelligence, computer equipment and a medium. The application extracts a text feature vector of a target text through a feature extraction model, predicts the text feature vector through a stress predictor, outputs a stress prediction vector, adds the stress prediction vector to the text feature vector to obtain a text stress vector, predicts the text feature vector through a pause predictor, outputs a pause prediction vector, adds the pause prediction vector to the text feature vector to obtain a text pause vector, predicts the text stress vector and the text pause vector through a prosody predictor, outputs a text prosody vector, matches the text prosody vector with a phoneme sequence of the target text, obtains a phoneme sequence with prosody labels, performs speech conversion on the phoneme sequence with prosody labels, obtains synthesized speech, and through the prediction of stress, pause and prosody, the expressiveness, naturalness and accuracy of the synthesized speech are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Maintaining connectivity in challenging network conditions with caption-assisted calls

A system is provided for managing and coordinating communication between STT / TTS systems and these systems during online conferences, and for mitigating connectivity issues that may occur during online conferences to provide a seamless and reliable conference experience with real-time captioning and / or presented audio. Initially, online conference communication is transmitted via a lossy connectionless protocol / channel. Then, in response to detected connectivity issues with one or more systems involved in the online conference (e.g., which may cause jitter or packet loss), instructions are dynamically generated and processed to enable one or more of the connected systems to utilize a more reliable connection / protocol (such as a connection-oriented protocol) to transmit and / or process the online conference content. Codecs are used at the system level when it is necessary to convert speech to text with associated speech attribute information and to convert text to speech.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

system

We provide the system. [Solution] A sensor means for detecting visitors, A server means that receives notifications from the aforementioned sensor means, activates a generative artificial intelligence, and performs an initial dialogue. The aforementioned generative artificial intelligence includes speech recognition means for converting speech into text, A natural language processing means for analyzing the aforementioned text and classifying the visitor's purpose, A response means for generating an automated response message according to the classified request and communicating it to the visitor, A notification means for notifying the user's terminal of the aforementioned request and response content, A system that includes this.
Owner:SOFTBANK GROUP CORP

A real-time voice conversion method and device, electronic equipment and medium

ActiveCN115910083BSpeech recognitionSpeech synthesisSpeech segmentationEngineering
The application provides a real-time voice conversion method and device, electronic equipment and medium. The method comprises the following steps: intercepting first voice data meeting voice segmentation conditions from voice data of a source speaking object recorded in real time; processing the first voice data to extract first semantic information; inputting the first semantic information into a pre-trained voice conversion model, and converting and processing effective information of historical voice data before the first semantic information and the first voice data through the voice conversion model to obtain target voice feature information corresponding to the first semantic information and a voice factor of a target speaking object; reconstructing the target voice feature information to obtain second voice data converted from the first voice data, thereby realizing low-delay streaming inference and low-delay and high-performance real-time voice conversion.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

A dam twin scene modeling method based on human-machine coupling

The application discloses a dam twin scene modeling method based on human-computer coupling, and particularly relates to the field of text-to-image technology in the artificial intelligence, first, the voice of human is converted into text through constructing a voice-to-text method, and then the text is transmitted into a model; the generator in the model can retain more key text information through text-image feature fusion, map the semantics into the image area, generate high-quality realistic images, and the generated images match the overall sentence semantics, achieving good text-image semantic consistency effect; the discriminator can better identify the image area more related to the text information through the attention mechanism, promoting the text-image semantic consistency of the generated image; the stronger discriminator in turn promotes the generator to generate higher-quality images; and the introduction of contrast learning promotes the authenticity of the semantic consistency of the generated image.
Owner:HOHAI UNIV

USB Voice Converter

1. Name of the product in this design: USB Voice Converter. 2. Purpose of this design: For converting USB devices into converters with voice on / off functionality. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: a 3D model.
Owner:李奕鑫

Offline large-model intelligent bidirectional translation system oriented to navigation communication guarantee

The invention relates to the technical field of ship equipment, discloses an off-line large-model intelligent bidirectional translation system oriented to navigation communication guarantee, and aims to solve the problem that navigation communication among different languages depends on networking requirements of translation software or depends on subjective judgment of accompanying translators in the prior art. According to the technical scheme, the communication content is extracted and translated by using the pre-established model, the subjectivity of manual translation is avoided, networking is not needed, and communication among different languages is realized. According to the invention, the effective voice is extracted, so that the influence of environmental noise is reduced; the voice is converted into the text, communication content can be visually displayed, and information omission is avoided. And the translation unit loaded with the text matching model integrated with the special term of the marine communication can translate special words and sentences in the marine communication, so that the translation accuracy is improved. And the audio recording unit and the storage unit can store communication contents and translation contents, so that subsequent redisk is facilitated.
Owner:HARBIN ENG UNIV

Voice conversion method, apparatus, medium, and device

The application discloses a speech conversion method, device, medium and equipment, the speech conversion method comprises the following steps: determining the task data input into different encoders in a speech conversion model according to a speech conversion type, and encoding the input task data through the encoders in the speech conversion model to obtain the encoding features output by each encoder; obtaining the public features corresponding to the encoding features output by a public classifier, and obtaining the adversarial features corresponding to the encoding features output by an adversarial classifier; performing decoding processing to obtain a target mel spectrum and a target high-frequency contour; and synthesizing a target audio based on the target mel spectrum and the target high-frequency contour through a vocoder. In one-time speech conversion, different representation styles are transmitted respectively, that is, according to the speech conversion type, speech conversion of timbre or timbre+pitch is realized respectively, so that the converted speech maintains the naturalness and expressiveness of the source speech.
Owner:SHENZHEN RAISOUND TECH

System

An object of a system according to an embodiment is to provide a foreign language learning environment corresponding to flexible time and level.SOLUTION: A system includes a voice text conversion part, a generation AI part, a text-to-speech part, a foreign language selection part, and a conversation field designation part. The speech-text conversion unit converts speech into text. The generated AI unit generates a response based on the text converted by the speech-to-text conversion unit. The text-to-speech unit converts the text generated by the generation and AI unit into voice. The foreign language selection unit selects a foreign language to be learned. The conversation field designation unit designates a field and a level in which a conversation is desired.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Speech synthesis method and apparatus, electronic device, and computer readable medium

The application discloses a speech synthesis method and device, electronic equipment and a computer readable medium, and relates to the technical field of speech synthesis. The method comprises the following steps: based on input text, obtaining a first synthesized speech according to a pre-acquired basic language speech synthesis model, and obtaining a second synthesized speech according to a pre-acquired target language speech synthesis model, wherein the similarity of the training speech of the target language speech synthesis model to the training speech of the basic language speech synthesis model is higher than a preset value; performing speech conversion on the second synthesized speech based on pre-acquired basic language training speech to obtain third synthesized speech; and obtaining target synthesized speech based on the first synthesized speech and the third synthesized speech. Therefore, the similarity of synthesized speech of different languages is further improved, and the target synthesized speech including bilingual or even multilingual speech has high timbre consistency, thereby improving the hearing effect.
Owner:VOICEAI TECH CO LTD

Voice processing method and device, equipment, storage medium and program product

The invention provides a voice processing method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring a response text for question voice of a target user; at least one of a local first speech synthesis model and a second speech synthesis model at the server side is selected at least based on the speech conversion requirement of the response text, and the speech synthesis model is used for executing a speech synthesis function on the response text; and acquiring response voice corresponding to the response text, wherein the response voice is generated after the at least one selected voice synthesis model is used for executing a voice synthesis function on the response text. Therefore, how to determine the response voice corresponding to the response text can be determined in combination with the network communication capability, the voice conversion requirement for the response text and the like, the TTS requirement of a user under the conditions of any network communication and different voice conversion requirements can be met, and the user experience of the user can be improved while the accuracy is met.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

system

The system according to the embodiment aims to efficiently record the contents of meetings and business negotiations and to automate task management and schedule adjustment. [Solution] A system according to an embodiment includes a speech recognition unit, a generation unit, a monitoring unit, and a calendar creation unit. The speech recognition unit converts speech into text. The generation unit analyzes the text created by the speech recognition unit and automatically creates minutes, ToDo lists, and a WBS. The monitoring unit monitors the progress of tasks created by the generation unit and issues an alert if there is no progress. The calendar creation unit automatically creates a calendar when the next meeting is required based on tasks determined by the monitoring unit to be no progress, and automatically sends it to relevant parties.
Owner:SOFTBANK GROUP CORP

A bone conduction speech conversion method based on spectral envelope mapping

The application discloses a bone conduction speech conversion method based on spectral envelope mapping, comprising the following steps: pre-processing and linear prediction analysis of the bone conduction speech signal, and calculating the LP filter coefficient; mapping the LSF coefficient of the air conduction speech signal corresponding to the bone conduction speech signal by using the trained neural network; controlling the minimum error of the original bone conduction speech signal and the synthesized air conduction speech signal according to the timbre weighting characteristic of the bone conduction speech characteristic; estimating the integer pitch of the bone conduction speech signal through the timbre weighting filter, obtaining the reference signal by passing the linear prediction residual signal through the timbre weighting filter, estimating the fractional pitch to obtain the adaptive codebook vector; obtaining the new reference signal by subtracting the adaptive codebook vector from the reference signal, searching for the optimal excitation in the fixed codebook; synthesizing the air conduction speech signal by combining the optimal excitation and the LP filter of the air conduction speech signal, and correcting the air conduction speech signal.
Owner:DALIAN UNIV OF TECH

Graphical user interface for voice cases for electronic devices

ActiveCN310087397SMedical recordKey pressing
1. Name of the product in this design: Graphical User Interface for Voice Cases in Electronic Devices. 2. Purpose of this design: An electronic device. 3. The key design features of this product are its graphical user interface. 4. The picture or photo that best illustrates the key design points: Design 1 front view. 5. The display screen panel adopts the conventional design, omitting the rear view, left view, right view, top view, and bottom view. 6. Design 1 is designated as the basic design. 7. Purpose of the graphical user interface: to convert speech into text medical records and provide relevant suggestions. 8. Human-computer interaction method of graphical user interface: Design 1 The main view is the initial startup interface, and the right side is the embedded AI assistant bar, which can edit medical records and obtain auxiliary diagnosis and treatment suggestions through natural language commands. After clicking the record button, the voice input is transcribed into text in real time, presenting the interface changes as shown in Figure 1. In the Design 1 interface change state diagram 1, after the voice input ends or the stop button is clicked and the notes button is clicked to upload the data, the Design 1 interface change state diagram 2 is displayed. Then, clicking the AI ​​assistant bar on the right will bring up a pop-up window, presenting the Design 1 interface state change diagram 3. Design 2's main view is the initial startup interface. After clicking the record button, the voice input is transcribed into text in real time, presenting the interface changes as shown in Figure 1 of Design 2. In the Design 2 interface change state diagram 1, after the voice input ends or the stop button is clicked, and the notes button is clicked to upload the data, the interface changes to the Design 1 interface change state diagram 2. Then, clicking the storage button in the upper right corner of the AI ​​assistant bar will present the Design 2 interface change state diagram 3. The main view of Design 3 is the initial startup interface. After clicking the record button, the voice input is transcribed into text in real time, presenting the Design 3 interface change state diagram 1. After the voice input ends or the stop button is clicked and the notes button is clicked to upload data, the Design 3 interface change state diagram 2 is presented. Clicking the clinic button at the bottom center, the 1.5x button in the middle, and the "+" button on the right presents the Design 3 interface change state diagram 3. The main view of Design 4 is the initial startup interface. After clicking the record button, the voice input is transcribed into text in real time, presenting the Design 4 interface change state diagram 1. After the voice input ends or the stop button is clicked and the notes button is clicked to upload data, the Design 4 interface change state diagram 2 is presented. Then, clicking the case generation button in the upper middle part presents the Design 4 interface change state diagram 3.
Owner:BEIJING UNITED FAMILY HOSPITAL CO LTD

Voice conversion method, voice conversion apparatus, electronic device, and storage medium

The application provides a speech conversion method, a speech conversion device, an electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: obtaining original speech data of a target speaker; performing segmentation processing on the original speech data to obtain first speech data and second speech data; performing encoding processing on the first speech data and the second speech data through a vector quantization coding network of a speech conversion model to obtain a first text vector, a first speech feature vector, a second text vector and a second speech feature vector; the first speech feature vector and the second speech feature vector are used for representing speech characteristics of the target speaker; performing splicing processing on the first text vector, the first speech feature vector, the second text vector and the second speech feature vector to obtain a target speech vector; and performing decoding processing on the target speech vector through a decoding network of the speech conversion model to obtain target speech data. The application can improve the speech conversion effect.
Owner:PING AN TECH (SHENZHEN) CO LTD