Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

570 results about "Voice transformation" patented technology

Techniques for determining conversational intent

The present disclosure relates to systems and methods for enhancing the interaction between users and automated agents, such as digital assistants, by employing Large Language Models (LLMs) to infer the intent of spoken language. The invention involves continuously monitoring ambient audio, converting speech to text, and utilizing LLMs to determine whether spoken language is intended for the automated agent. A structured prompt, including the converted text and specific instructions, is sent to the LLM, which is fine-tuned to process domain-specific prompts. The LLM provides a structured output in a standardized format, indicating the user's intent. The system may involve multiple prompts to perform separate tasks, such as identifying intent and generating additional context-specific data. This approach facilitates a more natural and intuitive user experience by eliminating the need for wake words and allowing seamless conversational interaction with virtual assistants across various platforms and devices.
Owner:SNAP INC

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Government affair service content navigation method and system based on large language model

The invention relates to the technical field of multi-modal large language models, in particular to a government affair service content navigation method and system based on a large language model, and the method comprises the following steps: receiving multi-modal information input by a user through texts, voices or pictures; the voice is converted into a text, and character information in the picture is analyzed by using an OCR (Optical Character Recognition) technology; fusing multi-modal data, and inputting the fused multi-modal data into a large language model for semantic understanding and context association analysis; matching items are retrieved in combination with a local government affair knowledge base, and an initial recommendation list is generated; the recommendation result is displayed through the intelligent assistant, and interaction optimization options are provided; the method has the beneficial effects that the current government affair service content navigation mode is optimized through the multi-modal capability, the retrieval enhancement capability and the content generation capability of the large language model, so that the navigation is more modal and more intelligent, and meanwhile, the privacy and authority of data reply are ensured.
Owner:INSPUR SOFTWARE CO LTD

Electronic medical record automatic generation method based on voice recognition

InactiveCN120690205ASpeech recognitionPatient-specific dataMedical recordSpeech segmentation
The invention relates to the technical field of electronic medical record generation, and discloses an electronic medical record automatic generation method based on voice recognition, which comprises the following steps: S1, initializing voice input, distributing a unique voice acquisition identifier for a medical session, and completing identifier generation, input, storage, association and identity verification; s2, voice information intelligent recognition: converting voice into a text by using a voice recognition engine, and ensuring semantic consistency through a context verification unit and a semantic analysis unit; and S3, performing multi-speaker processing based on intelligent interference detection and resolution, positioning an interference time period and an interference source through an interference detection unit, and realizing time period distribution and priority ranking of multi-speaker voices by using a voice segmentation protocol and a linear weighting model. And finally, extracting related information from the text generated by voice conversion, and filling the related information into a medical record template of a hospital. The method improves the efficiency and accuracy of electronic medical record generation, solves the problems of multi-speaker interference and semantic logic, and is suitable for medical informatization scenes.
Owner:THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)

Control method and artificial intelligence experiment system

The invention provides a control method, which is applied to an artificial intelligence experiment system, at least comprises a robot hardware platform, a controller, a wheeled robot chassis and an industrial mechanical arm, and the method comprises the following steps: starting an operation system and executing hardware self-inspection; loading a large language model service and a traditional AI model, wherein the traditional AI model comprises a speech recognition model, a target detection model and a semantic analysis model; a natural language instruction of a user is received, voice is converted into a structured text through a voice recognition model, an instruction intention is analyzed through a large language model and a semantic analysis model, robot control parameters are generated, and a multi-modal interaction control signal is obtained; based on the multi-mode interaction control signal, a camera is called to collect image data, the position and category of a target object to be grabbed by the robot are recognized through a target detection model, the moving path of the robot and the grabbing track of the mechanical arm are calculated in combination with laser radar data, and the wheeled robot chassis and the industrial mechanical arm are controlled to execute coordinated actions.
Owner:BEIJING ETERNAL CREATIVE TECH CO LTD

Speech enhancement method and system based on conditional stream matching and vocoder

The invention discloses a speech enhancement method and system based on conditional stream matching and a vocoder, and the method comprises the following steps: S1, constructing a Mel spectrum extraction module which is used for converting an input noisy speech into a noisy Mel spectrum; s2, constructing a condition flow matching noise reduction module which is used for processing the noisy Mel spectrum obtained in the step S1 and outputting an enhanced Mel spectrum; and S3, constructing a neural network vocoder module which is used for restoring the enhanced Mel spectrum obtained in the step S2 into a time domain voice waveform so as to obtain an enhanced voice signal. The speech enhancement method combining conditional stream matching and the vocoder is proposed for the first time, the conditional stream matching method is innovatively introduced into the Mel-frequency spectrum domain, an end-to-end speech enhancement system is constructed, and a complete processing flow from noise speech input to high-quality speech signal output is realized.
Owner:HANGZHOU DIANZI UNIV

Multi-scene self-adaptive guide robot interaction method and system

The invention provides a multi-scene self-adaptive guide robot interaction method and system in the technical field of guide robots. The method comprises the steps that S1, a guide robot obtains interaction voice and converts the interaction voice into an interaction text; s2, analyzing the interaction text to generate a reply text, converting the reply text into a reply voice, and playing the reply voice for interaction; s3, managing the dialogue content in the interaction process; s4, the guide robot moves based on the guide path in the interaction process, and automatically plays explanation voice when arriving at the guide point; step S5, when the interaction text carries the target point, planning an advancing path to the target point through a bidirectional progressive optimal search algorithm; and S6, displacement is carried out based on the advancing path, and operation information of the robot is displayed in real time. The method has the advantages that the efficiency and accuracy of path planning, the flexibility of voice interaction and the expansibility of the system are greatly improved, and then the user experience of guide robot interaction is greatly improved.
Owner:XIAMEN UNIV

Speaker embedding-based speaker adaptation method and system generated by using global style tokens and prediction model

Disclosed are a speaker embedding-based speaker adaptation method and system generated by using global style tokens and a prediction model. The speaker adaptation method performed by the speaker adaptation system, according to an embodiment, may comprise the steps of: generating a plurality of speaker embeddings representing the tone of a speaker from a speaker embedding by using a voice transformation model including a global style token mechanism; and predicting the final speaker embedding representing a new speaker through similarity comparison between a new speaker embedding predicted by using a prediction model for predicting a speaker embedding and the plurality of generated speaker embeddings.
Owner:INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY

Voice evaluation method and device, medium and program product

The invention provides a voice evaluation method, a voice evaluation device, a non-transitory storage medium and a computer program product. The voice evaluation method comprises the following steps: extracting a first person voice with a first tone based on a to-be-evaluated voice; converting the voice of the first person into voice with a second tone; and obtaining an evaluation result of the voice of the first person based on the converted voice with the second tone and the standard voice of the one or more accent.
Owner:NEW ORIENTAL EDUCATION & TECH GRP CO LTD

Method and device for converting lip language into voice, computer storage medium and terminal

The invention discloses a lip language-to-voice conversion method and device, a computer storage medium and a terminal, and aims to solve the problems that a lip language-to-voice conversion technology cannot be deployed on terminal equipment and voice quality cannot meet application requirements. The method is combined with a neural vocoder which reduces calculation complexity and resource requirements, system configuration requirements are reduced while system parameters are reduced, a design basis is provided for deploying a lip language-to-speech method on terminal equipment, and context-related visual feature sequences and user audio embedding vectors are fused, so that the user audio-to-speech conversion efficiency is improved, and the user audio-to-speech conversion efficiency is improved. A Mel spectrum acoustic feature sequence used for being converted into a voice waveform is obtained, and a user audio vector fused with the Mel spectrum acoustic feature sequence is obtained, so that the output voice waveform is more consistent with the real voice of a user; according to the embodiment of the invention, technical support is provided for deploying and applying the lip language-to-speech method meeting the speech quality requirement on the terminal equipment.
Owner:BEIJING WATERTEK INFORMATION TECH

Voiceprint characterization method and system for resisting voice conversion and voice synthesis based on deep learning

PendingCN120431939ASpeech analysisSpeaker verificationEngineering
The invention discloses a voiceprint characterization method and system for resisting voice conversion and voice synthesis based on deep learning. Relates to the field of speech recognition and biological feature security. Comprising the steps of 1, acquiring an input voice sample signal in an automatic speaker verification system, and extracting an FBANK feature of the voice sample signal; 2, performing depth feature extraction on the FBANK features by using a convolutional neural network (CNN) to obtain a depth feature vector; 3, processing the depth feature vector by using a recurrent neural network (RNN), and generating a forged identity vector of the voice; and 4, classifying the counterfeit identity vectors by using a linear discriminant analysis (LDA) module, and outputting a judgment result. According to the invention, the accuracy of counterfeit detection can be effectively improved in both clean and noise environments.
Owner:ZHEJIANG UNIV

Game assisting method and device

The invention provides a game assisting method and device.The method comprises the steps that in the process that a user plays a game, voice of the user is recognized, the voice is converted into a text, and a game image corresponding to the text is matched and recorded as a target game image; retrieving a preset game strategy knowledge base by utilizing the target game image to obtain target game strategy information; the target game image, the text and the target game strategy information are submitted to the large-scale multi-mode model for analysis, the instructive text is generated, the instructive text is synthesized into the voice and played, and a game assisting scheme which does not affect the game experience of the user in the game playing process of the user and efficiently and effectively guides the user is provided.
Owner:HAIMA CLOUD TIANJIN INFORMATION TECH CO LTD

Systems and methods for orchestrating interaction with an artificial intelligence application

Systems and methods for orchestrating interaction with an artificial intelligence (AI) application in a contact center environment receive, via an AI agent, a voice message from a user; convert the message from voice to text; generate an initial computational inference process based on the text message; determine whether or not all information required to execute the initial computational inference process is available to the processor; when a determination is made that all information required is available: execute the initial computational inference process; generate a text reply based on the initial computational inference process; convert the text reply to a voice reply; and send the voice reply to the user via the AI agent; when a determination is made that information is unavailable: generate a text query requesting the information; convert the text query to a voice query; and send the voice query to the user via the AI agent.
Owner:THE BANK OF NEW YORK MELLON

VITS-based singing voice conversion method

PendingCN120496546ASpeech analysisPitch shiftFundamental frequency
The invention discloses a singing voice conversion method based on VITS. The method comprises the following steps of: 1, respectively extracting a bottleneck feature, a fundamental frequency, a singing style feature and a linear spectrogram from the singing of a source singer by utilizing a Whisper superficial layer encoder, an automatic singing transcription model AST and an AutoVC-based pitch-energy encoder, and taking the bottleneck feature, the fundamental frequency, the singing style feature and the linear spectrogram as the input of a VITS-based singing voice conversion model; reconstructing the singing voice by using a VITS-based singing voice conversion model in combination with the voiceprint characteristics of the singer to obtain converted singing voice of the target singer; and 2, performing pitch adjustment according to the pitch difference between the source singer and the target singer through a pitch shifter, and generating a target singing sound with the characteristics of the target singer by taking the bottleneck characteristics, the F0 and the singing style characteristics as the input of a voice conversion model. The method has the characteristics of multi-level refined acoustic feature decoupling, explicit singing style migration, high-fidelity waveform generation and dynamic adaptive pitch conversion.
Owner:XIDIAN UNIV

Speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering

The invention discloses a speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering, and relates to the technical field of speaker segmentation. The method comprises the following steps: extracting at an extraction point of a target voice to obtain an extraction point voiceprint; selecting a reference point, and collecting a reference speaker and a reference voiceprint; if the voiceprint of the extraction point is not consistent with the reference voiceprint, judging whether multi-person talking overlapping exists or not according to the voiceprint perception density, and if the multi-person talking overlapping does not exist, judging whether the extraction point is used as a voice conversion point or not according to the health information and the environment information of the reference speaker; if yes, judging whether the extraction point is used as a voice conversion point or not according to the voice content of the extraction point and the background noise probability; and classifying the voice segments according to the voice conversion points to realize speaker segmentation. According to the invention, the resource utilization rate of speaker segmentation based on multi-scale feature fusion and voiceprint perception density clustering is improved.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Sign language-voice conversion system

The invention discloses a sign language-voice conversion system, and belongs to the technical field of auxiliary communication and wearable computing. The system comprises a wearable myoelectricity acquisition module used for acquiring double-arm myoelectricity signals when a user executes sign language; the mobile terminal module is wirelessly connected with the acquisition module and is used for receiving and preprocessing the signal and uploading the signal; the cloud processing module is used for receiving the signal, converting the signal into text information through a sign language recognition model, and further calling a voice synthesis service to convert the text into voice data; and the wearable audio output module is used for receiving and playing the voice data. Through an innovative end-to-end hardware system architecture, natural, accurate and real-time translation and voice output of sign language gestures are realized, communication barriers between hearing-impaired people and healthy hearing people are effectively solved, and the system has the advantages of flexible deployment, user friendliness and privacy protection.
Owner:宋飞 +1

Voice conversion authentic identification method and system based on knowledge distillation alignment

The invention discloses a voice conversion authentic identification method and system based on knowledge distillation alignment, and is applied to the technical field of voice authentic identification. The method comprises the following steps that a double-branch model used for voice authentic identification is constructed, pure voice serves as input of a teacher branch, noisy voice serves as input of a student branch, and the teacher branch and the student branch share a feature extraction network of the same structure; applying a speech enhancement technology at the front end of a student branch to generate an enhanced speech signal; deep features are extracted, and alignment of the deep features in a hidden space is restrained through a knowledge distillation loss function; carrying out dynamic weight fusion on the aligned deep features through a fusion weight; and training a classifier, performing joint optimization in combination with classification loss and knowledge distillation loss, and outputting a voice authentic identification result. According to the method, distribution alignment of pure and noise features is realized through knowledge distillation, forged traces in voice conversion are effectively recognized, and high detection precision is still kept in complex noise and unknown attack scenes.
Owner:ZHEJIANG UNIV

Communication method and system for realizing private call

The invention provides a communication method for realizing a private call, which comprises the following steps that: a wireless silencer preprocesses acquired voice to improve the voice quality; establishing communication connection between the earphone and the wireless silencer; the earphone receives the voice data sent by the wireless silencer; the microphone of the earphone is turned off, and the earphone is switched to an audio output mode; the wireless silencer is worn at the mouth of a speaker so as to collect voice sent by the speaker, and first audio information obtained by converting the voice is sent to the earphone; the earphone transmits the first audio information to a call device; the earphone receives the second audio information from the call device, converts the second audio information into sound and transmits the sound to human ears; in the conversation process, the wireless silencer shields the voice sent by the speaker so as to prevent the voice from being spread to the surrounding environment, and private conversation is achieved.
Owner:ZHONGKE HUAYI (SHENZHEN) INTELLIGENT TECHNOLOGY CO LTD

AI-based emotional text voice conversion method and device

The invention discloses an AI-based emotional text speech conversion method and device. The method comprises the following steps: acquiring speech segments and text records from historical data of a user; performing noise reduction processing and feature extraction according to the voice segments and the text records to obtain voice features; inputting the voice features into a pre-constructed emotional tendency model, and outputting emotional tendency and emotional intensity; according to the emotional tendency and the emotional intensity, adjusting a tone weight, a speech speed and a volume to obtain a speech parameter; extracting new voice features according to the voice parameters to perform scene emotion label matching, and determining voice adjustment parameters through a linear regression model; according to the emotional tendency and the new voice features, generating an emotional type through a pre-established emotional intention classification model, and calculating a voice parameter weight in combination with a pre-established emotional mapping table; and according to the voice adjustment parameter and the voice parameter weight, performing language synthesis to generate personalized voice. According to the method, personalized expression can be accurately generated according to the scene.
Owner:FUJIAN YUANZHI UNIVERSE CULTURE COMMUNICATION CO LTD

Voice conversion method and related equipment

The embodiment of the invention discloses a voice conversion method and related equipment. The related equipment can comprise a voice conversion device, electronic equipment, a computer program product and a computer readable storage medium. After at least one to-be-converted voice and a target timbre identifier corresponding to the to-be-converted voice are acquired, audio content features and acoustic features are extracted from the to-be-converted voice, and the target timbre feature corresponding to the to-be-converted voice is determined based on the target timbre identifier; extracting audio conversion features from the audio content features, extracting rhythm features from the acoustic features, fusing the target timbre features, the audio conversion features and the rhythm features to obtain target audio features, and then generating target voice corresponding to the target timbre identifier based on the target audio features; according to the scheme, the voice conversion accuracy can be improved. The embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi-agent-based old people information acquisition voice interaction system and method

The invention discloses a multi-agent-based old people information acquisition voice interaction system and method. The system comprises a voice acquisition module for forming an original voice signal; the voice recognition module receives an original voice signal, converts the voice into a text through compression and coding processing and back-end voice recognition service, and forms a user text instruction signal; the agent processing module receives a domain specific request signal, executes task processing and information retrieval by calling a knowledge base or a data service interface of a corresponding domain, and generates an integrated result signal containing structured data and a natural language answer; and the bimodal output module receives the integration result signal and synchronously generates a high-readability text display signal and a high-recognizability voice broadcast signal, the text display signal is rendered and presented through an interface suitable for aging, and the voice broadcast signal is output through an audio component. According to the invention, the problems of complex operation, dispersed service and difficult interaction when the old use the digital application can be solved.
Owner:BEIJING FUDI INTELLIGENT TECHNOLOGY CO LTD

Intelligent Technical Protocol Based Approach Leveraging AI-ML to Block Vishing Scammers

Systems and methods detect and prevent vishing attacks through an integrated framework combining SIP header customization, STIR / SHAKEN frameworks, AI / ML analysis, and real-time speech analysis using the Viterbi algorithm. The system begins with call initiation, embedding authentication information in the SIP header. The SIP data is transmitted and verified using STIR / SHAKEN frameworks, ensuring the authenticity of the caller's identity. Verified data is cross-referenced with third-party databases and analyzed by an AI / ML engine to detect anomalies. If potential fraud is detected, the call is blocked, and the customer is notified. Calls that pass initial checks are further analyzed using the Viterbi algorithm, which converts speech to text and identifies suspicious patterns. An anomaly pattern detector processes the converted text to detect vishing indicators, terminating the call if a match is found. This multi-layered approach ensures robust protection against vishing, enhancing the security and reliability of voice communications while safeguarding users from fraud.
Owner:BANK OF AMERICA CORP

Internet-of-things intelligent warehousing service robot control method and system

The invention discloses an Internet of Things intelligent storage service robot control method and system, and relates to the technical field of intelligent robot control, and the method comprises the steps: achieving path planning through combining a first fusion algorithm and a second algorithm; an instruction is generated through a first voice conversion technology, and intelligent interaction is realized in combination with a second language processing technology and a second voice synthesis technology; and comprehensive service management is realized through the first management system. According to the method, through the multi-sensor data fusion technology and the A * algorithm, the robot can perceive the environment in real time, the position of the robot can be accurately positioned, the route is dynamically adjusted according to the environment change, collision and jamming are effectively avoided, and therefore the operation efficiency and accuracy are improved; man-machine interaction is realized through cooperation of ASR, NLP and TTS technologies, and the operation convenience is improved; by constructing the knowledge base, the robot can learn and store knowledge and intelligently answer and make decisions according to the content of the knowledge base, and the intelligent level of homework is improved.
Owner:HUANENG BEIJING CO GENERATION

Audio communication method, audio conversion method, apparatus, electronic device, computer-readable storage medium, and computer program product

PCT designated stageWO2025237010A1Speech analysisComputer hardwareFeature coding
An audio communication method, an audio conversion method, a bitstream processing method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The audio communication method comprises: in response to a first communication request for an audio signal, acquiring, from a plurality of communication modes, a voice transformation mode for the audio signal (101); performing feature coding on the audio signal, so as to obtain a coded feature of the audio signal (102); acquiring a target timbre corresponding to the voice transformation mode, and determining a timbre feature of the target timbre (103); performing timbre conversion on the coded feature on the basis of the timbre feature, so as to obtain a target coded feature (104); and performing signal coding on the target coded feature, so as to obtain a target audio bitstream conforming to the target timbre, and transmitting the target audio bitstream to a decoding terminal (105).
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and apparatus for voice processing, device, storage medium, and program product

A method and apparatus for voice processing, a device, a storage medium, and a program product. The method comprises: acquiring a response text for a question voice of a target user (410); at least on the basis of a speech conversion requirement of the response text, selecting at least one of a first speech synthesis model on a local device and a second speech synthesis model on a server for executing a speech synthesis function on the response text (420); and acquiring a response voice corresponding to the response text, the response voice being generated after the speech synthesis function is executed on the response text by using the selected at least one speech synthesis model (430).
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Virtual digital human real-time interaction system based on AI copies and use method

The invention relates to the technical field of virtual digital humans, in particular to a virtual digital human real-time interaction system based on an AI (artificial intelligence) duplicate and a use method, comprising the following steps: step 1, acquiring original language signal data # imgabs0 # and text input data # imgabs1 # acquired by a microphone, inputting the acquired # imgabs2 # into a voice recognition module, and inputting the acquired # imgabs2 # into a voice recognition module; the voice is converted into a text T through an automatic voice recognition algorithm # imgabs3 # arranged in the voice recognition module, the calculation formula for converting the voice into the text T is # imgabs4 #, and text input data # imgabs5 # and the text T are combined to generate comprehensive text information T0; according to the method, AI technologies such as speech recognition, natural language understanding, semantic embedding, context reasoning and the like are fused, so that the virtual digital human can accurately recognize the intention of the user and naturally generate the response, and particularly has stronger semantic judgment and adjustment capability when processing multiple rounds of dialogues and ambiguous words; therefore, the accuracy and the natural fluency of virtual interaction are greatly improved, and the user experience is better.
Owner:SHENZHEN TONGZHU CLOUD TECHNOLOGY CO LTD

Efficient training and high-expressive-force voice conversion model based on acoustic model and vocoder decoupling architecture

The invention discloses an efficient training and high-expressive-force voice conversion model based on an acoustic model and vocoder decoupling architecture. The efficient training and high-expressive-force voice conversion model comprises an acoustic model and a vocoder. The acoustic model includes a speaker encoder, a content encoder, a normalized stream, a posterior encoder, a Mel decoder, and a discriminator. The method has the advantages that remarkable technical breakthroughs are realized in the aspects of improving the training efficiency of the voice conversion model, sound quality expression, emotion expression, interaction control and the like, a brand new solution is provided for a high-quality and high-controllability voice synthesis system, and the method has good practical value and industrial application prospect.
Owner:HAPPY ELEMENTS TECH (BEIJING) CO LTD

Voice interaction product test method and device based on artificial intelligence, and product

The invention discloses a voice interaction product test method and device based on artificial intelligence and a product, and the method comprises the steps: generating a corresponding question according to a selected question class through a first artificial intelligence large model; the questioning question is converted into questioning voice; the voice interaction product needing to be tested responds to the questioning voice and outputs corresponding answering voice; the answer voice is converted into a character answer; and evaluating the matching accuracy of the character answers and the corresponding questions by using the second artificial intelligence large model, and outputting the accuracy as a test result. The invention discloses a voice interaction product test method and device based on artificial intelligence and a product, and the method comprises the steps: generating a questioning question through a first artificial intelligence large model, evaluating the matching accuracy of a character answer and the corresponding questioning question through a second artificial intelligence large model, and outputting a test result. The system has the characteristics of high automation degree of voice interaction product testing, full testing and accurate testing effect.
Owner:BEIJING POLYTECHNIC

Telecommunications switch-type infrastructure for communications with call router and audio record server for computational source-to-target language conversion via application of artificial intelligence agents

A class-4 telecommunications switch hosts artificial-intelligence agents that convert live speech to text, translate the text between languages, and synthesize natural speech in real time. The switch proxies calls among public trunks, PBX / media gateways, and cloud ACD / CRM services, embedding diacritic-rich transcripts and user-specific language-model personalization. Deployable at the customer edge, in the PSTN core, or as SaaS, the system supports one-to-one, one-to-many, many-to-one, and many-to-many call patterns. FPGA, ASIC, or SoC accelerators minimize latency and bandwidth, cutting capital cost while improving global voice interoperability and cybersecurity.
Owner:CUNNINGHAM CHERYL EE LIN

Infusing knowledge graphs into automatic speech recognition

Systems and methods are described for converting speech to text. In one aspect, a method for converting speech to text includes generating an initial transcript of an audio or video file. At least one named entity may be extracted from the initial transcript using entity recognition. A subset of nodes of a knowledge graph that include the at least one named entity may be selected, where the subset of nodes of the knowledge graph correspond to a set of named entities. The method may further include encoding the set of named entities to generate a set of entity embeddings. The speech in the audio or video file may then be decoded using the set of entity embeddings to produce a final transcript of the audio file.
Owner:AMAZON TECH INC