Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

54 results about "Speech output" patented technology

Like speech input, speech output is a familiar and natural form of communication, so it is also an appropriate complement in a character-based interface. However, speech output also has its liabilities. In some environments, speech output may not be preferred or audible.

Information processing device and information processing method

A user is assisted in performing a voice operation appropriately.A situation determination section determines a situation. A state control section controls a voice command appropriate for the determined situation to put the voice command into a receivable state. For example, the user is informed of what the voice command in a receivable state is, by means of display or voice output. The user can utter a voice command without performing a user action to prevent false recognition, such as the utterance of a wake word. This reduces the troublesomeness and burden of the user.
Owner:SATURN LICENSING LLC

A large model-based speech generation method, device and medium

PendingCN122369426AMultiplexingAcoustics
The application discloses a large model-based speech generation method and device and medium, and belongs to the technical field of data processing. The method comprises the following steps: dividing target language text information into multiple speech generation units through a text division model; in response to the confidence that the target language text information is divided into multiple speech generation units through the text division model being lower than a first threshold value, dividing the target language text information into multiple speech generation units according to text structure features and parameter distribution features; in response to the speech generation units meeting a multiplexing condition, obtaining a first speech output result according to the speech generation units; in response to the speech generation units not meeting the multiplexing condition, calling a speech generation processing object to obtain a second speech output result, and combining the first speech output result and the second speech output result to generate a target speech result. The application can improve the overall processing efficiency, resource utilization rate and result multiplexing capability under a parameterized speech task.
Owner:SHANGHAI ZHONGAN XINKE INFORMATION TECH SERVICES CO LTD

system

We provide the system. [Solution] A recognition means for receiving travel conditions from users via voice, A data collection means for searching and obtaining information on tourist resources and accommodation resources via communication means based on the travel conditions of collected users, A schedule generation method that automatically generates and sequences travel itineraries using an AI model based on acquired information, A means for automatically making reservations for tourist attractions and accommodations according to the generated travel itinerary, A display means that notifies the user of automatically generated travel itinerary and reservation information through voice output and visual display, A system that includes this.
Owner:SOFTBANK GROUP CORP

An intelligent voice alarm system and method suitable for a power monitoring system

The application discloses an intelligent voice alarm system and method suitable for a power monitoring system, which comprises multiple functions of alarm event level definition, voice broadcast mode, specified area alarm mode, temporary shielding object, alarm event confirmation, mute function, voice broadcast sequence, voice text replacement, voice broadcast configuration, voice file configuration and system self-defined configuration; the system comprises an alarm event management module, a broadcast control module, a voice output module and a configuration management module; the alarm event management module generates an alarm to be broadcast according to alarm event level definition information and stores the alarm into a cache queue; the broadcast control module generates a voice broadcast request according to the voice broadcast sequence and the voice broadcast mode; and the voice output module receives the request and performs voice synthesis and broadcast. The application realizes fine control of the whole process of the alarm event through a modular architecture, supports deployment of a Windows system and a Linux system, and has high flexibility, expandability and cross-platform compatibility.
Owner:CHINA YANGTZE POWER

Identification and feedback system for communication and repair strategies

PendingUS20260171107A1Speech recognitionSets using external connectionHuman–computer interactionSpeech sound
Embodiments herein relate to ear-wearable device systems. In an embodiment, a method executed in a processor of a feedback system for providing information related to a user's communication strategies. The feedback system can include a microphone and a memory storage. The method can include monitoring a voice output of the user with the microphone. The method can include processing the voice output of the user by at least one of: counting a first number of instances in which the user implements a positive communication repair strategy; and counting a second number of instances in which the user implements a negative communication repair strategy. The method can include generating feedback indicative of the first number of instances, the second number of instances, or both the first number of instances and the second number of instances. Other embodiments are also included herein.
Owner:STARKEY LABORATORIES INC

system

Provide a system. 【Solution means】 Means for acquiring voice information, Means for converting the voice information into text information, Means for analyzing the text information and generating an instruction to acquire visual information based on its content, Means for acquiring visual information from a display, Means for extracting necessary information by analyzing the visual information, Means for outputting the extracted information as voice, Installed in an autonomous mobile body within a home and means for assisting in schedule management, A system including.
Owner:SOFTBANK GROUP CORP

Physical AI device and space operation support method for fully local space use

PendingJP2026110667ASpace roboticsSpace operations
In the operation of spacecraft, lunar rover, space habitation modules, and space robots, equipment status, work environment, worker status, and communication status are handled in a dispersed manner, making it difficult to provide consistent support in the event of communication delays or interruptions. [Solution] A fully local physical AI device for space that acquires equipment status data, environmental images, worker images / voices, and communication status data, locally generates, stores, and references mission-related information including observation events, target objects, work processes, safety constraints, estimated states, recommended actions, and candidate commands to be executed, and generates and displays natural language messages and machine-readable commands, and outputs them via voice output or to a control interface, without relying on a ground-based cloud.
Owner:松尾 信慎

A translation interaction device and electronic kit for use with smart wearable devices

ActiveCN224436892USpeech translationInteraction device
This utility model relates to the field of smart device technology, and more particularly to a translation interaction device and electronic kit for use with smart wearable devices. The translation interaction device includes a main body for use with an external smart wearable device. The main body includes a housing, a main control module disposed within the housing, and a display screen disposed on the surface of the housing. The main control module is communicatively connected to both the display screen and the external smart wearable device. The main control module includes a voice translation module, which includes a language receiving unit, a voice translation unit, and a voice output unit. The main control module also includes one or any combination of a recording module, a text translation module, an audio / video playback module, and a camera module. This utility model, through collaboration with smart wearable devices, can meet users' translation needs in different language environments, thereby expanding the functionality of external smart wearable devices.
Owner:深圳目渡科技有限公司

A voice service interaction method and device, a storage medium and an electronic device

ActiveCN121354564BSemantic understanding synchronizationStable multiple subtasksSemantic analysisSpeech recognitionSpeech inputHuman–computer interaction
The specification discloses a voice service interaction method and device, a storage medium and an electronic device, wherein the method comprises: in response to a service interaction request, generating a standardized instruction, the service interaction request containing unstructured voice data, and the standardized instruction being a standardized natural language text instruction; inputting the standardized instruction into a large language model to obtain composite feedback information; feeding back the composite feedback information to a client so that the client synchronously outputs voice feedback and renders a corresponding interaction component page; wherein the large language model contains a master sub-model, a fusion sub-model and a preset function component library, and the preset function component library contains a plurality of execution sub-models. The specification realizes complete link control from voice input to execution feedback, so that the model output can meet semantic understanding and drive business actions, and ensures that voice output is synchronized with client interface rendering.
Owner:CHONGQING ANT CONSUMER FINANCE CO LTD

Teleconference systems, communication terminals, teleconference methods, and computer-readable media

A telephone conferencing system capable of conducting telephone conferences smoothly is provided. According to the invention, a speech determination unit (2) determines whether the speech of each participant in a telephone conference is a statement or an interruption. A speech output control unit (4) performs control such that the speeches of the multiple participants are output to the communication terminals of the multiple participants respectively. When another participant speaks while one of the multiple participants is speaking, the speech output control unit (4) performs control to suppress the output of the other participant's speech. A counting unit (6) counts the number of speech conflicts for each participant. A count display control unit (8) performs control such that a display related to the count is displayed on the communication terminals of the multiple participants.
Owner:NEC PLATFROMS LTD

system

PendingJP2026101392ATherapiesInstrumentsHealth related informationBody shape
Provide a system. 【Solution means】An acquisition means for acquiring exercise data, An acquisition means for acquiring exercise data, A body shape analysis means for photographing and analyzing body shape data, A nutrition analysis means for analyzing diet data and determining the calorie and nutrient balance, A proposal generation means for generating a personalized exercise program and diet proposal based on these data, A notification means for notifying the user of the generated exercise program and diet proposal, An audio output means for providing health-related information via audio output, Means for providing a consumer health management robot for acquiring and analyzing data in a residential environment, A system including.
Owner:SOFTBANK GROUP CORP

Streaming audio generation system

Providing a user-friendly streaming audio generation system. [Solution] The streaming speech generation system 10 includes a speech recognition unit 20 that inputs the speaker's voice as text data to a generation AI 30, a generation AI 30 that generates a response sentence to the voice recognized by the speech recognition unit 20 as text data, a splitting unit 35 that receives the response sentence (text data) output by the generation AI 30 as streaming data, and sequentially divides it into short sentences and outputs them in real time, and a speech synthesis unit 40 that outputs the short sentences output by the splitting unit 35 as speech.
Owner:YANO YUAI OFFICE CO LTD +1

Digital human live broadcast interaction method and system, electronic device and storage medium

The application relates to a digital person live broadcast interaction method and system, electronic equipment and a storage medium. The method comprises the following steps: in the process of digital person live broadcast, extracting a plurality of historical comment information in a preset time period from a historical comment pool, and extracting a plurality of commodity portrait information corresponding to commodities from a commodity knowledge base; the commodity portrait information comprises commodity explanation rhetoric corresponding to the commodities; through a large language model, context understanding and context adaptation are performed on the plurality of historical comment information and the plurality of commodity portrait information, and a target intention of an audience is recognized; target commodity portrait information matched with the target intention is determined from the plurality of commodity portrait information, so that a target commodity explanation rhetoric is extracted from the target commodity portrait information; and the target commodity explanation rhetoric is used to generate and voice output a target reply text. The scheme provided by the application can meet the demand of the audience for specific commodity links, sizes and live broadcast room specific rhetoric.
Owner:HANGZHOU TEKAN TECHNOLOGY CO LTD

Adaptive text-to-speech output

The present invention relates to adaptive text-to-speech output. In some implementations, one or more computers determine a language proficiency of a user of a client device. The one or more computers then determine a text segment for output by a text-to-speech module based on the determined language proficiency of the user. After determining the text segment for output, the one or more computers generate audio data of a synthesized utterance that includes the text segment. The audio data of the synthesized utterance that includes the text segment is then provided to the client device for output. An improved user interface is provided through better text-to-speech conversion.
Owner:GOOGLE LLC

system

Provide a system. 【Solution means】 An input device that receives voice data or string data, A conversion device that converts the input voice data into string data, A natural language processing device that analyzes the converted string data to identify the intention, A search device that accesses the information system of a medical institution to check available appointment time slots and locations, A user interface that presents the search results and induces a selection, A notification device that finalizes and notifies an appointment based on the selected option, A voice output device that presents the generated options in voice, A system including the above.
Owner:SOFTBANK GROUP CORP

Voice response system

In voice response systems, we provide a voice response system that is more user-friendly for the user. [Solution] A voice response device that provides voice responses to input character information comprises a response acquisition means for acquiring multiple different responses to the character information, and a voice output means for outputting each of the multiple different responses in a different tone of voice. With such a voice response device, since multiple responses can be output in different tones of voice, even when a single solution to a character piece of information cannot be uniquely identified, different solutions can be output in different tones of voice in an easily understandable way to the user.
Owner:CASE CHARTER CO LTD

system

We provide the system. [Solution] A means for receiving image data captured by a user and analyzing the image data to identify item information, A means for converting voice instructions from the user into text data, and for analyzing the text data to understand the user's intent, A means for generating a process proposed based on the aforementioned item information and the user's intent, and providing the process to the user, A method for analyzing user voice input using artificial intelligence and outputting cooking instructions via voice in real time, A system that includes this.
Owner:SOFTBANK GROUP CORP

Voice modification

ActiveUS12670918B2TimbreSpeech sound
A computing system that receives an audio waveform representing speech from an individual and produces as output a modified version of the audio waveform that maintains the speaker's speech characteristics as well as prosody for specific utterances (e.g., voice timbre, intonation, timing, intensity). The system uses a bottleneck-based autoencoder with speech spectrograms as input and output. To produce the output audio waveform, the system includes a reconstruction error-based loss function with two additional loss functions. The second loss function is speaker “real vs fake” discriminator that penalizes for the output not sounding like the speaker. The third loss function is a speech intelligibility scorer that penalizes the output for speech that is difficult for the target population to understand. The produced modified audio waveform is an enhanced speech output that delivers speech m a target accent without sacrificing the personality of the speaker.
Owner:SRI INTERNATIONAL

An ai-driven dynamic audio fence control method

PendingCN122290622ANoiseEngineering
This invention discloses an AI-driven dynamic audio fence control method, specifically including the following steps: Step 1: Signal acquisition and preprocessing using a beamforming microphone array and a DSP processor; Step 2: AI feature extraction and speech noise classification of standardized multi-channel digital audio signals; Step 3: MVDR dynamic beamforming processing based on speech noise classification labels and noise feature masks; Step 4: 3D audio rendering of the beamformed single-channel target speech signal; Step 5: Effect monitoring and parameter iteration based on the target speech signal and feedback signals from the audio output module; Step 6: Pure target speech output playback based on the 3D rendered time-domain target speech signal. The advantages of this invention are: ensuring a continuous improvement in the target speech signal-to-noise ratio (SNR), ultimately transmitting pure target speech to the speaker, forming a personal sound bubble-like virtual audio fence around the target object.
Owner:JIANGSU COLLEGE OF INFORMATION TECH

Guidance device, method, apparatus and program product for energy storage system testing

The application is suitable for the technical field of energy storage system testing, and provides a guiding device, method, equipment and program product for energy storage system testing. The guiding device comprises a main controller, a memory, a voice output module, a human-computer interaction module and a communication interface module. The memory pre-stores a test procedure program containing a test step sequence and a parameter standard. The main controller controls the voice output module to output operation instructions to guide the operator according to the test procedure program. After receiving the operation confirmation instruction fed back by the human-computer interaction module, the main controller sends test instructions to external test equipment through the communication interface module and receives test data. Then, the test data is compared with the parameter standard to determine whether to control the test procedure to proceed or issue a warning. The application can effectively improve the test efficiency and consistency, reduce operation errors and enhance the safety of the test process by solidifying the test procedure and realizing human-computer cooperation.
Owner:SHAANXI GREEN ENERGY ELECTRONIC TECH CO LTD

system

PendingJP2026105500APattern recognitionMedicine
Provide a system. 【Solution means】 Means for acquiring position information, Means for analyzing the acquired position information, Means for modeling the normal behavior range based on the analysis, Means for detecting an abnormality by comparison with the normal behavior range, Means for generating an alarm when an abnormality is detected, Means for notifying the user of the alarm, Means for communicating with an external organization based on the alarm, Means for monitoring a moving object in a home environment, Means for controlling an image recording device when an abnormality is detected, Means for alerting with a voice output means when an abnormality occurs, A system including the above.
Owner:SOFTBANK GROUP CORP

Voice interaction system interaction interrupting method and device, medium and electronic equipment

The embodiments of the present disclosure relate to the technical field of human-computer interaction, and provide an interaction interrupting method and device of a voice interaction system, a medium and an electronic device. The method comprises: in response to detecting a voice input event, determining a current interaction state of the voice interaction system; based on the current interaction state, determining a target interrupting strategy of the voice interaction system, wherein the voice interaction system comprises a plurality of interaction states, the plurality of interaction states correspond to a plurality of interrupting strategies, the interrupting strategy is used to indicate an interrupting manner of a current action of the voice interaction system, the current action comprises a current voice output action and / or a current task execution action; and based on the target interrupting strategy, interrupting the current action of the voice interaction system. The embodiments of the present disclosure can adopt different interrupting strategies for different states of the system when a new voice instruction is received, improve the flexibility of the interrupting operation, and improve the user experience.
Owner:XG TECHNOLOGIES PTE LTD

Speech rehabilitation training control method based on voice initiation state evolution

The present application relates to the technical field of language rehabilitation, and particularly relates to a speech rehabilitation training adaptive control method and device based on vocalization starting state evolution and a storage medium. The speech rehabilitation training control method based on vocalization starting state evolution comprises the following steps: in the speech rehabilitation training process, a vocalization feature signal is acquired, the vocalization feature signal is used to represent state changes of at least two action links in a vocalization starting action chain; based on the vocalization feature signal, phase states of the at least two action links in the vocalization starting action chain are determined, the phase states are used to represent the advancing degree of the action links in the vocalization starting process; according to the phase states of the action links, a cooperative advancing state of the vocalization starting action chain is determined; and according to the cooperative advancing state, a speech rehabilitation training mode is adjusted, the speech rehabilitation training mode is defined by at least one non-speech output guide channel and / or participation weight configuration of at least one speech output guide channel.
Owner:THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV

Speech translation processing apparatus

A speech translation processing apparatus including a speech inputter and a speech outputter operated in cooperation with a wearable speech input / output apparatus worn on a user, includes a translation speech acquirer acquiring translation speech in a user language or the other language that is translated and generated on the basis of a spoken speech in the user language or the other language input through the wearable speech input / output apparatus or the speech inputter, and a translation speech output controller performing control such that the acquired translation speech is output from at least one of the speech outputter and the wearable speech input / output apparatus in an output mode according to a translation condition. According to such a configuration, it is possible to provide a user-friendly translation system.
Owner:MAXELL LTD

Information provision system, information provision method, and storage medium

An information provision system includes: a user state recognition unit that recognizes a state of a user in a mobile body; a user emotion estimation unit that estimates an emotional level of the user based on a predetermined index value, based on a recognition result of the state of the user; and an agent control unit that provides, using an anthropomorphic agent, the user with an HMI that performs speech output in a mode corresponding to an emotional level of the agent based on the predetermined index value. The agent control unit controls the emotional level of the agent to be a target level for changing the emotional level of the user to fall outside of a predetermined determination range, when the emotional level of the user estimated by the user emotion estimation unit falls within the predetermined determination range.
Owner:HONDA MOTOR CO LTD

Product processing equipment

The present invention provides a product processing device that allows the operator to confirm product information that is read aloud while viewing the displayed product information. [Solution] The product processing device 1 is configured such that when product information for a product is displayed on the display unit 5 by a product retrieval operation, the product information is read aloud by the voice output device 15A. The product processing device 1 comprises a display unit 5 that displays product information and a control unit 10 that controls the voice output device 15A. If pop-up screens PS1, PS2, and PS3 are displayed on the display screen SC1 of the product information displayed on the display unit 5 in a way that overlaps them, the control unit 10 waits for the pop-up screens PS1, PS2, and PS3 to disappear before having the voice output device 15A read aloud the product information.
Owner:ISHIDA CO LTD

Intelligent interactive general-purpose machine patient for war injury and trauma training

PendingCN122266228ACosmonautic condition simulationsEducational modelsSkeletal movementQuery string
This invention provides a general-purpose robotic wounded soldier for combat trauma training, belonging to the technical field of combat trauma training models. It includes: a robotic wounded soldier model; a human-computer interaction system for presetting initial combat trauma states; a wound simulation control module that generates query strings containing location, type, severity, and duration based on the preset states, calls a wound simulation knowledge graph to match corresponding wound state nodes, thereby obtaining descriptions of skin appearance, skeletal movements, language behavior, and physiological data, and uses this to drive the model to perform real-time morphological and appearance simulation, real-time voice output, and physiological data numerical output; and a training interaction module that collects real-time data on treatment actions and on-site environmental conditions, inputs this data into a state transition decision tree model integrating the influence of environmental and treatment conditions, and generates new wound state nodes that conform to the laws of medical evolution. This invention can accurately and in real-time simulate the physiological reactions and wound changes of wounded soldiers based on dynamic battlefield environments and treatment operations.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY