Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

195 results about "Speech output" patented technology

Like speech input, speech output is a familiar and natural form of communication, so it is also an appropriate complement in a character-based interface. However, speech output also has its liabilities. In some environments, speech output may not be preferred or audible.

Voice generation method and device based on pseudo-autoregression modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice generation method, device and equipment based on pseudo-autoregression modeling and a medium, and the method comprises the steps: obtaining a training sample containing a text sequence, a prompt voice segment and a target semantic token sequence; performing continuous fragment mask training on the text-to-semantic model to obtain a pseudo-autoregression trained text-to-semantic model; generating candidate speech output by using the text-to-semantic model and the initial semantic-to-acoustic model which are subjected to pseudo-autoregression training, and constructing a preference data pair; updating the semantics-to-acoustics model based on the preference data pair to obtain a preference optimized semantics-to-acoustics model; and generating target voice output based on the target text and the target prompt voice. According to the method, the time sequence modeling capability of the model is enhanced through pseudo-autoregression training, and the voice generation quality is directly optimized through the preference data pair, so that the voice alignment precision and the subjective listening feeling performance are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Sign language-voice conversion system

The invention discloses a sign language-voice conversion system, and belongs to the technical field of auxiliary communication and wearable computing. The system comprises a wearable myoelectricity acquisition module used for acquiring double-arm myoelectricity signals when a user executes sign language; the mobile terminal module is wirelessly connected with the acquisition module and is used for receiving and preprocessing the signal and uploading the signal; the cloud processing module is used for receiving the signal, converting the signal into text information through a sign language recognition model, and further calling a voice synthesis service to convert the text into voice data; and the wearable audio output module is used for receiving and playing the voice data. Through an innovative end-to-end hardware system architecture, natural, accurate and real-time translation and voice output of sign language gestures are realized, communication barriers between hearing-impaired people and healthy hearing people are effectively solved, and the system has the advantages of flexible deployment, user friendliness and privacy protection.
Owner:宋飞 +1

Sign language recognition method based on double-arm electromyographic signals

The invention discloses a sign language recognition method based on double-arm electromyographic signals, and belongs to the technical field of human-computer interaction and biological signal processing. The method comprises the steps that multi-channel electromyographic signals generated when sign language gestures are executed are synchronously collected through electromyographic arm rings worn on the left forearm and the right forearm of a user; performing preprocessing and feature extraction on the signal to obtain a time sequence feature sequence of left and right arms; the time sequence feature sequence is input into a pre-trained two-arm collaborative recognition model, and the model outputs a sign language gesture recognition result by fusing the spatial-temporal features of the left arm and the right arm; and finally, the recognized text information is converted into voice to be output. According to the method, the cooperation and time sequence relation of the double-arm electromyographic signals is creatively utilized, the problems that a traditional visual recognition method is greatly interfered by the environment, privacy is invaded, and double-hand linkage complex gestures cannot be effectively analyzed through single-arm electromyographic recognition are solved, and natural, accurate and real-time recognition and translation of the double-hand sign language gestures are achieved.
Owner:宋飞 +1

Voice interaction control method and device, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice interaction control method, device, equipment and medium, and the method comprises the steps: obtaining multi-mode perception data, and constructing a dynamic scene model; session feature parameters are extracted and input into a decision model to generate a strategy parameter set; a dialogue process and expected voice output are generated according to the strategy parameter set and voice characteristics of the voice interaction device; controlling the voice interaction device to execute voice output and collect actual voice output; monitoring an error between the actual voice output and the expected voice output, and performing compensation adjustment; and collecting feedback data after the voice interaction session is ended, and optimizing the model and the control parameters according to the feedback data. According to the method, closed-loop control of perception, decision, output and feedback is established, so that the voice interaction equipment can dynamically adjust voice output according to a real-time environment and user characteristics and continuously perform self-learning optimization, and the accuracy and naturalness of voice interaction are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Unmanned aerial vehicle natural language control and interaction method and system based on multi-modal large model

The invention discloses an unmanned aerial vehicle natural language control and interaction method and system based on a multi-mode large model, and belongs to the field of unmanned aerial vehicle interaction.The method comprises the steps that a current environment image and a user voice instruction of an unmanned aerial vehicle are obtained and converted into a text; performing intention classification on the text instruction information, and judging whether the current task is a flight control task or a question and answer interaction task; if the task is a flight control task, inputting the environment image and the text instruction into a multi-mode large model for vision-language joint reasoning, and generating a structured flight control instruction; if the task is a question-answer interaction task, performing knowledge retrieval by using the large model to generate question-answer information and outputting voice; adjusting the flight state of the unmanned aerial vehicle according to the structured instruction; and after the task is completed, the unmanned aerial vehicle is controlled to enter an obstacle avoidance self-stabilization mode and waits for a next instruction. According to the method, the mapping from the unstructured natural language to the bottom layer action of the unmanned aerial vehicle is realized, and the intelligent level of man-machine interaction and the flight safety in a complex environment are effectively improved.
Owner:NANJING FANMEILI ROBOT TECH CO LTD

Notification control apparatus, notification control method, and storage medium

A smartphone includes: a first biological information acquisition unit that acquires first biological information being biological information about a driver while driving a vehicle; a driving behavior acquisition unit that acquires driving behavior information indicating driving behavior of the driver; a danger level calculation unit that calculates a driving danger level indicating a danger degree while the driver is driving the vehicle, based on the first biological information and the driving behavior information; a speech output unit that performs speech output for the driver, in accordance with a degree of change in the driving danger level; a response acquisition unit that acquires response information indicating a response from the driver to the speech output; and an output control unit that controls the speech output unit based on the response information.
Owner:HONDA MOTOR CO LTD

Voice real-time question and answer processing method based on large model and domain knowledge base

The invention discloses a voice real-time question and answer processing method based on a large model and a domain knowledge base, which belongs to the technical field of computer data processing, and comprises the following steps: acquiring a voice input stream of a user, carrying out intelligent sound wave deconstruction on the voice input stream, generating a text stream, and carrying out dependency syntactic analysis and semantic role labeling; generating intention information and key information, dynamically accessing a domain knowledge base, executing predictive loading, generating a preloaded data subset, performing fusion processing by combining the intention information, the key information and the preloaded data subset, generating a fusion result, performing accuracy verification and correction, generating a natural language answer, and converting the natural language answer into voice output. The technical scheme of combining knowledge predictive loading driven by intention analysis, context fusion and answer traceability correction is adopted, and low-delay response, high-precision intention understanding and high-factuality answer of voice questions and answers can be achieved.
Owner:STATE GRID SHANDONG ELECTRIC POWER CO

Information processing system

The invention provides an information processing system. The information processing system comprises an image acquisition device used for acquiring front image data; the analysis server device is used for analyzing the collected image data; the voice output device is used for carrying out voice prompt on visually impaired people according to the analyzed image data; and the display device is used for performing augmented reality display according to the analyzed image data.
Owner:SOFTBANK GROUP CORP

Induced dialogue interaction system based on psychological intervention artificial intelligence model

The invention is applied to the technical field of dialogue interaction systems, and particularly discloses an induced dialogue interaction system based on a psychological intervention artificial intelligence model, which comprises a voice recognition module, a deep language module, a voice output module, a closed feedback module and an abnormal condition processing module, the voice recognition module is connected with the deep language module through wireless signals, and the deep language module is connected with the voice output module through wireless signals. According to the induced dialogue interaction system based on the psychological intervention artificial intelligence model, psychological theories such as a social emotion selection theory, a communication adaptation theory and the like are introduced in the system, and an old people age layering optimization deep language model is combined, so that the recognition and fitness reply ability of the model to common emotion expression of old people is enhanced; meanwhile, by means of intervention boundary setting, psychological intervention scenes of the old people are accurately adapted, and intervention safety and effectiveness are guaranteed.
Owner:SUZHOU UNIV

Voice interaction massage device

ActiveDE202025107998U1Gum massageVibration massagePenisSpeech sounds
Voice-interaction massage device for massaging at least one penis, comprising the following: a housing comprising a first massage chamber and a second massage chamber, each of the first massage chamber and the second massage chamber being configured to accommodate at least part of the penis; a massage module incorporated in the housing, wherein the massage module is configured to massage at least a part of the at least one penis that is inserted into at least one of the two massage chambers; a loudspeaker; and a first sensor configured to detect movements of the penis in the first massage chamber; a control unit that is electrically connected to the speaker and the first sensor; the control unit, as soon as the first sensor detects movements of the penis in the first massage chamber, activates the speaker to output an initial voice message.
Owner:SHENZHEN SHECLONE TECHNOLOGY CO LTD

Information processing device and information processing method

A user is assisted in performing a voice operation appropriately.A situation determination section determines a situation. A state control section controls a voice command appropriate for the determined situation to put the voice command into a receivable state. For example, the user is informed of what the voice command in a receivable state is, by means of display or voice output. The user can utter a voice command without performing a user action to prevent false recognition, such as the utterance of a wake word. This reduces the troublesomeness and burden of the user.
Owner:SATURN LICENSING LLC

System

A system is provided.SOLUTION: A learning system comprising: means for receiving an input of an alphabet; means for acquiring an English word based on the input alphabet; means for outputting the acquired English word by voice; and means for displaying an illustration corresponding to the English word.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

AI accompanying toy and method based on emotional acoustic wave algorithm

The invention discloses an AI accompanying toy and method based on an emotional acoustic wave algorithm, and relates to the technical field of man-machine interaction, a core processing layer establishes an independent virtual heart rate system for each AI role, and generates a personalized basic heart rate and an emotional intimacy initial value according to role attributes; dialogue content and AI emotional state are analyzed in real time based on data acquired by the input layer, and emotion type and intensity are identified; calculating a target heart rate and a spectrum modulation parameter according to the emotional state; dynamically adjusting parameters according to the interaction time, the dialogue depth and the emotion intimacy; the heart rate rhythm and spectrum modulation are applied to AI voice in real time; and the output layer presents the fused information of the emotional sound wave and the emotional state in a graph or color mode. Therefore, personification and life feeling of the AI role are greatly enhanced, man-machine interaction is upgraded from simple information exchange to accompanying experience with emotional depth and physiological resonance, and interaction naturalness and emotional attachment of the user are remarkably improved.
Owner:BEIJING LANGZHIWAN INTELLIGENT TECHNOLOGY CO LTD

Audio-visual multi-mode speech recognition method, model training method and electronic equipment

The invention provides an audio-visual multi-mode speech recognition method, a model training method and electronic equipment. The voice recognition method comprises the following steps: acquiring first video data and first audio data; a target model for audio-visual multi-modal speech recognition is obtained, the target model comprising a first target network, in which the first target network introduces a multi-scale packet sparse constraint comprising a plurality of norm constraints and a plurality of packets on the basis of a deep belief network (DBN), and the multi-scale packet sparse constraint comprises a plurality of norm constraints and a plurality of packets; the first target network divides hidden layer units of a restricted Boltzmann machine (RBM) into non-overlapping groups and performs feature extraction of the first voice by using a non-overlapping group lasso method; and inputting the first video data and the first audio data into a target model to obtain a voice recognition result of the first voice output by the target model. By introducing a multi-scale grouping sparse constraint and multi-head attention fusion mechanism, the accuracy and robustness of speech recognition in a complex scene are improved.
Owner:CHINA MOBILE INTERNET CO LTD +1

Voice large model question answering method based on retrieval enhancement generation

The invention belongs to the technical field of voice interaction and natural language processing crossing, and discloses a voice large model question answering method based on retrieval enhancement generation. Through ASR error correction and scenarized retrieval enhancement, the factual accuracy of answers is improved compared with that of a traditional scheme, the intention understanding accuracy of multiple rounds of dialogues is improved, and the problem of factual error and context disjunction in voice questions and answers is effectively solved; through hard constraint monitoring and dynamic resource scheduling, response delay is stably controlled within a set threshold value, which is reduced compared with an existing fusion scheme, and the real-time requirement of voice interaction is met; preference customization and hard constraint configuration of different scenes are supported, the adaptation score of the scenes such as medical treatment, vehicle-mounted and home furnishing is improved compared with a traditional scheme, and the differentiation requirements of professional users and common users can be met; through collaborative optimization of text generation and voice output, the naturalness score of a voice answer is improved, the problem that the text is smooth but the voice is stiff is reduced, and the interaction willingness of a user is improved.
Owner:NORTHEASTERN UNIV CHINA

A large model-based speech generation method, device and medium

PendingCN122369426AMultiplexingAcoustics
The application discloses a large model-based speech generation method and device and medium, and belongs to the technical field of data processing. The method comprises the following steps: dividing target language text information into multiple speech generation units through a text division model; in response to the confidence that the target language text information is divided into multiple speech generation units through the text division model being lower than a first threshold value, dividing the target language text information into multiple speech generation units according to text structure features and parameter distribution features; in response to the speech generation units meeting a multiplexing condition, obtaining a first speech output result according to the speech generation units; in response to the speech generation units not meeting the multiplexing condition, calling a speech generation processing object to obtain a second speech output result, and combining the first speech output result and the second speech output result to generate a target speech result. The application can improve the overall processing efficiency, resource utilization rate and result multiplexing capability under a parameterized speech task.
Owner:SHANGHAI ZHONGAN XINKE INFORMATION TECH SERVICES CO LTD

System

A system is provided.SOLUTION: A system comprising: means for collecting a moving image, a photo, and a chat of a deceased person; analyzing means for extracting a facial feature, a vocal feature, and a language pattern of the deceased person; generation and AI means for generating a digital clone of the deceased person based on the extracted features; means for providing an interface for a user to interact with the digital clone; and means for displaying or outputting by voice a reply generated by the generation and AI means to the user.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

system

We provide the system. [Solution] A recognition means for receiving travel conditions from users via voice, A data collection means for searching and obtaining information on tourist resources and accommodation resources via communication means based on the travel conditions of collected users, A schedule generation method that automatically generates and sequences travel itineraries using an AI model based on acquired information, A means for automatically making reservations for tourist attractions and accommodations according to the generated travel itinerary, A display means that notifies the user of automatically generated travel itinerary and reservation information through voice output and visual display, A system that includes this.
Owner:SOFTBANK GROUP CORP

Multimedia conference room sound system based on artificial intelligence

The present application relates to conference room sound control technical field, especially in kind based on artificial intelligence's multimedia conference room sound system, through personnel positioning module real-time acquisition of the image position and head posture of the participant, the automatic establishment of the space mapping relationship of personnel and channel is combined with microphone layout, the dynamic binding of microphone channel is realized. Voice activity detection based on audio data automatically identifies the main speaking channel, and differentiates the control of the channel gain through the sound output module, effectively suppresses the background noise of the non-speaking microphone. By extracting the behavior characteristics and interaction intention of the main speaker, the behavior characteristics and interaction intention are jointly modeled based on the reinforcement learning model, the prediction of the next speaker and the dynamic update of the main channel are realized, and the strategy parameters are continuously optimized based on the speech feedback. Reduce manual operation, improve the clarity of voice output and the natural fluency of conference interaction.
Owner:GUANGDONG RUIZHAO AUDIO EQUIPMENT CO LTD

Mind mapping sounding device for cardiac pacemaker implanted patient

ActiveCN224005568UElectrical appliancesMemory retentionHeart pacemakers
The utility model discloses a mind mapping sounding device for a cardiac pacemaker implanted patient, relates to the technical field of medical auxiliary equipment, and aims to solve the problems that in the prior art, traditional patient education depends on written materials or oral explanation, information is fragmented, the memory retention rate of the patient is low, and postoperative follow-up visit content is complex and is difficult for the patient to remember. Comprising a flexible map guide board and a control box, the upper end of the flexible map guide board is provided with a middle node, a first node and a second node, one side of the first node and one side of the second node are provided with a touch sensing area, the control box is provided with a prompt lamp and a voice interaction module, and the voice interaction module comprises a voice output module and a voice input module. A master control module, a storage module, a data interaction module and a follow-up time management module are further installed in the control box. Operation knowledge, postoperative care, follow-up visit nodes and the like are integrated through mapping, the cognitive level and follow-up visit compliance of a patient are remarkably improved, and the occurrence rate of complications is reduced.
Owner:SHANGHAI CHEST HOSPITAL

Information processing device, information processing method, and program

[Problem] To realize a technique capable of improving quality of life for a subject such as a patient who is impaired in vocalization. [Solution] An information processing device 1 includes a moving image data acquisition unit 11a, a preprocessing unit 11b, a feature amount acquisition unit 11c, a learning model generation unit 11d, an utterance content estimation unit 11e, and a substitute voice output unit 11f. The moving image data acquisition unit 11a acquires a moving image including lips of a subject. The preprocessing unit 11b extracts a lip portion in the moving image. The feature amount acquisition unit 11c acquires a feature amount of the extracted lip portion. The utterance content estimation unit 11e estimates an utterance content of the subject corresponding to the lip portion extracted by the preprocessing unit 11b, on the basis of a learning model that has been machine-learned in advance by associating a feature amount of a lip portion in a moving image for machine learning including the movement of lips during utterance with the utterance content uttered by the lip portion.
Owner:KEIO UNIV

Teaching system and teaching method

The invention relates to a teaching system and a teaching method. The teaching system comprises a voice acquisition module used for acquiring inquiry voice of a user; the main control module is connected with the voice acquisition module and is used for converting the inquiry voice into an inquiry text; the question and answer output module is connected with the main control module and used for outputting an answer text according to the inquiry text, and the main control module is further used for outputting an audio signal according to the answer text; and the sound amplification cabin module is arranged in the simulation patient robot, is connected with the main control module and is used for playing the audio signals. The system collects inquiry voice through the voice collection module and transmits the inquiry voice to the main control module, the question and answer output module generates an answer text after text conversion, and the answer text is played through the sound amplification cabin module of the simulation patient robot after voice synthesis, so that the space reality sense and immersion sense of voice output can be improved; and the user can carry out inquiry training close to a real clinical scene.
Owner:SHENZHEN UNIV

An intelligent voice alarm system and method suitable for a power monitoring system

The application discloses an intelligent voice alarm system and method suitable for a power monitoring system, which comprises multiple functions of alarm event level definition, voice broadcast mode, specified area alarm mode, temporary shielding object, alarm event confirmation, mute function, voice broadcast sequence, voice text replacement, voice broadcast configuration, voice file configuration and system self-defined configuration; the system comprises an alarm event management module, a broadcast control module, a voice output module and a configuration management module; the alarm event management module generates an alarm to be broadcast according to alarm event level definition information and stores the alarm into a cache queue; the broadcast control module generates a voice broadcast request according to the voice broadcast sequence and the voice broadcast mode; and the voice output module receives the request and performs voice synthesis and broadcast. The application realizes fine control of the whole process of the alarm event through a modular architecture, supports deployment of a Windows system and a Linux system, and has high flexibility, expandability and cross-platform compatibility.
Owner:CHINA YANGTZE POWER

Multi-modal video knowledge query system combined with agent task intention

The invention provides a multi-modal video knowledge query system combined with agent task intention, and relates to the technical field of multi-modal fusion, and the system comprises a feature extraction module which is used for receiving video clips and text problems, and carrying out the processing of a video frame sequence through a lightweight three-dimensional convolutional network to generate a video time sequence feature vector; performing semantic analysis on the text question to generate a text semantic feature vector representing question semantics; and the mapping module is used for analyzing the corresponding relationship between the space coordinates of the video key frame and the text semantics according to the relevance between the video time sequence feature vector and the text semantic feature vector, constructing a dynamic space semantic mapping representation, and determining an initial space reference position based on the mapping representation. Through multi-modal feature fusion, dynamic space-time constraint and external knowledge integration, the task intention of the agent is captured, the comprehensive answer fitting the scene is generated and output in the form of voice, and the accuracy, efficiency and interaction convenience of multi-modal video knowledge query are improved.
Owner:HANGZHOU LIGHT ELEPHANT TECH CO LTD +1

Registration apparatus, registration method, and non-transitory storage medium with audio output for products imaged and sensed by any of size and kind

The invention addresses the problem of improving labor of a registration operation of a product and enabling a checkout operator to recognize a recognition result. In order to solve the problem, the invention provides a registration apparatus (10) including: an image acquisition unit (11) that acquires an image obtained by imaging a placement surface of a table, on which a product is placed; an analysis unit (12) that recognizes the product included in the image, a registration unit (14) that registers the recognized product as a checkout target, and an output unit (13) that outputs a name of the recognized product by voice.
Owner:NEC CORP

Identification and feedback system for communication and repair strategies

PendingUS20260171107A1Speech recognitionSets using external connectionHuman–computer interactionSpeech sound
Embodiments herein relate to ear-wearable device systems. In an embodiment, a method executed in a processor of a feedback system for providing information related to a user's communication strategies. The feedback system can include a microphone and a memory storage. The method can include monitoring a voice output of the user with the microphone. The method can include processing the voice output of the user by at least one of: counting a first number of instances in which the user implements a positive communication repair strategy; and counting a second number of instances in which the user implements a negative communication repair strategy. The method can include generating feedback indicative of the first number of instances, the second number of instances, or both the first number of instances and the second number of instances. Other embodiments are also included herein.
Owner:STARKEY LABORATORIES INC

Voice control method, device, equipment, system, storage medium and program product

The invention provides a voice control method, device, equipment, system, a storage medium and a program product. According to the voice control method and device, under the condition that the control device cannot perform voice control on the household device, at least one of the voice input unit or the voice output unit of the other device associated with the control device is used for performing voice control, so that the error-tolerant rate and the reliability of voice control can be improved, and the use experience of a user can be improved.
Owner:DAIKIN INDUSTRIES LTD

System and method for communicating with a user with speech processing

A method and speech processing system for communicating with a user is provided. A speech signal may be received. The received speech signal may be processed by a first unified neural network to extract one or more of intents and entities. The one or more of intents and entities may be analyzed to generate a dialogue response. A second unified neural network may generate a speech output corresponding to the dialogue response for the user. In another example, a single unified neural network may process the received speech signal to extract one or more of intents and entities. The one or more of intents and entities may be analyzed, by the single unified neural network, to generate a dialogue response. The single unified neural network may generate a speech output corresponding to the dialogue response for the user.
Owner:TELEPATHY LABS GMBH

Digital human interaction method and device based on multiple modes and electronic equipment

The invention discloses a multi-modal-based digital human interaction method and device and electronic equipment, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining multi-modal data inputted by a user, carrying out the feature extraction and feature fusion of the multi-modal data, and obtaining the fusion feature data; performing sentiment analysis on the user by adopting a pre-trained sentiment analysis model based on the fused feature data to obtain a sentiment state and a sentiment intention corresponding to the user; according to the emotional state and the emotional intention, generating a text reply, an audio feature and an audio stream of the digital person and an action tag corresponding to the digital person; extracting a reference image from the visual data according to the action label, and generating a target image according to the reference image and the audio features; and performing video rendering based on the audio stream, the target image and the action sequence of the digital person to generate an interactive digital person. The emotional state and the emotional intention of the user can be deeply understood, and the expression, the action, the mouth shape and the voice output of the digital human are ensured to be harmonious.
Owner:BEIYIN FINANCIAL TECH CO LTD

Speech synthesis methods, devices, equipment and storage media

ActiveCN116612742BListening goals metImprove speech synthesisInternal combustion piston enginesSpeech synthesisSynthesis methodsAuditory feedback
This application discloses a speech synthesis method, apparatus, device, and storage medium. The method involves analyzing the original text to be synthesized to obtain a phoneme sequence; inputting the phoneme sequence into a configured speech synthesis model to obtain synthesized speech output by the model. The speech synthesis model is a final speech synthesis model after parameter adjustment of the basic speech synthesis model, using the scoring results of multiple candidate speech samples corresponding to the input test text synthesized by the basic speech synthesis model as reward signals. The scoring results of each candidate speech sample conform to the user's auditory perception goals. This application adds user auditory feedback signals (i.e., the scoring results as reward signals) to the training process of the speech synthesis model, guiding the speech synthesis model to optimize model parameters in a direction that better conforms to the user's auditory perception, making the synthesized speech more in line with the user's auditory perception goals and improving the speech synthesis effect.
Owner:IFLYTEK CO LTD