Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Speech recording" patented technology

An intelligent management method based on real-time speech recognition transcription

This invention discloses an intelligent management method based on real-time speech recognition and transcription, relating to the field of artificial intelligence technology. The method includes the following steps: Step S1, multimodal audio perception and environmental adaptive enhancement; Step S2, acoustic feature extraction and real-time transcription mapping; Step S3, semantic error correction compensation based on dynamic context weighting; Step S4, structured element extraction and logical association reconstruction; Step S5, intelligent management closed-loop decision-making and task distribution; Step S6, multi-source information backtracking and index construction; Step S7, intelligent management efficiency evaluation. This application can solve the problems of limited recognition accuracy and missing semantic logic under complex sound fields. By improving transcription accuracy through multimodal perception and dynamic semantic compensation, it achieves automated connection from speech recording to structured management decisions, significantly improving office management efficiency.
Owner:MUDANJIANG NORMAL UNIV

Issue reporting by a receiving device

ActiveUS12684198B2Speech recordingTelevision set
A technique is described for improved issue reporting by a receiving device such as a set-top boxes for satellite and cable television services. In an example embodiment, the receiving device generates an issue report based on internal operational logs, captured screens and / or video of a visual output, and a recording of the user's voice that includes a description of the issue they are experiencing. This issue report can be generated as an object file that can then be transmitted, via a communications network, to an issue reporting platform for processing, for example, by a technical support representative or an automated troubleshooting system.
Owner:DISH NETWORK LLC

System and method for a video avatar creation

A system trains a video synthesis model using a video training dataset comprising video samples of one or more persons. The system receives a video sample of the target person. The system trains a video custom synthesis model based on the video sample. The system generates, using both the video synthesis model and the video custom synthesis model, a video avatar that mimics visuals of the target person, wherein generating the video avatar further comprises: generating a preliminary video of a head of the target person with controlled gestures based on recorded gestures from the video sample and a target gesture script; and adding lip synchronization to the preliminary video by matching a voice recording of the target person to a plurality of lip movements based on words spoken by the target person in the preliminary video.
Owner:SIT AUTONOMOUS AG +1

Multi-modal fusion emotion recognition method and device based on graph neural network

The invention relates to the technical field of emotion recognition, and provides a multi-modal fusion emotion recognition method and device based on a graph neural network, and the method comprises the steps: obtaining the multi-modal input data of a to-be-recognized target person; respectively extracting modal features of the facial expression video data, the speaking recording audio data and the transcription text data to obtain a video feature vector, an audio feature vector and a text feature vector corresponding to each target statement moment; for each target statement moment, based on the video feature vector, the audio feature vector and the text feature vector, constructing a graph neural network structure; inputting the graph neural network structure into a graph neural network emotion recognition model to obtain an emotion category corresponding to the target statement moment; and determining an emotion recognition result based on the emotion category of each target statement moment. According to the embodiment of the invention, the multi-modal features are integrated by constructing the graph neural network structure, and emotion recognition is carried out by using the pre-trained model, so that the accuracy and robustness of emotion recognition can be improved.
Owner:ANHUI POLYTECHNIC UNIV MECHANICAL & ELECTRICAL COLLEGE

Voice separation and paragraph affiliation method and system for teleconference scene

The invention discloses a voice separation and paragraph affiliation method and system for a teleconference scene, and relates to the technical field of voice recognition. The method comprises the following steps: a conference platform triggers a voice record of a target conference window to generate a conference audio; according to the first engine, executing two-way voice separation processing on the conference audio, and constructing a voice logic map; according to a second engine, matching a target conference template, executing voice fragment attribution based on automatic conference record typesetting according to the voice logic map, and generating a target conference record; and carrying out platform database storage on the target conference record, and carrying out calling management by permission setting. The technical problems that in the prior art, due to the fact that voice interweaving in a multi-party conference is difficult to effectively distinguish spokesmen and speaking content logic structures of the spokesmen, conference record affiliation is not clear, and editing efficiency is low are solved, two-way semantic logic separation and structured affiliation of the voice content of the teleconference are achieved, and user experience is improved. And the conference recording accuracy and the automatic typesetting efficiency are improved.
Owner:BEIJING LIANXUN XINGYE TECH CO LTD

Hearing diagnostic method and system employing artificial intelligence

A system and method for diagnosing hearing impairments in individuals, particularly school-aged children, using voice sample analysis and data from electronic questionnaires. The system utilizes a deep neural network to process extracted acoustic features from voice recordings and correlate them with hearing performance data, generating a predicted audiogram. The system includes a microphone, signal acquisition and processing blocks, a neural network-based analysis engine, and a result interpretation interface.
Owner:INSTYTUT FIZJOLOGII I PATOLOGII SLUCHU

Speech separation and paragraph attribution method and system for teleconference scenarios

This invention discloses a method and system for speech separation and segment attribution in remote conferencing scenarios, relating to the field of speech recognition technology. The method includes: triggering speech recording in a target meeting window via a conferencing platform to generate meeting audio; performing two-way speech separation processing on the meeting audio using a first engine to construct a speech logic graph; matching a target meeting template using a second engine and performing speech segment attribution based on automated meeting record layout according to the speech logic graph to generate a target meeting record; and storing the target meeting record in a platform database for retrieval and management using access settings. This invention solves the technical problem in existing multi-party conferencing systems where speech interleaving makes it difficult to effectively distinguish speakers and the logical structure of their statements, leading to unclear meeting record attribution and low editing efficiency. It achieves two-way semantic logic separation and structured attribution of remote meeting speech content, improving the accuracy of meeting records and the efficiency of automatic layout.
Owner:BEIJING LIANXUN XINGYE TECH CO LTD

Voice modification detection using physical models of speech production

A computer may train a single-class machine learning using normal speech recordings. The machine learning model or any other model may estimate the normal range of parameters of a physical speech production model based on the normal speech recordings. For example, the computer may use a source-filter model of speech production, where voiced speech is represented by a pulse train and unvoiced speech by a random noise and a combination of the pulse train and the random noise is passed through an auto-regressive filter that emulates the human vocal tract. The computer leverages the fact that intentional modification of human voice introduces errors to source-filter model or any other physical model of speech production. The computer may identify anomalies in the physical model to generate a voice modification score for an audio signal. The voice modification score may indicate a degree of abnormality of human voice in the audio signal.
Owner:PINDROP SECURITY INC

Method for transferring at least one speech signal of a patient during a magnetic resonance imaging examination, and magnetic resonance imaging device

ActiveUS12455333B2Pulse automatic controlSensorsCarrier signalSpeech recording
Techniques are disclosed for transferring at least one speech signal of a patient during a magnetic resonance imaging examination, wherein the speech signal is recorded by a speech recording device of a wireless communication device assigned to the patient and transmitted at least as part of a communication signal to a receive device of the magnetic resonance imaging device. The communication signal is a modulated signal or is generated from a modulated signal, and to generate the modulated signal the speech signal is modulated onto a carrier signal. The modulated signal is generated by way of a modulation with reduction of the level of the carrier signal.
Owner:SIEMENS HEALTHINEERS AG

System and method for detecting cognitive decline using speech analysis

PendingJP2026031950ASurgeryKernel methodsPattern recognitionCognitively impaired
A system and method for detecting cognitive decline in a subject using a classification system for detecting cognitive decline in a subject based on speech samples.SOLUTION: The classification system is trained using speech data corresponding to audio recordings of speech from normal and cognitively impaired patients to generate an ensemble classifier that includes a plurality of component classifiers and an ensemble module. Each of the plurality of component classifiers is a machine learning classifier configured to generate a component output that identifies the sample data as corresponding to a normal patient or a cognitive patient. The machine learning classifier is generated based on a subset of the available features. The ensemble module receives the component outputs from all of the component classifiers and generates an ensemble output that identifies the sample data as corresponding to normal or cognitively impaired patients based on the component outputs.SELECTED DRAWING: None
Owner:JANSSEN PHARMA NV

Computer-implemented method for operating at least one vehicle and vehicle

The invention relates to a computer-implemented method (100) for operating at least one vehicle (10) with at least one recording device (24) for capturing a speech recording and at least one display device (14), comprising at least the following steps: determining (102) at least one target environmental parameter within a cabin of the vehicle (10) by means of at least one first trained learning algorithm, in particular a Large Language Model, from an audio signal of a speech recording provided by the recording device (24); generating (104) at least one visual signal by means of at least one second trained learning algorithm based on the at least one target environmental parameter; and controlling (106) the at least one display device (14) to display the at least one visual signal.The method (100) allows the user experience inside the vehicle (10) to be adapted to the user with increased accuracy.
Owner:DR ING H C F PORSCHE AG

Communication clue analysis reminding method and system

The invention relates to a communication clue analysis reminding method and system, and relates to the technical field of communication clue analysis, and the method comprises the steps: collecting audio recording information; performing preprocessing according to the audio recording information, and extracting speech recording information and environment recording information; extracting character voiceprint features and recording text information according to the speech recording information; inputting the character voiceprint features into a preset voice recognition model, and recognizing and acquiring character identity information; generating a communication clue probability value in combination with the character identity information, the recording text information and the environment recording information; and when the communication clue probability value is greater than a preset probability reference value, generating communication early warning information according to the character identity information and the recording text information, and outputting the communication early warning information. The method and the device have the effect of improving the early warning accuracy of the communication clues in the incoming call communication scene.
Owner:XINZHI DAOSHU (SHANGHAI) TECH CO LTD

Electronic device release effects, voice-over comments, graphical user interface

1. Name of the product in this design: Graphical User Interface with Release Effects, Voice Message, and Bullet Screen for Electronic Devices. 2. Purpose of this design: An electronic device. 3. The key design features of this product are its graphical user interface in electronic devices. 4. Images or photos that best illustrate the design points: Interface change state diagram 5. 5. Electronic devices are designed in a conventional way, so other views are omitted. 6. Purpose of the Graphical User Interface: The product interface serves as the interactive interface for publishing special effects voice-based bullet comments. After voice-based bullet comments are recorded, they are converted into text, triggering special effects keywords, generating animation effects, and enabling bullet comment interaction. In the main view interface, clicking the "Input Voice Bullet Comment" button at the bottom displays interface state 1, showing the voice bullet comment recording box and entering voice recording mode. In interface state 1, when the user's spoken words are recorded and converted into text displayed in the bullet comment box, interface state 2 is displayed. In interface state 2, after a 3-second pause in the user's speech recording, when the recording stops, interface state 3 is displayed, and keywords such as "explosion" and "heart" are highlighted in the recorded text. In interface state 3, clicking the "Send" button displays interface state 4, where a "bomb" graphic voice bullet comment is displayed in the player. In interface state 4, clicking the "bomb" special effect voice bullet comment displays interface state 5, where the bullet comment explodes with animation and the user's recorded voice is played.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Speech recording integrated control system, method, and program

We provide an integrated control system for recording speeches that enables a highly versatile platform and improves the evidentiary value of speech records when creating meeting minutes, various contracts, etc. [Solution] The system is characterized by comprising: an audio input unit that registers the speaker's first voice information in advance in a terminal device equipped with an audio recording function and inputs the speaker's second voice information and terminal information before recording; an integrated authentication unit that performs user authentication; a user authentication score calculation unit; a recording start activation unit that activates the recording start button based on the calculated score value; a speech recording unit that records audio data, transcribed text information, and speaker metadata after the recording start button is activated; an integration unit that integrates session data of the recorded audio information, text information, and metadata; and an integrated control unit that performs integrated control of the processes up to speech preparation, speaker authentication, speaker recording, and score calculation.
Owner:竹内祐树 +4

Synthesizing personalized speech through adaptive excitation signal generation

ActiveUS12620386B1Speech synthesisSpeech recordingLoudspeaker
A speech synthesis system is described and may include at least one microphone; a speaker; a sensing system, and memory storing processor-executable instructions, which when executed by the processor, cause the processor to: detect speech-related signals emanating from the subject; generate a variable excitation signal; shape the generated variable excitation signal according to previously stored speech recordings; and cause, from the speaker and based on the shaped variable excitation signal, produced speech content that approximates the matched one or more voice characteristics in the previously stored speech recordings.
Owner:INCENTMED IP LLC

Feedback loop combining augmented reality and virtual reality to facilitate interactions in the virtual world and the physical world

ActiveUS12718505B2EngineeringVirtual world
Combining augmented reality (AR), virtual reality (VR), and photogrammetry to facilitate interactions with the physical world is disclosed. Images may be collected through various means (e.g., images taken by drones, ground-based photography, satellite imaging, etc.) and used to construct a three-dimensional (3D) model representation of the real world. This virtual environment can then be experienced by users in either a computing system (e.g., a desktop computer, laptop computer, smart phone, tablet, etc.) or a VR interface (e.g., a headset). The users can create various annotations, such as voice recordings, images, text, etc., and place these annotations in the virtual 3D model of the real world. These annotations are linked to the physical world via a coordinate system. AR “explorers” and VR users can then interact with one another via geographically aligned AR and VR models.
Owner:AEROSPACE CORP

A communication lead analysis reminder method and system

The application relates to a communication clue analysis reminding method and system, and relates to the technical field of communication clue analysis, which comprises the following steps: collecting audio recording information; pre-processing the audio recording information to extract speech recording information and environment recording information; extracting a person's voiceprint feature and recording text information according to the speech recording information; inputting the person's voiceprint feature into a preset voice recognition model to identify and obtain person identity information; combining the person identity information, the recording text information and the environment recording information to generate a communication clue probability value; when the communication clue probability value is greater than a preset probability reference value, generating communication warning information according to the person identity information and the recording text information, and outputting the communication warning information. The application has the effect of improving the warning accuracy of communication clues in the incoming call communication scene.
Owner:XINZHI DAOSHU (SHANGHAI) TECH CO LTD

Voice processing method and device, equipment, storage medium and program product

PendingCN121237130ASpeech recognitionNeural learning methodsSpeech rateSpeech recording
The embodiment of the invention provides a voice processing method and device, equipment, a storage medium and a program product. The method comprises the steps that attribute information related to an object is determined based on voice collected in the voice recording process of the object, and the attribute information at least indicates the voice speed of the object; determining configuration information of a pause duration threshold value for the object in the voice recording process based on the attribute information; and controlling the voice activity detection of the voice recording process based on the configuration information of the pause duration threshold. Therefore, the speech speed of the object can be determined based on the collected speech, and the pause duration threshold value when speech activity detection is executed can be automatically adjusted based on the speech speed of the object. Different pause duration threshold values can be determined for different objects, the personalization and adaptability of voice activity detection performed on different objects can be improved, and the accuracy of the pause duration threshold values is improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Legal aid resource intelligent allocation and scheduling method, device and equipment and medium

The application relates to a legal aid resource intelligent allocation and scheduling method, device, equipment and medium. The method comprises the following steps: acquiring consultation data; the consultation data comprises original consultation text, voice recording and video file; legal entities are extracted based on the consultation data through a multi-task learning model to obtain three-dimensional classification labels; the three-dimensional classification labels comprise geographical location codes, case type codes and emergency level codes; the three-dimensional classification labels and lawyer characteristics are matched through a deep forest model to obtain a matching probability matrix; the lawyer load is monitored based on the current state of the lawyer to obtain a resource state heat map; and a matching scheme set is obtained through a reinforcement learning model based on the matching probability matrix and the resource state heat map; the matching scheme set comprises a legal aid scheme. The method can provide real-time feedback on the state of lawyers and increase the flexibility of recommended legal aid schemes.
Owner:GUANGDONG UNIV OF TECH