Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

62 results about "Speech patterns" patented technology

System and method for contextual analysis and metadata database generation for user-specific speech patterns

A system for contextual analysis and metadata database generation for user-specific speech patterns is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system splits the speech signal into a first set of audio frames, where each audio frame comprises an utterance of one or more words. The system determines a context associated with each word. In response, the system detects a context change between a first text and a second text. The system generates a contextually split set of frames by splitting the speech signal into a second set of audio frames according to the detected context changes.
Owner:BANK OF AMERICA CORP

Automated nonverbal analysis system

Examples relate to computer-implemented methods for analyzing communication in digital evaluation. A computing device accesses multimodal data comprising video and audio information of human subjects and configures a computational model using this data to identify patterns in communication that correlate with assessment metrics. The configuring implements processing techniques that preserve relationships between features across different modalities. When a video recording of a candidate is received, the computing device processes the video using the configured computational model to extract communication features. These features may include facial expressions, gestures, eye movements, posture, vocal tone, and speech patterns. The device generates an evaluation of the candidate based on the extracted communication features and outputs a representation of the evaluation.
Owner:LIGHT STEVEN PATRICK

Conference window conversation system based on voice sensor

The invention discloses a meeting window call system based on a voice sensor, and relates to the technical field of communication. Comprising a full-space voice acquisition module, a dynamic confidence evaluation module, a phase coupling enhancement module, a multi-dimensional speech mode recognition module, a dynamic gain adjustment and tone quality enhancement module and a full-link voice integrity verification module, an array type voice sensing unit is deployed based on audio reflection characteristics, a multi-band full-space acquisition matrix is constructed, voice signals of two parties meeting are acquired, and a basic audio data stream with a high signal-to-noise ratio is output. According to the invention, through array sensing acquisition, dynamic confidence evaluation, phase coupling and harmonic enhancement, multi-dimensional speech recognition and dynamic gain adjustment, accurate recognition and fidelity enhancement of low-volume speech are realized, and through combination of full-link integrity verification and mistaken killing backtracking, weak signals are ensured not to be missed and distorted, so that the accuracy of speech recognition is improved. And the privacy, the security and the information integrity of the meeting call are improved.
Owner:HANGZHOU HUA TING TECH CO LTD

System and method for dynamic audio slicing window selection based on context and speech patterns

A system for an audio slicing window selection for contextually splitting a speech signal is disclosed. The system identifies a first audio processing software algorithm that is assigned to a user. The system identifies a set of audio processing software algorithms and configures each of them with a respective audio slicing window. The system selects a second audio processing software algorithm, from among the set of audio processing software algorithms. The system selects one of the first and second audio processing software algorithms that is assigned an audio slicing window associated with the context of the speech signal. The system splits the speech signal using the selected audio processing software algorithm. The system determines whether the speech signal is split contextually. In response to determining that the speech signal is not split contextually, the selected audio processing software algorithm and / or the audio slicing window may be updated.
Owner:BANK OF AMERICA CORP

Systems and methods for generating conversational recommendations using non-serialized interpretations of serialized inputs

Systems and methods for generating dynamic conversational recommendations. Conversational recommendations include communications between a user and a system that may maintain and / or facilitate (e.g., via autocomplete functionality) a conversational tone, cadence, and / or speech pattern of a human during an interactive exchange between the user and the system. The system may use artificial intelligence applications to generate suggested dynamic conversational recommendations based on initial user inputs (e.g., such as autocomplete functionality).
Owner:CAPITAL ONE SERVICES LLC

Systems and methods for alternative content recommendations based on analyzing potential interpretations using supplemental inputs

Systems and methods for generating dynamic conversational recommendations. Conversational recommendations include communications between a user and a system that may maintain and / or facilitate (e.g., via autocomplete functionality) a conversational tone, cadence, and / or speech pattern of a human during an interactive exchange between the user and the system. The system may use artificial intelligence applications to generate suggested dynamic conversational recommendations based on initial user inputs (e.g., such as autocomplete functionality).
Owner:CAPITAL ONE SERVICES LLC

Real-Time AI-Driven Fraud Detection and Prevention System for In-Person Transactions

Systems and processes are disclosed for real-time fraud detection and prevention in in-person transactions. The invention utilizes an AI / ML engine to analyze customer application data for inconsistencies and unusual requests indicative of potential fraud. Concurrently, a real-time conversation analysis engine with speech recognition algorithms monitors interactions between bank associates and customers, identifying suspicious speech patterns, hesitations, and keywords associated with scams. By combining insights from application data and conversational analysis, the system generates a comprehensive risk assessment. When a high probability of fraud is detected, an alert notifies the bank associate, security personnel, and other relevant individuals. This proactive approach enables immediate action to prevent fraudulent transactions, reducing manipulation risks and minimizing financial losses. The system continuously learns from new data, adapting to evolving fraud tactics, thus providing robust, long-term protection for financial institutions and their customers.
Owner:BANK OF AMERICA CORP

System and method for contextual analysis and metadata database generation for user-specific speech patterns

A system for contextual analysis and metadata database generation for user-specific speech patterns is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system splits the speech signal into a first set of audio frames, where each audio frame comprises an utterance of one or more words. The system determines a context associated with each word. In response, the system detects a context change between a first text and a second text. The system generates a contextually split set of frames by splitting the speech signal into a second set of audio frames according to the detected context changes.
Owner:BANK OF AMERICA CORP

Systems and methods for secure tokenized credentials

Systems, devices, methods, and computer readable media are provided in various embodiments having regard to authentication using secure tokens, in accordance with various embodiments. An individual's personal information is encapsulated into transformed digitally signed tokens, which can then be stored in a secure data storage (e.g., a “personal information bank”). The digitally signed tokens can include blended characteristics of the individual (e.g., 2D / 3D facial representation, speech patterns) that are combined with digital signatures obtained from cryptographic keys (e.g., private keys) associated with corroborating trusted entities (e.g., a government, a bank) or organizations of which the individual purports to be a member of (e.g., a dog-walking service).
Owner:ROYAL BANK OF CANADA

A meeting window talk system based on a voice sensor

The application discloses a meeting window call system based on a voice sensor and relates to the technical field of communication.The meeting window call system comprises a full-space voice collection module, a dynamic confidence evaluation module, a phase coupling enhancement module, a multi-dimensional speech mode recognition module, a dynamic gain adjustment and sound quality enhancement module and a full-link voice integrity verification module.The full-space voice collection module is arranged in a meeting window space based on audio reflection characteristics, an array type voice sensor unit is arranged, a multi-frequency full-space collection matrix is constructed, voice signals of both parties in the meeting are collected, and basic audio data streams with high signal-to-noise ratio are output.The application realizes accurate identification and faithful enhancement of low-volume voice through array type sensing collection, dynamic confidence evaluation, phase coupling and harmonic enhancement, multi-dimensional speech recognition and dynamic gain adjustment, and ensures that weak signals are not missed and distorted through full-link integrity verification and false positive backtracking, thereby improving the privacy, security and information integrity of the meeting call.
Owner:HANGZHOU HUA TING TECH CO LTD

A computer-implemented system and method for deriving and generating a voice response to a user input

A computer-implemented system for providing text to speech (TTS) comprises a user communications device 104 having an input device 134 to receive text data from a text data source; and an audio playback device 134 for outputting an audio response. The user communications device 104 being configured to receive a text data stream from a text data source. In real time, as soon as said text data stream has started to be received, a speech emphasis handler module is used to segment the incoming text data into text data segments based on natural speech patterns and / or emphasis. The text data segments are passed, as they are created and in order, to a speech recognition module to transform each said text data segment into representative digital audio data. In real time, as the representative digital audio data is generated, digital audio data is fed to an audio buffer and stored. In real time and in order, the audio segments are delivered to the audio playback device 134. The text data source may be a Large Language Model (LLM) or Large Multimodal Model (LMM) API. The application may produce natural-sounding speech in real-time, without waiting for the entire text to be downloaded.
Owner:EVANS CELEVATORON

Multimodal emotion recognition-based ai communication support system for elderly and its method

PendingKR1020260139344ASpeech patternsSpeech sound
A multimodal emotion recognition-based AI communication support system and method tailored for the elderly are disclosed. Voice data of an elderly person of a certain age or older is collected, the speech patterns of the elderly person are learned using the voice data, and communication with the elderly person is supported based on the speech patterns.
Owner:INJE UNIVERSITY INDUSTRY ACADEMIC COOPERATION FOUNDATION

system

The system according to this embodiment aims to provide a personalized experience that responds to the individual needs and personalities of users. [Solution] The system according to the embodiment comprises a learning unit, an agency unit, and a clone creation unit. The learning unit learns the user's speech patterns and thought patterns. The agency unit functions as the user's agency based on the information learned by the learning unit. The clone creation unit creates a clone of the user.
Owner:SOFTBANK GROUP CORP

Systems and methods for determining actor status according to behavioral phenomena

Aspects relate to systems and methods for determining actor status according to behavioral phenomena. An exemplary system includes an eye sensor configured to detect an eye parameter as a function of an eye phenomenon, a speech sensor configured to detect a speech parameter as a function of a speech phenomenon, and a processor in communication with the eye sensor and the speech sensor; the processor is configured to receive the eye parameter and the speech parameter, determine an eye pattern as a function of the eye parameter, determine a speech pattern as a function of the speech parameter, and correlate one or more of the eye pattern and the speech pattern to a cognitive status.
Owner:GMECI LLC

A computer-implemented system for deriving and generating a voice response to a user input

PCT designated stage expiredWO2025114695A1Speech synthesisData streamApplication programming interface
A computer-implemented system for facilitating a text to speech application, comprising a user communications device 104 configured to communicate with a Large Language Model (LLM) or Large Multimodal Model (LMM) API, the user communications device (104) being configured to: in real time, as soon as a text data stream has started to be received, use a speech emphasis handler module to segment the incoming text data into text data segments based on natural speech patterns and / or emphasis, and pass the text data segments, in order, as they are created, to a speech recognition module to transform each said text data segment into representative digital audio data; in real time, as said representative digital audio data is generated, start feeding said digital audio data to an audio buffer and store it therein; and deliver, in real time and in order, said audio segments to an audio playback device.
Owner:EVANS CELEVATORON

Systems and methods for generating dynamic conversational responses using deep conditional learning

Methods and systems are described herein for generating dynamic conversational responses. Conversational responses include communications between a user and a system that may maintain a conversational tone, cadence, or speech pattern similar to a human during an interactive exchange between the user and the system. The interactive exchange may include the system responding to one or more user actions (which may include user inactions), and / or predicting responses prior to receiving a user action.
Owner:CAPITAL ONE SERVICES LLC

Location-based trivial question and answer competition with emotional response capability

PendingCN121999769APosition fixationBiological modelsEngineeringEmotional responsivity
One or more methods of using location-based trivial questions with emotional response capabilities in a vehicle include initiating one or more location-based trivial questions based on global positioning system (GPS) coordinates, and detecting active players using one or more microphones within the vehicle. Based on player names and seat locations, the method includes customizing trivial questions for one or more of the seat locations. The method may further include evaluating the player's emotion by analyzing the voice tones and the voice patterns, and based on the emotion on one or more trivial questions, and / or interrogating the player's name via the microphone (s) at the beginning of the trivial question, and / or configuring the number of channels and the processing chain topology based on the detected number of active players and their seat locations.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Location-based trivia contest with mood-responsive capabilities

A method, or methods, of using location-based trivia with mood-responsive capabilities in a vehicle includes initiating one or more location-based trivia questions based on global positioning system (GPS) coordinates and detecting active players using one or more microphones within the vehicle. Based on the player names and seat locations, the method includes customizing the trivia questions for one or more of the seat locations. The method may further include assessing a mood of the players by analyzing vocal tone and speech patterns and basing one or more of the trivia questions on the mood, and / or querying player names via the microphone(s) at a start of the trivia questions, and / or configuring the number of channels and processing chain topology based on the detected number of active players and their seat locations.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Advanced SIP-based caller identification and voicemail analysis system for fraud prevention in telecommunications

Systems and processes are disclosed for a multi-layered approach to fraud prevention by leveraging a machine learning engine integrated with the Session Initiation Protocol (SIP) to attempt caller identification before transitioning to a voice call, allowing for the potential blocking of unwanted calls. If SIP-based identification remains inconclusive, an anomaly detection engine employing the Viterbi algorithm analyzes the caller's speech patterns during voicemail messages. The Viterbi algorithm converts spoken language into text, identifying suspicious characteristics such as unusual speech patterns, inconsistencies, and keywords associated with scams. If suspicious characteristics are detected, the system automatically blocks callback attempts and notifies the customer of potential spam or unwanted calls. This proactive approach addresses both live and recorded fraudulent calls, enhancing the security of telecommunications by preventing fraudulent interactions before they can cause harm. The system continuously learns from new data, adapting to evolving fraud tactics, providing robust, long-term protection for telecommunications users.
Owner:BANK OF AMERICA CORP

Machine learning based self voice removal

A method includes receiving an audio signal, wherein the audio signal includes a speech component of a user and a noise component; filtering the audio signal with a self-speech filter that filters out the speech component using an intrinsic user vector, wherein the intrinsic user vector is determined based on speech input of the user during an offline training session and represents an intrinsic compressed representation of a speech pattern of the user; and outputting a filtered audio signal in which the speech component of the user has been substantially removed from the audio signal. The intrinsic user vector is obtained through machine learning using speech input from multiple speakers.
Owner:BOSE CORP

Personalized modification of audio and visual components of a virtual agent of a user interface

Systems, methods and / or computer program products personalizing user interactions with virtual agents of applications and / or services, using audio / visual components customized to appeal to user preferences. Upon initial interaction with virtual agents, AI ranking algorithms are triggered to adopt the highest ranked persona for the virtual agent based on previous positive interactions with the user, learned preferences, the user's state inferred from facial expressions, body language, tone. Personas can emulate voice signatures of popular characters, actors, celebrities and sports figures, and access a corpus of dialogue of the available personas to learn unique speech patterns, slang, tone, grammar, speed, and vocabulary. The corpus that comprises various personas of real and / or fictional individuals can include data of the mannerisms and visual likeness of the various characters and people, allowing avatars depicting the virtual agents to be animated in the likeness of the selected persona in real-time during conversational workflows.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Conversational avatar system

Systems and methods for conversational avatar systems are disclosed herein. The systems and methods may include receiving, via a computing system, an input comprising speech data; generating, via the computing system, a transcript of the speech data in real-time; analyzing, via the computing system, the transcript in real-time; generating, via the computing system, a response to the transcript in real-time based on the tagging; selecting, via the computing system, one or more avatar animation gestures based on tone and speech patterns of the generated response; synthesizing, via the computing system, an audible response based on the generated response; synchronizing, via the computing system, the one or more avatar animation gestures to the synthesized audible response to form a synchronized avatar animation; and rendering, via the computing system, the synchronized avatar animation.
Owner:2WAI INC

Multi-Modal Based Shortness of Breath Estimation

Multi-modal shortness of breath systems and methods are described. In aspects, one or more devices may be utilized to collect data associated with the user, such as audio data (e.g., speech pattern, breath, etc.) and motion data (e.g., walking, exercising, etc.) that overlaps in time with the audio data. Further, an assessment system may analyze both the audio data and the motion data collected by the one or more devices to provide an overall health and / or fitness metric for the user.
Owner:APPLE INC

Accent conversion for virtual conferences

One example method includes receiving, during a virtual conference hosted by a virtual conference provider, a first audio stream comprising speech having first speech patterns according to a first accent, the first audio stream received from a first client device associated with a first participant in the virtual conference; generating, by a first trained machine learning (“ML”) model, a second audio stream comprising the speech having second speech patterns according to a second accent; and outputting the second audio stream.
Owner:ZOOM COMMUNICATIONS INC

Systems and methods for alternative content recommendations based on analyzing potential interpretations using supplemental inputs

ActiveUS12682167B2User inputEngineering
Systems and methods for generating dynamic conversational recommendations. Conversational recommendations include communications between a user and a system that may maintain and / or facilitate (e.g., via autocomplete functionality) a conversational tone, cadence, and / or speech pattern of a human during an interactive exchange between the user and the system. The system may use artificial intelligence applications to generate suggested dynamic conversational recommendations based on initial user inputs (e.g., such as autocomplete functionality).
Owner:CAPITAL ONE SERVICES LLC

Text-to-speech and speech recognition for noisy environments

ActiveUS12315490B2Speech recognitionSpeech synthesisLombard effectNoise level
The present disclosure relates generally to speech processing. Humans change their speech patterns in noisy environments. The systems and devices described herein can compensate for noisy environments to be more human-like. Thus, the configurations and implementations herein can determine a sound profile for the sound environment where the user is listening. Based on the sound profile, the devices can determine a transform to apply to output speech from the device. This transform is applied to the wake word, speech recognition, and to the output speech to compensate for the noise level of the environment by mimicking the Lombard effect.
Owner:SPOTIFY

Multi-modality-based respiratory shortness estimation

The invention relates to breathing shortness estimation based on multi-modality. Multi-modal breathing shortness systems and methods are described. In aspects, one or more devices may be utilized to collect data associated with a user, such as audio data (e.g., voice pattern, breath, etc.) and motion data (e.g., walking, exercise, etc. In addition, the evaluation system may analyze both audio data and motion data collected by the one or more devices to provide overall health and / or fitness metrics for the user.
Owner:APPLE INC

A syntax-aware candidate matching based sentiment element extraction method and system

PendingCN122452549APart of speechData mining
The application provides a sentiment element extraction method and system based on syntax-aware candidate matching, which specifically comprises the following steps: performing syntax analysis on input text to obtain a syntax tree to extract part-of-speech tags and phrase structures; extracting a candidate attribute word set from the syntax tree based on a preset attribute part-of-speech pattern, retrieving attribute sample examples from an example library according to text syntax vector representations corresponding to candidate attribute word phrase structures, and using the attribute sample examples to guide a plurality of large language models to generate reliable triplets; extracting a candidate opinion word set from the syntax tree based on a preset opinion part-of-speech pattern, retrieving opinion sample examples from the example library according to text syntax vector representations corresponding to candidate opinion word phrase structures, and using the opinion sample examples to guide the plurality of large language models to generate reliable quadruplets; and fusing the reliable triplets and the reliable quadruplets to obtain a sentiment element extraction result.
Owner:MINJIANG UNIVERSITY

System and method for dynamic audio slicing window selection based on context and speech patterns

A system for an audio slicing window selection for contextually splitting a speech signal is disclosed. The system identifies a first audio processing software algorithm that is assigned to a user. The system identifies a set of audio processing software algorithms and configures each of them with a respective audio slicing window. The system selects a second audio processing software algorithm, from among the set of audio processing software algorithms. The system selects one of the first and second audio processing software algorithms that is assigned an audio slicing window associated with the context of the speech signal. The system splits the speech signal using the selected audio processing software algorithm. The system determines whether the speech signal is split contextually. In response to determining that the speech signal is not split contextually, the selected audio processing software algorithm and / or the audio slicing window may be updated.
Owner:BANK OF AMERICA CORP