Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "Speech analytics" patented technology

Speech analytics is the process of analyzing recorded calls to gather customer information to improve communication and future interaction. The process is primarily used by customer contact centers to extract information buried in client interactions with an enterprise. Although speech analytics includes elements of automatic speech recognition, it is known for analyzing the topic being discussed, which is weighed against the emotional character of the speech and the amount and locations of speech versus non-speech during the interaction. Speech analytics in contact centers can be used to mine recorded customer interactions to surface the intelligence essential for building effective cost containment and customer service strategies. The technology can pinpoint cost drivers, trend analysis, identify strengths and weaknesses with processes and products, and help understand how the marketplace perceives offerings.

A Human-Computer Voice Interaction Control Method and System Based on Smart TV

This application relates to a human-computer voice interaction control method and system based on a smart TV. The method includes acquiring voice data within a preset range, processing the voice data to obtain voice feature data carrying control commands; performing feature analysis on the voice feature data using a preset voice analysis model to extract wake-up keywords and compare them with a preset command library to obtain control command comparison results; sending a secondary confirmation request to the user based on the control command comparison results, and combining the confirmation voice information from the user feedback to perform command recognition evaluation and correct command deviation processing to obtain the corrected control command; performing function switching processing on the TV according to the correct control command, and optimizing the display effect by linking and adjusting related devices based on the program display requirements after the switch to obtain human-computer voice interaction control data. This application has the effect of improving the intelligence of voice interaction control of smart TVs.
Owner:GUANGZHOU XIANYOU INTELLIGENT TECH CO LTD

Intelligent Technical Protocol Based Approach Leveraging AI-ML to Block Vishing Scammers

Systems and methods detect and prevent vishing attacks through an integrated framework combining SIP header customization, STIR / SHAKEN frameworks, AI / ML analysis, and real-time speech analysis using the Viterbi algorithm. The system begins with call initiation, embedding authentication information in the SIP header. The SIP data is transmitted and verified using STIR / SHAKEN frameworks, ensuring the authenticity of the caller's identity. Verified data is cross-referenced with third-party databases and analyzed by an AI / ML engine to detect anomalies. If potential fraud is detected, the call is blocked, and the customer is notified. Calls that pass initial checks are further analyzed using the Viterbi algorithm, which converts speech to text and identifies suspicious patterns. An anomaly pattern detector processes the converted text to detect vishing indicators, terminating the call if a match is found. This multi-layered approach ensures robust protection against vishing, enhancing the security and reliability of voice communications while safeguarding users from fraud.
Owner:BANK OF AMERICA CORP

Voiceprint recognition and identity authentication method and system based on public safety management and medium

The invention discloses a voiceprint recognition and identity authentication method and system based on public safety management and a medium, and relates to the field of voiceprint recognition, and the method comprises the following steps: S1, carrying out the preprocessing of an original voice signal, extracting three types of features in parallel, and splicing the three types of features into a joint feature for representation; s2, coding the joint feature representation, outputting a high-level shared feature vector, and generating a voiceprint identification vector and a living body attribute vector through branches; s3, when high-security-level authentication is judged, dynamic challenge content is sent, response voice is received, and instant behavior characteristics are analyzed; and S4, calculating voiceprint matching similarity and living body attribute probability, fusing the voiceprint matching similarity, the living body attribute probability and the selectable dynamic evidence score, and generating a final authentication decision. By introducing the joint feature coding method based on the correlation constraint, the deep discrimination information can be better reserved, the defense capability on advanced speech synthesis and conversion attacks is enhanced, and the overall security level of the system is improved.
Owner:CHINA UNIV OF MINING & TECH

Intelligent call-out processing method and device based on artificial intelligence, and server

The invention relates to the technical field of intelligent outbound calls, in particular to an intelligent outbound call processing method and device based on artificial intelligence and a server. Real-time voice data input by a user is converted into text data based on a voice analysis model, intention analysis is conducted on the text data based on an intention analysis model, and the intention of the user is obtained. And performing sentiment analysis on the text data based on the sentiment analysis model to obtain a sentiment analysis result of the user, generating a dialogue strategy matched with the user according to a user intention analysis result and the sentiment analysis result, and finally executing dynamic adjustment operation according to the dialogue strategy and dialogue skill data to obtain verbal skill data matched with the user. According to the method, the accurate dialogue strategy is generated for the user not only through the intention of the user but also in combination with the sentiment analysis result of the user, so that the verbal skill data to be provided for the user is adjusted, the determination accuracy and precision of the verbal skill data needed by the user are improved, and the interaction accuracy of the intelligent outbound dialogue process is improved.
Owner:GUANGZHOU RES INTERESTING INFORMATION TECH CO LTD

AI-powered automated recruitment system for the intelligent identification and selection of talent.

ActiveDE202025106630U1InstrumentsEngineeringProcessing
An AI-powered automated recruitment system for intelligent talent identification and selection, consisting of: a processing unit configured to execute machine learning procedures for adaptive candidate evaluation; a storage unit that is operationally connected to the aforementioned processing unit and is configured to store candidate data, trained model parameters, and records of the recruitment process; a video interaction unit consisting of a camera and an audio recording unit configured for conducting real-time or asynchronous interviews, wherein the processing unit analyzes the captured audiovisual data using natural language processing, facial recognition and sentiment analysis to generate behavioral assessment metrics; a speech analysis unit consisting of a speech recognition processor and an acoustic characterization subsystem configured to receive and analyze telephone candidate responses, with the processing unit deriving linguistic competence, tonal clarity, and confidence indices from the responses; a technical evaluation unit configured to deliver interactive programming tasks to candidate terminals, with the processing unit evaluating code execution output and logic efficiency and detecting plagiarism through code similarity comparison techniques; a data analysis unit configured to combine the outputs of the aforementioned video interaction unit, speech analysis unit, and technical evaluation unit to generate predictive suitability scores for candidates and recruitment performance metrics; and a communication interface unit configured to establish secure data transmission connections between the system and external personnel databases via encrypted network protocols.
Owner:NELLIPUDI SOMA KIRAN KUMAR CUMMING

system

The system according to this embodiment aims to support workplace communication and employee mental health in real time. [Solution] The system according to the embodiment comprises an acquisition unit, an analysis unit, a decision unit, and a proposal unit. The acquisition unit acquires images and voices of employees. The analysis unit analyzes the employee's facial expressions, tone of voice, and language use based on the information acquired by the acquisition unit. The decision unit determines an action in accordance with the analysis results obtained by the analysis unit. The proposal unit proposes the action determined by the decision unit to the employee.
Owner:SOFTBANK GROUP CORP

Artificial intelligence-based method and system for cognitive assessment through speech analysis

PCT designated stageWO2026178554A1Data miningBiology
A method and non-transitory computer-readable medium for automated cognitive assessment and generation of cognitive assessment scores from patient oral cognitive assessment data using an artificial-intelligence ("AI") agent is provided. The method and computer-readable medium include delivering one or more oral cognitive assessments and storing collected audio data for analysis by a backend analysis server. Collected cognitive assessment data is processed by a backend analysis server using a feature-embedding model to extract derived features from collected audio data. The backend analysis server uses machine-learning classification to generate cognitive assessment scores from the features. Assessment scores may be transmitted to a clinician-accessible device for review. The method enables automated, scalable, and reproducible patient cognitive assessments on patient-accessible devices.
Owner:SENSORY SCIENCE INC

Artificial intelligence-based method and system for cognitive assessment through speech analysis

PendingUS20260248447A1Learning architectureBiology
A method and non-transitory computer-readable medium for generating cognitive assessment scores from patient speech using an artificial-intelligence (“AI”) agent is provided. The method and computer-readable medium include initiating a patient audio communication session, conducting one or more oral cognitive assessments including narrative recall, engaging in a context-dependent discussion phase, assessing responses, and generating scores from audiometric, lexical, prosodic, and semantic-coherence features using multi-modal machine learning architectures. Audio data is processed by a backend analysis server using automatic speech recognition to generate a time-aligned transcript, and the backend analysis server uses a feature-embedding model to extract prosodic, acoustic, lexical, and semantic-coherence features from the audio and corresponding transcript. The backend analysis server uses machine-learning classification to generate cognitive assessment scores from extracted features. Assessment scores, feature data, and session records may be transmitted to clinician-accessible locations for review. The method enables automated, scalable, and reproducible speech-based patient cognitive assessments on patient-accessible devices.
Owner:SENSORY SCIENCE INC

Subtitle rendering method, system and device based on AI technology and storage medium

The invention discloses a subtitle rendering method, system and device based on the AI technology and a storage medium, and the method comprises the steps: obtaining a to-be-rendered target file, and carrying out the voice analysis processing of the target file through employing the AI technology, and obtaining corresponding text information; obtaining statement core components and keywords in the text information so as to generate structured subtitle metadata; performing paging layout processing on the structured subtitle metadata by using preset layout parameters to obtain corresponding page data; on the basis of the page data, word-by-word rendering is carried out on structured subtitle metadata in the page data in combination with a timestamp; and according to a set display mode, performing paging rendering on the page data, and outputting and displaying the page data. According to the method, the target file is analyzed and processed through the AI voice, word-by-word rendering is carried out in combination with the layout parameters and the timestamps, and finally paging rendering output is completed in the set display mode, so that the subtitle rendering effect can be remarkably improved, and user requirements are met.
Owner:深圳牛学长科技有限公司

Communication systems and communication methods

This enables the use of the VOX function on multiple VOX communication devices. [Solution] The system comprises an extraction unit 22 that extracts voice data from received data, an analysis unit 23 that analyzes the voice data extracted by the extraction unit 22 as a voice analysis result, a calculation unit 26 that calculates priority information regarding the priority of the VOX function of the communication device based on the voice analysis result analyzed by the analysis unit 23, and a control unit 28 that determines whether to enable or disable the VOX function based on the priority information calculated by the calculation unit 26.
Owner:JVC KENWOOD CORP

System for assigning an operator to an interaction

PCT designated stageWO2026062516A1CommerceData setEngineering
System for assigning one of a plurality of operators to an interaction amongst one or more interactions each between a respective user and a user assistance centre comprising an artificial operator implemented on a computer, wherein each interaction comprises a real-time interaction between a user and the user assistance centre, the system comprising a controller configured to: obtain, during an information-collection phase of an interaction between a respective user and the user assistance centre, first information concerning a matter relating to the respective user; and determine an operator suitable to resolve the matter, based on said first information and on second information output by an AI-based operator assignment unit, the AI-based operator assignment unit being trained based on a dataset including at least the following information associated to each other: (i) past transcript information representing a transcript of an earlier interaction between a user and the user assistance centre; (ii) speech analysis information comprising parameters output from a speech analysis for the earlier interaction; and (iii) operator information identifying an operator assigned to said earlier interaction.
Owner:COVISIAN SPA

system

We provide the system. [Solution] A means for obtaining call content, electronic messages, and visitor images from a user's mobile device, A method using generative AI to analyze acquired information and detect the risk of fraud, Image recognition and voice analysis means for recording and identifying the audio and video of visitors, A means of alerting and notifying users or designated contacts when suspicious activity or visitors are identified, A means of exchanging information with users and providing instructions, A system that includes this.
Owner:SOFTBANK GROUP CORP

System

A system is provided.SOLUTION: A system comprising: data collection means for collecting facial data of a customer; data collection means for collecting voice data of the customer; facial expression recognition means for recognizing facial expressions of the customer from the collected facial data; voice analysis means for analyzing a tone of the customer from the collected voice data; emotion analysis means for determining an emotional state of the customer based on data obtained by the facial expression recognition means and the voice analysis means; and feedback generation means for suggesting a service improvement based on the determined emotional state.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

System

To provide a system for efficiently preparing a menu based on the request of a user, and for ordering necessary materials through a network.SOLUTION: A system includes a voice input part, an analysis part, a menu generation part, and a shopping support part. The voice input unit receives a voice of a user. The analysis unit analyzes the voice received by the voice input unit. The menu generation part generates a menu on the basis of the result analyzed by the analysis part. The shopping support unit orders necessary materials through the Internet based on the menu generated by the menu generation unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

system

The system according to the embodiment aims to generate and read out appropriate answers to questions from journalists in real time. [Solution] A system according to an embodiment includes a speech recognition unit, an analysis unit, a generation unit, and a reading unit. The speech recognition unit recognizes the speech of a reporter's question. The analysis unit analyzes the question recognized by the speech recognition unit. The generation unit generates an appropriate answer based on the question analyzed by the analysis unit. The reading unit reads out the answer generated by the generation unit.
Owner:SOFTBANK GROUP CORP

system

Provide a system. 【Solution means】 An information exchange means for detecting an incoming call, A voice analysis means for recognizing and analyzing the voice of the call partner, A determination means for evaluating a risk based on the analysis result, A moving means for transferring a call when it is determined that the risk is low, A storage means for issuing a warning and recording the call content when it is determined that the risk is high, A notification means for warning the user of the possibility of fraud detected by the voice analysis means, A management means for managing call history and making the risk assessment record viewable at a later date, A system including.
Owner:SOFTBANK GROUP CORP

system

A system is provided.SOLUTION: A system including means for acquiring posting data on an SNS, means for classifying the acquired posting data into a text, an image, and a moving image, means for analyzing the classified data by natural language processing, computer vision, and voice analysis, means for predicting a mental health risk from the analyzed data, means for generating a warning message and a corresponding message on the basis of a prediction result, and means for notifying a user of the generated message.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Generating data features from speech and sentiment analytics for enhanced predictive routing

A method for processing data for training a predictive routing model. The method includes receiving interaction data from previous interactions that includes audio data capturing a conversation and transcript data of the conversation. The method continues by performing speech analytics by processing the audio data to determine scores for speech metrics that include a measure of how much the agent or customer speaks during the conversation. The method continues by performing sentiment analysis to determine scores associated with sentiment metrics, the sentiment metrics including a measure of a sentiment based on classifying utterances appearing in the transcript data as being positive or negative. The method continues by performing feature engineering to generate feature data and generating a training dataset therefrom. The method continues by applying a machine learning algorithm to the training dataset to train a predictive routing model.
Owner:GENESYS CLOUD SERVICES INC

System

A system is provided.SOLUTION: A system comprising: a voice acquiring unit for analyzing a telephone conversation of an elderly person in real time; a voice analyzing unit for analyzing a voice received by a server using a generated AI model; a fraud detecting unit for comparing an analysis result with a past crime database and detecting a fraudulent act; a warning notifying unit for transmitting a warning to a user device based on a detection result of a fraudulent act; and an automatic notifying unit for notifying police and a social worker of a detection result.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Voice analysis-based agent routing

A computer system and method for improving call center routing through analysis of customer interactions including obtaining identifying information for a caller upon initiation of a call, identifying the caller as a repeat customer using the identifying information, retrieving historical interaction data associated with the repeat customer from a database, analyzing any combination of customer audio data, customer call log information, or customer feedback, utilizing an artificial intelligence algorithm to determine a current mood indicator of the customer, calculating a customer behavior score for the repeat customer based on the historical interaction data and the current mood indicator of the customer, and matching the repeat customer to a call agent, based on the customer behavior score.
Owner:WELLS FARGO BANK NA

Risk management and control method for power marketing business

The invention discloses a risk management and control method for power marketing business. Comprising the following steps: a provincial customer service center classifies work orders; the provincial customer service center obtains an emotion resonance index and intention ambiguity of the customer in real time through a voice analysis and text emotion analysis engine; based on the emotion resonance index, the intention ambiguity, the historical work order data and the customer behavior data, constructing a risk prediction model, and outputting the risk level of the work order; the risk level of the work order is primarily reviewed by provincial customer service center personnel, professional classification is carried out, and the risk appeal work order is automatically distributed; tracking and accessing the risk appeal work order; according to the invention, multi-modal emotion and intention analysis is fused, voice and text data are combined, a multi-dimensional risk prediction model is constructed, and intelligent classification and priority distribution of work orders are realized; through linkage of provincial-level first review and professional review, RPA automation and double telephone early warning, the risk identification accuracy and processing efficiency are improved.
Owner:HENAN TENGLONG INFORMATION ENG

System

An object of a system according to an embodiment is to detect a sign of power harassment at an early stage and automatically take appropriate measures.SOLUTION: A system includes a voice analysis unit, a text analysis unit, a warning unit, and a notification unit. The voice analysis unit analyzes the voice data and identifies an utterance that may be power harassment. The text analysis unit analyzes the text data and identifies a message that may be power harassment. The warning unit issues a warning based on the sign of power harassment identified by the voice analysis unit and the text analysis unit. The notification unit notifies a supervisor or a personnel department based on the warning issued by the warning unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Intelligent technical protocol based approach leveraging AI-ML to block vishing scammers

Systems and methods detect and prevent vishing attacks through an integrated framework combining SIP header customization, STIR / SHAKEN frameworks, AI / ML analysis, and real-time speech analysis using the Viterbi algorithm. The system begins with call initiation, embedding authentication information in the SIP header. The SIP data is transmitted and verified using STIR / SHAKEN frameworks, ensuring the authenticity of the caller's identity. Verified data is cross-referenced with third-party databases and analyzed by an AI / ML engine to detect anomalies. If potential fraud is detected, the call is blocked, and the customer is notified. Calls that pass initial checks are further analyzed using the Viterbi algorithm, which converts speech to text and identifies suspicious patterns. An anomaly pattern detector processes the converted text to detect vishing indicators, terminating the call if a match is found. This multi-layered approach ensures robust protection against vishing, enhancing the security and reliability of voice communications while safeguarding users from fraud.
Owner:BANK OF AMERICA CORP