Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

190 results about "Voice analysis" patented technology

Voice analysis is the study of speech sounds for purposes other than linguistic content, such as in speech recognition. Such studies include mostly medical analysis of the voice (phoniatrics), but also speaker identification. More controversially, some believe that the truthfulness or emotional state of speakers can be determined using voice stress analysis or layered voice analysis.

Man-machine physical twin system based on multi-modal video analysis and adaptive mapping and control method

The invention belongs to the technical field of man-machine interaction and robot control, and particularly relates to a man-machine physical twin system based on multi-modal video analysis and adaptive mapping and a control method, and the system comprises a multi-modal collection unit which is used for collecting original video streams of human body actions, expressions and voices, and compensating and repairing dynamic shielding; the semantic analysis unit comprises an action analysis module used for extracting a human skeleton key point set object interaction track from the original video stream, and an expression / voice analysis module used for outputting a facial action unit coefficient and a voice text with an emotion label; the self-adaptive mapping unit is used for establishing human body-robot joint kinematics mapping, distributing control bandwidth in real time according to task types and degrading non-key modes based on network states to guarantee action continuity, and non-contact action capture is achieved through a pure vision scheme by eliminating dependence on a wearable sensor; and the problems of failure and the like of a traditional visual scheme under shielding and illumination variation are solved.
Owner:AIMI (BEIJING) ROBOT CO LTD

Power supply service management method and system based on multi-source data

The invention relates to the field of power supply work order service management, in particular to a power supply service management method and system based on multi-source data. The method comprises the following steps: collecting a current fault repair work order based on a power supply service platform, carrying out deep work order text semantic analysis and standard fault instance modeling, and constructing a global fault instance map; recognizing a client real-time repair call audio stream, carrying out time window voice analysis one by one, and carrying out intelligent fault work order filling, thereby obtaining a real-time audio intelligent fault work order; performing fault demand decomposition on the real-time audio intelligent fault work order and the global fault instance map, performing power supply service resource dynamic configuration regulation and control, and constructing a service resource regulation and control engine; and performing global fault work order real-time processing based on the service resource regulation and control engine, performing overdue work order early warning, and generating a service timeliness early warning strategy. Through dynamic resource allocation and preventive work order creation, the power supply service efficiency, the power grid reliability and the intelligent level are improved.
Owner:ZHANGZHOU POWER SUPPLY COMPANY STATE GRID FUJIANELECTRIC POWER +1

Smart home control method and device for realizing user interaction based on smart mirror

The invention relates to the technical field of smart home, and discloses a smart home control method and device for realizing user interaction based on a smart mirror, and the method comprises the steps: confirming a plurality of user registration instructions, obtaining a registration voice stream based on the user registration instructions, carrying out the feature extraction of the registration voice stream, obtaining reference voiceprint features, summarizing the reference voiceprint features, and carrying out the collection of the reference voiceprint features. The method comprises the following steps: obtaining a plurality of reference voiceprint features, confirming to receive a pre-confirmed voice wake-up word, carrying out voice collection based on the voice wake-up word to obtain a user voice stream, obtaining a voiceprint label by using the user voice stream and the plurality of reference voiceprint features, carrying out voice analysis on the user voice stream to obtain a control instruction, and sending the control instruction to a server. And performing permission verification by using the voiceprint tag and the control instruction to obtain a permission verification result, and obtaining response feedback based on the permission verification result. According to the invention, the problem that the personalized intelligent service cannot be provided because the current user identity cannot be distinguished and the control authority cannot be verified according to the user identity in the smart home interaction scene can be solved.
Owner:DONGGUAN LAIMSEN TECH BUILDING MATERIAL CO LTD

Service compliance dynamic evaluation system based on real-time voice recognition

The invention relates to a service compliance dynamic evaluation system based on real-time voice recognition, and belongs to the cross technical field of artificial intelligence and service compliance management. The system comprises a voice signal enhancement acquisition unit, a voice analysis and risk initial judgment unit, a dynamic evaluation iteration unit and a compliance risk disposal unit. The voice signal enhancement acquisition unit processes low-quality voice and generates a voice data set; the voice analysis and risk initial judgment unit is used for transferring voice in real time, extracting a dialogue intention and identifying violation contents and interactive emotions; the dynamic evaluation iteration unit outputs compliance scores and violation details and is in butt joint with a manual labeling platform; and the compliance risk disposal unit generates a risk label, carries out linkage training and appeal review, and forms a compliance risk disposal scheme. According to the invention, full quality inspection of service scenes in multiple industries such as finance, operators and the like is realized, the voice transfer accuracy and compliance rectification rate are improved, the verbal skill violation complaint rate is reduced, and dynamic monitoring and closed-loop disposal of service compliance are achieved.
Owner:SHANGHAI RONGDA DIGITAL TECH CO LTD

Vehicle-mounted image generation device, in-vehicle infotainment system and vehicle

The invention relates to the technical field of vehicles, and discloses a vehicle-mounted image generation device, a vehicle machine system and a vehicle. The method comprises the following steps: at least acquiring first point cloud data and first image data through a data fusion module, and converting the two data into a bird's-eye view feature map through a preset conversion algorithm; performing semantic analysis on the user voice through a preset semantic analysis model by using a voice analysis module; generating a vehicle-mounted application image at least applied to head-up display and / or screen display according to the bird's-eye view feature map and the semantic analysis result by using an image generation engine through a preset compression model; and terminating the generation process of the vehicle-mounted application image in a preset time period after the automatic emergency braking signal is received through the emergency interruption module, and at least applying the graphic processing computing power of the generation process to a vehicle collision avoidance decision. Therefore, the problem that in an existing vehicle-mounted image generation technology, due to the fact that the computing power is too large, vehicle decision making is not timely under the emergency situation can be solved, and the driving safety of a user can be guaranteed.
Owner:CHINA FAW CO LTD +1

AI psychological counseling method and system integrating multi-modal sentiment analysis and own intelligence

The invention discloses an AI psychological counseling method and system based on multi-modal sentiment analysis and body intelligence fusion. The method comprises the steps that text, voice and image feature vectors of a patient are extracted; aligning feature dimensions through linear projection, adding space / time sequence position codes, inhibiting noise in combination with cross-modal attention and a gating weighting mechanism, hierarchically fusing text, image and voice features, and generating fusion features; based on facial actions, postures and voice analysis, identifying seven-estrus states of traditional Chinese medicine, outputting emotion probability prediction distribution of a patient, and performing element-by-element product on the emotion probability prediction distribution and the prediction distribution to obtain a video generation vector; a video frame is generated by using StyleGAN-V, the consistency of content and emotion is constrained through emotion driving loss, and a personalized virtual human image is configured. The system comprises a data acquisition module, a feature extraction module, a multi-modal fusion module, an emotion recognition module and a video generation module. According to the method, a closed loop of'emotion analysis-emotion quantification-video generation 'is constructed, and the accuracy and universality of remote psychological consultation are improved.
Owner:杨雪飞

Spoken English pronunciation quality evaluation method based on multi-mode speech feature analysis

ActiveCN121528247ASpeech analysisFeature extractionModal voice
The invention belongs to the technical field of speech analysis, and discloses a spoken English pronunciation quality evaluation method based on multi-modal speech feature analysis, which comprises the following steps: acquiring a spoken English speech signal of a target user, and performing multi-domain decomposition on the speech signal to obtain multi-modal speech features; performing time-frequency domain corresponding relation analysis and feature extraction on the multi-mode speech features to obtain a pronunciation detail feature sequence; performing multi-scale matching on the pronunciation detail feature sequence and a preset standard pronunciation template, constructing a multi-dimensional representation model based on a multi-scale matching result, and calculating a fine-grained quality score of a phoneme unit in each multi-dimensional representation model in combination with rhythm and rhythm parameters in the multi-modal speech features, the rhythm coherence score and the overall fluency score are fused to generate a comprehensive pronunciation quality evaluation result and a visual diagnosis report of pronunciation deviation; the oral English pronunciation evaluation method realizes comprehensive and refined evaluation of oral English pronunciation, and provides a scientific guidance basis for personalized language learning.
Owner:ZHANG ZHOU HALTH VOCATIONAL COLLEGE

Robot control method based on voice analysis

The invention relates to the field of voice analysis, in particular to a robot control method based on voice analysis, and the method comprises the steps: obtaining image information in a space region where a target robot is located, determining semantic tags, determining a potential association semantic tag group, and screening out a semantic guiding corpus group for the target robot; when voice control data is received, instruction fuzzy parameters of the voice control data are determined so as to judge instruction fuzzy tendency, optimization is carried out on the voice control data with the instruction fuzzy tendency, specifically, confidence centralized clusters and semantic discrete clusters are determined, semantic expansion is carried out on the semantic discrete clusters, and then an expansion text is obtained; and screening the expanded text based on the semantic oriented corpus group to obtain a confidence instruction text. According to the method, the semantic oriented corpus group is constructed in combination with the image information, the analysis of the voice control data with the instruction fuzzy tendency is guided, the analysis accuracy of the voice control data under the semantic fuzzy condition is improved, and the control instruction recognition precision is ensured.
Owner:厦门工学院

Gas cylinder filling method, early warning method, checking method, system and medium

The invention relates to a gas cylinder filling method, an early warning method, an assessment method, a system and a medium in the technical field of industrial safety production. The gas cylinder filling safety assessment method based on voice analysis comprises the following steps: acquiring acoustic data of an operator; and analyzing the acoustic data and calculating to obtain a knowledge proficiency score Spro. And acoustic data are extracted, and a psychological quality score Semo is obtained through calculation. Collecting motion posture data of the gas cylinder and calculating to obtain an operation stability score Ssta; and carrying out weighted fusion on the Stro, the Semo and the Ssta to obtain an assessment score Stotal. And setting a score line based on the assessment requirement, and if the Stotal is greater than or equal to the score line, determining that the operator passes the assessment. According to the invention, through innovative multi-mode assessment, the ability of the operator is changed from feeling-based ability to data-based ability, the final result can be known, and the specific weak link can be known according to the specific condition of each link, so that the psychological quality and emergency response ability of the operator can be effectively assessed.
Owner:ANHUI SPECIAL EQUIP INSPECTION INST

Self-adaptive voice semantic communication method based on hierarchical time sequence importance

The invention relates to a self-adaptive voice semantic communication method based on hierarchy time sequence importance, which belongs to the field of voice analysis, and comprises the following steps: a sending end performs priority division on a discrete feature matrix according to hierarchy and time sequence attributes of voice features, and screens the features according to importance scores in a feature matrix packaging and selecting link; the method comprises the following steps: interacting with wireless communication through a channel adaptive scheduling module, introducing a channel feedback mechanism, and dynamically adjusting transmission power and a resource allocation strategy by sensing current channel state information; and after completing signal demodulation, a receiving end inputs the acquired sparse feature flow into a voice restoration module, and globally reconstructs the received features by using voice priori knowledge of deep pre-training. According to the method, a hierarchical speech feature extraction technology based on a discrete codebook and a generative semantic repair technology are combined, so that the speech word error rate in a severe channel environment is reduced to the maximum extent, and the semantic intelligibility of a receiving end is improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

System and method for voice analysis

The invention provides a system for analysing vocal signals from an individual, said system comprising an acoustic sensor configured to have a measured frequency response in the ultrasonic range; mean
Owner:THE SEC OF STATE FOR DEFENCE IN HER BRITANNIC MAJESTYS GOVERNMENT OF THE UK OF GREAT BRITAIN & NORTHERN IRELAND

Intent inference in audiovisual communication sessions

In one aspect, a user's intent can be inferred based on voice analysis during a communications session, and prompts can be presented, or other actions taken, at least partly in response to the inferred intent. For example, a network microphone device (NMD) having one or more microphones can capture voice input and transmit the voice input to remote computing device(s) for a communication session (e.g., a videoconference). The NMD can analyze the voice input to detect one or more utterances. Based on the utterance(s), the NMD can cause a user prompt to be displayed via a display device communicatively coupled to the NMD. The particular prompt can depend at least in part on one or more context parameters associated with the communication session (e.g., a microphone state of one or more users, a screen share state of one or more users, or a recording status of the session, etc.).
Owner:SONOS INC

System

An object of a system according to an embodiment is to enable an elderly person to easily access necessary information and services.SOLUTION: A system includes a voice input unit, an analysis unit, a guidance unit, and a connection unit. The voice input unit acquires voice. The analysis unit analyzes the voice acquired by the voice input unit. The guidance unit provides appropriate information based on the content analyzed by the analysis unit. The connection unit connects to a necessary service based on the information provided by the guide unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

System

An object of the system according to the embodiment is to efficiently create minutes of a conference, provide opinions, and deal with foreign languages.SOLUTION: A system according to an embodiment includes a voice analysis unit, a minutes generation unit, an advice providing unit, and a foreign language handling unit. The voice analysis unit analyzes voice data during a conference. The minutes creating section creates minutes based on the voice data analyzed by the voice analyzing section. The advice providing unit provides advice or an opinion on the basis of the minutes generated by the minutes generation unit. The foreign language correspondence unit translates the minutes generated by the minutes generation unit into a foreign language.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

system

PendingJP2026105311ADialog systemVoice analysis
We provide the system. [Solution] A receiving method that accepts language selection, A generation means for generating learning content based on a generative AI model, An acoustic conversion means that presents the generated learning content as an audio output, A voice analysis method that converts user voice information into text data, A management system that continues the dialogue based on user responses and provides feedback, A dialogue system that includes this.
Owner:SOFTBANK GROUP CORP

Methods and system for distributing information via multiple forms of delivery services

A content distribution facilitation system is described comprising configured servers and a network interface configured to interface with a plurality of terminals in a client server relationship and optionally with a cloud-based storage system. A request from a first source for content comprising content criteria is received, the content criteria comprising content subject matter. At least a portion of the content request content criteria is transmitted to a selected content contributor. If recorded content is received from the first content contributor, the first source is provided with access to the received recorded content. The recorded content may be transmitted via one or more networks to one or more destination devices. Optionally, a voice analysis and / or facial recognition engine are utilized to determine if the recorded content is from the first content contributor.
Owner:GREENFLY

A breath-speech pause signal analysis system and method

The present application relates to the cross technical field of speech signal processing, respiratory physiological monitoring and artificial intelligence, and specifically discloses a respiratory-speech pause signal analysis system and method. The present application synchronously collects speech and respiratory signals, extracts language-adapted pause parameters and physiological indexes, performs time sequence alignment and correlation analysis, and then fuses a prediction model and a true-false recognition model for intelligent analysis, thereby solving the problems that traditional heart-lung function detection relies on professional equipment and cannot be remotely and non-contactly monitored, and that existing speech analysis lacks a physiological coupling mechanism, leading to the inability to identify synthetic speech and insufficient cross-language adaptation, and realizing the dual ability improvement of non-contact respiratory physiological state evaluation and speech fraud recognition.
Owner:ZHONGDE NUOHAO (BEIJING) EDUCATION TECH CO LTD

system

We provide the system. [Solution] A voice analysis method for acquiring voice data in real time and identifying the characteristics of the speaker, A learning method for analyzing past conversation history and learning conversation patterns, A means of providing and displaying interactive games to participants, A means of managing information to collect and organize family events and health information, A means of distributing information to notify user terminals of organized information. Includes system.
Owner:SOFTBANK GROUP CORP

Hearing device comprising an own voice estimator

Disclosed herein are embodiments of a hearing device including at least one first, outward-facing, input transducer configured to pick up first sounds from the environment of a user and a second, inward-facing, input transducer configured to pick up a second sounds at the eardrum of the user. The hearing device can further include a directional system including a) an own voice beamformer configured to provide an estimate of the user's own voice in dependence of the at least one first and the second electric input signals and configurable own voice beamformer weights; and b) an own voice analyzer configured to analyze at least one of the at least one first and said second electric input signals, or to analyze a signal or signals originating therefrom, and to provide an own voice beamformer weight control signal.
Owner:OTICON

A personalized sound medicine formula generation and evaluation method based on voice emotion recognition

PendingCN122290642APersonalizationSound therapy
This invention relates to a method for generating and evaluating personalized sound therapy formulas based on voice emotion recognition, belonging to the field of voice analysis, and more specifically to the field of digital music development and production. The method includes: using an AI sound therapy effect judgment model to intelligently determine the increase in the target user's happiness level after using each five-tone sound therapy formula based on the target user's current five-element attribute data and current five-organ attribute data, and the digital content of each five-tone sound therapy formula; selecting the five-tone sound therapy formula with the largest increase value as the target user's personalized sound therapy formula. This invention addresses the technical problems of difficulty in providing the most suitable personalized music therapy plan for different users and the insufficient precision of music therapy plans. It employs an AI-based traversal method to intelligently judge each sound therapy formula for the target user, and uses five-tone five-element sound therapy formulas to improve the precision of the sound therapy formulas, thereby solving the aforementioned technical problems.
Owner:SHANDONG SHANGYI HEALTHCARE TRADITIONAL CHINESE MEDICINE TECHNOLOGY DEVELOPMENT CO LTD +1

system

The system according to the embodiment aims to comprehensively manage the health condition and living environment of a user and provide optimal advice and reminders. [Solution] A system according to an embodiment includes a health data acquisition unit, a monitoring unit, an environmental data acquisition unit, an environment management unit, a voice analysis unit, a reminder setting unit, and a lifestyle analysis unit. The health data acquisition unit acquires health data. The monitoring unit monitors the user's health condition based on the data acquired by the health data acquisition unit. The environmental data acquisition unit acquires environmental data. The environment management unit manages the user's living environment based on the data acquired by the environmental data acquisition unit. The voice analysis unit analyzes voice. The reminder setting unit sets reminders based on voice commands analyzed by the voice analysis unit. The lifestyle analysis unit analyzes the user's lifestyle.
Owner:SOFTBANK GROUP CORP

A recording automatic generation method based on AI voice interaction

The application discloses a kind of based on AI voice interaction's record automatic generation method, belong to voice analysis technical field, comprising: obtaining original voice data, continuous audio stream is recorded from target environment by preset audio acquisition module, generates the voice segment set of preliminary division using time domain segmentation technique;For the voice segment set, extract pitch feature, speech rate variation and volume feature, generate voice feature set;The voice feature set is input into deep neural network model, determine emotional label by multidimensional feature mapping, generate voice unit set with emotional label;If the voice type of voice unit set is statement type, then use narrative template to generate text;If the voice type of voice unit set is inquiry type, then use inquiry template to generate text;Get formatted record.The based on AI voice interaction's record automatic generation method solves the problem that existing record generation mode is difficult to generate format standard and information-rich record.
Owner:GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)

system

PendingUS20260252649A1EngineeringVoice analysis
The system according to the embodiment comprises a collection unit, an analysis unit, and an estimation unit. The collection unit collects a person's utterances in real time. The analysis unit analyzes utterance data collected by the collection unit. The estimation unit estimates the person's intention based on information analyzed by the analysis unit.
Owner:SOFTBANK GROUP CORP

Children physiological education doll, method and program product

The invention belongs to the technical field of child dolls, and particularly discloses a child physiological education doll, a method and a program product, and the method comprises the steps: collecting the conversation voice of a child through the doll, preprocessing the conversation voice when the conversation voice triggers a conversation activation state, and uploading the preprocessed conversation voice to a cloud platform, the pre-trained large model is called through the cloud platform to carry out voice analysis and intelligent dialogue answering, corresponding physiological knowledge answering information is obtained and played to the child, and meanwhile, when the child touches the sensing device of the corresponding part of the doll, touch feedback voice information is played according to the corresponding part, so that early education of the physiological knowledge of the child is achieved. The child physiological knowledge education doll fills the blank of child physiological knowledge education dolls, dialogue education and touch feedback education of child physiological knowledge can be intelligently and visually achieved in the form of the dolls, and the child can effectively receive related physiological knowledge conveniently.
Owner:XINYI ANIMATION (HUIZHOU) CO LTD

Method and voice control device for voice control of a medical device

A method for voice control of a medical device (1) having the steps of: detecting a voice input intended to control the device containing an operator; analyzing the audio signal to provide a first voice analysis result; identifying a first voice command based on the first voice analysis result; determining a verification signal for the first voice command, including: analyzing the audio signal to provide a second voice analysis result; identifying a second voice command of the operator based on the second voice analysis result; comparing the first and second voice commands, wherein the verification signal confirms the first voice command only when the first voice command and the second voice command meet a consistency criterion; generating a control signal for controlling the medical device based on the verification signal if the first voice command is confirmed according to the verification signal, wherein the control signal is adapted to control the medical device according to the first voice command; and inputting the control signal into the medical device.
Owner:SIEMENS HEALTHINEERS AG

Supply telephone voice analysis and work order whole-process intelligent circulation method

The invention discloses a supply and service telephone voice analysis and work order whole-process intelligent circulation method, which comprises the following steps of S1, converting dialect voice of a user into a standard text in a high-precision manner by utilizing a high-precision dialect voice recognition model, and providing a reliable basis for subsequent processing; s2, intelligently identifying user appeal types by using the knowledge-enhanced brilliance large model, accurately extracting key information, and performing logic verification through a business rule engine to ensure that the information is accurate and compliant; s3, automatically matching a template according to the type, filling information, desensitizing sensitive data in real time, and finally automatically distributing the sensitive data to an execution unit through an encryption interface to realize full-process automatic closed loop; according to the method, the work order processing time can be greatly shortened, the labor cost is reduced, the fault processing accuracy and the service satisfaction degree can be improved, and core business support is provided for digital transformation of the state grid.
Owner:SHANGQIU POWER SUPPLY CO OF STATE GRID HANAN ELECTRIC POWER CO

Speech analysis method and device, computer equipment and readable storage medium

The invention relates to the technical field of voice processing, and provides a voice analysis method and device, computer equipment and a readable storage medium, and the method comprises the steps: obtaining a to-be-analyzed conversation voice in a preset business scene, and generating preprocessed voice data; converting the preprocessed voice data into text information, and extracting acoustic features and text features of the voice data; acquiring basic language understanding and pattern recognition capability according to the acoustic features and the text features, and optimizing the basic language understanding and pattern recognition capability according to a small amount of labeled sample data corresponding to the business scene; and dialogue feature representation is generated according to the optimized acoustic features and text features, the feature representation is analyzed, an intelligent decision corresponding to the business scene is generated, and analysis of dialogue voice is completed. Through deep coupling of acoustics and text features, shallow application of traditional speech recognition is broken through, and a general technical path of small data driven accurate decision is provided for the fields of finance, medical health, old-age care and the like.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

A system and method for detecting cognitive decline using speech analysis.

A system and method for detecting cognitive decline in a subject using a classification system for detecting cognitive decline in a subject based on speech samples. The classification system is trained using speech data corresponding to audio recordings of speech from normal and cognitively declining patients to generate an ensemble classifier including a plurality of component classifiers and an ensemble module. Each of the plurality of component classifiers is a machine learning classifier configured to generate a component output that identifies the sample data as corresponding to a normal patient or a cognitively declining patient. The machine learning classifiers are generated based on a subset of available features. The ensemble module receives the component outputs from all of the component classifiers and generates an ensemble output that identifies the sample data as corresponding to a normal patient or a cognitively declining patient based on the component outputs.
Owner:JANSSEN PHARMA NV