Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

252 results about "Recognition speech" patented technology

Speech recognition. Speech recognition is the inter-disciplinary sub-field of computational linguistics that develops methodologies and technologies that enables the recognition and translation of spoken language into text by computers. It is also known as automatic speech recognition (ASR), computer speech recognition or speech to text (STT).

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Interactive question and answer task processing method based on AI large model

The invention discloses an interactive question and answer task processing method based on an AI large model, and relates to the technical field of AI questions and answers, and the method comprises the following steps: performing correlation screening and function label labeling on high-confidence sub-queries in a retrieval result set, performing weight reduction on low-confidence sub-queries, constructing a cross-source consistency constraint vector, and generating an input data set; based on the input data set, a multi-source evidence consistency verification channel is constructed, entity-by-entity alignment and conflict detection are executed, the credibility interval of answers is calculated, meanwhile, voice emotions are recognized, and an answer data set is generated; and converting the answer data set into multi-modal feedback, monitoring user behaviors in real time, calculating behavior response strength indexes, dynamically adjusting a feedback form, and generating an interactive question and answer data set. The semantic consistency constraint of the multi-modal evidence is realized, and the robustness of answer credibility evaluation in the natural language processing task is improved.
Owner:张婧

Systems and methods for orchestrating interaction with an artificial intelligence application

Systems and methods for orchestrating interaction with an artificial intelligence (AI) application in a contact center environment receive, via an AI agent, a voice message from a user; convert the message from voice to text; generate an initial computational inference process based on the text message; determine whether or not all information required to execute the initial computational inference process is available to the processor; when a determination is made that all information required is available: execute the initial computational inference process; generate a text reply based on the initial computational inference process; convert the text reply to a voice reply; and send the voice reply to the user via the AI agent; when a determination is made that information is unavailable: generate a text query requesting the information; convert the text query to a voice query; and send the voice query to the user via the AI agent.
Owner:THE BANK OF NEW YORK MELLON

Classroom teacher teaching performance description method, model and system based on multi-modal data fusion and storage medium

The invention relates to the technical field of data processing, in particular to a classroom teacher teaching performance description method, model and system based on multi-modal data fusion and a storage medium. By introducing a multi-modal data fusion strategy, cooperative processing of classroom teacher visual information, audio information and text information is realized, and the accuracy and time sequence continuity of teacher target perception are significantly improved. Compared with a traditional single-mode method, the method not only can extract the posture, expression and action characteristics of the teacher from the visual mode, but also can recognize the voice emotion and the side language signal from the audio mode, and achieves the full-dimensional description of the teaching behavior of the teacher in combination with the text semantics. The objective of the invention is to solve the problem of how to perform multi-dimensional evaluation on teaching performance of a classroom teacher based on data of multiple modalities.
Owner:YUNNAN NORMAL UNIV

Contextually boosted aviation speech recognition

A variety of applications can include a system having a speech recognition system responsive to the speech input, where the speech recognition system can be configured to recognize the speech input using an aviation vocabulary including words extracted using state information of an aircraft or intent information of the aircraft associated with the received speech input. A control system can be implemented to automatically perform an action in the system in response to analysis of the recognized speech input, where the action is associated with flight of the aircraft.
Owner:CIRRUS DESIGN CORP D B A CIRRUS AIRCRAFT

Instant messaging-based proper noun speech recognition processing method and computer device

The invention discloses a proper noun speech recognition processing method based on instant messaging and a computer device. The method comprises the following steps: firstly, constructing annotation data sets of different scenes and a user-specific custom word list, and dynamically obtaining related data and a hot word list according to a to-be-recognized voice scene; afterwards, a training voice recognition model is subjected to fine tuning by using the annotation data set to obtain a first model, and after a recognition instruction is received, a hot word list is loaded for recognition to obtain a voice initial recognition text; and after the initial recognition text is obtained, dynamic optimization is carried out by utilizing a user-specific custom word list, and recognition errors are corrected. According to the method, matching data can be dynamically obtained, the model is combined with a scene to understand proper nouns, and the recognition difficulty caused by pronunciation and meaning complexity is reduced; in combination with targeted data and hot word information training, the proper noun recognition accuracy is improved, the defects of an existing model error correction mechanism are overcome, the accuracy of a final recognition result is remarkably improved, and powerful support is provided for speech recognition and subsequent application.
Owner:BEIJING VRV SOFTWARE CO LTD

Voice interaction method, device and equipment and computer storage medium

The invention relates to the technical field of artificial intelligence, and provides a voice interaction method and device, equipment and a computer storage medium. The method comprises the following steps: receiving a voice instruction input by a user; according to the voice instruction, obtaining an instruction text sequence corresponding to the voice instruction and historical interaction data corresponding to the user; the instruction text sequence and the historical interaction data are input into a semantic analysis model, enhanced semantic representation features are generated, and the enhanced semantic representation features are obtained based on fusion of text semantic features of the instruction text sequence, tag semantic features of the historical interaction data and a device topological relation; the equipment topological relation is obtained based on a pre-constructed knowledge graph of the household equipment; identifying intention information corresponding to the voice instruction and target equipment in the associated household equipment based on the enhanced semantic representation feature; and sending the equipment control instruction to the target equipment, wherein the equipment control instruction at least comprises the intention information and the equipment identification information of the target equipment.
Owner:CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1

Semantic recognition method, device and equipment based on improved BiLSTM-CRF model

The invention provides a semantic recognition method, device and equipment based on an improved BiLSTM-CRF model, and relates to the technical field of semantic recognition. The method comprises the steps of performing data extraction on to-be-recognized voice information to obtain a to-be-recognized semantic text, and inputting the to-be-recognized voice information into a preset emotion feature extraction model to obtain emotion features of the to-be-recognized voice information; preprocessing a semantic text to be recognized, and inputting the preprocessed semantic text to the improved BiLSTM-CRF model to obtain an initial semantic recognition result of the semantic text to be recognized; the improved BiLSTM-CRF model comprises a BiLSTM model, an attention model, a text convolutional neural network model and a CRF model; the text convolutional neural network model is used for performing convolution and pooling operation of different granularities on the preprocessed semantic text to be recognized; and fusing the emotion features with the initial semantic recognition result to obtain a semantic recognition result of the to-be-recognized voice information. The accuracy of the semantic recognition result of the voice information can be effectively improved.
Owner:STATE GRID HEBEI ELECTRIC POWER CO LTD BAODING POWER SUPPLY BRANCH CO +1

Systems and methods for real-time meeting summarization

Systems and methods are provided for processing electronic content and generating corresponding output. Electronic content is received from a meeting, including recognizable speech content. This content is then summarized into real-time summary output by processing and encoding the meeting content while selectively alternating between unidirectional attention and bidirectional attention that is applied to the meeting contents.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Speech recognition method and apparatus, electronic device, and storage medium

The application provides a speech recognition method and device, electronic equipment and storage medium, the method comprises the following steps: inputting a speech to be recognized into a speech recognition model to obtain a speech recognition result output by the speech recognition model; wherein, an attention layer of the speech recognition model is used to perform matrix multiplication on each sub-block attention matrix and a value matrix; each sub-block attention matrix is obtained by sequentially performing longitudinal division and horizontal truncation on a lower triangular attention matrix of the attention layer, or sequentially performing horizontal division and longitudinal truncation on the lower triangular attention matrix, and then removing zero matrices. Since there is no zero matrix in the sub-block attention matrix, the zero matrix will not participate in the matrix multiplication operation, thereby reducing unnecessary operation cost. Compared with the complete matrix multiplication operation of the entire lower triangular attention matrix and the value matrix in the traditional method, the multiplication operation of the zero element part is reduced, the overall operation speed is accelerated, and the training efficiency of the model is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Range hood control method and device and range hood

The invention provides a control method and device of a range hood and the range hood, and relates to the technical field of kitchen appliances, and the method comprises the following steps: responding to voice information collected by a first microphone and a second microphone; judging whether the voice information comes from a preset target area or not; wherein the target area is an area for implementing cooking; if yes, recognizing a voice instruction contained in the voice information, and controlling the range hood to operate based on the voice instruction. According to the control method and device of the range hood and the range hood, the target area is the area where cooking is carried out, so that the voice instruction from the area where cooking is carried out can be recognized, the influence of voice signals of other areas is filtered from the source of voice information, the interference of irrelevant noise is reduced, and therefore the user experience is improved. The voice wake-up rate and the recognition rate can be improved, and then the user experience is improved.
Owner:HANGZHOU ROBAM APPLIANCES CO LTD

Collaborative recitation method and device, computer storage medium and terminal

The application provides a collaborative recitation method and device, a computer storage medium and a terminal, and relates to the technical field of speech recognition and evaluation. The method comprises the following steps: inputting a to-be-recognized speech into a speech recognition model, determining the content corresponding to the to-be-recognized speech according to the output of the speech recognition model, and obtaining to-be-recognized content; inputting the to-be-recognized content into an intention recognition model, determining the target intention type corresponding to the to-be-recognized content according to the output of the intention recognition model, wherein the intention type output by the intention recognition model comprises at least one of the following: starting recitation, reciting, requesting help and completing recitation; and according to the target intention type, the collaborative recitation processing of the to-be-recognized content is realized. The speech recognition model and the intention recognition model in the scheme can assist students to be familiar with the content of the article and the poem in the liberal arts recitation scene, liberate manpower and improve the learning efficiency.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Intelligent evaluation method for project road performance

The invention relates to the technical field of data analysis and processing, and discloses an intelligent evaluation method for a project road presentation, which comprises the following steps of: acquiring multi-modal data of a road presentation video in real time, including speech text, voice audio, expression video and PPT content, and analyzing and processing the multi-modal data through a multi-modal fusion technology to ensure feature alignment; generating a core challenge question set based on the analysis data, and verifying the payment willingness deviation ratio of a target user for commercial logic, technical implementation and a financial model; executing three rounds of progressive questioning; when the multi-modal data analysis of the project road performance is carried out, the analysis deviation of voice, text and visual data can be recognized in real time by constructing a collaborative alignment mechanism of multi-modal feature fusion, and the high conformity of inquiry problem generation and real commercial risks is ensured; and meanwhile, a confidence-driven redundancy check process is established, expression data is actively fused to correct and output when the voice recognition credibility is insufficient, and the accuracy of dynamic questioning and the key vulnerability capturing capability are improved.
Owner:BEIJING ZHONGKE ZHIYUAN TECH CO LTD

An automatic speech recognition method and system for power grid communication scheduling

The application provides an automatic speech recognition method and system for power grid communication scheduling, and relates to the technical field of speech recognition. The method comprises the following steps: obtaining to-be-recognized speech data for power grid communication scheduling and performing segmentation processing to obtain a plurality of to-be-recognized speech paragraphs and inputting the to-be-recognized speech paragraphs into a first speech recognition model to generate a first speech recognition result; calculating a complexity feature factor of each to-be-recognized speech paragraph, and dividing the plurality of to-be-recognized speech paragraphs into first difficulty paragraphs and second difficulty paragraphs; extracting a plurality of groups of sub-paragraph speech data from the first difficulty paragraphs, inputting the plurality of groups of sub-paragraph speech data and the to-be-recognized speech paragraphs of the second difficulty paragraphs into a second speech recognition model to generate a second speech recognition result, and performing feature fusion on the first speech recognition result and the second speech recognition result to obtain a target speech recognition result of the to-be-recognized speech data. The application realizes the improvement of the recognition efficiency of power grid communication scheduling speech.
Owner:LANGFANG POWER SUPPLY COMPANY STATE GRID JIBEI ELECTRIC POWER COMPANY +1

Speech recognition method and device, storage medium and computer equipment

The invention discloses a voice recognition method and device, a storage medium and computer equipment, relates to the technical field of voice recognition, and can be applied to the field of digital medical treatment and finance. The main purpose is to solve the problem of low speech recognition efficiency. The method mainly comprises the steps of obtaining to-be-recognized voice and interaction scene information of the to-be-recognized voice; recognizing the voice to be recognized to obtain an initial conversion text and a multi-dimensional confidence feature of the initial conversion text; using the interaction context information as a scene matching basis, calculating a text credibility parameter of the initial conversion text together with the multi-dimensional confidence feature, and determining a target path for executing a text error correction task from three-level routing according to the text credibility parameter, an interaction category and an error correction routing strategy; and performing text error correction processing on the initial conversion text based on the target path to generate a target conversion text of the to-be-recognized voice. The method is mainly used for text error correction based on dynamic routing so as to improve the speech recognition efficiency.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Voiceprint recognition method and device, computer device, storage medium and program product

The application relates to a voiceprint recognition method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: obtaining to-be-recognized voice data, inputting the to-be-recognized voice data into a preset voiceprint recognition model, and obtaining a voiceprint recognition result of the to-be-recognized voice data. The method can apply a pre-trained voiceprint recognition model to recognize to-be-recognized voice data and obtain a voiceprint recognition result. Since the to-be-trained voiceprint recognition model references the convolutional layer parameters of a generative adversarial network model trained by data augmentation on a small sample voice data training set during training, the to-be-trained voiceprint recognition model references the knowledge obtained by training a large sample voice data set during training, which can further accelerate the convergence rate of the voiceprint recognition model training and improve the accuracy of the voiceprint recognition model recognition.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Speech recognition method and server

The application relates to a speech recognition method and a server. The method comprises the following steps: obtaining a to-be-recognized speech signal; recognizing each frame of the to-be-recognized speech signal according to an acoustic model of each language, and respectively outputting corresponding language phonemes and prediction probabilities; wherein the acoustic model of each language is respectively constructed according to shared hidden layer training; sequentially traversing a sentence decoding graph and a multi-language slot decoding graph connected with each other to obtain a corresponding path; wherein the sentence decoding graph is used for decoding phonemes entering a non-slot, and the slot decoding graph is used for decoding phonemes entering a slot; when it is determined that the path passes through the multi-language slot decoding graph in the speech decoding graph, screening the path according to the prediction probabilities of the language phonemes corresponding to each language, and determining the text information corresponding to the target path as a speech recognition result. The scheme provided by the application can accurately recognize mixed multi-language speech information.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

System with voice recognition and verbal command control

The invention relates to the field of digital information technology. A hardware and software system for controlling a robotic system (RS) includes an actuator and an RS control device. The RS control device includes a voice recognition unit which is capable of recognizing voice commands and is connected to a command vocabulary storage unit, a robot command identification unit for identifying commands on the basis of recognized vocabulary and information stored in the storage unit, a communication unit that transmits the command to the RS control device, and a feedback device. The claimed solution lends universality to the verbal command control of robotic systems.
Owner:AUTONOMOUS CLUSTER FUND PARK OF INNOVATIVE TECHNOLOGIES

Speech recognition method and device, electronic equipment and storage medium

ActiveCN117219063Breduce the number of elementsReduce the amount of decoding calculationsPrediction probabilitySpeech sound
Embodiments of the present application provide a speech recognition method and device, electronic equipment and storage medium, at least applied to the field of artificial intelligence and speech recognition, wherein the method comprises: performing vector coding processing on the audio feature vector of the speech to be recognized to obtain an audio coding vector; performing classification processing on the audio coding vector to obtain a prediction probability distribution of each predicted character in a preset vocabulary corresponding to each speech frame in the speech to be recognized; performing pruning processing on the audio coding vector based on the prediction probability distribution to obtain a pruned audio coding vector; and performing speech recognition on the speech to be recognized based on the pruned audio coding vector to obtain a speech recognition result. Through the present application, the decoding calculation amount in the speech recognition process can be reduced, the decoding efficiency can be improved, and thus the speech recognition efficiency can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1

Voice interaction method, device, server and computer readable storage medium

The application discloses a voice interaction method, comprising: receiving a voice request forwarded by a vehicle; encoding the voice request according to a preset model to obtain an encoded sequence matrix for slot recognition; dividing the encoded sequence matrix; performing slot recognition on the voice request according to the divided encoded sequence matrix; performing application program interface prediction on the voice request; selecting a predicted application program interface to perform application program interface parameter filling according to the result of slot recognition and the predicted application program interface; and outputting an execution result to the vehicle to complete voice interaction. The application encodes the voice request, performs slot recognition according to the divided encoded sequence matrix, performs application program interface parameter filling according to the result of slot recognition, and finally outputs an execution result and sends it to the vehicle, effectively improving the accuracy of recognizing slots composed of discontinuous words in the voice request and improving the voice interaction experience of users.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

System and method for identifying sentiment (emotions) in a speech audio input

In a system and method for enabling a user to identify the emotions of speakers during a telephone or online conversation, spoken audio input is pre-processed using a one-dimensional Mel Spectrogram and / or a two-dimensional Mel-Frequency Cepstral Coefficient (MFCC) matrix, reducing the two-dimensional matrix to a single dimension output, and identifying at least one emotion in the audio input using a convolutional or recurrent neural network.
Owner:VALENCE VIBRATIONS INC

Conversation system for voice control of motor vehicle

The invention relates to a dialogue system (100) for voice control of a motor vehicle (105), the dialogue system (100) comprising a device (110) for detecting a voice input (315, 325, 335, 345) of a person (310, 320, 330, 340) on the motor vehicle (105); and a processing device (120) arranged to recognize a language of the speech input (315, 325, 335, 345) on the basis of the speech marker and to provide a response (350) to the speech input (315, 325, 335, 345), the response (350) being provided in the recognized language.
Owner:BAYERISCHE MOTOREN WERKE AG

Systems and methods for improving data security by generating surrounding image sonification and dynamically adjusting graphical user interfaces

ActiveUS12554820B2Image analysisSpeech analysisSonificationEngineering
Systems, computer program products, and methods are described herein for improving data security by generating surrounding image sonification and dynamically adjusting graphical user interfaces. The present disclosure is configured to receive a data transmission request and at an authentication credential; identify a voice input, facial input data, and a physical characteristic input data; compare the voice input with a voice authentication, the facial input data with a facial authentication data, and the physical characteristic input data with a physical characteristic authentication; authenticate a user based on the comparison; receive an expected environment user input; receive at least one real-time image of the real-time geographic environment; analyze the at least one real-time image; generate a real-time geographic environment indication; transmit the real-time geographic environment indication; receive a real-time environment authentication; and authenticate the data transmission request based on the authentication of the user and the real-time environment authentication.
Owner:BANK OF AMERICA CORP

system

Provide a system. 【Solution means】 Means having a function of inputting basic information of a family, Mathematical means for calculating the admission permission points of a childcare institution based on the input information, Means for estimating the probability of admission based on the calculated points, Media generation means for visually presenting the estimated admission probability, Means for automatically generating necessary application records, Means for giving advice for proposing an optimal childcare institution selection policy, Means for immediately updating regional information and admission competition rate information of childcare institutions, Means for cooperating with a device capable of recognizing voices and having a conversation within a family, and providing the advice and information conversationally, A system including the above.
Owner:SOFTBANK GROUP CORP

system

The system of the embodiment aims to remotely monitor the health status of elderly people living alone and detect the risk of dementia and depression at an early stage. [Solution] A system according to an embodiment includes a voice recognition unit, a conversation recording unit, a health monitoring unit, and an alert unit. The voice recognition unit recognizes voices. The conversation recording unit records conversations based on the voices recognized by the voice recognition unit. The health monitoring unit monitors health conditions based on the conversations recorded by the conversation recording unit. The alert unit sends an alert for dementia or depression based on information monitored by the health monitoring unit.
Owner:SOFTBANK GROUP CORP

Multi-modal interactive experience platform for teenager education

The invention discloses a teenager education-oriented multi-mode interactive experience platform, and relates to the technical field of teenager education, interaction with a 5G intranet is carried out through a gigabit Ethernet, and user interaction comprises a 4KVR head display, a touch screen and a voiceprint recognition voice unit; the resource library is divided into four types of historical scenes and historical events and is updated quarterly; the generation module outputs multi-format content by using a GPT-4 lightweight model and verifies the multi-format content by an expert; intelligent adaptation is based on age, cognition and other portrait adaptation contents; a distributed architecture is adopted for storage, and a learning report is analyzed and output; safety comprises double content auditing and data desensitization; operation and maintenance are monitored for 7 * 24 hours, and remote operation and maintenance and automatic updating are supported; the problems that traditional education is single in form, poor in adaptation and insufficient in immersion are solved. Through multi-modal interaction, intelligent cognitive adaptation, safety control and closed-loop optimization, the interest and emotion resonance of teenagers are stimulated, and the attraction and effectiveness of cultural education are improved.
Owner:ANHUI POLYTECHNIC UNIV

Systems and methods for improving data security by generating surrounding image sonification and dynamically adjusting graphical user interfaces

ActiveUS20260030329A1Image analysisSpeech analysisSonificationEngineering
Systems, computer program products, and methods are described herein for improving data security by generating surrounding image sonification and dynamically adjusting graphical user interfaces. The present disclosure is configured to receive a data transmission request and at an authentication credential; identify a voice input, facial input data, and a physical characteristic input data; compare the voice input with a voice authentication, the facial input data with a facial authentication data, and the physical characteristic input data with a physical characteristic authentication; authenticate a user based on the comparison; receive an expected environment user input; receive at least one real-time image of the real-time geographic environment; analyze the at least one real-time image; generate a real-time geographic environment indication; transmit the real-time geographic environment indication; receive a real-time environment authentication; and authenticate the data transmission request based on the authentication of the user and the real-time environment authentication.
Owner:BANK OF AMERICA CORP

Automatic bus stop reporting system based on single-chip microcomputer control

The invention discloses an automatic bus stop reporting system based on single-chip microcomputer control. The automatic bus stop reporting system comprises a single-chip microcomputer control unit, a wireless transceiver module, a voice module, a display module, a temperature sensor and a clock module. The single chip microcomputer control unit is used for processing and controlling the whole station reporting system; the wireless receiving and transmitting module is used for receiving and transmitting wireless signals and identifying station names; the voice module is used for converting the station name information into voice broadcast; the display module is provided with a liquid crystal display screen, and the liquid crystal display screen is connected with a P0 port of the single chip microcomputer; the temperature sensor is connected with a P3.3 port of the single chip microcomputer; and the clock module is connected with P1.0 to P1.2 ports of the single chip microcomputer through a serial interface.
Owner:WUXI KAILAI MICROELECTRONICS CO LTD

Vehicle anti-theft and identity authentication method and system based on voiceprint recognition

The invention relates to the technical field of voice recognition, and particularly discloses a vehicle anti-theft and identity authentication method and system based on voiceprint recognition, and the method comprises the steps: obtaining to-be-recognized voice information for vehicle recognition, and obtaining to-be-recognized voiceprint feature information according to the to-be-recognized voice information, the to-be-recognized voiceprint feature information comprises scene voiceprint features and personnel voiceprint features; acquiring real-time scene voiceprint feature data of a preset time period before and after to-be-recognized voice information according to the scene voiceprint feature, and acquiring scene voiceprint similarity according to the real-time scene voiceprint feature data and the to-be-recognized voice information; and acquiring vocal cord fundamental frequency features, pronunciation comparison feature data and inter-frame vocal print time sequence relevance features according to the vocal print features of the personnel. According to the invention, through two-dimensional fusion of the scene voiceprint features and the personnel voiceprint features, the precision and reliability of vehicle identity recognition are greatly improved.
Owner:XINYOUXI TRAVEL TECHNOLOGY (HANGZHOU) CO LTD

Conversational control of an appliance

A method of operating an appliance includes obtaining a sound signal using a microphone of the appliance, analyzing the sound signal to identify a voice input including a request for a responsive action, determining an action confidence metric regarding performance of the responsive action, determining that the action confidence metric exceeds a predetermined action threshold, determining that the responsive action to the voice input is needed based on the action confidence metric exceeding the predetermined action threshold, and performing the responsive action.
Owner:HAIER US APPLIANCE SOLUTIONS INC