Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

126 results about "Speech interaction" patented technology

Voice interaction task execution method and device based on large model, equipment and medium

The invention discloses a voice interaction task execution method and device based on a large model, equipment and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: carrying out the voice recognition of voice features, obtaining a text character sequence, and carrying out the optimization of the text character sequence through a preset large language model, and obtaining a target text character sequence; performing entity recognition on the target text character sequence by using a preset large language model to obtain an entity recognition result, determining a relationship type among entities in the target text character sequence according to the entity recognition result, and constructing a knowledge graph according to the relationship type among the entities; generating an initial triple based on the knowledge graph, and optimizing the initial triple by using a preset large language model to obtain a target triple; and fusing the target triple with the initial knowledge graph, and executing a voice interaction task in the target voice interaction scene based on the updated knowledge graph. According to the method and the device, the efficiency and the accuracy of extracting the structured knowledge from the Chinese speech are improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Interactive method based on large model, training method, intelligent agent, device, and medium

An interactive method based on a large model, a training method, and an intelligent agent, which relate to fields of artificial intelligence, speech recognition, speech interaction, deep learning, large models, and application scenarios of knowledge search, autonomous driving, intelligent customer service, intelligent speech control, smart e-commerce, AI healthcare. The interactive method includes: acquiring a request speech; performing a speech recognition on the request speech to obtain a speech recognition feature representing a request semantics; and processing the speech recognition feature using the large model to obtain a response text, where the response text includes response words arranged in sequence, a target response word among the response words is determined by processing the speech recognition feature and an associated response word feature using an attention fusion layer of the large model, and the associated response word feature is related to an associated response word arranged before the target response word.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Sound duplicating and low-delay streaming speech synthesis method and system based on ultra-short sample

The invention provides a voice replication and low-delay streaming voice synthesis method and system based on an ultra-short sample, relates to the technical field of artificial intelligence, and is suitable for intelligent interaction, outbound service and multi-mode communication scenes. Deep personalized customization of intelligent voice interaction and real-time generation of ultra-low delay are realized, and bidirectional streaming interaction of a system level is supported, so that the fluency and response speed of dialogues are improved; an ultra-short sample sound duplicating module specially designed for processing ultra-short audio samples and a speech synthesis engine with ultra-low delay and bidirectional streaming output capability are integrated, and the capabilities of the ultra-short sample sound duplicating module and the speech synthesis engine are applied to a real-time and interactive intelligent speech interaction process to form a complete and efficient solution. The industrial pain point is directly solved in a targeted manner, and the method has important commercial application value.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Method for processing cross-modal question answerning based on large model, apparatus and storage medium

A method for processing cross-modal question answering based on large model, an apparatus, and a storage medium are suggested, which relates to the field of artificial intelligence technologies such as speech interaction processing, large models, machine learning and natural language processing. The specific implementation includes: performing an activity detection on a target speech input by a user; in response to detecting a pause in the inputting of the target speech, obtaining a first text corresponding to a first input speech before the moment of the pause in the target speech; performing a text response processing using a pre-trained speech question answering processing system based on the first text and the first input speech.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Speech interaction method, apparatus, and system

A speech interaction method includes receiving, by a server, a first play message, where the first play message includes an identifier of first audio content corresponding to a first non-speech instruction. The server determines a first intent and first slot information that correspond to the first non-speech instruction. In response to the first play message, the server instructs a playback device to play the first audio content. The server receives a first speech instruction input by a user into the playback device, where a second intent or second slot information or both in the first speech instruction are incomplete. The server determines, based on the first intent and the first slot information, the second intent and the second slot information that correspond to the first speech instruction, and the server, based on the second intent and the second slot information, instructs the playback device to play second audio content.
Owner:HUAWEI TECH CO LTD

Speech interaction method, speech interaction system and storage medium

A speech interaction method, a speech interaction system and a storage medium. The method includes: receiving interactive speech input by a user; determining an emotional tag corresponding to the interactive speech based on the interactive speech and interactive text corresponding to the interactive speech; determining, based on the emotional tag, a response text corresponding to the interactive text, and a first prosodic feature and a second prosodic feature corresponding to the response text. The first prosodic feature is used for characterizing the whole sentence prosodic feature of the response text, and the second prosodic feature is used for characterizing a local prosodic feature of each character in the response text; and generating and outputting a response speech corresponding to the interactive speech based on the response text, the first prosodic feature and the second prosodic feature.
Owner:NANJING SILICON INTELLIGENCE TECH CO LTD

Multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and medium

ActiveCN120954388ASpeech recognitionSpeech synthesisSpeech comprehensionModal voice
The invention discloses a multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and a medium, and relates to the technical field of multi-modal voice interaction.The method comprises the steps that a text token of a text training sample is obtained, a corresponding voice token is constructed, and pre-training data used for converting the text token into the voice token is obtained; in combination with the multi-modal input sample and the pre-training data, constructing fine-tuning training data for speech understanding and dialogue generation; constructing and pre-training a basic model by using the pre-training data; a multi-modal voice interaction large model is constructed based on a pre-training basic model, and fine tuning data is used for training, so that voice acoustic features can be regulated and controlled based on multi-modal input, and voice is output. According to the method, through alignment and staged training of the text token and the voice token, fine regulation and control of voice acoustic characteristics are realized, long voice continuity and interaction naturalness are improved, and a model is efficiently endowed with voice interaction capability of controllable timbre and emotion.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

VR large-space active noise reduction and directional speech enhancement system

The invention discloses a VR large-space active noise reduction and directional speech enhancement system, particularly relates to the technical field of speech enhancement, and aims to obtain a user attention direction by fusing multi-mode sensor data in user wearing equipment, guide a sound source positioning module to quickly lock a speech interaction main direction in a multi-sound-source interference environment, and improve the speech enhancement efficiency. Non-target speakers are further filtered through voiceprint feature recognition, it is ensured that enhanced voices have individual consistency, through deep anti-reverberation modeling, reflection residual sounds generated by multi-path propagation are effectively stripped, semantic fuzziness caused by pseudo sound source aliasing is avoided, a voice enhancement and active noise reduction module corrects the frequency spectrum phase while enhancing the amplitude spectrum, and the noise reduction effect is improved. The definition and naturalness of output voice are remarkably improved, the enhanced focusing performance of the system is monitored in real time by introducing a target locking drift coefficient and a beam pointing deviation coefficient, and adaptive reconstruction is triggered based on a closed-loop correction index.
Owner:HANGZHOU KAILIN CULTURE TECHNOLOGY CO LTD +1

Microphone voice recognition system and method based on multi-mode audio-visual fusion

The invention discloses a microphone voice recognition system and method based on multi-mode audio-visual fusion, and belongs to the technical field of artificial intelligence and voice interaction. Firstly, an audio module collects voice signals through a microphone, voice is converted into texts by means of a cloud voice recognition API, and words are further mapped into 300-dimensional semantic vectors through Word2Vec. A visual module extracts lip movement and log-Mel frequency spectrum features, after being subjected to Dlib detection and normalization processing, lip images are sent into a 3D CNN and a dense space-time CNN to extract space-time features, key areas are highlighted with the assistance of a space attention mechanism, and finally sequence visual features are extracted through bidirectional GRU. Meanwhile, a log-Mel spectrogram is generated from the audio signal, and the perception characteristic is enhanced through Mel filtering and logarithm processing. The audio word vector, the lip movement feature and the log-Mel feature are spliced into a multi-modal fusion vector, the multi-modal fusion vector is sent to a CTC decoder, and a text is predicted through Beam Search decoding. An Adam optimizer and a small-batch training strategy are used in the training process, and the model performance and generalization ability are improved.
Owner:ZHONGSHAN FENGXU ELECTRONIC IND CO LTD

Transform-based voice correction system and method

The invention provides a Transform-based voice correction system and method, and the system comprises a voice preprocessing module which is used for carrying out the noise suppression and feature extraction of an input southern Fujian dialect voice signal, carrying out the noise suppression through a frequency spectrum subtraction method, and extracting 40-dimensional MFCC features; the voice splicing module is used for solving the problem of voice input fragmentation; the dialect speech recognition module adopts an improved Transform encoder structure; and the character correction module adopts a character level Transform language model. According to the invention, more accurate and smooth voice interaction experience is provided for the user in the dialect area in southern Fujian. Meanwhile, the technical framework has good expandability, can adapt to other dialects and language variants, and provides important support for industrial application of the dialect voice technology.
Owner:ZHANGZHOU THIRD HOSPITAL (ZHANGZHOU LONGWEN HOSPITAL)

Airborne flight safety enhancement system and method based on multi-modal interaction

The invention discloses an airborne flight safety enhancement system and method based on multi-modal interaction, and relates to the field of crossing of avionics systems and artificial intelligence, and the method comprises the steps that an environment sensing module generates structured environment information; the airborne equipment monitoring module performs real-time health state analysis; the large language model module comprises a computing platform sub-module, a language model sub-module, a knowledge base sub-module, a voice recognition sub-module and a voice synthesis sub-module; and the output module is used for displaying the decision support information generated by the large language model module on a cockpit display device. The method has the advantages that through deep cooperation of multi-mode perception fusion, distillation large language model intelligent decision, natural voice interaction and airborne health monitoring, on the premise of ensuring airworthiness safety, the flight safety is remarkably improved, the pilot load is reduced, and the decision efficiency is optimized.
Owner:上海柘飞航空科技有限公司

Emotional voice interaction method and device based on humanoid robot

The invention provides an emotional voice interaction method and device based on a humanoid robot, and relates to the field of robot interaction. The method comprises the following steps: acquiring voice data of a target user through a humanoid robot, and extracting acoustic features and semantic features in the voice data; inputting the acoustic features and the semantic features into a cross-modal fusion emotion classification model, and outputting an emotion recognition result corresponding to the target user; determining an emotion recognition misjudgment risk, and correcting an emotion recognition result based on the emotion recognition misjudgment risk; inputting the corrected emotion recognition result and the semantic feature into an emotion voice understanding interaction model, and outputting an emotion voice interaction result; and controlling the humanoid robot to execute a corresponding multi-modal interaction response based on the emotion voice interaction result. The problem that the voice interaction experience is greatly reduced due to the fact that the voice interaction content output by an existing humanoid robot is difficult to fit the real emotional state of the elderly user in multiple aspects such as text expression and voice emotion is solved.
Owner:BEIJING SHENMOU TECH CO LTD

Shared space multi-user silent voice interaction method and device and medium

The invention relates to a shared space multi-user silent voice interaction method and device and a medium, and the method comprises the following steps: taking a frequency modulation continuous wave signal as an audio carrier and a sensing medium at the same time, directionally broadcasting an audio in a shared space, and capturing a silent voice signal in a multi-user scene through lip motion reflection; performing feature extraction and separation on the silent voice signals to obtain silent voice features of each user; on the basis of silent voice features, a lip movement analysis model driven by an integrated sliding window is adopted, silent voice content of a user is recognized in real time, and real-time interaction is achieved. Compared with the prior art, multi-user silent voices can be accurately recognized in a complex environment, and silent interaction in a privacy sensitive scene is achieved.
Owner:SHANGHAI JIAOTONG UNIV

User feedback for speech interactions

An interactive system may be implemented in part by an audio device located within a user environment, which may accept speech commands from a user and may also interact with the user by means of generated speech. In order to improve performance of the interactive system, a user may use a separate device, such as a personal computer or mobile device, to access a graphical user interface that lists details of historical speech interactions. The graphical user interface may be configured to allow the user to provide feedback and / or corrections regarding the details of specific interactions.
Owner:AMAZON TECH INC

A method and apparatus for echo cancellation and noise reduction

Embodiments of the present application disclose a method and device for echo cancellation and noise reduction. In the embodiments of the present application, a first near-end speech and a first far-end speech are obtained, and after pre-processing, delay detection and short-time Fourier transform, a fourth near-end speech and a fourth far-end speech are generated, and linear echo cancellation is performed to generate an error signal and a linear echo signal of the linear echo cancellation; the fourth near-end speech, the fourth far-end speech, the error signal and the linear echo signal of the linear echo cancellation are input into a two-stage neural network multi-task fusion model to be trained to determine a joint loss function; the two-stage neural network multi-task fusion model is generated according to the joint loss function, and then model compression is performed to generate a target two-stage neural network multi-task fusion model, which is used to realize echo cancellation and noise reduction. Through the above method, acoustic echo cancellation and noise reduction are realized under the condition of extremely low signal-to-echo ratio, and the speech interaction and communication performance of a terminal device are improved.
Owner:ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD

SEMANTIC ENTITY RECOGNITION METHOD AND DEVICE, LANGUAGE INTERACTION SYSTEM, VEHICLE AND MEDIUM

This application relates to a semantic entity recognition method and a semantic entity recognition device, a speech interaction system, a vehicle, and a computer-readable storage medium. The semantic entity recognition method includes: preprocessing speech instruction information to obtain a core-mentioned word contained within the speech instruction information; and searching a scenario-semantically directed structure graph based on the core-mentioned word to obtain a target entity object that matches the core-mentioned word, wherein the scenario-semantically directed structure graph is constructed from a scenario entity object in a current scenario and semantically associated words of different granularities of the scenario entity object.
Owner:BYD CO LTD

An ethical AI integration framework for smart home systems that use voice assistants

System (100) for integrating ethical artificial intelligence for smart home systems using voice assistants, wherein the system (100) comprises: a speech interaction module (104) that is used to receive and is configured to process voice-based user commands; an analysis module (106) configured to interpret the processed speech commands based on context parameters; Ethics assessment modules (108-1) configured to make automated decisions regarding bias, Assess transparency and data protection; a smart home control module (108-2) configured to perform automated actions based on ethically evaluated decisions; and a monitoring module (110) configured to monitor system behavior and ensure ongoing compliance with predefined ethical standards.
Owner:DR VISHWANATH KARAD MIT WORLD PEACE UNIV PUNE

Teaching speech interaction-oriented education metadata processing method and device, and storage medium

PendingCN121071590AData processing applicationsArtificial lifeEngineeringEducational metadata
The invention relates to a teaching speech interaction-oriented education metadata processing method and device, and a storage medium, and the method comprises the steps: carrying out the slicing and sentence segmentation of classroom teacher and student speech corpus data, and carrying out the sampling and aggregation to form a plurality of sentence blocks; based on a preset labeling instruction and the knowledge graph, performing word-level labeling of teaching behaviors and knowledge points on the sentence blocks by utilizing an intelligent agent, and generating a behavior container through majority voting; based on a preset discrimination instruction, the labeled sentence tag and the behavior definition of the tag, performing anti-fact verification by utilizing an intelligent agent, judging whether the sentence tag is related to the behavior definition or not, and if so, storing the sentence tag as a final labeling result in a preset graph database; and if not, triggering an error correction process, generating a new label by the intelligent agent based on the correction instruction, and selecting the label with the highest category overlapping degree as a result by matching with the positive example set and sorting similarity. And the correction result is fed back to the analysis process, so that iterative optimization update of the labeling judgment and correction module is realized.
Owner:SHANGHAI NORMAL UNIVERSITY +1

Speech recognition method and electronic device

PCT designated stageWO2026045574A1Speech recognitionPhonetic environmentAudio electronics
A speech recognition method and an electronic device. The method comprises: an electronic device acquiring audio which is input by a user (S801); on the basis of the audio, the electronic device detecting a speech environment when the user inputs the audio, and determining a speech interaction mode corresponding to the audio (S802); and the electronic device performing speech recognition on the audio on the basis of a speech recognition scheme corresponding to the speech interaction mode, so as to obtain a speech recognition result (S803). By detecting the environment in which the user inputs the audio, the electronic device can determine that the user is in a noisy, quiet, or another speech environment, and can then determine the speech interaction mode. The electronic device can execute corresponding speech recognition on the acquired audio by means of the speech recognition scheme corresponding to the speech interaction mode, thereby improving the accuracy of speech recognition.
Owner:HUAWEI TECH CO LTD

Speech interaction device and speech interaction method

A speech interaction device comprising a processor for controlling a speech interaction of a user, the processor configured to: recognize interaction contents of the user before interruption of the speech interaction; extract, on a basis of the interaction contents, a keyword in the speech interaction before the interruption as a pre-interruption keyword; to generate, on a basis of the extracted pre-interruption keyword, an icon image representing the interaction contents before the interruption; and generate a control command to display the generated icon image. The processor is configured to update, after the icon image is displayed, the icon image displayed to the user on a basis of a topic subject in which the user may have an interest and / or a driving load of the user who drives a vehicle during the interruption of the speech interaction.
Owner:NISSAN MOTOR CO LTD

A speech recognition method

The application provides a speech recognition method, which comprises: obtaining feature data of audio to be recognized; inputting the feature data into an acoustic model to obtain a time sequence label matrix corresponding to the feature data; decoding the time sequence label matrix through a first language model to obtain a plurality of decoding paths and corresponding probability scores thereof, and determining the decoding paths with the top N probability scores as N first decoding results, wherein N is a positive integer; determining a corresponding target intent field based on the N first decoding results, a previous round of speech interaction field and a current scene field; determining a corresponding second language model based on the target intent field, re-computing probability values of the decoding paths with the top N probability scores through the second language model, and generating second decoding results; and determining a speech recognition result of the audio to be recognized based on the second decoding results. The application improves the recognition efficiency while ensuring the speech recognition accuracy.
Owner:HUBEI QIGUANG TECHNOLOGY CO LTD

Speech interruption decision method and system based on multi-granularity semantic completeness prediction

PendingCN122313981ASemantic featureData mining
This invention discloses a speech interruption decision-making method and system based on multi-granularity semantic integrity prediction, belonging to the field of speech interaction technology. The method includes: extracting multi-granularity semantic features and combining them into a current multi-granularity semantic feature vector; calculating the offset of the current multi-granularity semantic feature vector with a standard semantic pause model to obtain a semantic offset feature vector; performing similarity matching and weighted voting between the semantic offset feature vector and a historical interruption decision case library to obtain the current predicted interruption confidence; and generating an interruption command when the current predicted interruption confidence exceeds a dynamic threshold. This invention solves the problems of existing technologies where speech interruption decisions rely on simple energy thresholds or semantic integrity probabilities, lack utilization of user expression habits and contextual experience, and have low decision accuracy by introducing a standard semantic pause model and a historical case analogy reasoning mechanism, thus achieving more intelligent and accurate speech interruption judgment.
Owner:GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

Cross-modal attention fusion method and device based on voiceprint features and language semantics

The application relates to the technical field of speech recognition, and discloses a cross-modal attention fusion method and device based on voiceprint features and language semantics, which comprises the following steps: performing a feature extraction operation on speech data to obtain voiceprint features; based on a sentiment perception mask mechanism and a context sentiment memory unit, fusing sentiment features into text data to obtain semantic features; constructing a cross-modal relationship graph of the voiceprint features and the semantic features; determining a time sequence dependency relationship between the voiceprint features and the semantic features based on the edge weights generated by the cross-modal relationship graph; projecting the voiceprint features and the semantic features into the same semantic space based on the time sequence dependency relationship, and performing a feature reconstruction fusion operation based on a reconstruction loss function to obtain voiceprint and semantic fusion features; and adjusting a dialogue strategy determined based on the voiceprint and semantic fusion features; and using the adjusted dialogue strategy to generate interactive response data with historical interactive data, so that a high-accuracy and personalized speech interaction experience is realized.
Owner:GUANGDONG GUANGXIN COMM SERVICES COMPANY

Automobile air-conditioner control method and air-conditioner control system

Provided in the present disclosure are an automobile air-conditioner control method and an air-conditioner control system. The air-conditioner control system comprises a speech interaction device, an infotainment host, a central operation module and a zone controller, wherein the speech interaction device collects user's speech, and sends the speech to the infotainment host; the infotainment host generates an air-conditioner control instruction on the basis of the speech, and sends the air-conditioner control instruction to the central operation module; the central operation module determines, on the basis of the instruction type of the air-conditioner control instruction, whether a vehicle meets an air-conditioner control condition, and if the vehicle meets the air-conditioner control condition, the central operation module sends the air-conditioner control instruction to the zone controller; and on the basis of the air-conditioner control instruction, the zone controller controls an automobile air conditioner to operate. In this way, an air-conditioner control instruction is accurately generated by identifying user's speech, whether a vehicle meets an air-conditioner control condition is verified, and on the basis of the rational air-conditioner control instruction that passes the verification, the automobile air-conditioner is controlled to operate, such that an air conditioner in a vehicle can be conveniently controlled by means of the user's speech, thereby avoiding the driving safety risk caused by means of manual adjustment.
Owner:CHINA FAW CO LTD

Voice noise reduction method and device, earphone, medium and program product

The invention discloses a voice noise reduction method and device, an earphone, a medium and a program product. The method comprises the following steps: acquiring voice sequence data; performing at least one self-attention operation on the voice sequence data; obtaining the output of the self-attention mechanism; noise reduction is carried out according to the output of the self-attention mechanism; the sg attention operation comprises the following steps: carrying out linear processing on the input sequence data by utilizing a preset query-key matrix and a preset value matrix to obtain a query-key vector and a value vector; the query-key matrix corresponds to a product between a preset query matrix and a preset key matrix; obtaining a target attention score according to the transpose of the voice sequence data and the query-key vector; carrying out Softmax operation on the target attention score to obtain an attention weight; and performing matrix multiplication on the attention weight and the value vector to obtain the output of the self-attention operation. In this way, the harsh requirement of the embedded device for high real-time performance in a voice interaction scene can be met.
Owner:BESTECHNIC SHANGHAI CO LTD

Multi-language natural voice interaction method and system for coffee robot

PendingCN121768397AMeet individual needsImprove consumer experienceSpeech recognitionEngineeringAcoustics
The invention relates to the technical field of voice services, in particular to a multi-language natural voice interaction method and system for a coffee robot, and the method comprises the steps: collecting a voice signal of a user, carrying out the recognition of the language type of the voice signal, obtaining the type information of the voice signal, activating a corresponding language recognition model according to the type information, and carrying out the recognition of the type information; performing semantic analysis processing on the voice signal, generating language text information corresponding to the voice signal, performing coffee customization demand analysis on a user based on the language text information to obtain coffee customization information of the user, initiating a voice interaction behavior to the user according to the coffee customization information to obtain an interaction feedback instruction of the user, and sending the interaction feedback instruction to the user. And a corresponding coffee beverage is produced according to the interaction feedback instruction and the coffee customization information.
Owner:SHENZHEN CHUANGJIE INTELLIGENT TECHNOLOGY CO LTD

Call method, and electronic device, storage medium and product

The present disclosure relates to the technical field of computers, and particularly relates to a call method, and an electronic device, a storage medium and a product. The call method comprises: in an interaction interface of a first object and a second object and on the basis of the input of the first object, determining a target scenario mode from among one or more scenario modes configured for the second object (S102), wherein each scenario mode is configured with a speech feature; and controlling the second object to perform speech interaction with the first object on the basis of the speech feature of the target scenario mode (S104).
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Speech interaction method and storage medium based on decision-making and large language model of hybrid training strategy

The present invention relates to a speech interaction method and storage medium based on a decision-making and large language model of a hybrid training strategy. The purpose of the present invention is to solve the problem that the existing large language model cannot provide accurate answers due to the lack of sufficient domain knowledge, and the use of two large language models will bring high computational costs and long response time. The process is: setting a specific answer format; constructing a decision data set; constructing a dialogue data set for dialogue questions and answers in a specific scenario; fine-tuning the large language model for the first time using full parameter fine-tuning based on the dialogue data set to obtain a large language model after the first fine-tuning; fine-tuning the large language model after the first fine-tuning for the second time using LoRA based on the decision data set to obtain a large language model after the second fine-tuning; connecting the speech recognition module and the speech synthesis module to the large language model after the second fine-tuning, processing the user's voice questions to be tested, and generating speech to interact with the user.
Owner:HARBIN INST OF TECH

Expression control method and system for driving lip, tooth and tongue linkage of humanoid robot based on anatomy, facial topological structure and language pronunciation rules

The invention discloses a humanoid robot expression control method and system based on anatomy and a facial topological structure, relates to the technical field of bionic robots and voice interaction, and aims to solve the problems that the mouth shape and pronunciation of an existing robot are not matched, the expression is rigid, and no natural micro-expression exists in a non-pronunciation state. A facial topological structure model matched with a robot is constructed based on human anatomy, a lip-tooth-tongue time sequence synchronous linkage control instruction is generated in combination with a language pronunciation rule, facial expressions, micro expressions and eye movement data matched with pronunciation are synchronously generated, meanwhile, an independent expression control instruction can be generated in a non-pronunciation state, and the control precision is improved. And finally, full-scene full-face natural expression control of the robot is realized. The humanoid reality and man-machine interaction naturalness of the humanoid robot are greatly improved, and the method can be widely applied to various bionic service robots and virtual digital humans.
Owner:孙量