Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1032 results about "Speech input" patented technology

Speech input is one of the most innovative browser technologies to appear in recent months. It’s easy to implement and there are several obvious uses: assistive dictation for those with impaired mobility. an alternative input option for mobile phones and tablets, and. any environment where a keyboard or mouse is impractical.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Enterprise intelligent finance and tax system fusing supervised learning and block chain

The invention provides an enterprise intelligent finance and taxation system fusing supervised learning and a block chain, and solves the core technical problems of insufficient data credibility, transparent contradiction between privacy protection and supervision, AI model training data island and the like of a traditional finance and taxation system. According to the system, an innovative technical fusion scheme is adopted, bank API, OCR recognition, voice input and other multi-source financial data are integrated through a multi-modal data fusion unit, and high-quality data fusion is achieved through confidence evaluation and an intelligent conflict resolution algorithm; an intelligent classification unit based on BERT deep learning integrates a pre-training model with a business rule engine and a gradient boosting decision tree, and accurate and automatic classification of financial transactions is achieved. A complete technical solution is provided for enterprise finance and taxation digital transformation, intelligent, automatic and credible processing of finance and taxation businesses is achieved on the premise that data safety and privacy protection are guaranteed, and the method has wide market application prospects.
Owner:张宏

Voice interaction system and method for customer service based on artificial intelligence

The invention relates to the technical field of voice recognition, in particular to a voice interaction system and method for customer service based on artificial intelligence, and the system comprises a voice input processing module, an intention classification and routing module, a context dynamic adjustment module, a user behavior learning module, a multi-level intention fusion module and a final result module. According to the method, a multi-dimensional feature system is constructed by extracting tone intensity, speech speed frequency and emotional fluctuation amplitude, intention categories, priority weights and confidence scores are generated to realize accurate acquisition of appeals, and dialogue history, context and emotional change dynamic reconstruction path nodes, switching rules and response time sequences are tracked during interaction. Historical behavior mining preference features, habit fusion intention relevance, emergency calculation of an optimal strategy, construction of service steps, resource allocation schemes and execution timelines, adjustment of an interactive interface, a service process and a feedback mechanism according to multi-dimensional analysis, guarantee of differentiated service experience, and improvement of response accuracy and user satisfaction.
Owner:NANJING XIUGUO INTELLIGENT TECH CO LTD

Wearable electronic devices, extended reality systems including neuromuscular sensors, and methods for generating text from speech input and modifying the generated text based on neuromuscular data

The disclosed system for interacting with objects in an extended reality (XR) environment generated by an XR system may include (1) neuromuscular sensors configured to sense neuromuscular signals from a wrist of a user and (2) at least one computer processor programmed to (a) determine, based at least in part on the sensed neuromuscular signals, information relating to an interaction of the user with an object in the XR environment and (b) instruct the XR system to, based on the determined information relating to the interaction of the user with the object, augment the interaction of the user with the object in the XR environment. Other embodiments of this aspect include corresponding apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
Owner:META PLATFORMS TECHNOLOGIES LLC

Intelligent voice interaction system and method based on streaming multi-mode fusion and equipment control protocol

PendingCN121260156ASpeech recognitionSpeech synthesisSpeech comprehensionEngineering
The embodiment of the invention discloses an intelligent voice interaction system and method based on streaming multi-mode fusion and an equipment control protocol, the system comprises a voice input processing module, a voice understanding and generating module and a voice synthesis module, the voice input processing module is used for converting an audio signal into a first token sequence, and the first token sequence is used for converting the audio signal into a second token sequence; the voice understanding and generating module is used for determining a response token sequence according to the first token sequence on the basis of a multi-modal Transform architecture so as to realize voice understanding and generation; and the voice synthesis module is used for synthesizing the response token sequence into an output audio so as to carry out at least one of the following adjustments on the converted audio of the response token sequence: emotion parameter adjustment, tone adjustment and rhythm adjustment. By adopting the embodiment of the invention, low-delay and high-naturalness intelligent voice interaction can be realized, multi-modal fusion and equipment control are supported, and the user experience is remarkably improved.
Owner:SHENZHEN HUANZHI TECHNOLOGY CO LTD

Real Time Dynamic Classification and Orchestration of Test Automated Components Leveraging Supervised Learning and Multi-Modal AI

This invention relates to systems and methods for real-time dynamic classification and orchestration of test automation components in distributed DevOps environments. The system features an Auto Identify Automation (AIA) engine that leverages supervised learning, Multi-Modal Artificial Intelligence (AI), and Generative AI technologies. It includes a Smart Scenario Designer interface that allows users to author test scenarios using handwriting and voice inputs, which are processed in real-time by AI-driven handwriting recognition, voice recognition, and Natural Language Processing (NLP). The system dynamically suggests relevant automated components via a smart bubble pane, facilitating rapid scenario creation. The architecture is tool-agnostic and scalable, with a Shared Workbench Engine that supports real-time collaboration and conflict resolution. The system continuously adapts and improves, ensuring that the automation suite remains consistent, up-to-date, and aligned with evolving software requirements, enabling efficient and user-friendly management of complex test automation processes.
Owner:BANK OF AMERICA CORP

Far-field single-channel speech enhancement method

The invention relates to the technical field of speech enhancement, in particular to a far-field single-channel speech enhancement method based on an MFSE (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error)). The method comprises the following steps: step 1, processing a far-field voice signal to obtain a complex spectrogram of a noise voice signal; 2, inputting the compressed complex spectrogram into a feature encoder, processing the output of the complex spectrogram by N MamAttention blocks, and then sending the processed complex spectrogram into an amplitude mask decoder and a phase decoder to respectively predict a clean compressed amplitude mask and a phase spectrum; step 3, preheating and training the MamAttention model, and performing supervised confrontation training by taking the MamAttention model as a generator and the multi-resolution discriminator as a discriminator; and step 4, inputting test voice into the trained model to realize far-field single-channel voice enhancement. According to the method, the supervised adversarial training strategy and the MamAttention model are combined, so that the problems of signal attenuation, noise and reverberation interference in far-field voice are effectively solved, and the voice quality is remarkably improved in a scene that the distance of a loudspeaker exceeds 5 meters.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Semantic processing system and method for home elevator intelligent customer service

The invention relates to the technical field of semantic processing, and discloses a semantic processing system and method for home elevator intelligent customer service, and the system comprises a data collection module, a semantic processing module, a response generation and optimization module, and a communication and log module. The method comprises the steps of collecting natural language request information input by a user through voice, performing semantic analysis and intention recognition on preprocessed information based on a deep learning model, and synchronously recording an interaction log of a whole link according to an intention recognition result and the emergency degree and priority of optimized response content. In order to solve the problems that a traditional keyword matching system is poor in semantic comprehension ability and prone to misjudgment of user intentions, the context and real intentions of natural language requests can be deeply understood by introducing a semantic analysis model based on deep learning and combining a confidence degree evaluation mechanism, and when the system recognizes the low confidence degree condition, the user intentions can be accurately judged. And multiple rounds of clear conversations can be automatically triggered or interaction channels can be switched.
Owner:SUZHOU FRANZ INTELLIGENT ELEVATOR CO LTD

Automatic voice response fault diagnosis method and system based on multi-source information fusion

The invention discloses an automatic voice response fault diagnosis method and system based on multi-source information fusion. The method comprises the following steps: receiving user voice input, synchronously obtaining user side intelligent electric meter data, meteorological environment information and historical service records, and extracting multi-modal features; semantic association is realized through an electric power service knowledge graph, and voice features, equipment data and environmental parameters are fused through space-time alignment; a machine learning dynamic decision tree engine is combined with a power consumption behavior analysis model to form a fault reasoning model, and a diagnosis result is effectively verified; and outputting the grading disposal scheme and triggering intelligent chemical order distribution. According to the invention, various information sources of the calling platform are fully utilized, a fault reasoning model with high accuracy is formed, the technical problems of inaccurate fault positioning and insufficient multi-source data collaboration of traditional voice response in power customer service are effectively solved, the diagnosis accuracy of customer power consumption problems is effectively improved, and the average processing time is shortened.
Owner:国家电网有限公司客户服务中心

Speech input based avatar face animation

Examples relate to systems and methods for generating an avatar animation. The systems and methods access an audio file comprising speech, spoken by a user, captured by a microphone of a user system, and receive input that selects an avatar associated with the user. The systems and methods process the audio file and the avatar, selected by the received input, by a generative machine learning model to generate an animation of the avatar having lips moving to represent the avatar speaking the speech of the audio file. The systems and methods generate a video comprising a depiction of the generated animation of the avatar speaking the speech of the audio file.
Owner:SNAP INC

Full-language voice interaction legal affair agent system and method for small and micro enterprises

The invention discloses a full-language voice interaction legal affair agent system and method for small and micro enterprises. The core of the system comprises a full-language voice input module, a language recognition module, a voice recognition engine routing module, a special voice recognition engine cluster, an enterprise-level multi-language natural language processing module, an enterprise legal affair agent core, a multi-mode output module and a feedback optimization module, and all the modules work cooperatively. And seamless connection from voice input to scheme output is realized. The process of the method comprises voice acquisition and language recognition, audio routing and text conversion, semantic understanding and intention recognition, commercial risk multi-dimensional trade-off analysis, structured scheme generation and multi-modal output, and meanwhile, the system is driven to continuously iterate through a feedback optimization mechanism. According to the method, the working efficiency and experience of an enterprise owner are remarkably improved, the legal consultation revolution of'moving without using hands' is achieved, the method has low threshold and high specialty, legal instruments can be directly generated, and the method has high intelligence and self-adaptive capacity.
Owner:齐洪建

Method for training speech synthesis model, speech synthesis method, and electronic device

A method for training a speech synthesis model includes obtaining training data; obtaining an initial speech synthesis model; training a semantic encoding network and a semantic decoding network in the speech synthesis model respectively based on a style sample speech, a timbre sample speech, an input sample text, and an output sample speech in training samples of the training data, to obtain a trained speech synthesis model.
Owner:BAIDU INT TECH (SHENZHEN) CO LTD

Term standardized query method and system based on large language model

The invention relates to the technical field of information retrieval, and discloses a term standardization query method and system based on a large language model, and the method comprises the steps: receiving original query, dialogue history and multi-modal input; analyzing the multi-modal input to generate an extended context; candidate aliases are extracted through natural language processing, and a domain term knowledge base is inquired in combination with a term list; adjusting the priority weight according to the user role and the real-time context; constructing a structured prompt including dialogue history, original query, extended context, candidate schemes and task instructions, and inputting a large language model to generate a replacement decision; updating the knowledge base based on user interaction data and model feedback; historical query records are stored to assist in follow-up reasoning. According to the invention, the picture, the PDF and the voice input are analyzed through the multi-modal context enhancement module, the extended context containing the term list and the semantic clue is generated, comprehensive context support is provided for term standardization, and query processing is ensured to adapt to various enterprise scenes.
Owner:江西博微新技术有限公司

Contextually boosted aviation speech recognition

A variety of applications can include a system having a speech recognition system responsive to the speech input, where the speech recognition system can be configured to recognize the speech input using an aviation vocabulary including words extracted using state information of an aircraft or intent information of the aircraft associated with the received speech input. A control system can be implemented to automatically perform an action in the system in response to analysis of the recognized speech input, where the action is associated with flight of the aircraft.
Owner:CIRRUS DESIGN CORP D B A CIRRUS AIRCRAFT

Digital human live broadcast voice interaction system fused with emotion calculation

ActiveCN121393436ASpeech recognitionLive voiceData stream
The invention relates to the technical field of digital human voice interaction, and discloses a digital human live voice interaction system fused with emotion calculation. The system constructs an emotional response time window by acquiring a user voice input data stream, the starting point of the emotional response time window is the end moment of the user voice input data stream, and the end point is obtained by subtracting the necessary duration of voice response synthesis from the preset maximum response cut-off moment. The system obtains an emotional state vector of the user in real time and judges whether the emotional state vector reaches an emotional intensity threshold value or not; and if so, predicting the total generation duration of the digital human voice response. And when the residual duration of the emotional response time window is equal to the total generation duration, taking the window starting point as a voice response starting moment, and controlling the digital human to start voice response generation. According to the system, the voice response matched with the emotional state of the user is generated through refined time window management and emotional state perception, invalid response and interaction delay are reduced, and the naturalness of digital human live broadcast voice interaction and the user experience are improved.
Owner:BEIJING ZHONGSHENGSHENG DIGITAL TECHNOLOGY CO LTD

Voice control in a healthcare facility

Systems for voice control of medical devices in a healthcare facility are disclosed herein. The systems employ continuous speech processing software, voice recognition software, natural language processing software, and other software to permit voice control of the medical devices. Systems are also provided for distinguishing which medical device from among multiple medical devices in a patient room is the particular medical device to be controlled by voice input from a caregiver or a patient.
Owner:HILL ROM SERVICES INC

Voice-driven information recording and checking method, device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice-driven information recording and checking method, device, equipment and medium. Key fields are extracted by applying a semantic analysis strategy to form preliminary structured data, an independent data system is connected to verify authenticity and generate data conflict information, and automatic error correction is executed based on the conflict information to generate updated structured data; and according to the user role matching output template, filling the template with the updated structured data to generate display content, and pushing the display content to target terminal equipment. Structured data generation is realized through voice input and semantic analysis, the data accuracy is improved by combining cross-system verification and automatic error correction, personalized display and real-time pushing are realized by utilizing a role matching template, and a closed loop of acquisition, verification and display is formed, so that the efficiency of information recording and management is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice Interface Integration System and Method for Mobile, Computer and Web Applications

A method for voice-based interaction between a user (e.g., a shopper) and online stores includes capturing an initial voice input from a user and converting the voice input to text. The text is analyzed and initial search terms are created such as product name, product description, product category, product brand name, product manufacturer and retailer. One or more searches are conducted using the initial search terms to identify one or more retailers meeting the search terms. One or more identified online retailers are accessed and searches are executed at the one or more identified retailers, generating search results. The identified online retailers have a user interface allowing purchases of the identified product or item. The search results are returned to the user.
Owner:OMNIBEK IP HLDG LLC

Digital human question and answer system

The invention provides a digital human question-answering system, and the system comprises a voice receiving and processing module which is used for receiving the voice input of a user in real time, splitting the input voice, extracting the acoustic features of each frame of voice in real time, and caching the acoustic features of continuous frames to form a streaming data queue; the voice feature reasoning module is used for inputting the streaming data queue into a pre-training question and answer model and outputting a question and answer text result and a voice synthesis instruction in real time; and the digital human synchronous driving module is used for generating a synchronous voice signal based on the question and answer text result and the voice synthesis instruction, mapping the synchronous voice signal to a mouth shape mapping library and an action template library, generating a mouth shape sequence and a limb action sequence matched with the voice time sequence, and completing digital human broadcasting. According to the method, user text or voice input is received, the question and answer result is returned after model reasoning, and the digital person is driven to complete voice broadcasting.
Owner:ZHONGCHUANG (WUHAN) TECH CO LTD

Complex transaction task execution and risk control method and system based on artificial intelligence large model (LLM), electronic equipment and storage medium

The invention discloses a complex transaction task execution and risk control method and system based on an artificial intelligence large model (LLM), electronic equipment and a storage medium, and relates to the field of financial science and technology. According to the method, a transaction demand input by a natural language (preferably voice) of a user is obtained, a structured task instruction sequence containing varieties, directions, number and trigger conditions is generated through large model semantic analysis, risk control verification is carried out according to account funds, position limitation and supervision rules, an executable transaction instruction is formed after the risk control verification is passed, and the transaction requirement of the user is met. And a manual confirmation link can be selected and sent to the transaction terminal for execution, and an execution result is returned for updating the session and risk state. A voice mode is preferably selected for feedback, and automatic transaction execution under complex and continuous conditions is supported. According to the method, natural language interaction is kept, meanwhile, automatic disassembly and rearrangement of complex transaction tasks are achieved, manual one-by-one operation is converted into automatic processing, and the execution efficiency and the risk control capability are improved.
Owner:GUIMIAO TECHNOLOGY (GUANGXI) CO LTD

Method and system for enhancing generated autism interview questions and answers based on multi-modal retrieval

The invention discloses a method for enhancing generated autism interview questions and answers based on multi-modal retrieval. The method comprises the following steps: step 1, constructing a multi-modal vector database in the autism field; 2, multi-modal data including voice input and text input of a user and audio and video recorded in advance are collected, and the user input is converted into vectors; 3, searching top-k vectors which are most similar to user input in a vector database through an algorithm, and filtering the retrieved vectors based on metadata; 4, if the relevant documents are not retrieved in the local documents, the retriever is used for retrieving in the exogenous data; and 5, designing a language model, outputting a reliable and traceable answer according to a retrieval result, and returning the reliable and traceable answer to the user.
Owner:EAST CHINA NORMAL UNIV

Automated speech recognition to support context-aware intent recognition

A computing system for determining a user intent from a speech input to effect a user intended action is provided. The computer system comprises a set of processing nodes and a controller module. Each processing node is capable of understanding only a subset of words directly relevant to a particular context. The processing nodes of the set are arranged to receive a same speech input, and each processing node attempts to interpret the input, based on its subset of words, to extract therefrom an output indicative of user intent. Each node is unable to interpret any portion of the input containing a word outside of its subset. The controller module receives the outputs from the set of processing nodes and determine a most likely user intent based on the outputs.
Owner:OAKSPIRE LTD

AI-supported multi-modal desktop interactive projection method and system

The invention discloses an AI-supported multi-modal desktop interactive projection method and system, and the method comprises the following steps: synchronously collecting visual signals, tactile data and voice input in a desktop projection environment, and obtaining a multi-modal information set after data fusion; calculating priority weight distribution of each mode according to the multi-mode information set, and obtaining a high-priority mode sequence; if the visual signals in the high-priority modal sequence are judged to be dominant, extracting a visual dominant feature group; performing cross validation on the visual dominant feature group and the tactile data, and judging information consistency after data fusion; according to an information consistency result, dynamically adjusting transparency control parameters of projection layering to obtain optimized hierarchical configuration; and updating the projection display content according to the optimized hierarchical configuration, and outputting the final projection content. According to the method, the interaction real-time performance and immersion of desktop projection are improved, and finally the technical effects of optimizing user experience and improving system intelligence are achieved.
Owner:TONGJI UNIV

Voice interaction method, user terminal, server side, equipment, medium and product

The invention discloses a voice interaction method, a user terminal, a server side, equipment, a medium and a product. The method comprises the steps that the user terminal acquires voice input of a user; the voice input is compressed and coded, first voice input is obtained, and the mode of compressed coding is dynamically selected according to the content of the voice input; performing voice recognition, text reasoning and voice synthesis on the first voice input to obtain a target synthesized voice; or, the first voice input is transmitted to the server side, the target synthetic voice obtained by the server side is received, in the transmission process, the first voice input is divided into a voice segment and a silent segment, and transmission of the voice segment is completed preferentially; and locally playing the target synthesized voice. Through the dynamic environment perception and self-adaptive noise reduction method, transmission delay is reduced, context understanding is enhanced through text reasoning, low-delay and high-naturalness real-time voice interaction is achieved, and user experience is improved.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

People and post matching method and system based on bidirectional collaborative filtering algorithm

The invention discloses a person and post matching method and system based on a bidirectional collaborative filtering algorithm, and the method comprises the steps: assisting a job seeker to clearly input job-seeking basic information and transmit job-seeking appeals in a flexible manner through an intelligent interaction interface integrating voice input and large language model analysis; the method comprises the following steps: collecting data through interaction information collection, resume information analysis and recruitment demand decomposition, carrying out data cleaning and label mining, generating a job seeker dominant label, a job seeker implicit label, a recruiter dominant label and a recruiter implicit label based on an ontology, preliminarily constructing an interestingness preference model of the job seeker and the recruiter, and establishing an interestingness preference model of the job seeker and the recruiter; and then calculating a job-seeking label data set and a recruitment position label data set based on an improved collaborative filtering algorithm, calculating a person-position bidirectional matching coefficient by integrating dominant demands and implicit demands of both parties, recommending a working position with a higher bidirectional matching degree for a job-seeker position, and improving the accuracy of position recommendation.
Owner:JIANGSU WANGXIN SOFTWARE CO LTD

Generative offer system with multilingual ai-driven voice assistant for real-time commerce

A system and method for generating and delivering personalized promotional offers in real time in response to user voice input. The system processes natural language queries and analyzes prior purchase history, inferred preferences, historical interactions, environmental context including device type, location, and time, and ongoing behavioral signals. Machine learning models adapt future offers based on accumulated user data to provide evolving personalization. The system supports multilingual speech recognition and can operate across smartphones, voice assistants, wearables, smart televisions, extended reality environments, and large language models. Offers may be delivered through audio, visual, or haptic interfaces to enable integration into commerce and advertising platforms. The combination of speech-triggered interaction, adaptive personalization, and contextual awareness enables dynamic, AI-driven promotional flows that operate consistently across multiple platforms and languages.
Owner:BYRD STEPHEN M

Sign language animation generation method and device based on semantic analysis, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a sign language animation generation method, device, equipment and medium based on semantic parse. The sign language animation generation method comprises the steps that voice input is received and recognized as text content, field semantic parse is conducted on the text content to generate a field semantic template, and the field semantic template is used for generating a sign language animation; and converting the domain semantic template into a sign language intermediate representation sequence, generating a three-dimensional sign language action sequence based on the sign language intermediate representation sequence, rendering the three-dimensional sign language action sequence into a virtual image sign language animation, and displaying the virtual image sign language animation. According to the invention, by fusing speech recognition, semantic analysis and three-dimensional action rendering, direct conversion from spoken language content to sign language animation is realized, and a complete visual expression link from speech to sign language is formed, so that a user can intuitively understand the speech content in a sign language form through a virtual image, and the user experience is improved. Therefore, the barrier-free performance of human-computer interaction and the accuracy of information transmission are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Rendering XR avatars based on acoustical features

In one embodiment, a method includes receiving a voice input having acoustic features from a first client system associated with a first user, determining emotions associated with the voice input based on one or more of the acoustic features by machine-learning models, determining facial features for a first extended-reality (XR) avatar representing the first user based on the emotions, and sending instructions for rendering the first XR avatar representing the first user to a second client system associated with a second user, wherein the first XR avatar is rendered with the determined facial features.
Owner:META PLATFORMS INC

Social robot semantic understanding method and system based on knowledge graph, electronic equipment and storage medium

The invention discloses a knowledge graph-based social robot semantic understanding method and system, electronic equipment and a storage medium, and the method comprises the steps: obtaining historical heterogeneous data composed of historical text input, historical voice input and historical image input, and coding the historical heterogeneous data to obtain a unified multi-modal feature vector representation; entities, relations and attributes in the multi-modal feature vector representation are extracted, and an initial knowledge graph is constructed; weighting and updating the initial knowledge graph by using a graph attention network and a graph convolution network to obtain a dynamically updated knowledge graph; dialogue information currently input by a user is acquired, and user intention and semantic association entities are identified by using the dynamically updated knowledge graph; based on the user intention and the semantic association entity, a semantic understanding result is generated in combination with the context of the dialogue information, and response content is generated based on the semantic understanding result. According to the method, the deep understanding capability of the social robot on complex semantics is remarkably improved.
Owner:UNIV OF CHINESE ACAD OF SCI

Reminding method and system for vehicle-mounted intelligent memorandum

The invention discloses a reminding method and system for a vehicle-mounted intelligent memo, and belongs to the technical field of automotive electronics, and the method comprises the steps: receiving memo content input by a user through a voice input module; automatically extracting position information, time information and associated user information in the memorandum content; based on the position information, the time information and the associated user information, whether the memorandum content meets important item judgment conditions or not is judged; determining a reminding strategy according to the judgment result; and executing the reminding strategy. According to the method and the system, active voice input of the user is combined with intelligent analysis, so that the problems of insufficient convenience, insufficient timeliness, insufficient accuracy and poor user experience in the prior art are solved, value-added services such as precise item classification reminding, family member collaborative reminding and tourism item intelligent generation strategy are realized, and the user experience is improved. The accuracy, the convenience, the intelligence and the safety of the vehicle-mounted memorandum are remarkably improved, meanwhile, the operation safety in the driving process is ensured, and important items are prevented from being missed.
Owner:VOYAH AUTOMOBILE TECH CO LTD