Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40981 results about "Human–computer interaction" patented technology

Human–computer interaction (HCI) researches the design and use of computer technology, focused on the interfaces between people (users) and computers. Researchers in the field of HCI observe the ways in which humans interact with computers and design technologies that let humans interact with computers in novel ways. As a field of research, human–computer interaction is situated at the intersection of computer science, behavioural sciences, design, media studies, and several other fields of study. The term was popularized by Stuart K. Card, Allen Newell, and Thomas P. Moran in their seminal 1983 book, The Psychology of Human–Computer Interaction, although the authors first used the term in 1980 and the first known use was in 1975. The term connotes that, unlike other tools with only limited uses (such as a hammer, useful for driving nails but not much else), a computer has many uses and this takes place as an open-ended dialog between the user and the computer. The notion of dialog likens human–computer interaction to human-to-human interaction, an analogy which is crucial to theoretical considerations in the field.

System and method for ai-driven multi-modal content generation and immersive interaction experiences

A system and method for creating complex, immersive, and interactive digital content is disclosed. The system integrates advanced artificial intelligence, multi-modal input processing, cloud-based shared environments, and immersive hardware to generate, optimize, and deliver rich interactive experiences. The platform supports content mashups, custom scenario generation, and adaptive AI behaviors, enabling the creation of unique and engaging digital environments across various media formats.
Owner:QOMPLX INC

Online social wager-based gaming system featuring dynamic cross-provider game filtering, persistent cross-provider voice-interactive group play, automated multi-seat group game reservation, and distributed ledger bet verification

An online social wager-based gaming system is disclosed, featuring dynamic cross-provider game filtering and persistent cross-provider voice-interactive group play. The system allows for automated multi-seat group game reservations and incorporates a distributed ledger for bet verification. Features include a Group Connect module for dynamic group formation and live audio across various game providers, enabling real-time play. A Professional Companion Connect module facilitates live webcam or microphone gambling sessions with professional companions, under contractual agreements. The User Games module enhances personalization by integrating user-uploaded images into the gaming environment. A cryptocurrency / blockchain-based bet tracking system ensures transparency and security in managing bet transactions through a controlled digital wallet system. This integrated platform aims to provide a cohesive, interactive, and personalized online casino experience, addressing limitations of traditional solitary online gaming by fostering social connections and enhancing user trust and engagement.
Owner:NOWAK DANIEL PATRYK

Dynamic agents with real-time alignment

An example may receive at least one input via at least one device. An example may use the at least one input to determine an entity identity. An example may use the entity identity to create an automated agent and load context data associated with the entity identity into at least one layer of a multi-layer memory of the automated agent. An example may cause the automated agent to machine-learn a supervision level via the context data. The machine-learned supervision level may indicate a level of supervision of the automated agent by an entity associated with the entity identity. An example may configure the automated agent to execute a task on behalf of the entity and in accordance with the machine-learned supervision level.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Techniques for determining conversational intent

The present disclosure relates to systems and methods for enhancing the interaction between users and automated agents, such as digital assistants, by employing Large Language Models (LLMs) to infer the intent of spoken language. The invention involves continuously monitoring ambient audio, converting speech to text, and utilizing LLMs to determine whether spoken language is intended for the automated agent. A structured prompt, including the converted text and specific instructions, is sent to the LLM, which is fine-tuned to process domain-specific prompts. The LLM provides a structured output in a standardized format, indicating the user's intent. The system may involve multiple prompts to perform separate tasks, such as identifying intent and generating additional context-specific data. This approach facilitates a more natural and intuitive user experience by eliminating the need for wake words and allowing seamless conversational interaction with virtual assistants across various platforms and devices.
Owner:SNAP INC

Collaborative knowledge fusion reinforcement learning method for sparse reward environment

The invention discloses a sparse reward environment-oriented collaborative knowledge fusion reinforcement learning method, and relates to the field of collaborative knowledge fusion reinforcement learning methods. By constructing a lightweight collaborative knowledge fusion model and a dynamic reward remodeling mechanism, the problems of low intelligent agent exploration efficiency and difficulty in strategy convergence in a sparse reward environment are solved. The method comprises the following steps: constructing a reinforcement learning framework comprising a policy network and a value network; designing an action space mutation supervision mechanism and a lightweight collaborative knowledge fusion model, and generating a smooth substitution action when a strategy is detected to be unstable; and a reward function is designed in combination with the task target and the dynamic constraint, and reward remodeling is realized by activating rewards through sub-target potential energy difference and knowledge fusion. According to the method, effective intermediate feedback can be provided for the agents in the sparse reward environment, the exploration efficiency is improved, the convergence time is shortened, the stability and cross-scene migration ability of the strategy are enhanced, and an effective solution is provided for sparse reward scenes such as robot control and multi-agent game.
Owner:CHANGCHUN UNIV OF TECH

Method of displaying user interfaces in an environment and corresponding electronic device and computer readable storage medium

Methods for displaying user interfaces in a computer-generated environment provide for an efficient and intuitive user experience. In some embodiments, user interfaces can have different immersion levels. In some embodiments, a user interface can have a respective immersion level based on its location in the three-dimensional environment or distance from the user. In some embodiments, a user interface can have a respective immersion level based on the state of the user interface. In some embodiments, a user interface can switch from one immersion level to another in response to the user's interaction with the user interface.
Owner:APPLE INC

Multi-modal natural language understanding and generating system and method

The invention discloses a multi-modal natural language understanding and generating system and method. The method comprises the following steps: constructing a cross-modal pre-training module, training a multi-modal encoder, and establishing a cross-modal association mapping space; mixing prompt fine tuning is carried out, and a complete blank filling template is constructed; according to the intention reasoning network, extracting multi-round dialogue intention representation of the user, and retrieving an external knowledge base for fine-grained reasoning; constructing a unified semantic representation framework, embedding the text, the image and the voice into a unified space, and generating a query vector of multi-modal intention perception; and the knowledge query module based on key value memory generates entity-level multi-modal replies and optimizes the semantic comprehension and generation capability of the dialogue model. According to the method, the multi-modal information understanding and generating capacity is improved, deep association and understanding of image and text information are achieved, downstream task adaptability is enhanced, task completion accuracy and efficiency are improved, unified semantic representation of the multi-modal information is achieved, and support is provided for information retrieval and utilization.
Owner:UNIV OF ELECTRONIC SCI & TECH OF CHINA CHENGDU COLLEGE

Adaptive Data System And A Method For Cognitive Data Processing

An adaptive data system (ADS) for cognitive data processing is disclosed. The ADS includes an adaptive semantic preprocessor, a trigger detector, a temporal batching engine, a symbolic encoder, and a dynamic cognitive transformer engine. The adaptive semantic preprocessor is configured to receive input data from one or more databases and identify cognitive data attributes comprising one or more contextual, semantic, and temporal attributes from the received input data. The trigger detector is configured to identify semantic divergence of the identified cognitive data attributes and provide a standardized data. The temporal batching engine is configured to provide a high-dimensional cognitive data from the standardized data. The symbolic encoder compresses the high-dimensional cognitive data. The dynamic cognitive transformer engine is configured to determine decision making rules, analyze the compressed high-dimensional cognitive data based on the decision making rules and provide recommendations based on an outcome of the analysis to a user.
Owner:DATAQUANTUM INC

User emotion recognition and psychological intervention system and method based on large language model

Aiming at the problems of insufficient language understanding depth, weak personalized dialogue generation ability, lack of continuous learning and long-term user state modeling and the like in the current emotion recognition and psychological intervention technology, the invention provides a user emotion recognition and psychological intervention method combined with a large language model (LLM). According to the method, the potential emotional state is identified by analyzing free text information input by a user by utilizing the powerful capabilities of a large language model in the aspects of natural language understanding, emotional modeling and text generation; constructing a multi-round dialogue context, and reasoning a psychological change trend of the user; in combination with a psychological knowledge base, personalized and mild psychological intervention dialogue content with a dredging effect is generated. The system supports recognition and classification of various emotional states such as depression, anxiety and alonity, is suitable for various interaction scenes (such as APPs, webpages and social robots), and can greatly improve the precision of emotion recognition and the timeliness and effectiveness of psychological intervention. The emotion recognition and psychological intervention method based on the large language model provides solid technical support for constructing an intelligent, continuous and personalized psychological health management system, and has wide application prospects and profound social significance.
Owner:CHANGCHUN UNIV OF TECH

Methods for adjusting and / or controlling immersion associated with user interfaces

In some embodiments, an electronic device emphasizes and / or deemphasizes user interfaces based on the gaze of a user. In some embodiments, an electronic device defines levels of immersion for different user interfaces independently of one another. In some embodiments, an electronic device resumes display of a user interface at a previously-displayed level of immersion after (e.g., temporarily) reducing the level of immersion associated with the user interface. In some embodiments, an electronic device allows objects, people, and / or portions of an environment to be visible through a user interface displayed by the electronic device. In some embodiments, an electronic device reduces the level of immersion associated with a user interface based on characteristics of the electronic device and / or physical environment of the electronic device.
Owner:APPLE INC

Robotic systems and methods

Systems and methods that allow robots to perform tasks for users are provided. A robot may comprise one or more robotic arms and / or a mobile base. The arms may be controlled by electric actuators and may have six or more degrees of freedom. The robot may have sensors which can be accessed remotely. Robots may have varying levels of autonomy, including, for example, full teleoperation (in which a human can have detailed control over the robot) or full autonomy (in which the robot can complete a task without any human intervention). Various entities can interact with individual robots or groups of robots over a network, locally, directly or in person, or any combination thereof. A management system can allow entities to control and / or monitor the robots.
Owner:ABRAMS DANIEL

Control method and equipment of intelligent robot with body and storage medium

The invention discloses a control method and equipment for an intelligent robot with a body and a storage medium, and belongs to the technical field of robots. The method comprises the steps of receiving a task instruction and environment perception data, performing fusion processing on the task instruction and the environment perception data, generating task semantic information associated with the task instruction, analyzing the task semantic information through an implicit planner, generating an implicit action mark, and based on the implicit action mark, generating an implicit action. And decomposing a to-be-executed task corresponding to the task instruction into a plurality of task sub-targets layer by layer, generating an action instruction sequence based on the task sub-targets, and executing a control action corresponding to the action instruction sequence. According to the method, based on the constraint of the implicit action mark and the action instruction sequence, the robot with the body can flexibly respond to environment parameter changes such as object position deviation or new task requirements in the task execution process, the coupling degree of the action instruction of the robot with the body and the predefined scene is reduced, and high stability of task execution is ensured.
Owner:YOUDI ROBOT (WUXI) CO LTD

Generative Outputs Confirming to User's Own Gameplay to Assist User

Generative models are disclosed to generate audio and visual outputs to a user when the user struggles with a particular aspect of a video game. The generative outputs can demonstrate what success at that aspect of the game looks like, doing so using the same playstyle, ability, and tactics as the user themselves to provide relevant and feasible assistance to the user.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

System and Method for Real-Time Team Intent Modeling Using Persistent Cognitive Machines with Federated Human Profiles

ActiveUS20260050745A1Memory architecture accessing/allocationDigital data information retrievalTeam compositionTeam learning
A system and method for real-time team intent modeling using persistent cognitive machines with federated human profiles which processes individual team member behavioral signals through geometric intent analyzers that generate high-dimensional vector representations of individual objectives and preferences. A team intent orchestrator aggregates individual vectors into collective representations within a dynamic geometric manifold that evolves based on team coordination patterns. Federated human profiles enable privacy-preserving knowledge sharing across teams through geometric abstraction techniques that preserve coordination utility while protecting individual privacy. The system implements proactive conflict detection through trajectory analysis that identifies potential coordination issues before performance impact, and provides real-time synchronization mechanisms that maintain team coordination coherence despite individual behavioral changes. Cross-team learning capabilities enable organizational intelligence development through pattern abstraction and context-aware adaptation of successful coordination strategies. The persistent cognitive architecture maintains coordination patterns across sessions and team composition changes, enabling continuous improvement through accumulated team experience.
Owner:ATOMBEAM TECH INC

Trust-Informed Engagement and Restraint (TIER) For Behavioral Governance in Conversational AI

The Trust-Informed Engagement and Restraint (TIER) system is a modular behavioral governance architecture for conversational AI that enables dynamic, real-time enforcement of trust-aligned interaction policies. The system comprises a Behavioral Governance Framework defining policy rules and an Enforcement Wrapper hosting Core Modules that operate externally to the AI model. These modules regulate behavioral traits including trust calibration, emotional tone, conversational containment, authority modulation, and cross-modal consistency. TIER continuously monitors interaction metrics and applies policy-driven constraints during live sessions without requiring model retraining or internal access. Unlike static filters or post-hoc moderation systems, TIER provides proactive, session-aware behavioral governance with auditable enforcement across domains. The architecture supports model-agnostic deployment in regulated and sensitive environments such as healthcare, finance, legal services, and intelligent assistance platforms.
Owner:ORTH JESSICA

Hallucination detection via multilingual prompt

Aspects of the present disclosure relate to detecting hallucinations in language model outputs. Embodiments include receiving a user query. Embodiments further include prompting a language processing machine learning model to generate responses to the user query in each language of a set of multiple languages. Embodiments further include receiving the responses from the language processing machine learning model in response to the prompting. Embodiments further include creating embedding representations of the responses. Embodiments further include calculating, based on the embedding representations, a degree of semantic similarity between the responses. Embodiments further include determining that a response of the responses contains a model hallucination based on comparing the degree of semantic similarity between the responses to a threshold.
Owner:INTUIT INC

Data processing method and apparatus, electronic device, computer readable storage medium and computer program product

The present application provides a data processing method and apparatus, an electronic device, a computer readable storage medium and a computer program product. The method comprises: acquiring historical interaction information and a predicted interaction text corresponding to the historical interaction information; extracting a first acoustic feature and a first semantic feature of the historical interaction information, and extracting a second semantic feature of the predicted interaction text; performing fusion mapping on the basis of the first acoustic feature, the first semantic feature and the second semantic feature to obtain a first paralanguage feature; denoising initial noise on the basis of the second semantic feature and the first paralanguage feature to obtain a second acoustic feature of the predicted interaction text; and on the basis of the second acoustic feature, generating a voice signal corresponding to the predicted interaction text.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Natural language processing

Techniques for generating tasks to be completed in order to perform an action responsive to a user input and, for a given task, shortlisting available components to those that are relevant for the task are described. The system processes a user input to determine tasks to be completed in order to perform an action responsive to the user input. The system determines a priority of the tasks and selects a top-ranked task. The system determines descriptions of processing performable by components that are semantically similar to the current task, and requests a description of the function the corresponding components would perform for the current task. Based on the received descriptions, the system selects one or more components to perform the task. Thereafter, the system causes the action to be performed and outputs a response to the user input.
Owner:AMAZON TECH INC

Intelligent voice recognition and natural language interaction method based on quadruped robot

The invention discloses an intelligent voice recognition and natural language interaction method based on a quadruped robot, and the method comprises the following steps: S1, collecting a user voice instruction, generating a standardized voice text, and extracting a semantic keyword set; s2, collecting multi-source sensing data of the quadruped robot and generating a structured state data tensor; s3, constructing a multi-modal collaborative modeling mechanism, and generating a multi-modal joint semantic embedding vector; s4, constructing a semantic map based on semantic embedding and generating an action chain plan structure; s5, executing each sub-action in the action chain, and performing path analysis and execution monitoring; s6, storing the interaction task as a multi-modal semantic behavior memory unit; and S7, performing similarity retrieval based on current semantic input and historical memory to realize behavior migration and action chain multiplexing. The method has the advantages of accurate semantic understanding, intelligent interaction response, high behavior migration capability and the like.
Owner:山东浪潮数据库技术有限公司

Generative and adaptive mediator for real-time interactions with conversational agents

A generative mediator engine can perform a requested interaction with a conversational agent of a target entity on behalf of a user. An internal conversational platform can identify intents for the requested interaction. An external artificial intelligence engine can perform intent discovery when an intent is not identified above a confidence threshold. A discovered intent unknown to the generative mediator engine can be received from the external artificial intelligence engine and used, with input requirements determined by the generative mediator for the requested interaction, by a dialog generator to generate a sample dialog for the requested interaction. User feedback can be received after review of action items and expected inputs identified from the sample dialog. The generative mediator engine can perform the requested interaction with the conversational agent on behalf of the user and without receiving user intervention during the requested interaction.
Owner:CISCO TECHNOLOGY INC

Generating a response for a communication session based on previous conversation content using a large language model

An example operation may include one or more of receiving interaction content from a communication session between a source device and a service provider device of a service provider, identifying a search criteria from the interaction content, retrieving a subset of vectors from a plurality of vectors stored in a vector database based on the search criteria of the interaction content, wherein the subset of vectors includes previous interaction content with the service provider, generating a response for the communication session based on execution of a large language model (LLM) on the subset of vectors, and outputting the response to at least one of the source device and the service provider device during the communication session.
Owner:THE TORONTO DOMINION BANK

Methods, systems, apparatuses, and devices for facilitating conversational interaction with users to help the users

A method for facilitating conversational interaction with users to help the users includes transmitting a conversational interaction interface for conversationally interacting with a user to a user device, receiving a request of the user through the conversational interaction interface from the user device, identifying an information based on the request, generating an input comprising the request and the information for a machine learning model based on the request and the information, processing the input using the machine learning model, generating a response for the request based on the processing of the input, transmitting the response through the conversational interaction interface for conversationally interacting with the user to the user device, and storing the machine learning model, the request, and the response.
Owner:NEXT LEAGUE EXECUTIVE BOARD LLC

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Psychological accompanying method based on multi-modal emotion recognition

The invention discloses a psychological accompanying method based on multi-modal emotion recognition, and belongs to the technical field of psychological health services, and the method comprises the steps: S1, synchronously collecting physiological signals, voice features, facial expressions and interactive behavior data of a user through a multi-modal sensor; s2, performing fusion analysis on the multi-modal data by using a deep learning model, and identifying a current emotional state and an emotional intensity level of the user; and S3, dynamically generating an adaptive psychological accompanying intervention scheme based on a preset emotion-intervention strategy mapping rule in combination with historical emotion data and personalized preferences of the user. According to the method, the physiological signals, the voice features, the facial expressions and the interactive behavior data are synchronously collected through the multi-modal sensor, the deep learning model is used for fusion analysis, and compared with single-modal recognition, the emotion state and the intensity level of the user can be judged more comprehensively and accurately, the emotion misjudgment risk is reduced, and a reliable basis is provided for subsequent intervention.
Owner:刘梓宸

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Dynamic agents with real-time alignment

An example may use an objective to retrieve first context data from at least one first memory layer of a multi-layer memory associated with an automated agent, and cause the first context data to be presented via at least one first conversational dialog element. An example may determine context feedback data in response to the first context data, and cause the context feedback data to be stored in at least one second layer of the multi-layer memory. An example may use the objective, the first context data, the context feedback data, and at least one workflow to configure a first prompt. An example may use the configured prompt and a machine learning model to generate a plan including one or more tasks executable by at least the automated agent to complete the objective. An example may cause the plan to be presented via at least one second conversational dialog element.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC