Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1622 results about "Human being" patented technology

Multi-modal articulation system for translating human gestures into control commands

A multi-modal articulation system translates human gestures into control commands for digital and automation devices. The system integrates a camera module, which captures facial and hand movements, with an image processing unit that extracts facial features, such as eyebrow movement, lip curvature, eyelid motion, and eye trajectory. These extracted features are analyzed in real time by an articulation recognition module that compares them against a database to recognize specific gestures. The system tracks and processes air-drawn hand gestures using a gesture trajectory tracking unit, which converts the motions into digital representations and classifies them using deep learning algorithms. The recognized gestures are converted into machine-readable control signals and executed on various connected devices via a relay control unit. The system incorporates adaptive algorithms to improve accuracy, account for environmental conditions, and stabilize hand tremors, providing a gesture-based control interface for applications ranging from smart home systems to multimedia devices.
Owner:AL GHAMDI RAYED +2

Multimodal intelligent agent system for dynamic environmental monitoring and human-centered support

A multimodal intelligent agent system for dynamic environmental monitoring and user-centered support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Hip assembly and kinematics of a humanoid robot

The present disclosure provides a humanoid robot with an arrangement of components that allows the robot to mimic the movements, functionality and capabilities of a human being. The robot includes a torso coupled to a waist, an arm assembly, and a head assembly. A pelvis is coupled to the waist and has left and right actuator mounts. Left and right hip assemblies are coupled to the respective actuator mounts. Each hip assembly includes a hip pitch actuator assembly, a hip roll actuator assembly, and a leg twist actuator assembly. The hip pitch actuator assembly has a portion positioned within the pelvis and is coupled to the actuator mount. The hip roll actuator assembly is coupled to the hip pitch actuator assembly, with a non-90 degree angle formed between their axes. The leg twist actuator assembly is coupled to the hip roll actuator assembly and positioned below extents of both the hip pitch and hip roll actuator assemblies.
Owner:FIGURE AI INC

Head and neck assembly of a humanoid robot

A head and neck assembly for a humanoid robot, including a head portion having an exterior surface defining an overall shape resembling a human head; a neck portion extending from the head portion; a head housing assembly enclosing the head portion and neck portion; an electronics assembly contained within the head housing assembly; and a head actuator assembly configured to move the head portion relative to a torso of the humanoid robot.
Owner:FIGURE AI INC

Human abnormal behavior monitoring method based on large-model multi-agent

The invention discloses a human abnormal behavior monitoring method based on a large-model multi-agent, which is executed by a modular multi-agent system deployed on a back-end server, obtains information through a monitoring camera, and comprises the following steps: obtaining a video stream from the monitoring camera by a sensing agent and extracting human body posture features; analyzing the key frame by a scene understanding agent by using a visual large model, and constructing a time sequence dynamic scene graph; the core reasoning agent evaluates the scene semantic conformity based on the pre-trained large model and performs abnormal preliminary judgment; performing fine-grained classification, interpretation generation and risk assessment on the abnormal behaviors; and the report and action agent generates an alarm and records event data. According to the invention, through multi-agent cooperative work and a large model technology, efficient and accurate monitoring of human abnormal behaviors is realized, and the intelligent level of the monitoring system and the abnormal behavior identification accuracy are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Compounds and combinations thereof for treating neurological and psychiatric conditions

This disclosure relates to a method of treating a neurological disorder, comprising administering a bupropion, or a combination of a bupropion and a dextromethorphan, to a human being in need thereof, monitoring the human being for Brugada pattern / syndrome, drug reaction with eosinophilia and systemic symptoms, or acute generalized exanthematous pustulosis, and reducing or discontinuing use of the bupropion, or the combination of the bupropion and the dextromethorphan, if the human being develops Brugada pattern / syndrome, subcutaneous tissue disorder, drug reaction with eosinophilia and systemic symptoms (DRESS), or acute generalized exanthematous pustulosis.
Owner:ANTECIP BIOVENTURES II LLC

Systems and methods for predicting and assessing motion of humans and / or objects in an environment

A method for predicting and assessing motion of a user in an environment. The method includes determining estimated positions of key points of the user at a first time and determining link lengths for linkages connecting the key points. The method further includes determining, for an upcoming time, predicted positions of the key points and predicted kinematics for the key points. The estimated positions, predicted positions, and kinematics are determined periodically over a time period. The method further includes classifying an activity being performed by the user over the time period and predicting a future motion path of the user for an upcoming time period. The method further includes generating a motion assessment for the user using the predicted future motion of the user and making the motion assessment available to a device or account associated with the user.
Owner:UNIVERSITY OF CINCINNATI

System and method for emotionally intelligent, personalized AI avatar-based health coaching using multi-domain data and adaptive behavioral intelligence

A programmatically generated AI avatar includes a customizable personality module, acting as the embodied interface for a powerful AI “mind” that delivers personalized coaching to improve user health, well-being, and longevity. The system uses machine learning, large language models, and biometric modeling to synthesize real-time, multi-modal health data—including sleep, nutrition, glucose, mood, and activity—and generate forward-prescribed KHAs. Unlike human coaches, it continuously adapts based on context and behavior, targeting the root cause: metabolic dysfunction—namely by restoring healthy, sustainable body composition through the preservation or building of lean muscle mass and reduction of excess fat. KHAs can also be shared with friends or programmatically generated AI avatars, allowing for coordinated action, emotional support, and accountability through social connection—further reinforcing positive behavior and adherence. The system's reinforcement learning engine incorporates both individual response data and anonymized population-level insights to optimize recommendations over time, learning which interventions are most effective for users with similar physiological and behavioral profiles. First validated with Olympic athletes—resulting in measurable improvements and medal-winning outcomes—this system offers a scalable, emotionally intelligent coaching engine that exceeds human capability, designed for the ultimate purpose of supporting sustainable health, resilience, and human thriving.
Owner:GOLD AND COMPANY

Remote operation system and control method for two mechanical arms based on VR head-mounted display

According to the double-arm mechanical arm teleoperation system based on the VR head-mounted display and the control method, high-precision and flexible teleoperation is achieved by accurately capturing movement tracks of the two hands and mapping the movement tracks to the double-arm mechanical arm, meanwhile, the content shot by the scene camera is transmitted to the VR head-mounted display in a streaming mode, and the real-time performance and immersion of operation are enhanced. In addition, the intelligent obstacle avoidance strategy designed by the invention effectively prevents the movement of the mechanical arm from being out of control in the operation process, and ensures the safety and reliability of the operation. The system has the remarkable advantages of high-precision operation, real-time performance, immersion, multi-scene adaptability, safety, reliability and the like, is suitable for interaction and task execution of a universal robot in industrial, household and other scenes, replaces human beings to execute precise operation in a high-risk environment, and has wide application prospects and important practical significance.
Owner:SUZHOU INST OF BIOMEDICAL ENG & TECH CHINESE ACADEMY OF SCI

Generating a modified digital image utilizing a human inpainting model

The present disclosure relates to systems, methods, and non-transitory computer-readable media that modify digital images via scene-based editing using image understanding facilitated by artificial intelligence. For example, in one or more embodiments the disclosed systems utilize generative machine learning models to create modified digital images portraying human subjects. In particular, the disclosed systems generate modified digital images by performing infill modifications to complete a digital image or human inpainting for portions of a digital image that portrays a human. Moreover, in some embodiments, the disclosed systems perform reposing of subjects portrayed within a digital image to generate modified digital images. In addition, the disclosed systems in some embodiments perform facial expression transfer and facial expression animations to generate modified digital images or animations.
Owner:ADOBE INC

Power system human resource allocation method and system based on multi-agent cooperation

A power system human resource allocation method based on multi-agent cooperation comprises the steps that S1, a human resource task collection module is constructed, and newly-added human tasks of a power system are collected in real time; s2, quantitatively decomposing each human task into a plurality of sub-tasks, calculating the complexity of each sub-task, and constructing a human task balance optimization model; s3, constructing an intelligent agent role allocation model according to the sub-task complexity, the intelligent agent capability matrix and the real-time load of the intelligent agent; and S4, performing iterative operation through a multi-objective optimization algorithm, solving the human task balance optimization model and the agent role allocation model, allocating the subtasks to the agents to obtain a task allocation matrix and estimated completion time, and executing the subtasks by the agents according to the allocation matrix to realize allocation of the human tasks. The design not only considers the balance of resource consumption and processing time of human tasks, but also allocates based on allocation efficiency and task completion efficiency.
Owner:ECONOMIC & TECH RES INST OF HUBEI ELECTRIC POWER COMPANY SGCC

Rich contact operation task-oriented robot skill learning and control method and device

The invention provides a robot skill learning and control method and device oriented to rich contact operation tasks. The method provided by the invention comprises the following steps: collecting human demonstration data; the human demonstration data at least comprises end effector pose data and stiffness matrix data; on the basis of the end effector pose data and the stiffness matrix data, a generalization motion track and a generalization stiffness track are obtained through time alignment, probability distribution modeling and new state constraint adaptation; constructing an obstacle model through environmental perception, taking the generalized motion trajectory and the generalized rigidity trajectory as reference sampling candidate trajectories, screening a collision-free trajectory in combination with motion consistency, rigidity consistency and smoothness cost, and performing iterative optimization to obtain an optimized motion trajectory and an optimized rigidity trajectory; and converting the optimized motion track and the optimized rigidity track into a robot joint torque instruction based on variable impedance control, and controlling the robot to execute a rich contact operation task based on the robot joint torque instruction.
Owner:BEIHANG UNIV

Long-tail human movement prediction method based on adaptive hierarchical learning

According to the long-tail human movement prediction method based on adaptive hierarchical learning, coarse-grained semantic grouping is introduced through a hierarchical tree structure based on the Maslow human motivation theory, and the optimization process is rebalanced in a framework-independent mode. Firstly, thinking chain cues are designed based on the Maslow human motivation theory, a hierarchical structure for city customization is constructed by applying a large language model, and knowledge in human mobile data is fully developed through hierarchical learning. Secondly, exploring hierarchical position prediction through Gumbel interference and adaptive weight, so as to fully capture complex space-time semantics; wherein Gumbel interference is used for rebalancing learning of head and tail positions, adaptive weight performs node-level adjustment on head categories and tail categories in each layer, and these components effectively promote exploration of mobility knowledge. The invention further comprises a method adaptive to the advanced mobile prediction algorithm, and the effectiveness and applicability of other methods in long-tail mobile data are improved by embedding the adaptive hierarchical learning loss into other mobile prediction methods.
Owner:ZHEJIANG UNIV

Humanoid robot imitation learning method and device, computer equipment and storage medium

The invention relates to a humanoid robot imitation learning method and device, computer equipment and a storage medium. The method comprises the steps of obtaining training data; according to the training data, training an initial model to obtain an imitation learning model; enabling a target humanoid robot to execute an action according to the imitation learning model; wherein the training data is data acquired by taking the simulation robot as a center when the simulation robot simulates an original object, and the training data comprises an actual joint position and an actual tail end force. According to the invention, the simulation learning model can be closer to the action habits of human beings, and the control of the humanoid robot can be more accurate and meticulous.
Owner:KEPLER ROBOT CO LTD

Psychotherapy and healing robot based on high human emotion fitting degree simulation analysis

The invention discloses a psychotherapy and healing robot based on high human emotion fitting degree simulation analysis, and the robot comprises a multi-mode perception layer which is used for collecting the interaction data of physiology, movement and environment; the multi-modal sensing layer comprises a heterogeneous data acquisition module, a spatial-temporal feature extraction network and an attention fusion mechanism module; the dynamic decision-making layer is used for generating an intervention strategy based on the interaction data; the dynamic decision-making layer comprises a reinforcement learning strategy engine and a hierarchical intervention selection tree; the generative interaction layer is used for generating a co-estrus response conforming to ethical specifications based on the intervention strategy; the generative interaction layer comprises an ethical constraint system and an emotional response generator; the brain science verification layer is used for monitoring neural feedback in real time through EEG and adjusting an intervention strategy; and the brain science verification layer comprises a neural feedback regulation module and a multi-mode feedback design module. Therefore, a precise and personalized psychological intervention decision closed loop is provided, and the defects of an existing AI psychological product in the aspects of emotion recognition, intervention strategies and effect quantification are overcome.
Owner:BEIJING PUJU HEALTH TECHNOLOGY CO LTD

Music reactive animation of human characters

Example methods for generating an animated character in dance poses to music may include generating, by at least one processor, a music input signal based on an acoustic signal associated with the music, and receiving, by the at least one processor, a model output signal from an encoding neural network. A current generated pose data is generated using a decoding neural network, the current generated pose data being based on previous generated pose data of a previous generated pose, the music input signal, and the model output signal. An animated character is generated based on a current generated pose data; and the animated character caused to be displayed by a display device.
Owner:SNAP INC

Planning, advice, and execution platform including techniques for improving advice through user interactions

Various embodiments described hereby include components of a planning, advice, and execution (PAE) system configured to deliver an advice, planning, and attainment experience that focuses on understanding clients as human beings and what they want to accomplish with their life. The PAE system, or one or more components thereof, may operate to provide technology-based solutions that continuously sync financial objectives with aspirations and values through the many moments of life. These technology-based solutions may empower humans to make financial decisions and attain life objectives, big or small, simple or complex, that make a real and lasting impact on their lives and future generations.
Owner:WELLS FARGO BANK NA

Machine behavior learning method based on emotion driving and human expert feedback

The invention discloses a machine behavior learning method based on emotion driving and human expert feedback, and the method comprises the following steps: a) fusing environment image features extracted by BLIP-2 and text instruction semantic features analyzed by GPT-4, and forming cross-modal input; b) training a basic VLA model through supervised fine tuning (SFT) by using human expert remote control trajectory data and cross-modal input to obtain a basic behavior strategy; c) combining an emotion recognition module with a multi-head self-attention mechanism, fusing emotional interaction dependency into a basic strategy, and generating high-order emotion driven behavior representation; and d) inputting the high-order emotional behavior representation into a reinforcement learning module, storing a track by using a Replay Buffer, carrying out optimization through human expert feedback preference learning, and outputting a final behavior strategy. Compared with an existing method, the method has the advantages of being high in multi-modal feature extraction capacity, high in emotion fusion degree, sufficient in expert feedback utilization and the like, and the response accuracy and interaction experience of a machine to human instructions and emotions can be improved to a certain degree.
Owner:EAST CHINA NORMAL UNIV +1

Robot control method based on tactile prediction pre-training

A robot control method based on tactile prediction pre-training comprises the following steps: acquiring and generating a human playing data set consisting of three-channel image tensors in an offline stage, and training a constructed conditional diffusion model comprising a tactile encoder, a tactile decoder and an action and visual encoder; in the online stage, the trained conditional diffusion model is integrated into a standard imitation learning strategy network, and an action instruction of the robot is generated according to the state of the robot, the current visual features and the tactile feature vectors extracted by the imitation learning strategy network. According to the method, a specific agent task is completed by training a deep neural network model, that is, a future tactile signal sequence is predicted according to historical information and future action intentions; the model is enabled to characterize generic haptic features contacting physical dynamic laws for further migration into downstream robot control tasks.
Owner:SHANGHAI JIAOTONG UNIV

Human-computer interaction method and system based on vision-language-action model

The invention discloses a human-computer interaction method and system based on a vision-language-action model, and belongs to the field of human-computer interaction. According to the method, an anchoring ring strategy is adopted to collect human teaching data to finely adjust the VLA model, and then the trained model is applied to an actual human-computer interaction scene. In the data acquisition stage, a first operator guides a master robot to execute task actions, and a slave robot synchronously moves and interacts with a second operator, and returns to a predefined initial position after each interaction; a teaching sample is formed by recording a robot state, an environment image and an instruction text, and a high-quality data set is generated through data enhancement. In the application stage, the real-time robot state, the environment image and the instruction text serve as input, an action instruction is generated through the VLA model, and the robot is driven to complete a cooperation task. According to the method, the data utilization efficiency and the model generalization ability are remarkably improved, the difference between simulation and the real environment is effectively overcome, and efficient, safe and natural man-machine cooperation is achieved.
Owner:ZHEJIANG UNIV

Service robot cross-modal user intention analysis method and system

The invention provides a service robot cross-modal user intention analysis method and system. The method comprises the following steps: firstly, acquiring and analyzing a voice instruction sequence and a visual flow containing a gesture, fusing a task target object and a gesture pointing direction to generate an intention hypothesis, and then projecting the intention hypothesis to a knowledge graph of an environment with a body; the method comprises the following steps: acquiring an intent hypothesis of a user, querying a candidate object set simultaneously matched with spatial constraints and semantic constraints of the intent hypothesis, when ambiguity exists in the candidate object set, generating an inquiry instruction for eliminating the ambiguity according to a relation path in a knowledge graph of the body environment, and finally analyzing and confirming a final user intent according to the response of the user to the inquiry instruction; according to the technical scheme provided by the invention, the ability of the robot to cope with uncertainties such as unknown reference and environment shielding in a real scene is enhanced, the interactive experience closer to the human cognitive level is realized, and the interactive efficiency and satisfaction of the user in a complex environment are also improved.
Owner:BEIJING ZHONGHE HUICHUANG TECH CO LTD

Self-adaptive safety planning method based on multi-modal intention perception and physiological signals

According to the self-adaptive safety planning method based on the multi-modal intention perception and the physiological signals, through fusion of multi-modal information such as vision, IMU and the physiological signals, accurate intention perception and motion trail prediction are achieved, so that balance is achieved between safety margin and efficiency, and the safety performance of the system is improved. Through fusion of human behavior intention prediction, physiological state feedback and environmental perception data, in combination with advanced planning of intention prediction and immediate adjustment of physiological feedback, sudden actions and environmental changes are coped with. The behavior of the robot or the mechanical arm is dynamically adjusted through feature fusion and a human body intention classification algorithm, the collision probability is effectively reduced, and the man-machine cooperation efficiency and the user comfort degree are optimized while safety is guaranteed.
Owner:CHINA THREE GORGES UNIV

Mechanical arm control method based on visual language model and human feedback

The invention provides a mechanical arm control method based on a visual language model and human feedback. The method comprises the following steps: a) acquiring scene information: acquiring states of a mechanical arm and objects in a scene in real time through a camera; b) generating a control code according to the user instruction: decomposing the user instruction into a plurality of subtasks by utilizing a large language model, and combining and calling a predefined API (Application Program Interface) to generate the control code; c) verifying the generated control code: verifying the code through a visual language model, and judging that the control code completes a user instruction; d) executing the control codes: sequentially executing the control codes according to the sequence of the subtasks, and judging whether the step is successfully executed or not; and e) fault repair and man-machine interaction: feeding back reasons that tasks are not completed according to expectations, and regenerating control codes according to user feedback. According to the invention, the mechanical arm can understand and execute the natural language instruction of the user, and can communicate with the user when an abnormal condition occurs, so that the mechanical arm can complete various tasks in accordance with the expectation of the user.
Owner:EAST CHINA NORMAL UNIV +1

Brain-like decision-making method, device and equipment for body intelligence and storage medium

The invention relates to the technical field of artificial intelligence, and discloses a brain-like decision-making method, device and equipment for body intelligence and a storage medium, and is applied to a robot. The method comprises the following steps: receiving a human language instruction, and carrying out semantic understanding on the human language instruction to obtain a task description; collecting multi-modal data, and generating an environment model based on the multi-modal data; and generating a pulse event sequence by adopting a pulse neural network model based on the task description and the environment model, and generating a motion strategy of the robot based on the pulse event sequence. Through application of the bionic neuromorphic computing architecture, the decision-making speed of the robot in a complex scene is increased by 50 times, the power consumption is reduced to 1 / 10, the real-time response capability of the robot is greatly improved, the robot can quickly cope with various emergencies, for example, when encountering a sudden obstacle, the robot can quickly do an avoidance action, and the robot can be prevented from being damaged. And collision accidents are avoided.
Owner:JIANGXI INST OF FASHION TECH

LLM driven multimodal human-robot interaction planning

A computer-implemented method for controlling a robot collaborating with a human in an environment of the robot comprises: obtaining, by at least one sensor, multimodal information on the environment of the robot including information on a human acting in the environment; converting, by a first converter, the obtained multimodal information into text information; estimating, by an intent estimator, an intent of the human based on the text information; determining, by a state estimator, a current state of the environment including the human based on the text information; planning, by a behavior planner, based on the current state of the environment and the estimated intent of the human, a behavior of the robot including at least one multimodal interaction output for execution by the robot, and generating control information including text information on the at least one multimodal interaction output; converting, by a second translator, the generated text information into multimodal actuator control information; and controlling at least one actuator of the robot based on the multimodal actuator control information.
Owner:HONDA MOTOR CO LTD

Multi-scene internet-of-things off-site dynamic inspection self-adaptive monitoring system

The invention discloses a multi-scene internet of things off-site dynamic inspection self-adaptive monitoring system, and relates to the technical field of intelligent inspection and path planning, the system is composed of a plurality of functional modules, and the system comprises a moving object detection module for detecting a moving object by adopting a Gaussian mixture background modeling and optical flow method; meanwhile, a light-weight MobileViT-XXS model is deployed, and a moving object target is tracked; the moving object classification and identification module is used for judging the type of a moving object through a multi-modal classification decision tree: when the moving object is judged to be a non-living body, carrying out ROI extraction and optimal frame interception on the image saliency of the moving object, and generating EXIF metadata through super-resolution reconstruction; when the judgment result is a biological object, biological category judgment is carried out: when the judgment result is an animal, geo-fencing data matching is carried out, and a spatial-temporal feature fusion model is adopted; and when human beings are judged, face quality evaluation is carried out through multi-frame fusion super-resolution.
Owner:XIAMEN SHUYIDA INFORMATION TECHNOLOGY CO LTD

Automated nonverbal analysis system

Examples relate to computer-implemented methods for analyzing communication in digital evaluation. A computing device accesses multimodal data comprising video and audio information of human subjects and configures a computational model using this data to identify patterns in communication that correlate with assessment metrics. The configuring implements processing techniques that preserve relationships between features across different modalities. When a video recording of a candidate is received, the computing device processes the video using the configured computational model to extract communication features. These features may include facial expressions, gestures, eye movements, posture, vocal tone, and speech patterns. The device generates an evaluation of the candidate based on the extracted communication features and outputs a representation of the evaluation.
Owner:LIGHT STEVEN PATRICK