Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

122 results about "Multimodal interaction" patented technology

Multimodal interaction provides the user with multiple modes of interacting with a system. A multimodal interface provides several distinct tools for input and output of data. For example, a multimodal question answering system employs multiple modalities (such as text and photo) at both question (input) and answer (output) level.

Intelligent digital human training method and system based on multi-modal interaction

The invention discloses an intelligent digital human training method and system based on multi-modal interaction, and belongs to the technical field of semantic indexing.The method specifically comprises the steps that voice, vision and text data are analyzed and converted into high-dimensional feature vectors through a modal exclusive encoder, the high-dimensional feature vectors are projected to a unified semantic space through a cross-modal semantic mapping model, and the high-dimensional feature vectors are obtained; generating a semantic primitive containing a modal identifier, a core semantic tag and a feature weight; semantic primitives are used as nodes, directed edges and edge weight table association strength are established based on semantic similarity, typical scene node connection weights are strengthened, and a mesh map containing intra-modal hierarchy and inter-modal cross association is formed; constructing a double-layer index on the basis of the mesh map; semantic primitives are extracted from newly added data, the position of a new node in an association graph is determined through a graph matching algorithm, an association edge with an existing node is automatically established, and a lower-layer modal exclusive index is synchronously updated.
Owner:JIANGXI INST OF FASHION TECH

Multimodal interaction method, apparatus, controller, system, automobile, and storage medium

The application discloses a multimodal interaction method, device, controller, system, automobile and storage medium. The method comprises the following steps: determining a current interaction dialogue according to a current interaction voice at an interaction time; determining a current scene image and current scene data corresponding to a current interaction interface corresponding to the interaction time; adopting a multimodal recognition model to perform multimodal recognition on the current interaction dialogue, the current scene image and the current scene data, and determining a target control instruction; and executing the target control instruction to complete a human-computer interaction operation. The method can guarantee the output efficiency and accuracy of the target control instruction, reduce the complexity of customizing templates, rules and associations and other complex control logics in the development process, and improve the adaptability and generalization ability of voice interaction.
Owner:BYD CO LTD

Multi-modal interaction control method and system of multifunctional teaching assisting robot and robot

The invention relates to the technical field of intelligent education and Internet of Things fusion, in particular to a multi-modal interaction control method and system of a multifunctional teaching-assistant robot and the robot. According to the system, a tablet AI processor runs an Android system as a core, a display module, a man-machine interaction module, an audio input module, an audio output module, an image acquisition module, a temperature and humidity sensor module, a WIFI or Bluetooth module, an Ethernet module and an Internet of Things module are connected, and the Internet of Things module supports RS485 wired access and Zigbee wireless access at the same time; the tablet AI processor executes voice interaction, video call, face recognition, environment monitoring, network interaction and equipment linkage, and provides a face recognition alignment and depth feature matching algorithm and an annular microphone array beam forming and sound source direction estimation algorithm, so that multi-modal interaction and multi-equipment linkage in a teaching scene are realized. The convenience and the safety are improved; and the equipment access cost is reduced.
Owner:SHENZHEN YUXIN DIGITAL TECH CO LTD

Automatic generation method of rejection defense document based on multi-agent collaboration

The invention discloses a method for automatically generating a rejection defense document based on multi-agent collaboration, and belongs to the technical field of artificial intelligence. The method comprises the steps of collecting and preprocessing original multi-modal interaction data related to a payment refusing case, processing the original multi-modal interaction data into a text form to obtain an original interaction text, and storing the original interaction text in a database; performing cleaning processing on the original interaction text in the database by adopting Non-Agent to obtain semantic intermediate representation with consistent format; and based on the structured input, generating an anti-distinguishing reason by a multi-agent collaborative anti-distinguishing framework to obtain an anti-distinguishing document of the current payment refusing case. According to the invention, based on structured layering, responsibility constraint and parallel scheduling, the key defects of'hallusion caused by input noise ', 'single LLM responsibility overload' and'end-to-end time delay 'in the prior art are overcome together, so that the automatic defense signal generation system realized based on the method is remarkably improved in the aspects of compliance and fact accuracy; and practical advantages are embodied in business generalizability and engineering efficiency.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Interactive Page System

A web page embeds an AI-driven widget that converts static content into an interactive experience. Executable code builds a page interaction index from the page's DOM, including text, selectors, and positional metrics for DOM nodes. In response to natural-language input, pipelines perform summarization, stepwise explanations, voice-guided form completion with rule-based validation, and on-page product scanning to create a dynamic, user-tunable comparison table. Results are rendered as in-place overlays with interactive back-references that highlight source DOM nodes in the viewport. Speech recognition and text-to-speech enable multimodal interaction. Optional privacy gating redacts sensitive data or routes processing to local models. The system improves webpage usability by binding AI outputs to precise DOM regions and providing unified, context-aware assistance within the page.
Owner:BOLOURI RAMIN

Multi-mode fused immersive interaction system for virtual exhibition hall

ActiveCN121879587AResolve semantic ambiguity issuesImprove understanding accuracyInput/output for user-computer interactionImage data processingDefuzzificationComputer graphics (images)
The invention relates to the technical field of virtual reality interaction, and particularly discloses a multi-modal fused virtual exhibition hall immersive interaction system, which comprises the following steps of: acquiring real-time multi-modal interaction data of a user and performing synchronous preprocessing; inputting the orientation semantic words into an interval type-2 fuzzy logic system to generate a three-dimensional space semantic membership field; according to the gesture pointing vector and the hand shaking amplitude, a pointing conical area is constructed, two-dimensional Gaussian distribution is established in a layered mode, and a gesture pointing probability field is generated; taking a semantic membership field, a gesture pointing probability field and a gaze point Gaussian kernel density field as independent evidences, introducing a wall boundary and a floor channel as spatial topology constraints, and fusing by adopting an evidence theory combination rule to generate a three-dimensional intention probability distribution field; carrying out gravity center defuzzification processing on the distribution field to extract a navigation target area, and planning a navigation path by taking the current position of the user as a starting point and taking a nearest floor channel entrance as a passing point; according to the method, the problems of multi-mode fuzzy intention understanding and space adaptation are solved.
Owner:XINZHIHANG MEDIA TECH GRP CO LTD

Intelligent conference multi-modal interaction optimization method and system based on large model

The invention provides an intelligent conference multi-modal interaction optimization method and system based on a large model, and the method comprises the steps: obtaining conference multi-modal data, and constructing a locked image frame set based on an image content snapshot locking mechanism; inputting each piece of modal data of the multi-modal data stream into a pre-trained large model to generate embedded vectors corresponding to the modal data, performing anchor point extraction by using similarity calculation between each pair of embedded vectors, and screening through the large model to obtain a semantic anchor point set; when the conference is carried out, each new speech is converted into an embedded vector, and then the similarity between the embedded vector and each semantic anchor point is calculated to obtain an anchor point reference frequency set; and according to the anchor point reference frequency set and the anchor point state set, generating a real-time interaction suggestion, and based on the interaction suggestion, feeding back an updated anchor point state and an updated anchor point priority, and updating the interaction suggestion. According to the invention, automatic identification, structured representation and process state perception of the conference content are realized.
Owner:GUANGZHOU HUIYI INFORMATION TECHNOLOGY CO LTD

Intelligent exhibition hall multi-mode interactive digital human system and implementation method

The invention relates to the technical field of computer data processing, and discloses an intelligent exhibition hall multi-modal interaction digital human system and an implementation method, which are used for solving the problem that multi-source data of multi-modal interaction lacks a unified clock and verifiable timestamp alignment mechanism in a traditional method. According to the method, a unified time domain and a time version are established on an edge side, terminal access is restrained, and an acquisition time mark, an access time mark and a serial number are written in data or a state; performing gating shunting according to a time version, performing de-duplication and out-of-order rearrangement based on a serial number and double time marks, and generating and solidifying a session window evidence index; locking a transaction window boundary according to the evidence index, generating a participation source list and a transaction number, establishing a fragment reference relationship and generating an alignment voucher; in the linkage stage, phase division issuing is carried out according to a preparation phase, an execution phase and a confirmation phase, backward reading verification is carried out according to an action sequence number, a transaction log is archived, and alignment, rechecking and playback of an interaction link are achieved.
Owner:SUZHOU CHUANGJIE MEDIA EXHIBITION CO LTD

A Low-Light Scene Analysis Method Based on Multimodal Feature Fusion and Clustering

This invention discloses a low-light scene analysis method based on multimodal feature fusion and clustering, belonging to artificial intelligence technology. It constructs a single-branch feature extraction network based on Transformer; a multimodal feature interaction and fusion module to achieve multimodal feature interaction and fusion; a multimodal fusion feature clustering module to cluster the features after multimodal interaction and fusion, utilizing the semantic and spatial distances between different features to achieve clustering while simultaneously downsampling the features, mitigating the problem of loss of detailed edge information in structured downsampling; and a multi-scale feature aggregation and decoding module to receive feature information from the encoding network and classify each feature pixel according to the semantic distance of the multi-scale features. This invention can fully utilize visible light and thermal image information and can be applied to scene analysis and navigation of unmanned systems in diverse and complex low-light scenes.
Owner:CHINA UNIV OF MINING & TECH

Interaction story machine based on ROS and large model and control method thereof

The invention discloses an interactive story machine based on an ROS and a large model and a control method of the interactive story machine. The story machine comprises an input module, an ROS system, a large model module, an audio output unit and a color ink screen. According to the invention, through cooperative work of software and hardware modules, an ROS system is used as an intelligent scheduling center, and meanwhile, a large model module is introduced, so that user interaction input modes and personalized demands of different user types can be understood and responded, and personalized story contents are dynamically generated; for the hearing-impaired user type, based on a barrier-free interaction mode, generating image prompt information and a smooth sign language action sequence, and driving a color ink screen to perform visual presentation through an optimized rendering instruction; and for a non-hearing-impaired user type, synchronously generating a story text fragment, a matched image illustration and voice information based on a multi-mode interaction mode, and providing an immersive story reading environment for different types of child users.
Owner:GUANGDONG UNIV OF TECH

A method and system for matching the dialogue intent of intelligent NPCs in multimodal interaction

This invention relates to the field of natural language processing technology, specifically to a method and system for matching the intent of intelligent NPC dialogues in multimodal interaction. The method involves real-time acquisition of multimodal data, generating multimodal semantic features through submodal preprocessing; integrating the semantic features of each modality using a cross-modal fusion module based on semantic association, generating a candidate intent set based on semantic context representation and combined with a semantic parsing module and predefined intent templates; tracking changes in user intent in real time through a continuous semantic learning mechanism and an interactive memory module, combined with a Bayesian update method, dynamically adjusting the confidence level of each intent in the candidate intent set, and filtering the final intent; constructing an NPC semantic cognition model, and performing semantic consistency analysis to perform semantic checks on user input and NPC dialogue state; combining the final intent and a decision engine to generate a dialogue strategy and output synchronized response content; this invention improves the accuracy of dynamic matching of intelligent dialogue intents.
Owner:JIANGSU COLDPLAY INFORMATION TECH CO LTD

Toys (AI Smart Plush Toys)

1. Name of the product in this design: Toy (AI Intelligent Plush Toy). 2. Purpose of this design: Intelligent plush toys for AI-powered multimodal interaction, emotional companionship, and educational purposes. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: a 3D model.
Owner:TONGDA SMART TECH (XIAMEN) CO LTD

A digital reading platform construction system supporting multi-modal content interaction

PendingCN122451027ADigital readingEngineering
The application discloses a kind of digital reading platform construction systems of supporting multimodal content interaction, it is related to digital reading technical field, the present application includes multimodal content acquisition module, multimodal interaction analysis module and digital reading platform construction module, the present application is classified and is regularized by from multiple channels acquisition heterogeneous reading resources, standardization multimodal content feature database is constructed by mode extraction feature, pre-processing and semantic analysis are carried out to user multimodal interaction input, user interaction intent is identified and interaction execution instruction is generated by progressive threshold matching and scene weighting score mechanism, based on feature database, build hierarchical platform framework, configure control component and establish multimodal content display and interaction response logic, the present application improves the resource management efficiency of digital reading platform, interaction accuracy, immersive experience and running reliability.
Owner:CHINA FOCUS LTD

A medical care-patient bidirectional communication translation method based on multi-modal interaction

This invention belongs to the field of medical wristband technology, specifically disclosing a translation method for bidirectional communication between medical staff and patients based on multimodal interaction, including the following steps: Step S1, the multimodal interaction unit receives multimodal input signals; Step S2, the multimodal interaction unit fuses the multimodal input signals to generate comprehensive input information; Step S3, the comprehensive input information is identified and translated based on a medical-specific translation engine; Step S4, the context-aware module analyzes the dialogue content, identifies the scenario, and dynamically adjusts the translation results; Step S5, the adjusted translation results are synchronously output bidirectionally and broadcast via voice. This invention achieves three core breakthroughs in clinical cross-language communication through a translation mechanism combining multimodal signal fusion and context awareness: enabling real-time, accurate, and private barrier-free communication between doctors and patients, reducing the risk of medical errors due to language barriers, improving consultation efficiency and patient satisfaction, and reducing the risk of misdiagnosis.
Owner:HUAZHOU PEOPLES HOSPITAL

Cockpit multi-modal interaction control method, device, equipment, medium and product

PendingCN122300532AImprove experienceAvoid inoperable modalsInteraction controlData pack
This invention belongs to the field of intelligent cockpit technology and discloses a cockpit multimodal interaction control method, device, equipment, medium, and product. The method includes: acquiring user interaction capability characteristic data and historical interaction data, wherein the historical interaction data includes the interaction modal used in each interaction, associated scene data, and interaction performance; dividing the scene data into sub-scenes; for different sub-scenes, determining the priority of the corresponding interaction modal based on the usage frequency and interaction performance of the interaction modal in that sub-scene, obtaining an initial interaction priority matrix; judging the user's capability adaptability to each interaction modal based on the user interaction capability characteristic data, and correcting the initial interaction priority matrix based on the capability adaptability; acquiring current scene data, and determining the highest priority interaction modal in the current sub-scene as the current standby primary modal based on the user interaction priority matrix. This invention can meet the diverse interaction needs of different users.
Owner:CHERY AUTOMOBILE CO LTD

A multimodal interaction control method for a large field of view display system

The application discloses a multimodal interaction control method of a large field of view display system, and belongs to the technical field of reality augmentation and intelligent human-computer interaction. The method comprises the following steps: acquiring environment state data in real time; synchronously collecting gesture, eye movement and voice multimodal interaction data of a user; dynamically determining a multimodal fusion decision strategy adapting to a current condition based on the environment state data; according to the strategy, performing collaborative analysis and fusion decision on the multimodal data through a multimodal attention fusion network to generate a final control instruction; and executing the instruction to control display output. The application dynamically adjusts a strategy through environment perception, and combines an intelligent fusion algorithm to realize efficient and robust collaboration of gesture, eye movement and voice modes, effectively overcomes the limitations of a single mode in a complex and changeable environment, and significantly improves the naturalness, accuracy and reliability of human-computer interaction in high-load and high-safety requirement scenes such as aviation.
Owner:SUZHOU LIPAI TECH CO LTD

Personalized recommendation method for virtual digital humans in enterprise publicity

The invention relates to the technical field of enterprise propaganda, and discloses a personalized recommendation method for virtual digital humans in enterprise propaganda. According to the method, multi-modal interaction data such as visual attention data and voice feedback data in the interaction process of a user and a virtual digital human are collected in real time, and an original interaction flow is generated; performing multi-dimensional fusion analysis on the original interaction flow, analyzing a dependency relationship among different dimensions, identifying a hidden association between a user interest mode and a virtual digital human performance feature, labeling an analysis result as an initial interest index, and generating an interest labeling data set with confidence; optimizing model adaptability and dynamically updating a model state by utilizing a fusion analysis result; matching the real-time interaction data with the dynamic interest evolution model, and identifying recommendation candidates and opportunities to obtain recommendation contexts; and mining a resource library based on a matching result, identifying a deep recommendation strategy and an adjustment signal, and tracing the strategy to identify a core factor and an optimization direction, thereby realizing accurate and adaptive personalized recommendation.
Owner:ANHUI RUIXUAN SUPPLY CHAIN TECH CO LTD

Construction method and device of attack input, equipment and storage medium

The invention discloses an attack input construction method and device, equipment and a storage medium, and relates to the technical field of large model security, and the disclosed attack input construction method comprises the following steps: obtaining an original attack instruction of a multi-modal large model; based on the semantic information of the original attack instruction, extracting a main intention and behavior logic represented by the original attack instruction; based on the main intention and behavior logic, multiple frames of target images are generated, and the multiple frames of target images are used for describing the process of realizing the main intention; based on the multiple frames of target images, target attack input is constructed, and the target attack input is used for attacking the multi-modal large model. According to the method, the target attack input is used for attacking the large model, the attack effectiveness and success rate are effectively improved, the method is suitable for complex attack scenes, the potential of multi-modal interaction is fully played, stable threats can be formed for the large model with certain defense capability, and the security of the large model can be conveniently improved in an auxiliary mode.
Owner:BEIJING QIHOOD TECHNOLOGY CO LTD

Supply chain multi-mode interactive question and answer method based on large language model

The invention discloses a supply chain multi-modal interactive question and answer method based on a large language model, and the method specifically comprises the steps: S1, collecting text, voice, image and table data in a supply chain scene, and extracting semantic features to form a multi-modal semantic vector; s2, extracting inventory, transportation, production and order time sequence data, and inputting the data into the improved self-organizing mapping neural network to generate a semantic state topological structure; s3, executing concept drift detection and locally reconstructing nodes, and outputting a stable supply chain semantic state vector; s4, fusing the supply chain semantic state vector and the user context information to generate context state enhanced representation; s5, calculating a multi-modal correlation weight to realize semantic alignment and unified coding; and S6, inputting the large language model and combining with knowledge graph reasoning to generate text, voice or chart answers. According to the method, supply chain multi-modal information intelligent fusion and semantic question and answer accurate generation are realized, and the decision-making efficiency and the intelligent interaction level are remarkably improved.
Owner:江西博微新技术有限公司

Multi-mode interaction control method and system based on docking station, medium and product

The invention discloses a multi-mode interaction control method and system based on a docking station, a medium and a product, and relates to the field of docking stations. The method comprises the following steps: acquiring page layout information of a display screen of the docking station, and generating structured user interface metadata; generating a preset collaborative activation area based on the user interface metadata; if the behavior of the cursor of the pointer equipment in the preset cooperative activation area meets a cooperative control triggering condition, triggering a cooperative control mode, and mapping a two-dimensional coordinate of the cursor into a target coordinate under a display screen coordinate system of the docking station; sending an instruction of the target coordinate to a docking station, so that the docking station displays a proxy cursor on the target coordinate; if it is determined that the cursor of the pointer device triggers the preset interaction triggering event, the proxy cursor is controlled to execute the preset interaction triggering event, so that the docking station executes operation corresponding to the preset interaction triggering event, and cross-device interaction is completed. By implementing the technical scheme provided by the invention, cross-device interaction is completed, and the user experience is improved.
Owner:SHENZHEN SINOBRY ELECTRONICS LTD

A vehicle machine testing method

This invention discloses a vehicle infotainment system testing method. By having a controller uniformly receive modal simulation parameters, generate precise multimodal interaction commands, and coordinate the synchronous actions of the multimodal interaction simulation device and the audio / video acquisition device, a fully automated closed-loop testing process is constructed. This application fundamentally eliminates the randomness and uncertainty caused by manual operation. Since the test input commands (such as voice content, touch coordinates, and timing) are precisely reproduced by the device according to preset parameters, and the entire execution, acquisition, and judgment process is automatically driven by the controller, it ensures that the test conditions, execution steps, and triggering logic remain consistent at any time and for any number of tests. This application's embodiments make performance test results (such as response latency) objectively comparable, providing a reliable and consistent technical foundation for accurately evaluating vehicle infotainment system performance and identifying performance regression between versions.
Owner:VOYAH AUTOMOBILE TECH CO LTD

Intelligent scale adaptive operation guiding method based on multi-modal interaction

The invention relates to an intelligent scale self-adaptive operation guiding method based on multi-mode interaction, and the method comprises the steps: collecting user voice through a voice interaction module after goods on an intelligent scale are changed, and extracting an effective instruction; analyzing the effective instruction through an instruction understanding engine, and comparing the effective instruction with a pre-stored template of a local action mapping library; when the effective instruction is matched with the pre-stored template of the action mapping library, generating a dynamic guide instruction through an SOP generator, and broadcasting the dynamic guide instruction to a user step by step; and when the effective instruction is not matched with the action mapping library pre-stored template, performing semantic association on the effective instruction through an instruction understanding engine to match the most similar action mapping library pre-stored template, then generating a dynamic guide instruction through an SOP generator, and broadcasting the dynamic guide instruction to a user step by step. According to the method, the effective instruction can be compared with the pre-stored template of the action mapping library through the instruction understanding engine, the optimal instruction is analyzed, the user operation is guided, and the method has better operation efficiency.
Owner:TAIHENG PRECISION MEASUREMENT & CONTROL (KUNSHAN) CO LTD

Method and system for cooperative interaction of multi-mode blind area outside vehicle

The invention discloses an out-of-vehicle multi-mode blind area cooperative interaction method and an out-of-vehicle multi-mode blind area cooperative interaction system. Starting blind area sensing when a preset starting condition is met; the camera and the millimeter-wave radar collect target position, speed, attitude and track information; the edge calculation unit performs multi-source fusion and behavior recognition, and outputs a target behavior intention and / or a risk result; the strategy decision-making module is used for matching an interaction strategy, controlling laser projection equipment to project a warning or guiding graph in a target area, and controlling a directional sound wave array to output a directional voice prompt; and continuously monitoring the target state and updating the strategy, and stopping output and waiting when a termination condition is met. The system comprises a sensing layer module, a multi-mode interaction terminal, a strategy decision module and a template adaptation module, and template parameters are used for generating projection and voice output parameters and can be linked with a vehicle light and / or steering system through a vehicle-mounted bus.
Owner:YIXIAN INTELLIGENCE

A method and system for AI character interaction in a smart lollipop

This invention discloses an AI character interaction method and system for a smart lollipop. The smart lollipop includes a smart handle and multiple candy heads with different physical shapes. The smart handle is communicatively connected to a user terminal. The candy heads are detachably installed at the front end of the smart handle. Each type of candy head contains a built-in character identification module. The method includes: after the current candy head is plugged into the smart handle, reading the character identification module in the current candy head to load the current AI character; then collecting the user's voice data and analyzing the emotion category and intent information; inputting the analysis results into the personalized interaction model corresponding to the current AI character, which can output a response voice that matches the style of the current AI character. This invention couples hardware form, emotion analysis, and dynamic response technologies, upgrading the existing smart lollipop from a one-way playback tool into an AI interactive terminal with emotion perception, character evolution, and personalized dialogue capabilities, realizing multimodal interaction and emotional companionship.
Owner:AMES (GUANGDONG) FOOD TECH CO LTD +1

A multi-modal fusion-based intelligent home central control system and method

This invention discloses a smart home central control system and method based on multimodal fusion. The control system includes a multimodal interaction interface layer, a device execution layer, and a central controller. The central controller is communicatively connected to both the multimodal interaction interface layer and the device execution layer. The central controller includes an instruction receiving and parsing unit, a context awareness unit, a multimodal fusion decision engine, a scheduling and execution unit, and a feedback unit. The instruction receiving and parsing unit receives and parses interactive instructions; the context awareness unit acquires various dynamic data to provide the context information required for decision-making; the multimodal fusion decision engine is trained based on a hierarchical hybrid decision architecture, which includes a rule engine layer, a probabilistic inference layer, and a reinforcement learning layer. This invention achieves a more accurate, robust, and personalized smart home control experience through a dedicated fusion algorithm, a hierarchical conflict arbitration strategy, hardware and software co-optimization, and an adaptive feedback mechanism.
Owner:GUANGXI UNIV FOR NATITIES

An abnormal behavior recognition method, device and system based on multi-modal interaction

PendingCN122290048AHuman bodyFeature extraction
This invention provides a method, device, and system for abnormal behavior recognition based on multimodal interaction. It simultaneously extracts RGB sequences and human skeleton sequences from surveillance videos, performing feature extraction and alignment respectively. Second-order interaction features between the two modalities are calculated using compact bilinear pooling, and a lightweight Transformer encoder is used for global spatiotemporal correlation modeling to achieve two-stage deep fusion. Simultaneously, a feature pyramid module with a three-level structure (local, regional, and global) is designed to extract multi-scale spatiotemporal features, which are then dynamically fused using an adaptive attention mechanism. Finally, multi-level features are aggregated for classification decisions. This invention integrates the complementary advantages of appearance and structural information, achieving deep cross-modal interaction and multi-scale understanding, significantly improving the accuracy and robustness of behavior recognition in complex scenarios, and is particularly suitable for public security monitoring scenarios such as rail transit.
Owner:SHENYANG ERYISAN ELECTRONICS TECH CO LTD

Kitchen safety cooperative control system based on multi-mode interaction

The invention relates to the technical field of intelligent instrument data interaction, in particular to a kitchen safety cooperative control system based on multi-mode interaction, and is characterized in that an intelligent screen in the system is in communication connection with a system server, and the system server is in communication connection with a gas meter; the intelligent screen is in communication connection with the door and window control system, the safety valve, the alarm, the sensor and the range hood, more humanized operation experience is provided for a user by integrating multiple interaction modes such as visual sense, auditory sense and tactile sense, and the multi-mode interaction function is achieved. Through a hybrid architecture of edge computing and cloud computing, real-time performance and expansibility are both considered, the local quick response requirement is met, and strong data analysis capability is achieved. Through an adaptive learning algorithm of the system server, a decision model of the system server can be continuously optimized according to historical data, and performance and reliability are gradually improved.
Owner:ZENNER METERING TECH (SHANGHAI) LTD

A multi-modal interaction and dynamic-based cognitive impairment digital therapy system

PendingCN122290871Aincrease engagementObjective quantitative assessment of cognitive functionPersonalizationPhysical medicine and rehabilitation
This invention discloses a digital therapy system for cognitive impairment based on multimodal interaction and dynamics. Through an interaction and data acquisition layer, multidimensional data is collected via gamified tasks in a gamified training environment, and preprocessed to obtain preprocessed multidimensional data. An intelligent analysis and decision-making layer is responsible for modeling, evaluating, and optimizing the preprocessed multidimensional data to generate rehabilitation strategies. An application and presentation layer provides dynamically adjusted training content, enabling objective, continuous, and fine-grained quantitative assessment of cognitive function. Based on the patient's real-time abilities and performance, truly personalized rehabilitation training content is dynamically generated and adjusted, enhancing the fun and patient participation in the training process, and improving rehabilitation efficacy and efficiency.
Owner:THE FIRST REHABILITATION HOSPITAL OF SHANGHAI

Companion robot interaction method and system

PendingCN122287691AImprove interactive playabilityhigh degree of personalizationPersonalizationRobot hardware
This application discloses a companion robot interaction method and system, relating to the field of multimodal interaction technology and applied to a cloud service platform. The method includes: acquiring interaction events reported by a user through a companion robot hardware terminal or mobile terminal; converting the interaction events into interaction commands; and sending the interaction commands to the companion robot hardware terminal, so that the companion robot hardware terminal can trigger game plots or develop personality attributes according to the interaction commands, and interact with the user according to the game plots or personality attributes. Through the interwoven personality attribute development and game plot-triggered interaction process, the playability and personalization of human-computer interaction are both enhanced.
Owner:SHENZHEN MUSI INTERACTIVE ENTERTAINMENT TECHNOLOGY CO LTD