Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

182 results about "Multimodal interaction" patented technology

Multimodal interaction provides the user with multiple modes of interacting with a system. A multimodal interface provides several distinct tools for input and output of data. For example, a multimodal question answering system employs multiple modalities (such as text and photo) at both question (input) and answer (output) level.

Intelligent digital human training method and system based on multi-modal interaction

The invention discloses an intelligent digital human training method and system based on multi-modal interaction, and belongs to the technical field of semantic indexing.The method specifically comprises the steps that voice, vision and text data are analyzed and converted into high-dimensional feature vectors through a modal exclusive encoder, the high-dimensional feature vectors are projected to a unified semantic space through a cross-modal semantic mapping model, and the high-dimensional feature vectors are obtained; generating a semantic primitive containing a modal identifier, a core semantic tag and a feature weight; semantic primitives are used as nodes, directed edges and edge weight table association strength are established based on semantic similarity, typical scene node connection weights are strengthened, and a mesh map containing intra-modal hierarchy and inter-modal cross association is formed; constructing a double-layer index on the basis of the mesh map; semantic primitives are extracted from newly added data, the position of a new node in an association graph is determined through a graph matching algorithm, an association edge with an existing node is automatically established, and a lower-layer modal exclusive index is synchronously updated.
Owner:JIANGXI INST OF FASHION TECH

AI-based digital media interface design optimization method

The invention discloses an AI-based digital media interface design optimization method, and relates to the technical field of interface design, and the method comprises the following steps: collecting multi-modal interaction input data of a user in real time, and constructing a user behavior sequence tensor with multiple time steps and multi-modal interaction dimensions; carrying out joint modeling on the user focus state vector, the user behavior sequence tensor and the user real-time feedback vector; generating a recommended new layout scheme based on an output result of the interactive intention weight model; when the probability is higher than a preset threshold value, triggering an interface pushing mechanism; after the interface is rearranged, data are fed back according to subsequent interaction behaviors and feedback data of the user; the technical problem that the layout or position of the interface element cannot be dynamically adjusted according to the behavior of the user in the traditional interface design is solved.
Owner:SHANDONG ZIMO CREATIVE DESIGN CO LTD +1

Customer service interaction method and system fusing AI digital employee and multi-agent decision

The invention relates to a customer service interaction method and system fusing AI digital employees and multi-agent decision, and the method comprises the steps: receiving a multi-mode interaction request from a user, converting the multi-mode interaction request into interaction text data in a unified format, and forwarding the interaction text data to a multi-agent decision unit. And performing intention recognition and sentiment analysis on the interactive text data through the multi-agent decision-making unit to obtain a recognition analysis result. And based on the identification analysis result, guiding the interaction process of the user in combination with the historical interaction content so as to determine the interaction task type, and feeding back the interaction task of the corresponding type to the AI digital employee. And in response to the interaction task, calling the AI digital employee to execute the interaction task according to the rules and knowledge in the knowledge base, and generating a task execution result. According to the task execution result and the recognition analysis result, reply content based on the multi-modal interaction request is generated, the reply content is fed back to the user side, and the flexibility and accuracy of the interaction process are improved.
Owner:CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD

Dialogue Agent interaction method based on multimodal intention understanding

The invention relates to the technical field of man-machine interaction, in particular to a dialogue Agent interaction method based on multi-modal intention understanding, which comprises the following steps: S1, collecting multi-modal data in a user interaction process in real time, and calculating a time synchronization deviation value of each modal data source; s2, constructing a space-time fusion feature vector; s3, analyzing a dominant action instruction and a recessive behavior clue in the space-time fusion feature vector; s4, generating a multi-level intention analysis tree; s5, when the corrected confidence coefficient of any node in the intention analysis tree is lower than a set threshold value, activating a targeted sensor to complementarily collect data; and S6, analyzing a tree drive response decision according to the finally confirmed intention. According to the method, high-precision identification and response control of the dialogue Agent on the user intention in a complex scene are realized by constructing a multi-modal interaction method with space-time consistency fusion capability, an explicit and implicit intention analysis mechanism and an adaptive modal clarification strategy.
Owner:ZHONGKE JUXIN INFORMATION TECH BEIJING CO LTD

Intelligent voice telephone robot system and method based on multi-modal interaction and dynamic decision

The invention belongs to the field of intelligent information system management, and particularly discloses an intelligent voice telephone robot system and method, voice and image multi-mode data are collected through a microphone and a camera, and after preprocessing, voice, emotion and semantic features are fused through an improved Transform architecture to achieve accurate recognition of user intentions; a double-layer decision network based on reinforcement learning is combined with a dynamic reward function to generate an optimal response strategy; and realizing rapid task migration and parameter optimization of the model by adopting a meta-learning mechanism. The system also has the functions of adaptive noise robustness, multi-language interaction, user portrait dynamic updating, man-machine collaboration and the like. Compared with a traditional scheme, the intention recognition accuracy, the task completion rate and the scene adaptability are remarkably improved, the interaction experience is effectively improved, and the method can be widely applied to the fields of customer service, intelligent marketing and the like.
Owner:BEIJING XINJIACHUN TECHNOLOGY CO LTD

Self-adaptive training method and system for cognitive function of old people based on multi-modal interactive feedback

The invention discloses an elderly cognitive function adaptive training method and system based on multi-modal interaction feedback, and relates to the technical field of smart medical treatment. The method comprises the steps that basic information of a user is collected for initial cognitive ability evaluation, a user cognitive portrait is constructed according to an evaluation result, and an initial training task with the corresponding difficulty is allocated; collecting multi-modal interaction data in real time according to the initial training task; carrying out fusion analysis on the multi-modal interaction data by utilizing a machine learning model to obtain a quantized real-time state index; based on the real-time state index and the performance data of the current task, dynamically adjusting a subsequent training task through an adaptive decision rule engine; all-dimensional data of each training task is recorded, a visual cognitive competence development trend report is generated through longitudinal comparative analysis, and a machine learning model and a self-adaptive decision rule engine are continuously optimized and trained by utilizing accumulated user data to form an optimized training closed loop. The cognitive function training effect of the old people can be improved.
Owner:JILIN ACAD OF TRADITIONAL CHINESE MEDICINE

Multi-modal task interaction assisting method and system and computer equipment

The invention relates to the technical field of man-machine interaction, and discloses a multi-modal task interaction assisting method and system and computer equipment, and the method comprises the steps: collecting multi-modal interaction data of a user executing a multi-dimensional interaction task; determining a user behavior feature vector based on the multi-modal interaction data, and constructing a task completion quantitative model; determining the weight of each operation step by adopting a dynamic weight distribution strategy based on the multi-dimensional interaction task completion degree quantification model, and dynamically adjusting a preset adaptive coefficient by adopting a preset adjustment mechanism based on the multi-modal interaction data and the weight of each operation step; generating a dynamic task guiding strategy by adopting a dynamic path planning algorithm based on the weight of each operation step and a dynamically adjusted preset adaptive coefficient; and executing the dynamic task guiding strategy to obtain a multi-modal execution result. According to the method, a complete closed-loop process is formed from data acquisition, analysis, decision making to execution, intelligent services can be provided for middle-aged and elderly users, and the user experience and the task execution efficiency are improved.
Owner:BEIJING RENSHENG INTELLIGENT TECHNOLOGY CO LTD

Multi-modal interactive virtual teaching method, system, equipment and medium

The invention discloses a multi-mode interactive virtual teaching method and system. The method comprises the following steps: acquiring multi-source data; extracting voice data acoustic features, and inputting the voice data acoustic features into a learning model to obtain a text triple; key terms are extracted from the text data, a traceable operation chain is generated, semantic analysis is carried out, and an operation scheme is output in combination with a knowledge graph; performing abnormal state recognition on the image data through a target detection model, positioning abnormal equipment in combination with a character recognition model, performing action mapping through gesture recognition and a spatial constraint rule, and outputting a corresponding instruction; fusing the three types of outputs to generate a scheduling event chain; and performing semantic analysis, evaluation and optimization on the scheduling event chain, interacting with personnel, updating operation suggestions and providing an operation analysis result. According to the method, manual rechecking requirements are reduced through multi-modal data collaboration, a new man-machine interaction database and a new man-machine interaction standard in the power industry can be formed through multi-modal interaction rules and event chain construction, and data intelligent driving is achieved while the training efficiency is improved.
Owner:GUANGXI POWER GRID CORP

Multi-modal fusion man-machine interaction control method, system and equipment and storage medium

The embodiment of the invention provides a multi-mode fusion man-machine interaction control method, system and device and a storage medium, and relates to the technical field of intelligent driving, and the method comprises the steps: obtaining vehicle driving data and driver state data; based on the vehicle driving data and the driver state data, calculating fusion weights of a visual mode, an auditory mode and a tactile mode to obtain a multi-mode fusion weight matrix; generating a multi-modal interaction signal according to the multi-modal fusion weight matrix, wherein the multi-modal interaction signal comprises a visual signal, an auditory signal and a tactile signal; and triggering corresponding visual warning, auditory warning and tactile warning according to the multi-mode interaction signal. In this way, multi-mode fusion is conducted on vision, hearing and touch, the output intensity of visual warning, hearing warning and touch warning is adjusted according to the weight of each mode of vision, hearing and touch, a driver is helped to make the most urgent operation at present according to the warning of each mode, and the driving safety of man-machine interaction in the driving process is improved.
Owner:FAW HAIMA AUTOMOBILE CO LTD +1

Response method, device and equipment for customer service robot

The invention provides a response method, device and equipment for a customer service robot, and relates to the technical field of customer service robots, and the method comprises the steps: obtaining multi-mode request data inputted into a current session with the customer service robot by a user, and a conversation state of the current session; according to the request data of different modals and the modal features of the request data of the corresponding modals, constructing a multi-modal heterogeneous graph; fusing the different modal features based on the request data of the different modalities, the corresponding modal features, the dialogue states and the multi-modal heterogeneous graph to obtain multi-modal fusion features; inputting the multi-modal fusion feature into a pre-constructed user intention response model based on a causal relationship to obtain a response strategy of the customer service robot to the multi-modal request data; wherein the user intention response model is used for representing the causal relationship between the multi-modal interaction data input by the user and the response strategy of the customer service robot. The response accuracy of the customer service robot can be improved, and the user interaction experience is improved.
Owner:BEISEN CLOUD COMPUTING CO LTD

Digital cultural tourism management system based on multi-source data analysis

The invention discloses a digital cultural tourism management system based on multi-source data analysis, and the system comprises a multi-source heterogeneous data collection module which is used for collecting tourist behavior data, environment data, cultural resource data and third-party platform data in real time; the user demand intelligent analysis module is used for constructing a dynamic user portrait through the multi-modal interaction data and identifying dominant and implicit demands; the culture knowledge graph construction module is used for generating a reasonable multi-dimensional knowledge network based on culture resource attributes and historical data; the travel route dynamic generation module is used for generating a personalized touring route in combination with the user portrait, the real-time environment and the resource state; according to the digital cultural tourism management system based on multi-source data analysis disclosed by the invention, the utilization rate of cultural resources is greatly improved, and the decision response speed is increased; the cultural cognition depth of tourists is improved, and the residence time of the tourists is prolonged; the damage rate of high-sensitivity cultural relics is reduced; and the special group culture acquisition efficiency is improved in a breakthrough manner.
Owner:HAINAN VOCATIONAL COLLEGE OF SCI & TECH

Multimodal interaction method, apparatus, controller, system, automobile, and storage medium

The application discloses a multimodal interaction method, device, controller, system, automobile and storage medium. The method comprises the following steps: determining a current interaction dialogue according to a current interaction voice at an interaction time; determining a current scene image and current scene data corresponding to a current interaction interface corresponding to the interaction time; adopting a multimodal recognition model to perform multimodal recognition on the current interaction dialogue, the current scene image and the current scene data, and determining a target control instruction; and executing the target control instruction to complete a human-computer interaction operation. The method can guarantee the output efficiency and accuracy of the target control instruction, reduce the complexity of customizing templates, rules and associations and other complex control logics in the development process, and improve the adaptability and generalization ability of voice interaction.
Owner:BYD CO LTD

Robot multi-mode interaction control method and related device

The invention discloses a robot multi-mode interaction control method and a related device, and the method comprises the steps that a robot controller obtains user interaction information, and the user interaction information comprises user voice information and / or user action information; analyzing the user interaction information, and determining a shopping guide service type and a target vehicle part; acquiring first component information of the target vehicle component from a preset vehicle knowledge base according to the shopping guide service type; according to the shopping guide service type and the first component information, a target action sequence is determined from a preset action library, the target action sequence comprises joint actions and voice actions, the voice actions are used for outputting voice shopping guide information, and the voice shopping guide information is determined according to the first component information; and sending a first control instruction to the target robot, wherein the first control instruction is used for indicating the target robot to execute the target action sequence. According to the invention, the functionality and intelligence of the humanoid robot as an intelligent shopping guide robot can be improved.
Owner:SHANGHAI FOURIER INTELLIGENCE CO LTD

Multi-modal large model-based AI recruitment candidate comprehensive ability evaluation method

The invention discloses an AI recruitment candidate comprehensive ability evaluation method based on a multi-modal large model, and relates to the technical field of artificial intelligence recruitment and multi-modal data processing. The method comprises the following steps: S1, constructing a post digital twinborn environment; S2, collecting candidate multi-modal interaction data; S3, training and optimizing a multi-modal large model; and S4, evaluating the comprehensive ability of the candidate and outputting a result. According to the method, a high-simulation post digital twinborn environment is constructed, multi-modal interaction data acquisition is combined, and deep analysis and evaluation are performed by using a multi-modal large model, so that the method not only considers the traditional resume and interview performance of candidates, but also improves the evaluation efficiency by simulating a real working scene. The method comprehensively evaluates the actual operation capability, the communication cooperation capability, the problem solving capability and other multi-dimensional capabilities of the candidates, and compared with a traditional recruitment mode, the method can more accurately predict the scene adaptation speed and the comprehensive capability performance of the candidates after the candidates enter the job.
Owner:SHANGHAI DAOAN INFORMATION TECHNOLOGY CO LTD

Multi-modal interaction control method and system of multifunctional teaching assisting robot and robot

The invention relates to the technical field of intelligent education and Internet of Things fusion, in particular to a multi-modal interaction control method and system of a multifunctional teaching-assistant robot and the robot. According to the system, a tablet AI processor runs an Android system as a core, a display module, a man-machine interaction module, an audio input module, an audio output module, an image acquisition module, a temperature and humidity sensor module, a WIFI or Bluetooth module, an Ethernet module and an Internet of Things module are connected, and the Internet of Things module supports RS485 wired access and Zigbee wireless access at the same time; the tablet AI processor executes voice interaction, video call, face recognition, environment monitoring, network interaction and equipment linkage, and provides a face recognition alignment and depth feature matching algorithm and an annular microphone array beam forming and sound source direction estimation algorithm, so that multi-modal interaction and multi-equipment linkage in a teaching scene are realized. The convenience and the safety are improved; and the equipment access cost is reduced.
Owner:SHENZHEN YUXIN DIGITAL TECH CO LTD

Interactive museum display guiding explanation system based on intelligence

The invention relates to the technical field of museum intelligent service, and discloses an intelligent-based interactive museum display guide explanation system, which comprises an interactive perception module, a user portrait construction module, a content matching module, a path planning module and a feedback optimization module. The interactive perception module processes user behaviors and display environment data and constructs a multi-modal interactive perception model; the user portrait construction module integrates the user basic information and the behavior preference data to form a user feature portrait set; the content matching module combines the multi-modal interactive perception model and the user feature portrait to generate a personalized explanation content candidate sequence; the path planning module optimizes and determines an optimal guiding path; and the feedback optimization module evaluates the explanation effect and dynamically adjusts the content. According to the system, personalized explanation and intelligent guidance are realized, the interactivity and accuracy of museum service are improved, and the visiting experience of the user is optimized.
Owner:SANMENXIA GUO STATE MUSEUM

Automatic generation method of rejection defense document based on multi-agent collaboration

The invention discloses a method for automatically generating a rejection defense document based on multi-agent collaboration, and belongs to the technical field of artificial intelligence. The method comprises the steps of collecting and preprocessing original multi-modal interaction data related to a payment refusing case, processing the original multi-modal interaction data into a text form to obtain an original interaction text, and storing the original interaction text in a database; performing cleaning processing on the original interaction text in the database by adopting Non-Agent to obtain semantic intermediate representation with consistent format; and based on the structured input, generating an anti-distinguishing reason by a multi-agent collaborative anti-distinguishing framework to obtain an anti-distinguishing document of the current payment refusing case. According to the invention, based on structured layering, responsibility constraint and parallel scheduling, the key defects of'hallusion caused by input noise ', 'single LLM responsibility overload' and'end-to-end time delay 'in the prior art are overcome together, so that the automatic defense signal generation system realized based on the method is remarkably improved in the aspects of compliance and fact accuracy; and practical advantages are embodied in business generalizability and engineering efficiency.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Interactive Page System

A web page embeds an AI-driven widget that converts static content into an interactive experience. Executable code builds a page interaction index from the page's DOM, including text, selectors, and positional metrics for DOM nodes. In response to natural-language input, pipelines perform summarization, stepwise explanations, voice-guided form completion with rule-based validation, and on-page product scanning to create a dynamic, user-tunable comparison table. Results are rendered as in-place overlays with interactive back-references that highlight source DOM nodes in the viewport. Speech recognition and text-to-speech enable multimodal interaction. Optional privacy gating redacts sensitive data or routes processing to local models. The system improves webpage usability by binding AI outputs to precise DOM regions and providing unified, context-aware assistance within the page.
Owner:BOLOURI RAMIN

End-to-end 3D target detection method and system based on multi-modal feature fusion

The invention relates to an end-to-end 3D target detection method and system based on multi-modal feature fusion. The method comprises the following steps: acquiring environment image data and environment point cloud data, and generating a position code; respectively carrying out image mapping and point cloud mapping on the 3D reference points based on the transformation matrix to generate query embedding; in a Transform decoder, information fusion is carried out through a self-attention mechanism and a trans-attention mechanism; carrying out image-to-point cloud fusion and point cloud-to-image fusion on the basis of the multi-modal interaction features, realizing feature bidirectional interaction, and obtaining fusion features; and multi-modal consistency loss is introduced, a total loss function is constructed and obtained, and 3D target detection is completed by using a target detection model. A multi-layer perceptron is used for generating position codes, so that implicit alignment of images and point cloud features can be realized; cascade bidirectional fusion is introduced, interaction between point cloud features and image features is enhanced, the feature expression ability is improved, and the precision of 3D target detection is improved.
Owner:NINGBO PRESCHOOL TEACHERS COLLEGE

Multi-modal interaction method and system for cluster robot control

The invention belongs to the technical field of robot interaction, and particularly relates to a multi-mode interaction method and system for cluster robot control. Firstly, eye movement data of an operator is obtained, eye movement features of the operator are extracted from the eye movement data, the extracted eye movement features are input into a trained intention recognition model, and the decision intention of the operator for the cluster robot is obtained; generating a control instruction of the robot according to the obtained decision intention, and enabling the robot to execute the control instruction; and in the process of executing the control instruction by the robot, acquiring a hand image of an operator, detecting key point positions of a hand in the hand image, identifying a gesture of the operator according to the detected key point positions, adjusting the control instruction according to a gesture command represented by the identified gesture, and enabling the robot to execute the adjusted control instruction. According to the method, the intuition and convenience of eye movement control are kept, and the accuracy and efficiency of gestures are fully utilized, so that an operator can execute complex tasks more efficiently.
Owner:ZHENGZHOU UNIV

Multi-mode fused immersive interaction system for virtual exhibition hall

ActiveCN121879587AResolve semantic ambiguity issuesImprove understanding accuracyInput/output for user-computer interactionImage data processingDefuzzificationComputer graphics (images)
The invention relates to the technical field of virtual reality interaction, and particularly discloses a multi-modal fused virtual exhibition hall immersive interaction system, which comprises the following steps of: acquiring real-time multi-modal interaction data of a user and performing synchronous preprocessing; inputting the orientation semantic words into an interval type-2 fuzzy logic system to generate a three-dimensional space semantic membership field; according to the gesture pointing vector and the hand shaking amplitude, a pointing conical area is constructed, two-dimensional Gaussian distribution is established in a layered mode, and a gesture pointing probability field is generated; taking a semantic membership field, a gesture pointing probability field and a gaze point Gaussian kernel density field as independent evidences, introducing a wall boundary and a floor channel as spatial topology constraints, and fusing by adopting an evidence theory combination rule to generate a three-dimensional intention probability distribution field; carrying out gravity center defuzzification processing on the distribution field to extract a navigation target area, and planning a navigation path by taking the current position of the user as a starting point and taking a nearest floor channel entrance as a passing point; according to the method, the problems of multi-mode fuzzy intention understanding and space adaptation are solved.
Owner:XINZHIHANG MEDIA TECH GRP CO LTD

Intelligent conference multi-modal interaction optimization method and system based on large model

The invention provides an intelligent conference multi-modal interaction optimization method and system based on a large model, and the method comprises the steps: obtaining conference multi-modal data, and constructing a locked image frame set based on an image content snapshot locking mechanism; inputting each piece of modal data of the multi-modal data stream into a pre-trained large model to generate embedded vectors corresponding to the modal data, performing anchor point extraction by using similarity calculation between each pair of embedded vectors, and screening through the large model to obtain a semantic anchor point set; when the conference is carried out, each new speech is converted into an embedded vector, and then the similarity between the embedded vector and each semantic anchor point is calculated to obtain an anchor point reference frequency set; and according to the anchor point reference frequency set and the anchor point state set, generating a real-time interaction suggestion, and based on the interaction suggestion, feeding back an updated anchor point state and an updated anchor point priority, and updating the interaction suggestion. According to the invention, automatic identification, structured representation and process state perception of the conference content are realized.
Owner:GUANGZHOU HUIYI INFORMATION TECHNOLOGY CO LTD

Multi-modal interaction method based on intelligent networked automobile vehicle-mounted interaction practical training platform and practical training platform

The invention provides a multi-modal interaction method based on an intelligent networked automobile vehicle-mounted interaction practical training platform and a practical training platform, and the method comprises the steps: responding to an interaction instruction of an automobile vehicle-mounted human-computer interaction interface, and obtaining multi-modal data based on the interaction instruction; constructing user interaction space coordinates based on the voice interaction data and the gesture interaction data; performing interference analysis and screening on the voice interaction data, the gesture interaction data and the touch interaction data based on the vehicle driving state to obtain target interaction data; fusing the target interaction data based on the user interaction space coordinates to obtain fused feature data; and classifying and identifying the fused feature data to obtain a target interaction instruction, and triggering the controlled object to execute an instruction operation based on the target interaction instruction. According to the method, the interference among interaction modes such as voice, gestures and touch in a complex real scene is reduced, the condition of signal confusion is reduced, and the problem of high identification difficulty caused by different user characteristics during simultaneous practical training of multiple persons is solved.
Owner:广东合赢教育科技股份有限公司

Barrier-free intelligent medical guide method and platform based on multi-modal interaction

The invention provides a barrier-free intelligent medical guide method and platform based on multi-modal interaction, and relates to the technical field of multi-modal interaction, and the method comprises the steps: obtaining user time sequence state data according to a bimodal state collection unit, carrying out the detection of a medical guide support state, and outputting a target medical guide support mode; calling a target modal medical guide interaction container; traversing a medical resource navigation map to position an initial interaction node by adopting medical guide demand data input by a user; the interaction response reply is converted into a modal inquiry response to carry out user initial demand intention reply; according to the update feedback demand of the user, executing multiple rounds of demand intention analysis, and outputting a department navigation suggestion label set; and packaging the demand interaction record and the department navigation suggestion label set into a medical guide demand alarm and sending the medical guide demand alarm to the target auxiliary medical guide. The technical problem that the user experience and the service efficiency are affected due to the fact that the prior art generally depends on single-mode interaction and cannot provide comprehensive support for special groups is solved.
Owner:BEIJING NANSHI INFORMATION TECH CO LTD

Intelligent exhibition hall multi-mode interactive digital human system and implementation method

The invention relates to the technical field of computer data processing, and discloses an intelligent exhibition hall multi-modal interaction digital human system and an implementation method, which are used for solving the problem that multi-source data of multi-modal interaction lacks a unified clock and verifiable timestamp alignment mechanism in a traditional method. According to the method, a unified time domain and a time version are established on an edge side, terminal access is restrained, and an acquisition time mark, an access time mark and a serial number are written in data or a state; performing gating shunting according to a time version, performing de-duplication and out-of-order rearrangement based on a serial number and double time marks, and generating and solidifying a session window evidence index; locking a transaction window boundary according to the evidence index, generating a participation source list and a transaction number, establishing a fragment reference relationship and generating an alignment voucher; in the linkage stage, phase division issuing is carried out according to a preparation phase, an execution phase and a confirmation phase, backward reading verification is carried out according to an action sequence number, a transaction log is archived, and alignment, rechecking and playback of an interaction link are achieved.
Owner:SUZHOU CHUANGJIE MEDIA EXHIBITION CO LTD

Video anomaly detection method based on anomaly feature enhancement of multi-modal interaction

The invention discloses a video anomaly detection method based on abnormal feature enhancement of multi-modal interaction, and the method comprises the following steps: S1, obtaining a monitoring region video, a public region video and a pre-annotation anomaly data set, and taking the videos as input video data; s2, processing the input video data; S21, segmenting the input video data into video frames; s22, obtaining a language text description of the video according to the input video data, and taking the language text description as subsequent text feature information; s3, processing the visual feature information and the text feature information into features of the same dimension, and inputting the features into CLIP for feature space modal matching; s4, abnormal visual features are extracted twice through CLIP, the obtained features have higher enhancement and key anomalies, abnormal information in the video is extracted in a visual and language mode interaction mode, more attention is paid to key abnormal features during model detection, and the detection accuracy is improved. And enhanced and comprehensive key anomalies are obtained through secondary extraction.
Owner:GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY

A Low-Light Scene Analysis Method Based on Multimodal Feature Fusion and Clustering

This invention discloses a low-light scene analysis method based on multimodal feature fusion and clustering, belonging to artificial intelligence technology. It constructs a single-branch feature extraction network based on Transformer; a multimodal feature interaction and fusion module to achieve multimodal feature interaction and fusion; a multimodal fusion feature clustering module to cluster the features after multimodal interaction and fusion, utilizing the semantic and spatial distances between different features to achieve clustering while simultaneously downsampling the features, mitigating the problem of loss of detailed edge information in structured downsampling; and a multi-scale feature aggregation and decoding module to receive feature information from the encoding network and classify each feature pixel according to the semantic distance of the multi-scale features. This invention can fully utilize visible light and thermal image information and can be applied to scene analysis and navigation of unmanned systems in diverse and complex low-light scenes.
Owner:CHINA UNIV OF MINING & TECH

Interaction story machine based on ROS and large model and control method thereof

The invention discloses an interactive story machine based on an ROS and a large model and a control method of the interactive story machine. The story machine comprises an input module, an ROS system, a large model module, an audio output unit and a color ink screen. According to the invention, through cooperative work of software and hardware modules, an ROS system is used as an intelligent scheduling center, and meanwhile, a large model module is introduced, so that user interaction input modes and personalized demands of different user types can be understood and responded, and personalized story contents are dynamically generated; for the hearing-impaired user type, based on a barrier-free interaction mode, generating image prompt information and a smooth sign language action sequence, and driving a color ink screen to perform visual presentation through an optimized rendering instruction; and for a non-hearing-impaired user type, synchronously generating a story text fragment, a matched image illustration and voice information based on a multi-mode interaction mode, and providing an immersive story reading environment for different types of child users.
Owner:GUANGDONG UNIV OF TECH

A method and system for matching the dialogue intent of intelligent NPCs in multimodal interaction

This invention relates to the field of natural language processing technology, specifically to a method and system for matching the intent of intelligent NPC dialogues in multimodal interaction. The method involves real-time acquisition of multimodal data, generating multimodal semantic features through submodal preprocessing; integrating the semantic features of each modality using a cross-modal fusion module based on semantic association, generating a candidate intent set based on semantic context representation and combined with a semantic parsing module and predefined intent templates; tracking changes in user intent in real time through a continuous semantic learning mechanism and an interactive memory module, combined with a Bayesian update method, dynamically adjusting the confidence level of each intent in the candidate intent set, and filtering the final intent; constructing an NPC semantic cognition model, and performing semantic consistency analysis to perform semantic checks on user input and NPC dialogue state; combining the final intent and a decision engine to generate a dialogue strategy and output synchronized response content; this invention improves the accuracy of dynamic matching of intelligent dialogue intents.
Owner:JIANGSU COLDPLAY INFORMATION TECH CO LTD

Toys (AI Smart Plush Toys)

1. Name of the product in this design: Toy (AI Intelligent Plush Toy). 2. Purpose of this design: Intelligent plush toys for AI-powered multimodal interaction, emotional companionship, and educational purposes. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: a 3D model.
Owner:TONGDA SMART TECH (XIAMEN) CO LTD