Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

36 results about "Visual memory" patented technology

Visual memory describes the relationship between perceptual processing and the encoding, storage and retrieval of the resulting neural representations. Visual memory occurs over a broad time range spanning from eye movements to years in order to visually navigate to a previously visited location. Visual memory is a form of memory which preserves some characteristics of our senses pertaining to visual experience. We are able to place in memory visual information which resembles objects, places, animals or people in a mental image. The experience of visual memory is also referred to as the mind's eye through which we can retrieve from our memory a mental image of original objects, places, animals or people. Visual memory is one of several cognitive systems, which are all interconnected parts that combine to form the human memory. Types of palinopsia, the persistence or recurrence of a visual image after the stimulus has been removed, is a dysfunction of visual memory.

Multi-class anomaly detection method and system based on pre-training visual language model in training data scarcity scene

The invention provides a multi-class anomaly detection method and system based on a pre-training visual language model in a training data scarcity scene, and relates to the technical field of anomaly detection. A first pre-training visual language model is used for obtaining feature representation of a small number of normal sample images in a text space; using a second pre-training visual language model to obtain global features and block features of a small number of normal sample images, constructing and training an adaptive prompt vector generator, in the training process, updating parameters of the adaptive prompt vector generator to obtain a trained adaptive prompt vector generator, and obtaining a training result of the adaptive prompt vector generator; according to the method, the ability of a pre-training visual language model is effectively combined, a training prompt vector generator is guided, a self-adaptive prompt vector for an anomaly detection task is generated, a one-to-many visual memory warehouse and a one-to-many prompt vector warehouse are constructed, a one-to-many training normal form is adapted, and the detection performance of anomaly detection in a training data scarcity scene is improved.
Owner:SUN YAT SEN UNIV

Instruction aware memory device for video understanding

The invention provides an instruction perception memory device for video understanding. The instruction perception memory device comprises a text-visual memory library module and a cross attention module, the text-visual memory library module is used for storing and retrieving cross-modal features and supporting video analysis, the text-visual memory library module is integrated with a multi-modal large language model, and video data are processed in an incremental mode, so that the limitation of a memory and a context length is overcome; and the cross attention module is used for fusing the text and the visual features and generating cross-modal representation. By introducing a text-visual memory library and a cross attention module, early fusion and long-term memory management of video and text information are realized. The fine-grained time dependency relationship in the video can be effectively captured, and the performance of the model in a long video understanding task is improved, so that the aim of improving the accuracy and efficiency of video understanding is fulfilled.
Owner:LANZHOU UNIV

Target tracking method and system based on traffic scene

The invention discloses a target tracking method and system based on a traffic scene, and relates to the field of intelligent traffic management.The method comprises the steps that target tracking is achieved through a denoising learning process in a diffusion model, and visual memory information and search area information are used as condition input; according to the method, the real position of the target is predicted by introducing a noise bounding box corresponding to the target, the de-noising process is decomposed into a plurality of de-noising blocks in order to improve the model prediction efficiency and realize real-time tracking, each de-noising block realizes one de-noising process, and finally accurate target position prediction is realized.
Owner:INTELLIGENT INTER CONNECTION TECH CO LTD

Unmanned aerial vehicle visual language navigation method based on multi-mode historical information fusion

The invention provides an unmanned aerial vehicle visual language navigation method based on multi-mode historical information fusion. The method comprises the following steps: acquiring an image frame sequence acquired by an unmanned aerial vehicle before a current moment, a trajectory point sequence of flight of the unmanned aerial vehicle, a task language instruction and an observation value of the unmanned aerial vehicle at the current moment; processing the image frame sequence by utilizing visual Slot aggregation to obtain a visual memory feature sequence; processing the flight trajectory point sequence of the unmanned aerial vehicle to obtain trajectory embedding features; processing the task language instruction by using a Bert language encoder to obtain language semantic features; processing the language semantic feature, the visual memory feature sequence, the track embedding feature and the observation value at the current moment to obtain a fusion feature; and processing the fused features by using a large language model to obtain navigation decision information. According to the embodiment of the invention, the historical image, the track point and the task language instruction are fused, so that the context understanding capability of the unmanned aerial vehicle and the robustness of the navigation decision are remarkably improved.
Owner:BEIJING INST OF TECH

Real-time non-graphical autonomous navigation system for robot

The invention relates to a real-time non-graphical robot autonomous navigation system, and belongs to the field of robots. The system comprises a main control module, an environment sensing module and a hybrid navigation module, the main control module is responsible for man-machine interaction, task flow control and cooperation of other modules, and storing and managing a visual semantic database; the environment sensing module is responsible for controlling the robot to collect image data and depth data under the condition of no prior map, identifying objects in a scene, extracting semantic information of the objects, and integrating identification results into a visual semantic database; and the hybrid navigation module analyzes the instruction sent by the main control module by using a large language model, matches visual memory, determines a navigation target, and completes autonomous navigation of the robot through a hybrid navigation strategy. According to the method, the problem of low navigation task accuracy caused by the sparsity of global information in map-free navigation is solved.
Owner:GUANGDONG UNIV OF TECH

SAM2 small sample segmentation method based on semantic-visual dual-memory fusion

The invention discloses an SAM2 small sample segmentation method based on semantic-visual dual-memory fusion, and the method comprises the steps: constructing semantic query memory, visual query memory and query-related support visual memory, and fusing the semantic query memory and the visual query memory through a memory refinement module guided by the query-related support visual memory. SAM2 dense matching and decoding module end-to-end training are combined. The problems of single memory and foreground-background confusion of an existing method are solved, target semantic consistency and fine-grained modeling ability are enhanced, segmentation precision and generalization ability are remarkably improved in a complex scene, and the method is suitable for the fields of medical image analysis, automatic driving and the like.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Test system for motor couriers

The invention is a test simulation developed to assess whether motorcycle courier workers are psychologically and psychomotorally competent to use motorcycles and do it as a job, wherein it comprises the process steps of: displaying test applications to the candidate via a screen, virtual reality glasses or similar, the test applications being prepared to prevent people from working as couriers without a license, to introduce a selectivity for becoming a motorcycle couriers and to determine whether individuals are suitable for this profession; the candidate performing the required actions within the time given for each test application and giving the candidate points according to the predetermined scoring system in accordance with the evaluation of the tested skill; evaluating the candidate's attention level, selective attention, motor speed, visual motor level, conceptual scanning, complex attention level, executive functions, visual memory, reaction speed, impulsivity and reflex characteristics, in line with the candidate's score from the test applications.
Owner:PSIKOTEKNO YAZILIM TICARET ANONIM SIRKETI

Organic compounds

The disclosure relates to the use of phosphodiesterase 1 (PDE1) inhibitors for preventing or treating chemobrain, e.g., inhibiting or mitigating one or more symptoms of chemobrain, e.g., feeling of mental fogginess, deficit in attention, disorganized behavior or thinking, confusion, impaired concentration, difficulty finding the right word, difficulty learning new skills, difficulty multitasking, memory impairment (e.g., short-term memory problems, long-term memory problems, impaired verbal memory, impaired visual memory), and / or fatigue.
Owner:INTRA CELLULAR THERAPIES INC

Driver target attention prediction method, system, device, medium and product

The invention discloses a driver target attention method, system and device, a medium and a product, and relates to the field of target detection, and the method comprises the steps: constructing a driver target attention prediction model; the driver target attention prediction model comprises a target recognition model, a driving event classification model and an attention prediction model; the attention prediction model comprises an environment model fusing multi-layer visual memory and a target-level visual attention prediction model; obtaining visual behavior data and driving data of a driver in real time; and inputting the driver visual behavior data and the driving data into the driver target attention prediction model to obtain a visual attention prediction target. According to the method, the adaptive capacity of the driver visual target attention prediction model can be improved, the accuracy and robustness of model prediction are improved, the prediction result better fits the visual search mechanism of the driver, and the practicability of the driver visual target attention prediction model is improved.
Owner:BEIJING INST OF TECH

Multi-scene-oriented security robot intelligent inspection method and system

The invention discloses a multi-scene-oriented security robot intelligent inspection method and system, and relates to the field of intelligent security, and the method comprises the steps: collecting machine body vibration data generated by interaction between a robot and the ground in an environment with good visibility, and building a standard fingerprint band and a historical image; in a low-visibility environment, the robot collects a fuzzy visual image and vibration data in real time, generates a current inspection fingerprint and matches the current inspection fingerprint with a standard fingerprint band; and when deviation is detected, calling a visual memory library and determining the closest reference landmark by utilizing a deep learning visual matching network, so that the steering angle of the robot is adjusted, and the robot is enabled to be close to the central path of the fingerprint zone again. The method can be applied to industrial control software, is deployed on security robots of multiple models to realize process control of inspection, solves the problem that the security robots are unstable in inspection positioning in a low-visibility environment, and realizes path self-correction and intelligent inspection control based on vibration fingerprints and depth vision matching.
Owner:ANLIZHI INTELLIGENT ROBOT TECH (BEIJING) CO LTD

A method for assessing the risk of Alzheimer's disease based on visual memory assessment

The present invention discloses a method for assessing the risk of Alzheimer's disease based on visual memory evaluation. Step 1: Collect eye movement images of Alzheimer's disease patient groups and normal people. Step 2: Generate a fixation heat map as a test picture for each eye movement image to generate a sample set. Step 3: Use a convolutional neural network to extract the features of the fixation heat map to form a test sample feature vector. Step 4: Obtain the common internal feature representation of each test picture. Step 5: Extract the common external modality representation. Step 6: Use the common feature representation for binary classification, that is, use the finally extracted common feature representation as the input of a fully connected network to perform binary classification processing on the eye movement data to be tested corresponding to Alzheimer's disease patients and normal people respectively. Compared with the prior art, the eye movement data features used in the present invention are more objective and convenient, and can effectively improve the accuracy of model detection and solve the problems of clinical data loss and insufficiency.
Owner:SHANGHAI WULIDUO TECH CO LTD

Method and system for analyzing ad characteristic information based on brain medical image

The application discloses an AD characteristic information analysis method based on brain medical images, which comprises the following steps: data preprocessing is performed on brain MRI images to be analyzed to obtain standard images of brain regions; spatial characteristic analysis of the whole brain and key brain regions; whole brain image characteristic analysis; cognitive characteristic analysis of the whole brain and key brain regions; and aggregation analysis results are obtained to form corresponding explanation data. The application further discloses an AD characteristic information analysis system based on brain medical images. The method for analyzing cognitive index characteristics such as brain age, logical memory score, visual memory score and long-time delay memory score from brain MRI images by using deep learning technology can provide analysis results of brain MRI images and other important medical indexes for doctors, and can provide effective arguments and explanations for subsequent scientific research analysis and interpretation.
Owner:SHANGHAI TONGJI HOSPITAL +1

Method and system for multi-modal collaborative generation of visual memory compression and privacy protection

The present application relates to the technical field of privacy protection, in particular to a kind of visual memory compression and privacy protection multi-modal collaborative generation method and system.The method comprises: the visual memory sample is divided into recent memory layer, intermediate memory layer and long-term memory layer, generates each clustering cluster after weighting clustering, extracts the weighted average feature of each clustering cluster and generates cluster prototype vector after abstract coding;Perform anonymization operation, inject noise disturbance into cluster prototype vector in combination with differentially private budget decaying with time, generate abstract memory representation after privacy enhancement;Receive task instruction through cloud big model, generate structural sketch, generate personalized response through local small model;Calculate importance score for each memory unit in the abstract memory representation after privacy enhancement, perform retention or elimination decision for the abstract memory representation after privacy enhancement, update abstract preference memory bank;It can maintain multi-modal collaborative generation continuity, reduce overhead, enhance security.
Owner:DUKE KUNSHAN UNIVERSITY

Logistics goods sorting method, device and system based on computer vision

The invention discloses a logistics cargo sorting method based on computer vision, and relates to the field of logistics, and the method comprises the following steps: obtaining multi-modal visual data of cargoes; obtaining a multi-modal feature vector set based on the multi-modal visual data; inputting the multi-modal feature vector set into a preset adaptive visual memory network to output a goods identification result and a corresponding identification confidence; and when the recognition confidence coefficient is larger than or equal to a preset threshold value, fusion features are generated based on the multi-modal feature vector set, and a three-dimensional space model of the goods is constructed through the fusion features so that the three-dimensional space model can be used for path planning of a mechanical arm. According to the logistics goods sorting method based on the computer vision, the multi-mode visual data of the goods are obtained, the self-adaptive visual memory network is matched, high-precision goods recognition can be achieved, and the sorting error rate is remarkably reduced; and meanwhile, an accurate three-dimensional space model is provided, and the grabbing efficiency and safety of the mechanical arm are improved.
Owner:HEFEI XINCHUANG ZHONGYUAN INFORMATION TECH CO LTD

Emergency evacuation design and rescue method for large-scale comprehensive building

ActiveCN122241161BSpatial structureFire house
The application discloses an emergency evacuation design and rescue method for large-scale comprehensive buildings, and particularly relates to the field of emergency evacuation and rescue design, and is used for solving the problem that the rescue path planning and people flow guidance are disconnected due to the complex internal space structure of the existing comprehensive buildings and the lack of visual cognitive support in traditional evacuation design; high-frequency stay areas are identified by collecting the crowd moving tracks of multiple comprehensive buildings, consensus space memory features are extracted and a space memory anchor network is constructed, evacuation path nodes are labeled in the network and weak visual connection sections are identified, new memory anchor points are formed by arranging micro fire stations, the guide utility coefficients are calculated in combination with fire drill data and are converted into dynamic conflict weights, when a fire event occurs, rescue path search is performed in the anchor network with the target of minimizing the cumulative dynamic conflict weights, and the evacuation guidance optimization based on visual memory features and the dynamic cooperative design of the rescue path are realized.
Owner:TIANJIN FIRE SCI & TECH RES INST OF MEM

Visual memory in autonomous agents

Embodiments described herein provide a method for an autonomous agent (which may be a contextualized, homed agent) to learn new visual objectives, and identify the same objective in the same "runtime" / session / interaction, and can train a long-term visual memory model on the objective. The visual memory system stores templates of the target and compares the images to these templates. Internal or external (to proxy) signals may determine target (s) to compare. The visual memory system takes new and unidentified targets as "templates" and returns whether any region of the image may include a match to any stored template. Single sample visual learning is thus provided.
Owner:SOUL MACHINES LTD

A method for extending long video understanding capability of a multi-modal large language model through a visual memory mechanism

This invention discloses a method for extending the long-video understanding capability of a multimodal large language model through a visual memory mechanism. The method includes: segmenting the long video into consecutive segments and initializing a visual memory bank; iteratively encoding each segment and generating contextual memory and local memory for each segment using a dual-path compression mechanism, where the former is used to convey historical information and the latter is stored in the memory bank; retrieving relevant local memory segments from the memory bank according to textual instructions; and inputting the retrieved segments and instructions into the multimodal large language model to generate an answer. This invention effectively extends the model's long-video understanding capability without increasing the GPU memory burden, balancing temporal coherence and local details, and achieving fast and flexible memory retrieval.
Owner:XIAMEN UNIV

EYE TRACKING BASED BLOCK DESIGN TESTING SYSTEM FOR VISUAL ATTENTION AND VISUAL MEMORY EVALUATION

This invention relates to the field of cognitive testing techniques and sensor-based measurement systems to create a block design test-based visual attention and visual memory assessment system that can be used in various age groups. This system integrates testing media in the form of physical blocks, digital simulations, or a combination of both, as well as an eye-tracking data acquisition module to record gaze direction, eye movements, and attention duration in real-time during the process of observing and arranging blocks. The eye-tracking data obtained is then processed to analyze visual attention patterns, focus shifts, and their relationship to the process of remembering and arranging patterns that represent the user's visual memory ability. Unlike conventional methods that rely on subjective observations by the examiner, this system is capable of conducting objective, measurable, and data-based assessments of visual behavior.Furthermore, this system enables more efficient testing by storing results digitally for monitoring, evaluation, and research. This invention is intended to support a variety of applications, from monitoring cognitive development in children, evaluating cognitive performance in adults, to early detection of cognitive decline in the elderly.
Owner:ELECTRONIC ENGINEERING POLYTECHNIC INSTITUTE OF SURABAYA

Task understanding method and device for images and model training method and device

The embodiment of the invention provides a task understanding method and device for an image and a model training method and device. The task understanding method comprises the following steps: extracting a plurality of fine-grained features of a group of images to be processed, and storing the fine-grained features in a visual memory library; meanwhile, performing global visual coding on each image in a group of images, and mapping global visual features to a feature space of the large model through a pre-trained visual language adapter to obtain initial global understanding; thirdly, inputting the initial global understanding and the user task instruction into a large model, determining to-be-reviewed contents through the large model, and retrieving and extracting first fine-grained features related to the to-be-reviewed contents from a visual memory library through a reviewing module; in this way, the first fine-grained feature can be input into the large model, and the understanding content based on the multiple images and aiming at the user task instruction is output through the large model. And when the image and the user task instruction contain privacy data, privacy protection processing needs to be performed on the privacy data in the processing process.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD +1

A target robust detection method based on a brain cognitive model

The present invention discloses a target robust detection method based on a brain cognitive model, which relates to the technical field of target detection. First, paired real images and standard images are input; secondly, a primary perception feature is extracted from the input real image by an image feature extraction module that simulates the visual information attention perception function; then, a target perception feature is extracted from the primary perception feature by an image feature extraction module that simulates the visual spatial attention function; subsequently, a target memory feature is extracted from the standard image by a standard visual memory bank module that simulates the visual memory function; then, the target perception feature and the target memory feature are input into a perception-standard feature association and fusion module that simulates the visual and memory association function, and using domain adaptation technology, the target perception feature and the target memory feature are interactively fused and learned to form a fusion feature; finally, the fusion feature is input into a prediction and inference module that simulates the visual prediction and inference function, and the position prediction of the target is completed using Kalman filtering.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

AI-based endoscope image tumor marking method

The invention discloses an AI-based endoscopic image tumor labeling method, which relates to the technical field of medical equipment, and comprises the following steps: during use, a white light acquisition module and a fluorescence acquisition module perform image acquisition at the same time, at the same place and at the same focal length; the image storage module stores white light image information collected by the white light collection module and fluorescence image information collected by the fluorescence collection module, a lightweight target recognition algorithm of channel convolution and space convolution separation is constructed based on a Mask R-CNN algorithm, ultra-high-definition fluorescence image data is rapidly detected and recognized, and the tumor tissue boundary is accurately recognized. According to the invention, the error caused by the fact that a doctor observes a fluorescent endoscope image and a white light endoscope image for contrast reference in a traditional scheme to match a tumor range by means of visual memory is solved, and real-time guidance and accurate excision of tumor excision are realized.
Owner:HEFEI DVL ELECTRON CO LTD

Visual memory puzzle device

ActiveCN224207350USmall footprintAvoid falling and losingIndoor gamesSoftware engineeringMetal sheet
The utility model discloses a visual memory puzzle device, which relates to the technical field of educational toys, and comprises a box body, a box cover, a magnetic plate and a plurality of puzzle blocks, the plurality of puzzle blocks are placed on the inner side of the box body, ferromagnetic metal sheets are arranged in the puzzle blocks, the box cover is rotatably connected to the upper end of the box body in a damping manner, and a buckle is connected between the box body and the box cover. And the magnetic suction plate is mounted on the box cover. Through the matching effect of the jigsaw blocks with the built-in ferromagnetic metal sheets and the magnetic attraction plates arranged on the box cover, after the box cover is opened, a plurality of jigsaw blocks can be spread on the inner side of the box body, the needed jigsaw blocks can be found, the jigsaw blocks are attracted to the inner side of the box cover one by one, the occupied area of the jigsaw is reduced, and meanwhile the jigsaw puzzle is convenient to use. The jigsaw blocks can be effectively prevented from falling off and being lost, after jigsaw is completed, the magnetic attraction plate can be pulled out through the handle, at the moment, the jigsaw blocks originally attracted to the inner side of the box cover through the magnetic attraction plate lose attraction force and automatically fall into the box body, and manual taking down one by one is not needed.
Owner:YIBIN XITONG HEALTH TECHNOLOGY CO LTD

An ai agent-based hierarchical visual memory method and system

This invention discloses a hierarchical visual memory method and system based on AI agents, belonging to the field of artificial intelligence agent technology. It addresses the technical problem that existing AI agent text memory systems cannot efficiently process continuous visual input and cannot convert visual data into contextual experience that the agent can reason about, resulting in visual data being unusable stably by the AI ​​agent. The method includes: receiving visual input data and generating corresponding visual observation objects according to preset rules, and writing the visual observation objects into a first visual memory layer; filtering and controlling the visual observation objects through promotion gating; synthesizing the filtered visual observation objects into contextual memory records, and writing the contextual memory records into a second visual memory layer; grouping the contextual memory records according to preset indicators, generating a corresponding natural language summary for each group of contextual memory records, and writing the natural language summary into a third visual memory layer.
Owner:BEIJING MIANBI INTELLIGENT TECH CO LTD

A visual memory parking method based on a deep learning network, medium, device and vehicle

The application discloses a visual memory parking method based on a deep learning network, a medium, equipment and a vehicle, and comprises the following steps: acquiring a preloaded map of a parking lot; acquiring a feature map of a target vehicle at a current time and a real vehicle pose of the target vehicle on the preloaded map at a previous time; predicting a virtual vehicle pose of the target vehicle on the preloaded map at the current time based on the real vehicle pose of the target vehicle on the preloaded map at the previous time; cutting out a local map of the target vehicle on the preloaded map based on the predicted virtual vehicle pose of the target vehicle on the preloaded map at the current time; aligning the feature map and the local map under the deep learning network, and outputting a real vehicle pose of the target vehicle in the preloaded map at the current time; and parking based on a parking path corresponding to a target parking space in the cloud memory and the selected target parking space. Through the above method, the cruise process of memory parking can be realized only by visual observation and the vehicle's own odometer, the complex road topological information and semantic lane are avoided, only the memorized route and the light semantic map need to be maintained, the storage space is reduced, the maintenance cost is reduced, and the data between platforms is easier to migrate.
Owner:SHENZHEN DEEPROUTE AI CO LTD

Sam2 small sample segmentation method based on semantic-visual dual memory fusion

The application discloses a SAM2 small sample segmentation method based on semantic-visual double memory fusion, constructs semantic query memory, visual query memory and query related support visual memory, fuses the semantic query memory and the visual query memory through a memory refinement module guided by the query related support visual memory, and combines SAM2 dense matching and decoding module end-to-end training. The method solves the single memory and foreground-background confusion problems of the existing method, enhances target semantic consistency and fine-grained modeling capability, significantly improves segmentation precision and generalization capability in a complex scene, and is suitable for medical image analysis, automatic driving and the like.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Elliptical billiards aiming training article and method

The invention discloses an elliptical billiard aiming training device, comprising a training disc and a training ring; the training disc comprises a pointer, a disc body and rivets; the elliptical billiard aiming training is mainly completed by respectively intersecting the pointer and ellipses 1 and 2 engraved on the disc body to form aiming points to strengthen visual memory; the training ring comprises a ring body and legs, is put on a target ball, and is used to verify the standing distance of a billiard player and adjust the change of the minor axis of the ellipse; the invention can easily find an exact aiming point in the pattern of the aiming training device according to the standing distance between the billiard player and the target ball and from the extension of the goal line at any angle, has a wide adaptability, high accuracy, strong operability, is simple and clear, and is easy to master, and can make up for the shortcomings of aiming methods such as the tail-finding method, the coincidence method, the symmetry method and the double coincidence method; the aiming process is safe and reliable, and has great promotion value.
Owner:刘强

Composition containing diamine and / or polyamine for preventing deterioration of brain function or improving brain function

To provide a composition for preventing deterioration of brain function or improving brain function.SOLUTION: The present invention provides a composition containing diamine and / or polyamine, wherein brain function is at least one selected from the group consisting of comprehensive memory, verbal memory, visual memory, cognitive function speed, motor speed, processing speed, comprehensive attention, cognitive flexibility, response time, and executive function.SELECTED DRAWING: None
Owner:TOYOBO CO LTD

Visual memory in autonomous agents

Embodiments described herein provide a method by which an autonomous agent (which may be a situationalized and materialized agent) learns new visual objects, recognizes those objects within the same “runtime” / session / interaction, and enables the training of a long-term visual memory model for those objects. The visual memory system stores templates of objects and compares images to these templates. The choice of which objects to compare may be determined by internal or external signals (of the agent). The visual memory system stores new and unrecognized objects as “templates” and returns whether any region of an image is likely to contain a match with any of the stored templates. In this way, one-shot visual learning is achieved.
Owner:SOUL MACHINES LTD

Multi-modal video reasoning method and system based on hierarchical multi-agent

The invention discloses a hierarchical multi-agent-based multi-modal video reasoning method and system, and belongs to the technical field of artificial intelligence and multi-modal video understanding. The method comprises the following steps: receiving an input video and a natural language query, and analyzing a complex query into a plurality of logically coherent sub-questions through a question decomposition agent; distributing the sub-questions to multi-source answer generation agents, wherein the multi-source answer generation agents comprise an answer generation agent based on Web, an answer generation agent based on time sequence memory and a video-language answer generation agent; the intelligent agent based on time sequence memory models video long-range dependence through a visual memory library and a query memory library, and adopts a cross-attention mechanism to fuse time sequence features; performing consistent voting and fusion on the multi-source answers through the decision agent to generate a final answer; and carrying out training optimization on the system based on a cross entropy loss function.
Owner:LANZHOU UNIV