Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

16 results about "Visual memory" patented technology

Visual memory describes the relationship between perceptual processing and the encoding, storage and retrieval of the resulting neural representations. Visual memory occurs over a broad time range spanning from eye movements to years in order to visually navigate to a previously visited location. Visual memory is a form of memory which preserves some characteristics of our senses pertaining to visual experience. We are able to place in memory visual information which resembles objects, places, animals or people in a mental image. The experience of visual memory is also referred to as the mind's eye through which we can retrieve from our memory a mental image of original objects, places, animals or people. Visual memory is one of several cognitive systems, which are all interconnected parts that combine to form the human memory. Types of palinopsia, the persistence or recurrence of a visual image after the stimulus has been removed, is a dysfunction of visual memory.

Unmanned aerial vehicle visual language navigation method based on multi-mode historical information fusion

The invention provides an unmanned aerial vehicle visual language navigation method based on multi-mode historical information fusion. The method comprises the following steps: acquiring an image frame sequence acquired by an unmanned aerial vehicle before a current moment, a trajectory point sequence of flight of the unmanned aerial vehicle, a task language instruction and an observation value of the unmanned aerial vehicle at the current moment; processing the image frame sequence by utilizing visual Slot aggregation to obtain a visual memory feature sequence; processing the flight trajectory point sequence of the unmanned aerial vehicle to obtain trajectory embedding features; processing the task language instruction by using a Bert language encoder to obtain language semantic features; processing the language semantic feature, the visual memory feature sequence, the track embedding feature and the observation value at the current moment to obtain a fusion feature; and processing the fused features by using a large language model to obtain navigation decision information. According to the embodiment of the invention, the historical image, the track point and the task language instruction are fused, so that the context understanding capability of the unmanned aerial vehicle and the robustness of the navigation decision are remarkably improved.
Owner:BEIJING INST OF TECH

Real-time non-graphical autonomous navigation system for robot

The invention relates to a real-time non-graphical robot autonomous navigation system, and belongs to the field of robots. The system comprises a main control module, an environment sensing module and a hybrid navigation module, the main control module is responsible for man-machine interaction, task flow control and cooperation of other modules, and storing and managing a visual semantic database; the environment sensing module is responsible for controlling the robot to collect image data and depth data under the condition of no prior map, identifying objects in a scene, extracting semantic information of the objects, and integrating identification results into a visual semantic database; and the hybrid navigation module analyzes the instruction sent by the main control module by using a large language model, matches visual memory, determines a navigation target, and completes autonomous navigation of the robot through a hybrid navigation strategy. According to the method, the problem of low navigation task accuracy caused by the sparsity of global information in map-free navigation is solved.
Owner:GUANGDONG UNIV OF TECH

SAM2 small sample segmentation method based on semantic-visual dual-memory fusion

The invention discloses an SAM2 small sample segmentation method based on semantic-visual dual-memory fusion, and the method comprises the steps: constructing semantic query memory, visual query memory and query-related support visual memory, and fusing the semantic query memory and the visual query memory through a memory refinement module guided by the query-related support visual memory. SAM2 dense matching and decoding module end-to-end training are combined. The problems of single memory and foreground-background confusion of an existing method are solved, target semantic consistency and fine-grained modeling ability are enhanced, segmentation precision and generalization ability are remarkably improved in a complex scene, and the method is suitable for the fields of medical image analysis, automatic driving and the like.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Multi-scene-oriented security robot intelligent inspection method and system

The invention discloses a multi-scene-oriented security robot intelligent inspection method and system, and relates to the field of intelligent security, and the method comprises the steps: collecting machine body vibration data generated by interaction between a robot and the ground in an environment with good visibility, and building a standard fingerprint band and a historical image; in a low-visibility environment, the robot collects a fuzzy visual image and vibration data in real time, generates a current inspection fingerprint and matches the current inspection fingerprint with a standard fingerprint band; and when deviation is detected, calling a visual memory library and determining the closest reference landmark by utilizing a deep learning visual matching network, so that the steering angle of the robot is adjusted, and the robot is enabled to be close to the central path of the fingerprint zone again. The method can be applied to industrial control software, is deployed on security robots of multiple models to realize process control of inspection, solves the problem that the security robots are unstable in inspection positioning in a low-visibility environment, and realizes path self-correction and intelligent inspection control based on vibration fingerprints and depth vision matching.
Owner:ANLIZHI INTELLIGENT ROBOT TECH (BEIJING) CO LTD

Method and system for analyzing ad characteristic information based on brain medical image

The application discloses an AD characteristic information analysis method based on brain medical images, which comprises the following steps: data preprocessing is performed on brain MRI images to be analyzed to obtain standard images of brain regions; spatial characteristic analysis of the whole brain and key brain regions; whole brain image characteristic analysis; cognitive characteristic analysis of the whole brain and key brain regions; and aggregation analysis results are obtained to form corresponding explanation data. The application further discloses an AD characteristic information analysis system based on brain medical images. The method for analyzing cognitive index characteristics such as brain age, logical memory score, visual memory score and long-time delay memory score from brain MRI images by using deep learning technology can provide analysis results of brain MRI images and other important medical indexes for doctors, and can provide effective arguments and explanations for subsequent scientific research analysis and interpretation.
Owner:SHANGHAI TONGJI HOSPITAL +1

Logistics goods sorting method, device and system based on computer vision

The invention discloses a logistics cargo sorting method based on computer vision, and relates to the field of logistics, and the method comprises the following steps: obtaining multi-modal visual data of cargoes; obtaining a multi-modal feature vector set based on the multi-modal visual data; inputting the multi-modal feature vector set into a preset adaptive visual memory network to output a goods identification result and a corresponding identification confidence; and when the recognition confidence coefficient is larger than or equal to a preset threshold value, fusion features are generated based on the multi-modal feature vector set, and a three-dimensional space model of the goods is constructed through the fusion features so that the three-dimensional space model can be used for path planning of a mechanical arm. According to the logistics goods sorting method based on the computer vision, the multi-mode visual data of the goods are obtained, the self-adaptive visual memory network is matched, high-precision goods recognition can be achieved, and the sorting error rate is remarkably reduced; and meanwhile, an accurate three-dimensional space model is provided, and the grabbing efficiency and safety of the mechanical arm are improved.
Owner:HEFEI XINCHUANG ZHONGYUAN INFORMATION TECH CO LTD

A method for extending long video understanding capability of a multi-modal large language model through a visual memory mechanism

This invention discloses a method for extending the long-video understanding capability of a multimodal large language model through a visual memory mechanism. The method includes: segmenting the long video into consecutive segments and initializing a visual memory bank; iteratively encoding each segment and generating contextual memory and local memory for each segment using a dual-path compression mechanism, where the former is used to convey historical information and the latter is stored in the memory bank; retrieving relevant local memory segments from the memory bank according to textual instructions; and inputting the retrieved segments and instructions into the multimodal large language model to generate an answer. This invention effectively extends the model's long-video understanding capability without increasing the GPU memory burden, balancing temporal coherence and local details, and achieving fast and flexible memory retrieval.
Owner:XIAMEN UNIV

EYE TRACKING BASED BLOCK DESIGN TESTING SYSTEM FOR VISUAL ATTENTION AND VISUAL MEMORY EVALUATION

PendingIDS00202607934APerceived cognitive abilitiesVisual perception
This invention relates to the field of cognitive testing techniques and sensor-based measurement systems to create a block design test-based visual attention and visual memory assessment system that can be used in various age groups. This system integrates testing media in the form of physical blocks, digital simulations, or a combination of both, as well as an eye-tracking data acquisition module to record gaze direction, eye movements, and attention duration in real-time during the process of observing and arranging blocks. The eye-tracking data obtained is then processed to analyze visual attention patterns, focus shifts, and their relationship to the process of remembering and arranging patterns that represent the user's visual memory ability. Unlike conventional methods that rely on subjective observations by the examiner, this system is capable of conducting objective, measurable, and data-based assessments of visual behavior.Furthermore, this system enables more efficient testing by storing results digitally for monitoring, evaluation, and research. This invention is intended to support a variety of applications, from monitoring cognitive development in children, evaluating cognitive performance in adults, to early detection of cognitive decline in the elderly.
Owner:ELECTRONIC ENGINEERING POLYTECHNIC INSTITUTE OF SURABAYA

Visual memory puzzle device

ActiveCN224207350USmall footprintAvoid falling and losingIndoor gamesSoftware engineeringMetal sheet
The utility model discloses a visual memory puzzle device, which relates to the technical field of educational toys, and comprises a box body, a box cover, a magnetic plate and a plurality of puzzle blocks, the plurality of puzzle blocks are placed on the inner side of the box body, ferromagnetic metal sheets are arranged in the puzzle blocks, the box cover is rotatably connected to the upper end of the box body in a damping manner, and a buckle is connected between the box body and the box cover. And the magnetic suction plate is mounted on the box cover. Through the matching effect of the jigsaw blocks with the built-in ferromagnetic metal sheets and the magnetic attraction plates arranged on the box cover, after the box cover is opened, a plurality of jigsaw blocks can be spread on the inner side of the box body, the needed jigsaw blocks can be found, the jigsaw blocks are attracted to the inner side of the box cover one by one, the occupied area of the jigsaw is reduced, and meanwhile the jigsaw puzzle is convenient to use. The jigsaw blocks can be effectively prevented from falling off and being lost, after jigsaw is completed, the magnetic attraction plate can be pulled out through the handle, at the moment, the jigsaw blocks originally attracted to the inner side of the box cover through the magnetic attraction plate lose attraction force and automatically fall into the box body, and manual taking down one by one is not needed.
Owner:YIBIN XITONG HEALTH TECHNOLOGY CO LTD

An ai agent-based hierarchical visual memory method and system

This invention discloses a hierarchical visual memory method and system based on AI agents, belonging to the field of artificial intelligence agent technology. It addresses the technical problem that existing AI agent text memory systems cannot efficiently process continuous visual input and cannot convert visual data into contextual experience that the agent can reason about, resulting in visual data being unusable stably by the AI ​​agent. The method includes: receiving visual input data and generating corresponding visual observation objects according to preset rules, and writing the visual observation objects into a first visual memory layer; filtering and controlling the visual observation objects through promotion gating; synthesizing the filtered visual observation objects into contextual memory records, and writing the contextual memory records into a second visual memory layer; grouping the contextual memory records according to preset indicators, generating a corresponding natural language summary for each group of contextual memory records, and writing the natural language summary into a third visual memory layer.
Owner:BEIJING MIANBI INTELLIGENT TECH CO LTD

Sam2 small sample segmentation method based on semantic-visual dual memory fusion

ActiveCN121600514BReduce error enhancementsreduce mismatchCharacter and pattern recognitionNeural learning methodsImaging analysisComputer vision
The application discloses a SAM2 small sample segmentation method based on semantic-visual double memory fusion, constructs semantic query memory, visual query memory and query related support visual memory, fuses the semantic query memory and the visual query memory through a memory refinement module guided by the query related support visual memory, and combines SAM2 dense matching and decoding module end-to-end training. The method solves the single memory and foreground-background confusion problems of the existing method, enhances target semantic consistency and fine-grained modeling capability, significantly improves segmentation precision and generalization capability in a complex scene, and is suitable for medical image analysis, automatic driving and the like.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Visual memory in autonomous agents

Embodiments described herein provide a method by which an autonomous agent (which may be a situationalized and materialized agent) learns new visual objects, recognizes those objects within the same “runtime” / session / interaction, and enables the training of a long-term visual memory model for those objects. The visual memory system stores templates of objects and compares images to these templates. The choice of which objects to compare may be determined by internal or external signals (of the agent). The visual memory system stores new and unrecognized objects as “templates” and returns whether any region of an image is likely to contain a match with any of the stored templates. In this way, one-shot visual learning is achieved.
Owner:SOUL MACHINES LTD

Multi-modal video reasoning method and system based on hierarchical multi-agent

The invention discloses a hierarchical multi-agent-based multi-modal video reasoning method and system, and belongs to the technical field of artificial intelligence and multi-modal video understanding. The method comprises the following steps: receiving an input video and a natural language query, and analyzing a complex query into a plurality of logically coherent sub-questions through a question decomposition agent; distributing the sub-questions to multi-source answer generation agents, wherein the multi-source answer generation agents comprise an answer generation agent based on Web, an answer generation agent based on time sequence memory and a video-language answer generation agent; the intelligent agent based on time sequence memory models video long-range dependence through a visual memory library and a query memory library, and adopts a cross-attention mechanism to fuse time sequence features; performing consistent voting and fusion on the multi-source answers through the decision agent to generate a final answer; and carrying out training optimization on the system based on a cross entropy loss function.
Owner:LANZHOU UNIV

A method and system for intelligent inspection of security robots for multiple scenarios

ActiveCN121541646BPattern recognitionVisual matching
This invention discloses an intelligent inspection method and system for security robots in multiple scenarios, relating to the field of intelligent security. The method includes: in environments with good visibility, collecting vibration data generated by the robot's interaction with the ground to establish a standard fingerprint strip and historical images; in low-visibility environments, the robot collects blurred visual images and vibration data in real time, generating the current inspection fingerprint and matching it with the standard fingerprint strip; when a deviation is detected, the system calls upon a visual memory library and uses a deep learning visual matching network to determine the closest reference landmark, thereby adjusting the robot's turning angle to bring it back to the center path of the fingerprint strip. This invention can be applied to industrial control software and deployed on various models of security robots to achieve process control during inspections. It solves the problem of unstable positioning of security robots during inspections in low-visibility environments, realizing path self-correction and intelligent inspection control based on vibration fingerprint and deep visual matching.
Owner:ANLIZHI INTELLIGENT ROBOT TECH (BEIJING) CO LTD

Text-guided cross-scale collaborative gating remote sensing target detection method

The invention relates to the technical field of remote sensing images, in particular to a text-guided cross-scale collaborative gating remote sensing target detection method, which comprises the following steps of: generating fine-grained text description corresponding to remote sensing image contents; respectively extracting image multi-scale features and text hierarchical semantic features; constructing a language guide query vector based on the text hierarchical semantic features; performing dynamic fusion on the image multi-scale features, the text hierarchical semantic features and the language guide query vectors to obtain visual memory features and text enhancement features after cross-modal fusion; and based on the language guide query vector, the visual memory features and the text enhancement features, iteratively optimizing a prediction bounding box through a decoder, and outputting a final remote sensing target detection result. The method aims at overcoming the limitations of the existing remote sensing image visual detection technology in the aspects of semantic description coarseness, cross-modal feature alignment noise sensitivity, weak spatial relationship modeling and the like.
Owner:NORTHWESTERN POLYTECHNICAL UNIV