A scene fingerprint-based visual memory processing method and system
Patent Information
- Application Number
- CN202610918939.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明实施例提供了一种基于场景指纹的视觉记忆处理方法及系统,用于解决如下技术问题:现有视觉记录与场景变化检测方案以画面视觉差异为判断核心,冗余数据多、易受干扰误判且无法表征复杂语义变化,无法适配AI Agent的视觉语义记忆需求
本发明以场景语义变化为记录判断标准,通过场景指纹比对来识别重复语义场景,仅在语义状态发生实质变化时生成视觉记忆记录,相较于逐检测事件写入模式可大幅减少记忆条目数量,在降低存储资源占用的同时,有效缩减后续记忆检索、语义蒸馏的运算量,提升AI Agent记忆调用的响应速度。
Smart Images

Figure CN122842012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a visual memory processing method and system based on scene fingerprints. Background Technology
[0002] In the field of video processing, keyframe extraction and scene change detection are relatively mature technologies. Existing video processing solutions mostly use pixel features, optical flow features, histogram features, or image similarity as the core judgment criteria. They select representative frames from continuous video streams through video segmentation, content clustering, motion entropy calculation, etc., and mainly serve video indexing, video browsing, content summary generation, and compressed storage of monitoring video clips.
[0003] With the development of AI agent technology, visual memory has become a core foundation for agents to perceive the physical environment, perform semantic reasoning, and interact. AI agents need to persistently record continuous visual input to support subsequent scene recall, behavioral pattern analysis, and task decision-making. Currently, AI agents mainly use two modes for visual memory recording: one is a frame-by-frame or per-detection-event writing mode, where a memory record is generated each time a visual detection result is output; the other is an alarm-triggered writing mode, where visual data is recorded only when a security alarm event is triggered.
[0004] However, the event-by-event write mode generates massive amounts of redundant data, significantly increasing the computational cost of subsequent memory retrieval and semantic reasoning. Write modes triggered solely by alarms lose a significant amount of everyday scene information; this mode only records abnormal alarm events and cannot retain normal environmental states, making it difficult for the AI agent to build a complete environmental awareness. Furthermore, traditional pixel-level or basic feature-level change detection algorithms have weak anti-interference capabilities, easily misjudging meaningless screen fluctuations as scene changes, generating a large number of invalid memory records. Some solutions rely solely on the number of targets to determine scene changes, failing to distinguish the semantic differences arising from varying spatial positions, postures, and object interactions with the same number of targets, thus failing to accurately identify substantial semantic changes in the scene. Traditional keyframe output, on the other hand, lacks semantic structure and cannot directly adapt to the agent's needs.
[0005] Ultimately, existing video processing technologies judge changes based on whether there are differences in visual imagery, while AI Agent visual memory judges changes based on whether the semantics of the scene have undergone changes worth recording. The technical goals and evaluation dimensions of the two are fundamentally different, which makes it impossible for existing solutions to directly adapt to the adaptive recording and deduplication requirements of AI Agent visual memory. Summary of the Invention
[0006] This invention provides a visual memory processing method and system based on scene fingerprints to solve the following technical problems: existing visual recording and scene change detection schemes take visual differences in the image as the core of judgment, resulting in a lot of redundant data, easy interference and misjudgment, and inability to represent complex semantic changes, which cannot meet the visual semantic memory requirements of AI agents.
[0007] The embodiments of the present invention adopt the following technical solutions: On one hand, embodiments of the present invention provide a visual memory processing method based on scene fingerprints, the method comprising: starting a real-time visual detection thread to acquire a continuous frame stream output by a camera, and performing target detection on the continuous frame stream to obtain shared detection state data of each target; Parallel startup of event-driven recording paths and periodic snapshot recording paths; The event-driven recording path listens for behavioral events, and visual memory records are constructed based on the listening results and the shared detection state data, and written into the visual memory library. The shared detection state data is read through the periodic snapshot recording path to generate a scene fingerprint representing the semantics of the current scene; Perform semantic consistency comparison between the current scene fingerprint and the historical scene fingerprint; If the comparison result is semantically inconsistent, the change type of the current scene is determined, and a visual memory record is generated based on the change type and written into the visual memory bank.
[0008] In one feasible implementation, a real-time visual detection thread is initiated to acquire a continuous frame stream output by the camera, and target detection is performed on the continuous frame stream to obtain shared detection state data for each target, specifically including: The real-time visual detection thread acquires continuous frame streams from the camera. The camera performs target detection on each frame in the continuous frame stream using a target detection algorithm, and tracks at least one detected target. Identify the real-time pose of each target and classify it to obtain the target pose type; The target detection results, target tracking results, and target pose type are summarized into the shared detection state data.
[0009] In one feasible implementation, the event-driven recording path and the periodic snapshot recording path are started in parallel, specifically including: After acquiring the shared detection status data of each target, the event-driven recording path and the periodic snapshot recording path are started simultaneously through parallel threads; The shared detection status data is simultaneously input into two paths for visual memory writing operations.
[0010] In one feasible implementation, the event-driven recording path listens for behavioral events, and a visual memory record is constructed based on the listening results and the shared detection state data, specifically including: In the event-driven recording path, behavioral events in the shared detection status data are monitored; wherein, the behavioral events include at least behavioral transition events and security events, as well as alarm rules corresponding to the behavioral events; Match the detected behavioral events with their corresponding priorities; If a high-priority behavioral event is detected, a visual memory record is constructed based on the shared detection state data and written into the visual memory library.
[0011] In one feasible implementation, the shared detection state data is read through the periodic snapshot recording path to generate a scene fingerprint characterizing the semantics of the current scene, specifically including: In the periodic snapshot recording path, the shared detection status data is read at preset time intervals; Key image information is extracted from the shared detection state data; wherein, the key information includes at least: target category, detection confidence, bounding box, tracking ID, pose and motion data, and physical region data; Filter low-value data from the key image information; wherein, the low-value data includes low-confidence targets, purely moving areas, briefly appearing targets, and objects within non-interested areas; The filtered key image information is subjected to coarse-grained spatial quantization processing, and the center point of each target is mapped to the corresponding spatial region. Based on the spatial region, feature descriptions are performed on the key image information to generate a scene fingerprint that represents the semantics of the current scene.
[0012] In one feasible implementation, based on the spatial region, the key image information is characterized by feature description to generate a scene fingerprint representing the semantics of the current scene, specifically including: Feature normalization is performed on human targets and non-human targets respectively to generate corresponding target description fragments; among them, the human target description fragment includes anonymous tracking label, spatial region, posture and action and duration; the non-human target description fragment includes target category, number of targets and spatial region. Extract the spatial interaction relationships between different targets and generate relationship description fragments; All target description fragments and relationship description fragments are sorted and concatenated according to a preset sorting rule to obtain the initial scene fingerprint; Perform empty scene verification on the initial scene fingerprint to generate the final scene fingerprint that represents the semantics of the current scene.
[0013] In one feasible implementation, the initial scene fingerprint is subjected to empty scene verification to generate a final scene fingerprint for characterizing the semantics of the current scene, specifically including: If the initial scene fingerprint does not correspond to an empty scene, then the initial scene fingerprint is used as the scene fingerprint representing the semantics of the current scene; If the initial scene fingerprint corresponds to an empty scene, then a preset number of periodic snapshots are continuously acquired; when all acquired periodic snapshots generate an initial fingerprint for an empty scene, a scene clear fingerprint is generated as the scene fingerprint representing the semantics of the current scene.
[0014] In one feasible implementation, semantic consistency comparison is performed between the current scene fingerprint and the historical scene fingerprint, specifically including: Perform a semantic consistency comparison between the current scene fingerprint and the previous scene fingerprint stored on the same visual acquisition device; If the comparison result is semantically consistent, then skip this memory write.
[0015] In one feasible implementation, if the comparison result is semantically inconsistent, the change type of the current scene is determined, and a visual memory record is generated and written into the visual memory bank based on the change type. Specifically, this includes: The change type of the current scene is determined based on the scene fingerprint; the change type includes changes in empty scene, changes in the number of objects, changes in posture, changes in spatial region, and changes in relationships; If the change type is an empty scene change, then a delayed confirmation process is performed. If multiple consecutive detections determine that it is an empty scene change, then the scene semantics are confirmed to have changed, and a visual memory record is generated and written into the visual memory bank. If the change type is any other than the change in an empty scene, then the change in scene semantics is directly confirmed, and a visual memory record is generated and written into the visual memory bank.
[0016] On the other hand, embodiments of the present invention also provide a visual memory processing system based on scene fingerprints, the system comprising: The state detection module is used to start a real-time visual detection thread to acquire continuous frame streams output by the camera, and to perform target detection on the continuous frame streams to obtain shared detection state data of each target. The visual memory recording generation module is used to launch the event-driven recording path and the periodic snapshot recording path in parallel; it listens for behavioral events through the event-driven recording path, and constructs visual memory records based on the listening results and the shared detection state data, and writes them into the visual memory library; it reads the shared detection state data through the periodic snapshot recording path and generates a scene fingerprint that represents the semantics of the current scene. The visual memory recording and writing module is used to compare the current scene fingerprint with the historical scene fingerprint for semantic consistency. If the comparison result is semantically inconsistent, the change type of the current scene is determined, and a visual memory record is generated and written into the visual memory library based on the change type.
[0017] Compared with the prior art, the visual memory processing method and system based on scene fingerprinting provided in this invention have the following beneficial effects: This invention uses scene semantic changes as the recording and judgment standard, identifies repetitive semantic scenes through scene fingerprint comparison, and generates visual memory records only when the semantic state changes substantially. Compared with the event-by-event writing mode, it can significantly reduce the number of memory entries, reduce storage resource consumption, effectively reduce the amount of computation for subsequent memory retrieval and semantic distillation, and improve the response speed of AI Agent memory retrieval.
[0018] Scene fingerprinting integrates multi-dimensional information such as object category, number of targets, spatial distribution area, posture and action state, duration and interaction relationship between objects. It breaks through the limitation of traditional solutions that only compare the number of targets or pixel features. It can accurately distinguish the semantic differences of complex scenes such as multi-person interaction, target displacement, and behavior state transformation, and output memory records with complete semantic structure. It can directly support AI Agent to carry out scene recall, behavior pattern analysis and task decision-making.
[0019] In response to the special case of a scene transitioning from a target to an empty scene, this invention sets up a multi-round snapshot lag confirmation mechanism. The scene clearing fingerprint is only generated after multiple consecutive snapshots detect an empty scene. This effectively reduces the problem of repeated state flipping caused by instantaneous missed detections in computer vision, avoids generating a large number of invalid state switching records, and ensures the continuity and stability of scene state recording.
[0020] This invention adopts a dual-mode recording architecture that combines event-driven and periodic snapshots. High-priority alarm events, security events, and behavior transformation events can be triggered for immediate writing, ensuring zero-latency recording and no omission of critical events. In daily routine scenarios, periodic snapshots are used to complete semantic change detection and retain complete environmental state information. The combination of the two satisfies the real-time requirements of security scenarios and covers the memory collection needs of daily scenarios. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1A flowchart of a visual memory processing method based on scene fingerprinting provided in an embodiment of the present invention; Figure 2 A flowchart for scene fingerprint generation is provided in an embodiment of the present invention; Figure 3 A schematic diagram of a fingerprint in a complex scene provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a scene fingerprint-based visual memory processing device provided in an embodiment of the present invention. Detailed Implementation
[0022] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.
[0023] This invention provides a visual memory processing method based on scene fingerprints, such as... Figure 1 As shown, the scene fingerprint-based visual memory processing method specifically includes steps S101-S105: S101. Start the real-time visual detection thread to acquire the continuous frame stream output by the camera, and perform target detection on the continuous frame stream to obtain the shared detection status data of each target.
[0024] Specifically, a real-time visual inspection thread acquires continuous frame streams from the camera. A target detection algorithm is used to detect targets in each frame of the continuous frame stream, and at least one detected target is tracked.
[0025] Furthermore, the real-time pose of each target is identified and classified to obtain the target pose type. The target detection results, target tracking results, and target pose type are then aggregated into shared detection state data.
[0026] As one feasible implementation, the method provided by this invention is executed by the visual recording module of a visual perception plugin or an AI Agent host system. The visual perception plugin or visual recording module starts a real-time visual detection thread to perform target detection, multi-target tracking, and pose / motion classification processing on the continuous frame stream output by the camera, and continuously updates and maintains the shared detection state.
[0027] S102, Parallel startup of event-driven recording path and periodic snapshot recording path.
[0028] Specifically, after acquiring the shared detection state data for each target, the event-driven recording path and the periodic snapshot recording path are started simultaneously via parallel threads. Then, the shared detection state data is simultaneously input into both paths for visual memory writing operations.
[0029] S103. Record path listening behavior events through event-driven recording, construct visual memory records based on the listening results and shared detection state data, and write them into the visual memory library.
[0030] Specifically, in the event-driven logging path, behavioral events in the shared detection status data are monitored. These behavioral events include at least behavioral transition events and security events, as well as the corresponding alarm rules.
[0031] Furthermore, the detected behavioral events are matched with their corresponding priorities. If a high-priority behavioral event is detected, a visual memory record is constructed based on the shared detection state data and written into the visual memory bank.
[0032] As a feasible implementation method, the event-driven recording path monitors preset alarm rules, security events and behavior transition events in real time. When a high-priority event is hit, a visual memory record is immediately constructed based on the current shared detection state and written into the visual memory bank.
[0033] S104. Read shared detection state data by periodically snapshotting the path and generate a scene fingerprint that represents the semantics of the current scene.
[0034] Specifically, shared detection status data is read from the periodic snapshot recording path at preset time intervals. Key image information is extracted from the shared detection status data; the key information includes at least: target category, detection confidence, bounding box, tracking ID, pose and motion data, and physical region data.
[0035] Furthermore, low-value data in key image information is filtered out; low-value data includes low-confidence targets, purely moving areas, briefly appearing targets, and objects in non-interested areas.
[0036] Furthermore, coarse-grained spatial quantization is performed on the filtered key image information, mapping the center points of each target to their corresponding spatial regions. Based on these spatial regions, feature descriptions are generated for the key image information, producing a scene fingerprint that characterizes the semantics of the current scene. Specifically, this includes: Feature normalization is performed on human and non-human targets separately to generate corresponding target description fragments. The human target description fragment includes anonymized tracking labels, the spatial region it belongs to, its posture and action, and its duration. The non-human target description fragment includes the target category, the number of targets, and the spatial region it belongs to. Spatial interaction relationships between different targets are extracted to generate relationship description fragments. All target description fragments and relationship description fragments are sorted and concatenated according to a preset sorting rule to obtain the initial scene fingerprint.
[0037] Then, empty scene verification is performed on the initial scene fingerprint to generate the final scene fingerprint used to represent the semantics of the current scene, specifically including: If the initial scene fingerprint does not correspond to an empty scene, then the initial scene fingerprint is used as the scene fingerprint representing the semantics of the current scene. If the initial scene fingerprint corresponds to an empty scene, then a preset number of periodic snapshots are continuously acquired; when all acquired periodic snapshots generate an initial fingerprint for an empty scene, a scene clear fingerprint is generated and used as the scene fingerprint representing the semantics of the current scene.
[0038] As a feasible implementation method, Figure 2 A flowchart for scene fingerprint generation is provided as an embodiment of the present invention, such as... Figure 2 As shown, the detection status is first read to obtain data including the object category, confidence level, bounding box, tracking ID, pose, object category, and physical region within the current frame or the nearest window.
[0039] Then, low-value inputs are filtered out, eliminating low-confidence targets, purely moving areas, briefly appearing targets, and objects outside the area of interest.
[0040] Then, spatial quantization is performed to divide the image into coarse-grained regions, such as left / center / right, near / middle / far, or mapped according to the actual room area, with the target center point mapped to the corresponding area.
[0041] Then, object normalization is performed to generate descriptive fragments for people, pets, and important items. Person fragments may include anonymous tracking tags, spatial regions, gestures, and duration buckets; pet and object fragments may include category, quantity, spatial regions, and relationships.
[0042] Furthermore, relation extraction is performed on the normalized objects. If there are relationships such as proximity, contact, occupation, or being near a piece of furniture between the objects, relation fragments are generated. For example, "person - near - medicine box" or "cat - located - doorway".
[0043] Finally, the above information is concatenated into a scene state code, sorted by object category, spatial region, tracking stability, and relationship fragments to generate a normalized string.
[0044] If a scene changes from having a target to being empty, a scene clearing fingerprint can only be generated if the number of consecutive snapshots that are empty reaches a threshold; otherwise, a clearing fingerprint will not be generated.
[0045] In one embodiment, Figure 3 A schematic diagram of a fingerprint in a complex scene provided as an embodiment of the present invention, such as... Figure 3 As shown, in a complex scene, an elderly person is sitting on a sofa on the left side of the living room, a child is standing on the right side, and a cat is at the door. The character fragments are: "Left: Sitting: Elderly Person," "Middle: Walking: Child." The pet / object fragments are: "Cat, located at the doorway," "Medicine box, located on the table." The relationship fragments are: "Elderly Person - near - Sofa," "Person - near - Medicine Box." After normalizing and sorting these fragments, the scene state code is obtained. The above fingerprint discards precise coordinates and pixel details but retains the semantic changes required for agent memory. The scene fingerprint remains unchanged when the bounding box jitters slightly; it changes when the person changes from sitting to standing, from left to right, or when the number of people near the medicine box or exhibits changes.
[0046] S105. Perform semantic consistency comparison between the current scene fingerprint and the historical scene fingerprint; if the comparison result is semantically inconsistent, determine the change type of the current scene, generate a visual memory record based on the change type, and write it into the visual memory bank.
[0047] Specifically, the current scene fingerprint is compared semantically with the previous scene fingerprint stored on the same visual acquisition device. If the comparison result is semantically consistent, the current memory write is skipped.
[0048] If the comparison result is semantically inconsistent, the change type of the current scene is determined based on the scene fingerprint; the change type includes changes in empty scene, changes in the number of objects, changes in pose, changes in spatial region, and changes in relationships.
[0049] If the change type is an empty scene change, then a delayed confirmation process is performed. If multiple consecutive detections determine that it is an empty scene change, then the scene semantics are confirmed to have changed, and a visual memory record is generated and written into the visual memory bank.
[0050] If the change type is any other than the change in an empty scene, the scene semantics are directly confirmed to have changed, a visual memory record is generated and written to the visual memory bank, and the fingerprint of the last stable scene and the last write time are updated.
[0051] This invention uses scene fingerprints, rather than pixel differences, to determine whether to write to visual memory. Scene fingerprints include object category, quantity, spatial buckets, pose / action, duration buckets, and object relationships. They are not written when the same semantic scene appears repeatedly; they are only recorded when the semantic fingerprint changes, significantly reducing redundant writing. The proposed delayed confirmation mechanism reduces the occurrence of "person / no one" switching caused by brief missed detections. Furthermore, this method aligns with the Agent's memory objective, outputting semantic changes that can be written into contextual memory, rather than simply video keyframes.
[0052] In addition, embodiments of the present invention also provide a visual memory processing system based on scene fingerprints, such as... Figure 4 As shown, the scene fingerprint-based visual memory processing system 400 specifically includes: The state detection module 410 is used to start a real-time visual detection thread to acquire a continuous frame stream output by the camera, and to perform target detection on the continuous frame stream to obtain shared detection state data of each target. The visual memory recording generation module 420 is used to launch the event-driven recording path and the periodic snapshot recording path in parallel; it listens for behavioral events through the event-driven recording path, and constructs a visual memory record based on the listening results and the shared detection state data, and writes it into the visual memory library; it reads the shared detection state data through the periodic snapshot recording path and generates a scene fingerprint that represents the semantics of the current scene. The visual memory recording and writing module 430 is used to compare the current scene fingerprint with the historical scene fingerprint for semantic consistency; if the comparison result is semantic inconsistency, the change type of the current scene is determined, and a visual memory record is generated according to the change type and written into the visual memory library.
[0053] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0054] The foregoing has described specific embodiments of the present invention. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0055] The above description is merely an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present invention should be included within the protection scope of the present invention.
Claims
1. A visual memory processing method based on scene fingerprints, characterized in that, The method includes: A real-time visual detection thread is started to acquire continuous frame streams output by the camera, and target detection is performed on the continuous frame streams to obtain shared detection status data of each target. Parallel startup of event-driven recording paths and periodic snapshot recording paths; The event-driven recording path listens for behavioral events, and visual memory records are constructed based on the listening results and the shared detection state data, and written into the visual memory library. The shared detection state data is read through the periodic snapshot recording path to generate a scene fingerprint representing the semantics of the current scene; Perform semantic consistency comparison between the current scene fingerprint and the historical scene fingerprint; If the comparison result is semantically inconsistent, the change type of the current scene is determined, and a visual memory record is generated based on the change type and written into the visual memory bank.
2. The visual memory processing method based on scene fingerprints according to claim 1, characterized in that, A real-time visual detection thread is initiated to acquire continuous frame streams output by the camera, and target detection is performed on the continuous frame streams to obtain shared detection state data for each target, specifically including: The real-time visual detection thread acquires continuous frame streams from the camera. The camera performs target detection on each frame in the continuous frame stream using a target detection algorithm, and tracks at least one detected target. Identify the real-time pose of each target and classify it to obtain the target pose type; The target detection results, target tracking results, and target pose type are summarized into the shared detection state data.
3. The visual memory processing method based on scene fingerprints according to claim 1, characterized in that, Parallel startup of event-driven recording paths and periodic snapshot recording paths, specifically including: After acquiring the shared detection status data of each target, the event-driven recording path and the periodic snapshot recording path are started simultaneously through parallel threads; The shared detection status data is simultaneously input into two paths for visual memory writing operations.
4. The visual memory processing method based on scene fingerprints according to claim 1, characterized in that, The event-driven recording path listens for behavioral events, and visual memory records are constructed based on the listening results and the shared detection state data, specifically including: In the event-driven recording path, behavioral events in the shared detection status data are monitored; wherein, the behavioral events include at least behavioral transition events and security events, as well as alarm rules corresponding to the behavioral events; Match the detected behavioral events with their corresponding priorities; If a high-priority behavioral event is detected, a visual memory record is constructed based on the shared detection state data and written into the visual memory library.
5. The visual memory processing method based on scene fingerprints according to claim 1, characterized in that, The shared detection state data is read through the periodic snapshot recording path to generate a scene fingerprint representing the semantics of the current scene, specifically including: In the periodic snapshot recording path, the shared detection status data is read at preset time intervals; Key image information is extracted from the shared detection state data; wherein, the key information includes at least: target category, detection confidence, bounding box, tracking ID, pose and motion data, and physical region data; Filter low-value data from the key image information; wherein, the low-value data includes low-confidence targets, purely moving areas, briefly appearing targets, and objects within non-interested areas; The filtered key image information is subjected to coarse-grained spatial quantization processing, and the center point of each target is mapped to the corresponding spatial region. Based on the spatial region, feature descriptions are performed on the key image information to generate a scene fingerprint that represents the semantics of the current scene.
6. The visual memory processing method based on scene fingerprints according to claim 5, characterized in that, Based on the spatial region, feature descriptions are performed on the key image information to generate a scene fingerprint representing the semantics of the current scene, specifically including: Feature normalization is performed on human targets and non-human targets respectively to generate corresponding target description fragments; among them, the human target description fragment includes anonymous tracking label, spatial region, posture and action and duration; the non-human target description fragment includes target category, number of targets and spatial region. Extract the spatial interaction relationships between different targets and generate relationship description fragments; All target description fragments and relationship description fragments are sorted and concatenated according to a preset sorting rule to obtain the initial scene fingerprint; Perform empty scene verification on the initial scene fingerprint to generate the final scene fingerprint that represents the semantics of the current scene.
7. The visual memory processing method based on scene fingerprints according to claim 6, characterized in that, Perform empty scene verification on the initial scene fingerprint to generate a final scene fingerprint that represents the semantics of the current scene, specifically including: If the initial scene fingerprint does not correspond to an empty scene, then the initial scene fingerprint is used as the scene fingerprint representing the semantics of the current scene; If the initial scene fingerprint corresponds to an empty scene, then a preset number of periodic snapshots are continuously acquired; when all acquired periodic snapshots generate an initial fingerprint for an empty scene, a scene clear fingerprint is generated as the scene fingerprint representing the semantics of the current scene.
8. The visual memory processing method based on scene fingerprints according to claim 1, characterized in that, Perform semantic consistency comparison between the current scene fingerprint and historical scene fingerprints, specifically including: Perform a semantic consistency comparison between the current scene fingerprint and the previous scene fingerprint stored on the same visual acquisition device; If the comparison result is semantically consistent, then skip this memory write.
9. The visual memory processing method based on scene fingerprints according to claim 1, characterized in that, If the comparison result indicates semantic inconsistency, the type of change in the current scene is determined, and a visual memory record is generated and written into the visual memory bank based on the type of change. Specifically, this includes: The change type of the current scene is determined based on the scene fingerprint; the change type includes changes in empty scene, changes in the number of objects, changes in posture, changes in spatial region, and changes in relationships; If the change type is an empty scene change, then a delayed confirmation process is performed. If multiple consecutive detections determine that it is an empty scene change, then the scene semantics are confirmed to have changed, and a visual memory record is generated and written into the visual memory bank. If the change type is any other than the change in an empty scene, then the change in scene semantics is directly confirmed, and a visual memory record is generated and written into the visual memory bank.
10. A visual memory processing system based on scene fingerprinting, characterized in that, The system includes: The state detection module is used to start a real-time visual detection thread to acquire continuous frame streams output by the camera, and to perform target detection on the continuous frame streams to obtain shared detection state data of each target. The visual memory recording generation module is used to launch the event-driven recording path and the periodic snapshot recording path in parallel; it listens for behavioral events through the event-driven recording path, and constructs visual memory records based on the listening results and the shared detection state data, and writes them into the visual memory library; it reads the shared detection state data through the periodic snapshot recording path and generates a scene fingerprint that represents the semantics of the current scene. The visual memory recording and writing module is used to compare the current scene fingerprint with the historical scene fingerprint for semantic consistency. If the comparison result is semantically inconsistent, the change type of the current scene is determined, and a visual memory record is generated and written into the visual memory library based on the change type.