Scene difference information storage method and system
By collecting and analyzing scene image sets in the intelligent environmental perception system, identifying and storing difference feature vectors, the problems of low historical data utilization and insufficient monitoring accuracy in the existing system are solved, and efficient scene difference analysis and query are achieved.
Patent Information
- Application Number
- CN202511180269.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing intelligent environmental perception systems have shortcomings in historical data utilization and monitoring accuracy. It is difficult to quickly and accurately query and utilize key information about environmental changes from historical data. Traditional big data processing methods lack an in-depth understanding of high-dimensional spatial semantic information, which affects the query processing accuracy and efficiency of scenario tasks.
The spatial intelligent machine collects scene image sets and spatial structure dynamic information, identifies target objects, and uses difference recognition algorithms to identify difference information in multiple dimensions. The preset world model is used to generate difference feature vectors, which are associated with the original scene data and stored to avoid redundant storage and improve query efficiency.
It achieves real-time and accurate scene difference analysis, significantly reduces data storage volume, improves the query and reasoning efficiency of subsequent tasks, and reduces data processing complexity.
Smart Images

Figure CN120726401A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent perception systems, and in particular to a method and system for storing scene difference information. Background Art
[0002] Current intelligent environmental perception systems typically rely on real-time image processing technology to monitor environmental changes. However, these systems suffer from widespread issues such as poor utilization of historical data and low monitoring accuracy, making it difficult to quickly and accurately retrieve and utilize key information about environmental changes from historical data. Furthermore, while traditional big data processing methods, such as access log analysis or the Event Sourcing paradigm, can record events and provide historical traceability, they typically only process structured or semi-structured data and lack a deep understanding of high-dimensional spatial semantics. This makes complex scene analysis and real-time reasoning difficult, hindering the accuracy and efficiency of query processing for scenario-based tasks. Summary of the Invention
[0003] In response to the above technical problems, the present invention provides a method and system for storing scene difference information, which performs difference analysis on complex scenes and stores difference information through high-dimensional spatial data. It can achieve real-time and accurate scene difference analysis and significantly reduce the amount of data storage, which is conducive to efficient query and reasoning of subsequent tasks.
[0004] According to a first aspect of the present invention, a method for storing scene difference information is provided, comprising the following steps: S100, based on the scene image set of the target area collected by the pre-deployed spatial intelligent machine within a preset time period and the acquired spatial structure dynamic information, a number of target objects are identified from the scene image set based on a number of preset category objects; the said number of preset category objects include scene background, people, objects, robots and IoT devices, scene behavior events and scene macro laws.
[0005] S200, for any target object, a preset difference recognition algorithm is used to identify the difference information of the target object in several dimensions based on the spatial structure dynamic information; the difference information is the semantic information corresponding to the difference of the target object in any dimension; wherein, the difference recognition algorithms corresponding to the scene background, personnel, objects, robots and IOT devices, scene behavior events and scene macro-laws are semantic segmentation and scene recognition algorithms, human posture estimation and behavior recognition algorithms, target detection and tracking algorithms, device state recognition algorithms, event detection algorithms and statistical analysis algorithms.
[0006] S300: Input the difference information of each target object in different dimensions into a preset world model, and output a target difference feature vector corresponding to the scene image set through the preset world model.
[0007] S400: associating the target difference feature vector corresponding to the scene image set with the preset original scene data and storing them in the spatial intelligent machine.
[0008] According to a second aspect of the present invention, a system for storing scene difference information is provided, the storage system comprising: The first recognition module is used to identify a number of target objects from the scene image set based on a number of preset category objects based on a set of scene images of the target area collected by a pre-deployed spatial intelligent machine within a preset time period and the acquired spatial structure dynamic information; the said number of preset category objects include scene background, people, objects, robots and IoT devices, scene behavior events and scene macro-laws.
[0009] The second recognition module is used to identify the difference information of any target object in several dimensions by using a preset difference recognition algorithm based on the spatial structure dynamic information; the difference information is the semantic information corresponding to the difference of the target object in any dimension; among which, the difference recognition algorithms corresponding to the scene background, personnel, objects, robots and IOT devices, scene behavior events and scene macro-laws are semantic segmentation and scene recognition algorithms, human posture estimation and behavior recognition algorithms, target detection and tracking algorithms, device state recognition algorithms, event detection algorithms and statistical analysis algorithms.
[0010] The processing module is used to input the difference information of each target object in different dimensions into a preset world model, and output the target difference feature vector corresponding to the scene image set through the preset world model.
[0011] The storage module is used to associate and store the target difference feature vector corresponding to the scene image set and the preset original scene data in the spatial intelligent machine.
[0012] The present invention has at least the following beneficial effects: 1. Based on the scene image set of the target area within a preset time period collected by the pre-deployed spatial intelligent machine, several target objects are identified from the scene image set. Through object decomposition, it is beneficial to perform difference analysis on each independent object separately, which can improve the quality of difference analysis.
[0013] 2. By using the corresponding difference recognition algorithm for each target object, the difference information of the target object in several dimensions can be identified. By using different difference recognition algorithms for different target objects, the difference information of different target objects in multiple dimensions can be accurately identified, making the identified difference information more comprehensive, which is conducive to improving subsequent accurate queries.
[0014] 3. By inputting the difference information of each target object in different dimensions into the preset world model, and outputting the target difference feature vector corresponding to the scene image set through the preset world model, the characteristics of the world model that can uniformly quantify the difference information of multiple dimensions and efficiently process high-dimensional spatial data are referenced, which is conducive to real-time and accurate scene difference analysis of the complex scenes in this application.
[0015] 4. By associating and storing the target difference feature vector corresponding to the scene image set and the preset original scene data in the spatial intelligent machine, the scene change information is stored in the form of a difference feature vector, which effectively avoids the repeated storage of a large amount of similar or redundant original scene data, significantly reduces the data storage space occupied, and in subsequent scene queries, only the differentiated information needs to be retrieved and inferred, which greatly reduces the complexity of data query and processing, and improves the efficiency of the system's real-time response and query. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 This is a flowchart of a method for storing scene difference information provided in the first embodiment of the present invention; Figure 2 This is a structural diagram of a system for storing scene difference information provided in the second embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] Example 1 The first embodiment of the present invention provides a method for storing scene difference information, such as Figure 1 As shown, the storage method includes the following steps: S100, based on the scene image set of the target area collected by the pre-deployed spatial intelligent machine within a preset time period and the acquired spatial structure dynamic information, identifies several target objects from the scene image set based on several preset category objects; it can be understood as: the spatial intelligent machine collects images in real time through camera data, and integrates local sensors and its own environmental perception and spatial modeling functions to continuously monitor the local spatial structure dynamic information to understand the physical space.
[0020] Specifically, the preset category objects include scene background, people, objects, robots and IOT devices, scene behavior events and scene macro laws.
[0021] Specifically, the spatial structure dynamic information includes semantic information of dynamic changes of each target object in the target area. It can be understood that the spatial intelligent machine performs difference analysis on continuous image frames to obtain the dynamic changes of each object.
[0022] As mentioned above, by identifying several target objects from the scene image set, the decomposition of different objects is achieved, which is conducive to performing difference analysis on each independent object separately, thereby improving the accuracy of difference analysis. The obtained spatial structure dynamic information can grasp the macro-dynamic changes of the target area, which is conducive to semantic analysis of dynamic events in the scene.
[0023] S200, for any target object, a preset difference recognition algorithm is used to identify the difference information of the target object in several dimensions based on the spatial structure dynamic information; wherein, the difference recognition algorithms corresponding to the scene background, personnel, objects, robots and IOT devices, scene behavior events and scene macro-laws are semantic segmentation and scene recognition algorithm, human posture estimation and behavior recognition algorithm, target detection and tracking algorithm, device state recognition algorithm, event detection algorithm and statistical analysis algorithm.
[0024] Specifically, the difference information is the semantic information corresponding to the difference of the target object in any dimension; for example, when information of three dimensions, namely, the position movement of a person, the posture of bending over and standing up, and the action of picking up an object, is detected, the corresponding semantic information is moving to the object and bending over to pick up the object.
[0025] Specifically, semantic segmentation and scene recognition algorithms are used to identify structural changes in the scene background; human posture estimation and behavior recognition algorithms are used to identify changes in the position, posture and movement of people; target detection and tracking algorithms are used to detect changes in the position, form, quantity and spatial position of objects, such as the movement of tables and chairs; device state recognition algorithms are used to detect the state, position changes and functional state differences of robots and IOT devices in real time; event detection algorithms are used to identify the occurrence or changes of various scene events in space, such as gatherings, abnormal behaviors, special events, etc.; statistical analysis algorithms and machine learning algorithms are used to identify and predict abnormal changes in macro-behavioral laws or patterns in scenes, such as trends in changes in crowd density, abnormal frequency of equipment use, etc. Those skilled in the art are aware of the specific implementation methods of each of the above algorithms and will not be repeated here.
[0026] As described above, by adopting a corresponding difference recognition algorithm for each target object to identify the difference information of the target object in several dimensions, and by adopting different difference recognition algorithms for different target objects, the difference information of different target objects in multiple dimensions can be accurately identified, making the identified difference information more comprehensive, which is conducive to improving subsequent accurate queries.
[0027] S300: Input the difference information of each target object in different dimensions into a preset world model, and output a target difference feature vector corresponding to the scene image set through the preset world model.
[0028] Specifically, step S300 includes the following steps: S301, for any target object, based on the difference information of the target object in several dimensions, generates an initial difference feature vector corresponding to the target object through a preset world model; it can be understood that the preset world model extracts the difference information of each target object in different dimensions and uniformly expresses it as a feature vector.
[0029] S302, merging the initial difference feature vectors corresponding to each target object through the preset world model, and outputting the target difference feature vector corresponding to the scene image set; it can be understood that the preset world model is embedded in the corresponding spatial intelligent machine.
[0030] As mentioned above, by introducing the world model, it is possible to uniformly quantify the difference information of multiple dimensions, and the world model has the characteristic of efficiently processing high-dimensional spatial data, making the extraction of difference feature vectors more unified and efficient, which is conducive to real-time and accurate scene difference analysis of the complex scenes in this application.
[0031] S400: associating the target difference feature vector corresponding to the scene image set with the preset original scene data and storing them in the spatial intelligent machine.
[0032] Specifically, the preset original scene data includes a scene image set corresponding to a target difference feature vector, a reading of a preset sensor, and an original scene description text.
[0033] As mentioned above, due to the acquisition of the aforementioned difference feature vector, when storing information, it is only necessary to store the original scene data and store the scene change information in the form of a difference feature vector, instead of storing the global scene data, which effectively avoids the repeated storage of a large amount of similar or redundant original scene data, significantly reduces the data storage space occupied, and in subsequent scene queries, only the differentiated information needs to be retrieved and inferred, which greatly reduces the complexity of data query and processing, and improves the efficiency of the system's real-time response and query.
[0034] Furthermore, the method further comprises the following steps: S10, when receiving a user complex task query request, the user complex task query request is decomposed into several sub-query tasks through the pre-deployed central machine; the user complex task query request refers to a request including at least two query tasks; for example, when a person is detected entering the target area and approaching a fixed item box, the person's actions and purpose are inferred.
[0035] S20, for each sub-query task, extract the target difference feature vector corresponding to the sub-query task and the original scene data associated with the extracted target difference feature vector from each pre-deployed spatial intelligent machine, and obtain the semantic information corresponding to the extracted target difference feature vector.
[0036] In a specific embodiment, step S20 includes the following steps: S21, obtaining semantic information of the sub-query task and converting it into a task feature vector corresponding to the sub-query task; the task feature vector has the same vector dimension as the target difference feature vector.
[0037] S22, calculating the vector similarity between the task feature vector corresponding to the sub-query task and each stored target difference feature vector, and extracting the target difference feature vector corresponding to the maximum vector similarity.
[0038] In another specific embodiment, several keywords of the subquery task and several semantic keywords corresponding to the target difference feature vector are obtained, and by comparing the same situations of the several keywords and the several semantic keywords, the target difference feature vector corresponding to the most identical keywords is extracted.
[0039] Specifically, the spatial intelligent machines are provided in a plurality and are deployed in a distributed manner, and each spatial intelligent machine communicates with the central machine.
[0040] S30: Semantically integrate the semantic information corresponding to the acquired target difference feature vector with the original scene data associated with the extracted target difference feature vector to obtain the task reasoning result text. For example, if a person is detected entering the monitoring space and approaching a fixed item box, the corresponding target difference feature vector is found based on the semantic similarity of the historical data. If the person's arm is extended in front of the fixed item box and the door of the fixed item box is opened and closed, then the person's intention is to take the item from the fixed item box.
[0041] As described above, through the coordinated operation of the central machine and several spatial intelligent machines, the central data can decompose the complex query tasks and then request historical target difference feature vector information and related original scene data from different spatial intelligent machines respectively. The central machine then integrates the content of all the returned data to provide complete and efficient overall reasoning results.
[0042] Example 2 The second embodiment of the present invention provides a storage system for scene difference information, such as Figure 2 As shown, the storage system includes: The first recognition module 100 is configured to identify a plurality of target objects from a set of scene images of a target area within a preset time period, based on a set of scene images and spatial structure dynamic information acquired by a pre-deployed spatial intelligent machine. The plurality of preset categories of objects include scene background, people, objects, robots and IoT devices, scene behavior events, and scene macro-regularities. The second recognition module 200 is used to identify, for any target object, difference information of the target object in several dimensions using a preset difference recognition algorithm based on the spatial structure dynamic information; the difference information is semantic information corresponding to the difference of the target object in any dimension; wherein the difference recognition algorithms corresponding to the scene background, people, objects, robots and IoT devices, scene behavior events, and scene macro-regulations are semantic segmentation and scene recognition algorithms, human posture estimation and behavior recognition algorithms, target detection and tracking algorithms, device state recognition algorithms, event detection algorithms, and statistical analysis algorithms respectively; The processing module 300 is used to input the difference information of each target object in different dimensions into a preset world model, and output the target difference feature vector corresponding to the scene image set through the preset world model; The storage module 400 is used to associate and store the target difference feature vector corresponding to the scene image set and the preset original scene data in the spatial intelligent machine.
[0043] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0044] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A method for storing scene difference information, characterized in that: The storage method comprises the following steps: S100, based on a set of scene images of a target area collected by a pre-deployed spatial intelligent machine within a preset time period and the acquired spatial structure dynamic information, identifying a plurality of target objects from the scene image set based on a plurality of preset categories of objects; the plurality of preset categories of objects include scene background, people, objects, robots and IoT devices, scene behavior events, and scene macro-regularities; S200: For any target object, based on the spatial structure dynamic information, a preset difference recognition algorithm is used to identify difference information of the target object in several dimensions; the difference information is semantic information corresponding to the difference of the target object in any dimension; wherein the difference recognition algorithms corresponding to the scene background, people, objects, robots and IoT devices, scene behavior events, and scene macro-regulations are semantic segmentation and scene recognition algorithms, human posture estimation and behavior recognition algorithms, target detection and tracking algorithms, device state recognition algorithms, event detection algorithms, and statistical analysis algorithms, respectively; S300, inputting the difference information of each target object in different dimensions into a preset world model, and outputting the target difference feature vector corresponding to the scene image set through the preset world model; S400: associating the target difference feature vector corresponding to the scene image set with the preset original scene data and storing them in the spatial intelligent machine.
2. The method for storing scene difference information according to claim 1, characterized in that: The preset original scene data includes a scene image set corresponding to a target difference feature vector, a reading of a preset sensor, and an original scene description text.
3. The method for storing scene difference information according to claim 2, wherein: The method further comprises the steps of: S10, upon receiving a user complex task query request, decomposing the user complex task query request into a plurality of sub-query tasks through a pre-deployed hub; the user complex task query request is a request including at least two query tasks; S20, for each sub-query task, extracting a target difference feature vector corresponding to the sub-query task and original scene data associated with the extracted target difference feature vector from each pre-deployed spatial intelligent machine, and obtaining semantic information corresponding to the extracted target difference feature vector; S30 , semantically integrating the semantic information corresponding to the acquired target difference feature vector and the original scene data associated with the extracted target difference feature vector to obtain a task reasoning result text.
4. The method for storing scene difference information according to claim 1, wherein: Step S20 also includes the following steps: S21, obtaining semantic information of the sub-query task and converting it into a task feature vector corresponding to the sub-query task; the task feature vector has the same vector dimension as the target difference feature vector; S22, calculating the vector similarity between the task feature vector corresponding to the sub-query task and each stored target difference feature vector, and extracting the target difference feature vector corresponding to the maximum vector similarity.
5. The method for storing scene difference information according to claim 1, wherein: There are several spatial intelligent machines and they are deployed in a distributed manner, and each spatial intelligent machine communicates with the central machine.
6. The method for storing scene difference information according to claim 1, wherein: The spatial structure dynamic information includes dynamically changing semantic information of each target object within the target area.
7. The method for storing scene difference information according to claim 1, wherein: Step S300 includes the following steps: S301, for any target object, generating an initial difference feature vector corresponding to the target object through a preset world model according to difference information of the target object in several dimensions; S302: Merge the initial difference feature vectors corresponding to each target object through a preset world model, and output the target difference feature vector corresponding to the scene image set.
8. A storage system for scene difference information, characterized in that: The storage system includes: A first recognition module is configured to identify a plurality of target objects from a scene image set based on a plurality of preset categories of objects, based on a set of scene images of a target area collected by a pre-deployed spatial intelligent machine within a preset time period and the acquired spatial structure dynamic information; the plurality of preset categories of objects include scene background, people, objects, robots and IoT devices, scene behavior events, and scene macro-regularities; The second recognition module is used to identify the difference information of any target object in several dimensions using a preset difference recognition algorithm based on the spatial structure dynamic information. The difference information is the semantic information corresponding to the difference of the target object in any dimension. The difference recognition algorithms corresponding to the scene background, people, objects, robots and IoT devices, scene behavior events and scene macro-regulations are semantic segmentation and scene recognition algorithms, human posture estimation and behavior recognition algorithms, target detection and tracking algorithms, device state recognition algorithms, event detection algorithms and statistical analysis algorithms. A processing module is used to input the difference information of each target object in different dimensions into a preset world model, and output the target difference feature vector corresponding to the scene image set through the preset world model; The storage module is used to associate and store the target difference feature vector corresponding to the scene image set and the preset original scene data in the spatial intelligent machine.
Citation Information
Patent Citations
Target scene generation method and device, server, and storage medium
CN113887129A
Space-time context extraction method and system for real world model training
CN120451882A
Method for object- and scene related storage of image-, sensor- or sound sequences, involves storing image-, sensor- or sound data from objects, where information about objects is generated from image-, sensor- or sound sequences
DE102012014022A1
Method and apparatus for video frame sequence-based object tracking
US20040240542A1
Real-Time Digital Video Identification System and Method Using Scene Information
US20080313152A1
Cited By
First-view-angle drilling method and device based on memory enhancement and storage medium
CN121544793A