Method and device for object- and scene-related storage of image, sensor and / or sound sequences
By focusing on object and scene-related storage of image, sensor, and sound sequences, the method addresses the challenge of high data volumes, achieving reduced storage needs and improved data quality through efficient preprocessing and condensation.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- WASCHULZIK THOMAS DR
- Filing Date
- 2012-07-14
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for storing image, sensor, and sound sequences face challenges with increasing data volumes, leading to higher costs and storage requirements, while improvements in object recognition and parallel processing create a need for more efficient data preprocessing and condensation.
The method focuses on storing image, sensor, and sound data by recognizing objects and scenes, storing only modified information such as deviations from reference objects, and using multiple sensors to enhance data usability and reduce storage needs.
This approach significantly reduces storage requirements and costs by improving data quality and compression, enabling efficient object and scene-based processing and retrieval, with enhanced image and audio data usability.
Abstract
Description
[0001] The invention relates to a method and a device for object- and scene-related storage of image, sensor and / or sound sequences, in particular in microscopy, optical quality assurance, surveillance and security technology and in similar fields where the storage of image, sensor and / or sound sequences is carried out for later processing, analysis and / or visualization and playback.
[0002] It is known from the prior art that image and / or sound data are stored for processing in applications such as microscopy, optical quality assurance, surveillance and security technology, and the film industry. In the MPEG-4 standard, image sequences are stored, and storage is optimized by saving only individual images and the differences between them and subsequent images. With the ongoing improvement of image analysis algorithms and more cost-effective methods for parallel data processing, e.g., through multicore CPUs, increasingly better possibilities arise for recognizing objects contained in video information and determining the 2.5D or 3D shape and texture information of these objects from the video data.This makes it possible to directly assign image information to the captured real-world objects and to perform scene rendering for playback based on the information stored with the objects and scenes. A disadvantage is that in many applications, the increasing number of sensors leads to larger data streams and significantly higher costs for data transmission and storage. However, improvements in object recognition and the availability of increasingly affordable parallel processing create new conditions where it becomes more attractive to further preprocess and condense sensor information, also to improve its usability. Early segmentation and recognition of the objects contained in the data stream leads to earlier data refinement, as evaluation programs can access the data stream at a much higher level of abstraction.Scenes can be created from the objects detected in the sensor data, containing the detected objects, their spatial relationships, lighting, sensor positions, and temporal changes. If needed, the raw data can also be reconstructed from the objects, allowing for the detection of other objects without any detrimental information loss during the initial object detection and compression step. Thanks to the large number of available processing cores, it's possible, for example, to select cores specifically for detecting a subset of the total number of objects to be detected. The data stream is provided to all or several processing cores in parallel, and the algorithms running on these cores then search within this data stream for the objects for which the algorithms were developed or for which their parameters are suitable.There is a protocol for communication between the processes running on the information processing units, through which the detected objects, along with their position information and the corresponding image information, are reported back to the other information processing units. Based on this information, the other search processes can adjust their search areas within the sensor data. This allows them to either skip search areas where the sensor data has already been sufficiently explained, as no further objects are likely to be found there, or to process search areas more intensively because other objects are more likely to be found near certain objects.This approach makes it very easy to combine libraries for the recognition of different objects, which can recognize large numbers of objects, since only the protocol of communication between the computer cores needs to be standardized and the algorithms responsible for recognizing a subset of objects can be implemented largely independently of each other.
[0003] WO 97 / 13372 A2 discloses video encoding and decoding processes that enable the compression and decompression of digitized video signals representing display motion in video sequences spanning multiple frames. The encoder process employs object- or feature-based video compression to enhance the accuracy and versatility of encoding interframe motion and interframe image features. Video information is compressed relative to objects or features of arbitrary configurations, rather than fixed, regular pixel arrangements as in conventional video compression methods. This reduces error components, thereby improving compression efficiency and accuracy. The decoding process decompresses the encoded video information to reconstruct the objects or features of arbitrary configurations.
[0004] Based on this prior art, the invention aims to create a method and a device for object- and scene-related storage of image, sensor and / or sound sequences with a significant reduction in storage requirements for large amounts of data and with low or controllable information loss, while simultaneously improving system properties, so that objects can be recorded, stored and compared with databases more easily and in better quality, and at the same time a reduction in manufacturing costs for systems with large data volumes and multiple sensors can be achieved.
[0005] The method is intended for use in various technical fields, such as microscopy, optical quality assurance, surveillance and security technology, film technology, and similar areas, where image, sensor, and / or sound sequences are stored for later processing, analysis, and / or visualization. In particular, for object- and scene-related video information storage, the method is intended to improve image quality and reduce storage space requirements, so that objects and scenes enriched with extensive information are directly usable for interaction.
[0006] At the same time, the task should be expanded to include combining the use of multiple sensors with the networking of information processing of data streams and scene- and object-based processing to improve the scope and information content of the recorded data and the usability of the recorded information, thus also reducing the costs for acquiring and storing the data.
[0007] To solve these problems, the invention proposes a method for object- and scene-related storage of image, sensor, and / or sound sequences, characterized in that image, sensor, and / or sound data of objects, such as technical and / or natural objects, landscapes, rooms, material / tissue structures, humans, animals, plants, microorganisms, liquids, gases, or the like, are stored, and information about the objects is generated wholly or partially from the image, sensor, and / or sound sequences. A scene or sequence of scenes formed from these objects then contains information about the spatial relationships of the objects and / or the recording sensors to one another, as well as the environmental conditions, lighting, background, or background noise.The storage of image, sensor, and / or audio data is carried out in such a way that this data is checked for a possible, complete or partial, correlation with the objects. If a correlation is found, only the modified information—a new texture, a changed shape, a changed positional relationship, or a changed lighting—is stored, in addition to the correlation itself. This allows for the complete or partial reconstruction of image, sensor, and / or audio sequences of the recorded situation from the stored data. Finally, the detected objects are not always stored in their entirety. Instead, after the initial storage as a reference object, only the reference to the object, along with the deviations of the detected object from the reference object, is transferred or stored. The stored data can be updated over time if it is determined that certain properties have changed permanently.This creates a reference object that changes over time, and its changes can be tracked. The found objects are evaluated, and it is checked whether they are useful to the application.
[0008] Object deviations can include, for example, the object's orientation in space, a different projection direction, or the object's illumination, such as intensity, structure, direction, spectral composition, or the object's size or distortion—that is, a change in the object's shape or texture. The smaller the deviations between the reference object and the detected object, the higher the compression achieved. To achieve high compression, the change can be transmitted as a reference to the object's most closely matching state. This is useful, for example, in time sequences when only the object's changes since the last transmission need to be transmitted. This represents a significant compression advantage compared to current algorithms that transmit changes to the raw data.When an object is rotated, only the new rotation angle needs to be transmitted, e.g., the overall rotation; not all changed pixels need to be transmitted. If data compression is the primary goal, then the object hypothesis is accepted, for example, if this hypothesis results in data compression. In this case, it is helpful to continuously search for new potential objects, which may appear more frequently, using, for example, shape criteria. When a new, useful object is found, it is stored in the object database and used to compress the image and / or audio data stream. If, for example, the identification of objects or microbes and their temporal tracking—that is, the movement of the objects or microbes in time and space—is the primary focus in microscopy, then the recognition hypothesis is also stored in the data stream.If the focus in surveillance and security technology is on generating alarms, these alarms are triggered when a specific detection threshold is reached. The achieved compression factor for a portion of the data stream that can be attributed to an object then serves as a possible indicator of the object detection quality. To simplify object detection and tracking, high-quality and detailed object capture is emphasized in all situations where new objects can enter the context. A scene object is a set of scene objects with a hierarchical structure (e.g., scenes unfolding simultaneously in the rooms of a building) and a spatial and temporal relationship to one another (e.g., the scene sequences of a film). Furthermore, a scene has a spatial and temporal coordinate system into which a set of objects is arranged.A scene object consists of a set of sub-objects with a hierarchical structure and a set of properties (either static or with models of temporal change). The sensors and / or supporting systems (such as lighting) used to record the scene can be integrated into the scene description. The properties of a scene object are determined by its position in space, its spatial extent, its internal structure (possibly hierarchical), its external structure, its sensor properties, and its effectors. The internal structure of scene objects results from components, which can themselves be scene objects, and a deformation model of the object. The object's deformation model includes models of the object's movement, e.g.,Deformation of the human body during breathing, where the surface change as a whole is modeled using control points and interpolation, but not by modeling individual joints. For example, 2- or 3-dimensional polynomial meshes or splines are used to model the surfaces. The type of model is determined either from the recorded sensor data or obtained via object recognition if a subset of the sensor data can be successfully assigned to an object or object type. If no model could yet be assigned to the sensor data due to object recognition, a surface model is first created for the sensor data using control points. Then, after the data has been recorded, the sensor data is examined to determine whether it contains joints at specific locations, relative to the position and orientation within the object and the degrees of freedom with angles of movement.The object's deformation model also encompasses the gait of a living being or an artificial object. The motion model, along with other parameters or calculation methods, can also be obtained via object recognition. The external structure of a scene object consists of its surface, texture, and reflection properties. These properties are determined from sensor data using image processing methods or, after successful object recognition, extracted from the model knowledge available about the recognized object. The sensor properties of a scene object include the sensor orientation and the sensor type, for example, acoustic, optical, or geographic.
[0009] The sensor types differ in terms of spatial and temporal resolution, wavelengths or frequency ranges, bit depth, and sensor data formats. Determining the properties of a scene object with respect to the effectors depends, for example, on sound emission, radiation emission (e.g., for illumination or distance measurement), or the spatial changes of other objects due to their movement or deformation, or changes in other object properties.
[0010] Application areas for the targeted use of sensors or imaging techniques to simplify object recognition include, in optical microscopy and electron or particle beam microscopy, systematically changing the optical or electronic focal plane to capture all planes of an object and use them for recognition. Furthermore, the integration of other sensors (e.g., distance sensors, energy-dispersive X-ray spectroscopy, particle beam ablation, and mass spectrometers), multispectral imaging, increasing the acquisition frequency, systematically moving the camera or slide, varying the acquisition speed, or dynamically changing the resolution can be used when new objects enter the observed area, for example, through addition, cell division, or movement.
[0011] A significant reduction in storage requirements for large datasets is achieved by combining multiple sensors with different characteristics or by using information from previously recorded scenes. Information from sensor data acquired at a later time, such as details or spectral information, can be used to render earlier points in time on a virtual camera trajectory, provided the object is assumed to have remained unchanged during that period. To avoid misinformation, information acquired at a different time, for example in surveillance technology, can be highlighted using false colors so that the viewer can recognize that this information was not present in the original sensor data.The ability to capture, store, and compare objects with databases more easily and in higher quality is achieved by retrieving data from different recording times and locations. For example, in surveillance technology, if a person moves from one room to another, information about that person's face, captured in one room, can be retrieved for that person in the other. Data from different perspectives, with varying levels of detail, exposure times, and spectral ranges—such as the normal visible spectrum combined with infrared or ultraviolet information—are combined for scene rendering. This ensures that the best available and plausible information is always used.In this way, a significant improvement in data quality can be achieved with reasonable financial expenditure.
[0012] The advantage is that, due to the amount of data stored and the objects and / or scenes generated from the data stream, the image quality is improved by eliminating noise from individual camera images. This allows details that are no longer discernible due to the resolution limitations of a sensor or the camera's viewing angle in a specific camera setting to be superimposed onto the generated image from other images. This detailed information is then collected and stored in the reference object to be available for further processing. Details that are no longer detectable due to the limited intensity resolution of the camera sensor under the current image setting and lighting conditions are superimposed onto the reference object.
[0013] Preferably, to reduce storage space, redundant information contained in various image, sensor, and / or audio data is stored only once, as needed, in the reference object and / or in a scene of the object. The other data can then be reconstructed from knowledge about the objects and scenes, the lighting, the background, the orientation of the objects, and similar data, so that only deviations that are inconsistent with the model need to be stored in addition to the model itself. This ensures that improving the model of the depicted objects simultaneously increases the compression factor. The more high-quality information is available, the less redundant information needs to be stored.For this reason, the technical setup is designed to ensure the highest possible quality of the object's model in the reference object right from the start, when an object enters a scene. This additional effort during the initial recordings pays off through the higher compression factor later on and the improved data quality over the entire period. The objects and / or scenes of objects, enriched with extensive information, can be used directly for interaction, allowing a viewer to rotate objects within a scene and view them from different perspectives, in different resolutions, or in different spectral ranges. Alternatively, this information can be exported for external use (archiving, manipulation, comparison).The reference objects and / or scenes of reference objects, enriched with extensive information, can be used directly for interaction. By rotating the objects, 3D images can be viewed in different positions, meaning a viewer can see scenes of the objects from a different perspective, one not captured in the original image. In microscopy, for example, spatial relationships that were not discernible from the original image perspective can be reconstructed based on the model knowledge and the information gathered from the reference objects.Through object recognition, the results of which are stored in the reference object, a process for the automatic annotation of objects and scenes is performed. This is used, for example, to improve search functionality within image, sensor, and / or sound data, and also allows objects, scenes, or sounds to be stored with varying degrees of lossy compression depending on the type and significance of the object or scene. Based on the data collected in the objects, higher-dimensional representations can also be easily derived from the originally lower-dimensional sensor information. This allows for the addition of further spectral information in additional dimensions, or for the creation of three-dimensional camera images, e.g., for 3D films, from two-dimensional camera images using the information stored in the objects.
[0014] Furthermore, it is advantageously designed that, to improve compression, the information from object and scene recognition is used to store objects, scenes, or sounds with varying degrees of lossy compression depending on the type and significance of the object or scene. If one or more virtual camera / sensor trajectories are defined within a scene during recording planning, the data compression can also take into account that all information captured by the virtual camera is stored with a higher level of detail. The virtual camera trajectory can be defined in time and space. It can also contain additional information about the resolution, depth of field, frequency ranges in which sensor information is to be captured, and other aspects of the desired sensor information.This simplifies the recording process, as all sensors can independently optimize their settings and data stream for the planned trajectory, thus relieving the operator of this task. Scene aspects not captured by the virtual camera are either not saved or saved with a lower level of detail. Similarly, optical or other properties of the virtual camera can be defined and taken into account during data compression. It is advantageous if annotations created manually at one point or obtained through object recognition and assigned to an object are then used throughout the entire recording period. This makes automatic scene annotations simpler and more effective. For example,Building components detected by sensors can be assigned to the corresponding objects in the building plans. Alternatively, the name of a person, identified by face or access key, can be assigned to an object in the scene and retrieved upon user request (e.g., by pointing at an image of the person in a rendered scene), along with access permissions loaded from databases based on the person's identification. By linking objects and their information about their size, shape, surface texture, and annotations to the scene within their spatial context, alarms can be easily implemented. These alarms might be triggered, for example, if a large piece of luggage is left unattended against a supporting column or if a person in a shop puts merchandise into their clothing pocket.The flexibility of using the captured data for different applications is achieved by rendering the data not only from the sensor's perspective, as it was captured by a sensor. The viewer can move through a scene and view it from different angles. Across multiple images and sequences, the objects within a scene are digitized, enriched with image information, and a spatial model of the objects, including their surface properties such as texture, is built. This model can be augmented by object recognition if additional information from other images or other sources (annotations) is available for the recognized object. Based on this model knowledge, the scene can then be rendered from different perspectives.The user can manipulate the scene and move objects relative to each other, thus, for example, changing the script of a film. This can be done either manually or based on motion models underlying the objects, such as when a person walks across the room. The parameters for these motion models can also be generated from the image data. The user can modify objects within the scene, such as their surface texture. This allows certain aspects of scenes to be retouched by manipulating specific properties, such as the shape of an object via a reference object. This makes it easy and cost-effective to implement significant artistic changes for all subsequent scenes. The user can also replace objects within the scene, for example, by...Users can replace actors with themselves or people from their surroundings and personalize feature films. They can also create their own films as directors using the captured objects. The objects derived from the recordings can form the basis for other applications, such as games, making it very easy and cost-effective to create film-like computer games. However, in surveillance situations, users may be prohibited from precisely the aforementioned manipulations in order to collect data that can be used to support testimony in court. In scientific contexts, the information obtained can be used for model building.The information collected about an object can also be used to gather very detailed information about a perpetrator in a surveillance situation, so that a 2.5-dimensional image of his face and / or a movement model is created, which can be used for investigative purposes.
[0015] It is also preferred that a sensor image be divided into segments with different relevance or semantic relationships, so that a different compression level is used depending on the relevance when employing lossy compression methods, for example, low relevance and a high compression level, or high relevance and a low compression level with little or no information loss. The division of a sensor image is achieved, for example, by means of motion analysis or online object recognition, into a static background and a dynamic component of moving objects, so that the compression of object-related information also depends on the object or object type.
[0016] Increased throughput or easy separation of different object classes, scene components, or semantic relationships is achieved by storing the different compression levels or semantic relationships on separate storage media. This allows objects, scene components, or objects with a semantic relationship to be easily modified, replaced, or deleted without affecting other objects or being impacted by their data volume. Storing the different compression levels of object classes on separate storage media, such as disks, is advantageous for increasing throughput and enabling easy separation later. For example, the background or an object class can then be easily deleted without affecting the other objects.
[0017] Furthermore, it is advantageously designed that the compression factor is dynamically controlled for different object and scene categories to optimize the use of available storage media. If sufficient storage space is still available, a high level of detail can be achieved initially. If it becomes apparent that storage space is running low, the level of detail for the various object or scene categories can be easily reduced retrospectively, ensuring the best possible results for the application. Compression can also be enhanced with suitable application-specific filters. These filters can, for example, eliminate noise or details whose spatial size is smaller than a predefined threshold. Other filters can remove specific frequencies from the audio or spectral information from the image data.The filters and filter parameters depend on detected objects, the spatial relationships of objects within a scene, the camera's current viewing angle relative to a planned virtual camera trajectory (for later playback of the scene), the size of the spatial details, the frequency of the image information (elimination of noise or invisible aspects), the frequency of the audio information, and the current lighting conditions. Controlling the compression factor is advantageous for optimally utilizing available data transmission paths or storage media, such as main memory and disks. The ring buffer in a video surveillance system can then be configured so that data about moving objects in the room is stored for much longer than data about the background or static objects.
[0018] Backgrounds or static objects can then be inserted during playback from an object and scene memory or from other sequences, e.g., with false color coding, so that it is clear that the data is not the original but rather information generated from other sensor data. The cost of data storage can be significantly reduced through a variety of methods, such as reducing redundancy. Image information about an object contains considerable redundancy, as the same information is stored repeatedly. MPEG encoding exploits the fact that the similarity between successive images is used for compression.By leveraging the similarity of information belonging to an object for compression across different scenes and sensors, a completely new dimension of compression emerges. This makes the use of massively parallel camera systems economically viable, as conventional storage methods are extremely costly and resource-intensive due to data overload. In the new method presented here, compression is improved by increasing the number of sensors, as the quality of the object model is enhanced. This improved model allows for even better assignment of sensor data to the object, thereby improving both the quality of the scene digitization and the achievable compression ratio.Furthermore, more information about the surface properties of the objects, their motion patterns, and the lighting conditions can be obtained, thereby improving the model and allowing a larger portion of the information captured by the sensors to be correctly attributed to the objects, the lighting conditions, or sensor noise. The noise in the sensor data can then be reduced, improving image quality and decreasing storage requirements. This reduces the data stream that needs to be stored to reconstruct the original sensor information. A distinction can also be made between primary and auxiliary sensors. The data from the primary sensors is stored in such a way that their sensor data stream can be reconstructed, while the auxiliary sensors are used only to, for example, better capture the lighting conditions or to record additional object details from different perspectives.The cost of data storage can be further reduced through context-specific compression levels. By segmenting objects and assigning sensor data to specific objects, object recognition is significantly simplified, and its quality can be improved. Scenes can then be easily and reliably annotated, and data compression can be controlled according to the context. For example, data acquisition in a surveillance situation where only the lighting changes and no objects move can be dramatically reduced. However, after a person has entered the room, it often needs to be re-digitized to determine if anything has been removed or added.If no changes were detected, the data can be discarded after the information has been stored and re-digitized.
[0019] Content-based control of storage location is achieved by allowing decisions to be made for each individual object, depending on the scene being displayed, specifying the level of detail and location for storage. Selecting specific storage locations ensures that a guaranteed bandwidth is available for digitizing and storing certain objects, and that selected information, such as background or lighting details, can be selectively deleted, copied, or modified without affecting or being affected by other information. A wide range of parameters is available to control data compression according to the specific usage situation.
[0020] It is still preferred that a library of frequently occurring objects be established to improve object recognition and segmentation, that the object library contain models and / or special algorithms for object identification, and that Gestalt laws such as the law of Prägnanz, proximity, similarity, continuity, closure, common motion, continuous line, common region, simultaneity, or connected elements be used for this purpose. In addition, methods that evaluate the spatial proximity of characteristic features can also be used.
[0021] Another advantage of the method for object- and scene-related storage of image, sensor, and / or sound data is that the acquisition of sensor information is influenced by the quality of previous digitization. This means that images of highly detailed aspects are captured that have not yet been recorded, or images of less detailed aspects are captured if these have already been digitized very accurately and the aspects match the reference model. The acquisition of object aspects can also depend on the result of object recognition. For example, one would want to capture as many details as possible of a wanted criminal, while only a few details would be captured of a known chair that is of little interest to a particular application.
[0022] The fundamental innovation of this invention lies in the fact that the detected object and scene become the focus of storage and processing, rather than the sensor data stream as before. This shift in the central reference point for storage opens up entirely new possibilities and features for working with the stored data. For example, data from multiple cameras can be stored for a single object. Furthermore, decisions regarding the storage and processing of information can be made based on knowledge about the object.
[0023] The device for carrying out the method is characterized in that it has one or more cameras for recording in parallel or sequentially, wherein the data from several sensors can be integrated into a scene model and wherein the cameras for automatically capturing the data quantities have a wide range of viewing angles, focus planes, resolution levels or a wide spectral bandwidth.The cameras are arranged in such a way that the sequence of objects entering a monitored area can be recorded with a high level of detail. Multiple information processing units analyze the data stream from one or more cameras in parallel, and the results of this analysis are exchanged between these units for object recognition. Information about the cameras' orientation in space is also used, and data compression or preprocessing is performed prior to object recognition to correlate the information emanating from a single object. The parallel use of multiple cameras allows for the simultaneous recording of different viewing angles, zoom levels, focus settings, and spectral ranges.As the number of sensors / cameras used in parallel increases, so does the attractiveness of this method, since all sensors / cameras can integrate their information into the same scene representation and mutually benefit from the objects already identified in the data streams of the other sensors by combining multiple sensors with different properties. Examples of sensors include microphones, cameras that simultaneously determine the distance of pixels, cameras with flexible focus, autofocus, apertures, variable lenses with respect to focal length, magnification, and depth of field, radar sensors, ultrasonic sensors, laser scanners, microwave radiometers, distance sensors, and receivers for coded recognition signals of objects (e.g., RFID) or sub-objects.Sensors capable of identifying the chemical composition of objects, such as energy-dispersive X-ray spectroscopy or mass spectrometry, can also be used. For this purpose, microscopically small samples of the objects are taken using special sensors and then analyzed in detail for their composition. The sensor data is initially stored in the sensor object's data stream. The sensor data stream is then analyzed, and the sensor data is assigned to the objects. Thus, image data points captured by a person are no longer assigned to the capturing sensor but to the person captured. During this transition, the sensor data stream can be updated to indicate where the sensor data has been assigned, so that the complete sensor data stream can be reconstructed.To assign sensor data to objects, all available information is used whenever possible to ensure a fast, reliable, and accurate assignment. This could be, for example, a transmitter that indicates the position, movement, and / or orientation of a scene object in space, or the object's position in the scene model relative to other scene objects, or sensor information that enables object recognition, leading to an update of the scene object's position, orientation, texture, color, shape, or similar attributes within the scene.This can also include information about the movement and / or deformation of objects in space, which leads to the expectation that the next sensor data from an object is coming from a different direction or with a different orientation. It can also include information from registration algorithms indicating that sensor data assigned to a scene object is now located in a different sensor data area, or that the scaling of the sensor data has changed (e.g., when the scene object approaches the camera). From the information about the orientation and movement of the sensor object, as well as changes in the sensor's properties (e.g., a change in the camera's zoom factor), the new relative position of the scene object can then be determined. This position is then stored, if necessary, after comparing the information with that of other sensors in the scene. If an object is deliberately moved, e.g.,For targeted, high-dimensional data collection during films or games, or when entering or leaving a monitored area, the data is collected using technical measures, allowing the measurements to be reliably assigned to the individual object. This approach yields the highest data quality. For example, the object's weight can be determined to ascertain whether a person has gained more weight during their visit than would be expected based on the weight of purchased items recorded at the checkout. This could trigger an intelligent alarm to investigate the individual further regarding a possible theft. Similarly, an unexpected weight loss could indicate the planting of a bomb.
[0024] The present method according to the invention and the apparatus for carrying out the method can be used wherever large amounts of data need to be processed and a significant reduction in storage requirements for large sensor data sets leads to a reduction in costs, with minimal or controllable information loss and a simultaneous improvement in system properties. In microscopy, the method according to the invention enables the image to be divided into several parts with different relevance. Depending on the relevance, a different compression level is used when employing lossy compression methods: a high compression level for low relevance and a low compression level for high relevance with little or no information loss. The image is divided into objects by means of online object recognition.The different compression levels are stored on separate storage media to increase throughput. Improved recording and distribution of the captured image data enhances search functionality, and the compression levels are controllable, allowing for optimal utilization of available buffer memory, such as main memory and disk space. This method enables the use of virtual reality techniques, providing users with new perspectives. 3D scenarios and 2D or 3D films can be calculated from 3D object information and current 2D images or 3D images with definable resolution. The present invention further achieves a significant reduction in storage requirements for large image datasets in microscopy, with minimal or controllable information loss, while simultaneously improving system performance.This reduces the manufacturing costs for systems with large data volumes by saving on storage media. The slightly higher costs for the CPUs will become less significant in the future due to the availability of inexpensive multicore CPUs. The image analysis algorithms are designed to fully exploit the advantages of multicore CPUs. The improved performance increases the attractiveness of the systems, as only smaller amounts of data need to be processed after compression. Recording and content-based search are improved, and tracking within the sensor image streams is facilitated. Overall, the shift from sequence-based image storage to object- and scene-based storage with a high compression ratio represents a major innovation.In security and surveillance technology, as well as in the film industry, object recognition is simplified by implementing technical measures at potential access points to the monitored area. These measures enable the highest possible quality and most comprehensive initial object capture. This can be achieved, for example, by installing camera systems in entrance doors, barriers, or in the entrance area, or by using structured lighting to capture the surface structure of moving object parts and their typical angles of movement. Further methods for improving the initial object information can be achieved through targeted variations in focus, magnification, lighting, intensity, direction, recording frequency, and resolution.Additionally, when new objects enter the monitored area, a search can be performed in the existing object database to load information about the incoming object. Using object- and scene-based storage of image and / or audio data in security and surveillance technology reduces the required storage capacity. This is because the system stores less redundant information about the objects or the scene itself, including deviations of sensor data from the current model, information about the objects in the space, their relationships to one another, changes in their position, and lighting conditions. Furthermore, the transition to object-based information improves the quality of the image and audio data.It is not only possible to use the information contained in a short video sequence, but also to correlate data from different cameras, such as the time of recording, viewing direction, lighting conditions, and similar data. For each camera, the position, viewing direction, and focus in space, degrees of freedom of movement of the camera in space, degrees of freedom of changes in the optical properties of the cameras in space, and the position and properties of the light sources in space are stored. Objects include static objects, such as landscapes, buildings, and furnishings, as well as dynamic objects, such as people, animals, moving objects, liquids, and gases.For each object, the object's prototype, spectral and intensity information in high-dimensional format, and references to other information retrieved from other databases via object recognition or an identification card are stored. Furthermore, typical movement patterns of objects within the rooms can be recorded, and deviations from these patterns can be used to trigger intelligent alarms. Storing data on an object-specific basis allows for much simpler and more accurate comparison and correlation of objects with databases. This is relevant for object-specific alarms, such as when a large object is placed against a load-bearing column of the building, or for tracking the temporal progression of objects like clothing or a suitcase being picked up or put down, or when the shape or weight of a bag changes unexpectedly.The objects can also be equipped with motion models so that a person can be recognized by the system, even if they are walking, standing, or sitting. The parameters of the motion model, such as position, degrees of freedom and orientation of joints, typical directions of movement, and speeds of movement, can either be stored in the object model or also generated from the sensor data. Object- and scene-specific storage also allows for higher-quality recording and storage of the acquired data. Data storage can be more compact, and redundancy can be reduced. Data acquisition is selectively controlled to complete the entire dataset by utilizing the degrees of freedom of the cameras, viewing direction, zoom, focus, spectral properties, lighting, and obstacles.To improve system performance, it is advantageous to capture high-quality images of objects as they enter and leave the monitored area. The better the initial information about an object, the easier it is to later register it with additional image data. Tracking objects through a monitored area should be accomplished using varying cameras and viewing angles. This allows for early assessment of an object. The earlier high-quality information is available, the less processing time is required for data registration, and the more effectively the data can be aggregated. High-quality sensors only need to be installed in the entry and exit areas and, if necessary, at particularly relevant locations. In other areas, high-quality object information within the sensor's range can be extracted from stored object data.The operating personnel and the automated algorithms have access to better information over a longer period. Special entry barriers are designed that allow comprehensive digitization of entries into and exits from the monitored area. These barriers guide an entering person or object past a sensor or group of sensors in a targeted manner, enabling comprehensive recording and, if possible, detection. The weight of a person or object can also be recorded upon entering and exiting the secured area. The system can then focus on tracking the object and periodically verifying the sensor data. By using the inventive method and the device for implementing the method in security and surveillance technology, the costs for image data storage can be reduced by approximately 70%.A new 2.5D wanted poster can be easily and securely compared with other wanted posters. Alarms are triggered when specific objects are removed from or introduced into the monitored area. This new method for object- and scene-related storage of image, sensor, and / or sound sequences significantly improves the scope and information content of the data obtained, while reducing the costs of data acquisition and storage.
[0025] The invention is not limited to the application examples mentioned above, but can be used in any area requiring high data storage capacity. In particular, it also includes variants that can be formed by combining the features described in connection with the present invention. All features mentioned in the preceding description are further components of the invention, even if they are not specifically highlighted and mentioned in the claims.
Claims
[1] Methods for object- and scene-related storage of image, sensor and / or sound sequences, where objects are found in the image, sensor and / or sound sequences and information about the objects is generated, a scene is formed from these objects, which contains information about the spatial relationships of the objects and / or the recording sensors to each other, where image, sensor and / or sound data are stored in such a way that the image, sensor and / or sound data are checked for a possible full or partial assignment to the objects found, where, in the case of a possible match, only the changed information is stored in addition to this match, Finally, the found objects are not always stored as a whole, but after the initial storage as reference objects, only the reference to the found object along with the deviations of the The found object is transferred from or stored to the reference object, thereby creating a spatial model of the objects, and where 3-dimensional polygon meshes or splines are used to model the surfaces. [2] Method according to claim 1, wherein, based on the amount of data stored and the objects and / or scenes of objects generated from the data stream, the image quality is improved by eliminating the noise of individual camera images in such a way that the details which are no longer recognizable due to the resolution limitation of a sensor or the viewing angle of a camera in a certain camera setting are superimposed from other shots into the generated image, and this detail information is collected and stored in the reference object in order to be available for further processing. [3] Method according to one of claims 1 and 2, wherein the details which can no longer be detected due to the limited intensity resolution of the camera sensor at the current image setting and illumination are superimposed from the reference object. [4] Method according to any one of claims 1 to 3, wherein, in order to reduce the storage space requirement, the redundant information contained in various image, sensor, and / or sound data is stored only once as required in the reference object and / or in a scene of the object, the other data is reconstructed from knowledge about the objects and scenes of the objects, the lighting, the background, the orientation of the objects and similar data, so that only the deviations that are not consistent with the model need to be stored in addition to the model. [5] Method according to any one of claims 1 to 4, wherein the objects and / or scenes of objects enriched with extensive information are directly usable for interactions, so that a viewer can rotate objects in a scene and view them from other perspectives, in other resolutions or in other spectral ranges, or this information is to be exported for external use. [6] Method according to any one of claims 1 to 5, wherein a viewer of objects and / or scenes of objects can view them from a perspective that was not captured in the original version, and wherein the positional relationships of the objects to each other and the calculated image data for this new perspective, which were not recognizable from the original perspective of the recording, are determined on the basis of the model knowledge and the information collected in the reference objects. [7] Method according to any one of claims 1 to 6, wherein an automatic annotation of objects and / or scenes of objects is performed, which is used both to improve the search functionality within the image and / or sound data, and also allows the storage of objects, scenes or sounds in different lossy compression depending on the type and importance of the object or scene. [8] Method according to any one of claims 1 to 7, wherein the division of a sensor image into segments with different relevance or with different semantic relationships is carried out, so that a different compression level is used depending on the relevance when using lossy compression methods, and that the division of a sensor image into a static background and a dynamic component of moving objects is carried out by means of motion analysis or online object recognition, so that compression of the object-related information also depends on the object or object type. [9] Method according to any one of claims 1 to 8, wherein, to increase throughput or to easily separate different object classes, scene components or semantic relationships, the different compression levels or semantic relationships are stored on different storage media, so that objects, scene components or semantic relationships can be easily changed, replaced or deleted without affecting the other objects or being affected by their data volume. [10] Method according to one of claims 1 to 9, wherein, in order to make targeted use of the storage media available for the recordings, a dynamic control of the compression factor is carried out for the different object and / or scene categories, so that filters and filter parameters depend on recognized objects, on positional relationships of objects in a scene, on the positional relationship of the viewing angle in relation to a planned virtual camera trajectory, on the size of the spatial details, on the frequency of the image information, on the frequency of the sound information or on the current lighting situation. [11] Method according to any one of claims 1 to 10, wherein a library of frequently occurring objects is established to improve the recognition and segmentation of objects and that the object library contains models and / or special algorithms for identifying the library objects. [12] Method according to any one of claims 1 to 11, wherein the sound information emanating from an object is stored in that object and that the sound information is also synthesized according to the viewing angle when rendering a scene, so that the spatial assignment of the sound information when played back from other viewing angles can correspond to the actual spatial conditions in the recordings. [13] Method according to any one of claims 1 to 12, wherein the acquisition of sensor information is controlled by the quality of the previous digitization or by the recognized objects or the semantic relationships in a scene, in order to selectively obtain details or aspects relevant to the application. [14] Device for carrying out the method according to claims 1 to 13, wherein the device has one or more cameras for recording in parallel or sequentially, wherein the data from several sensors can be integrated into a scene model, wherein the cameras for automatically capturing the data quantities have a wide range of viewing angles, focus planes, resolution levels or a wide spectral bandwidth and wherein the cameras are arranged relative to each other in such a way that the recording of objects entering a monitored area is initially provided with a high level of detail. [15] Device for carrying out the method according to claim 14, wherein several information processing units analyze the data stream from one or more cameras in parallel, wherein the results of which are exchangeable between the information processing units for object recognition, and wherein the information about the orientation of the cameras in space is used to provide data compression or preprocessing prior to object recognition in order to associate the information emanating from an object with each other.
Citation Information
Patent Citations
Feature-based video compression method
WO1997013372A2