Systems and methods for multi-modal neural symbolic scene understanding
By integrating image and sound information through a multimodal neural symbol scene understanding system, the problem of insufficient sensor modal data fusion is solved, enabling efficient detection and reasoning of events in complex environments and generating control commands to cope with different scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2022-02-28
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies struggle to effectively integrate data from multiple sensor modalities to achieve scene understanding, particularly in their insufficient ability to detect and infer complex events within an environment.
A multimodal neural symbol scene understanding system is adopted, which captures image and sound information through sensors such as cameras and microphones, extracts data features using encoders, and combines spatiotemporal inference engines and metadata to determine the scene and output control commands.
It enables event detection and reasoning in complex environments, improving the accuracy and robustness of scene understanding, and can generate control commands to meet the needs of different scenarios.
Smart Images

Figure CN114972727B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image processing using sensors such as cameras, radar, microphones, etc. Background Technology
[0002] The system can be capable of performing scene understanding. Scene understanding can refer to the system's ability to infer the objects and the events they participate in based on the semantic relationships between objects and other objects in the environment and / or the geospatial or temporal structure of the environment itself. The basic goal of scene understanding tasks is to generate statistical models that can predict (e.g., classify) high-level semantic events given some observations of the context within a given scene. Observation of the scene context can be enabled by using sensor devices placed in various locations, which allow sensors to obtain contextual information from the scene in the form of sensor modalities, such as video recordings, acoustic patterns, ambient temperature time-series information, etc. Given such information from one or more modalities (e.g., sensors), the system can classify events initiated by entities in the scene. Summary of the Invention
[0003] According to one embodiment, a system for image processing includes: a first sensor configured to capture at least one or more images; a second sensor configured to capture sound information; and a processor communicating with the first and second sensors, wherein the processor is programmed to receive one or more images and sound information, extract one or more data features associated with the image and sound information using an encoder, output metadata to a spatiotemporal inference engine via a decoder, wherein the metadata is derived using the decoder and one or more data features, one or more scenes are determined using the spatiotemporal inference engine and the metadata, and control commands are output in response to the one or more scenes.
[0004] According to a second embodiment, a system for image processing includes: a first sensor configured to capture a first set of information indicating an environment; a second sensor configured to capture a second set of information indicating the environment; and a processor communicating with the first and second sensors. The processor is programmed to receive the first and second sets of information indicating the environment, extract one or more data features associated with image and sound information using an encoder, output metadata to a spatiotemporal inference engine via a decoder, wherein the metadata is derived using the decoder and the one or more data features, one or more scenes are determined using the spatiotemporal inference engine and the metadata, and control commands are output in response to the one or more scenes.
[0005] According to a third embodiment, a system for image processing includes: a first sensor configured to capture a first set of information indicating an environment; a second sensor configured to capture a second set of information indicating the environment; and a processor communicating with the first and second sensors. The processor is programmed to receive the first and second sets of information indicating the environment, extract one or more data features associated with the first and second sets of information indicating the environment, output metadata indicating the one or more data features, determine one or more scenes using the metadata, and output control commands in response to the one or more scenes. Attached Figure Description
[0006] Figure 1 A schematic diagram of the monitoring setup is shown;
[0007] Figure 2 This is an overview system diagram of a wireless system according to embodiments of the present disclosure;
[0008] Figure 3A This is the first embodiment of a computational pipeline;
[0009] Figure 3B This is an alternative implementation of a computational pipeline that utilizes the fusion of sensor data;
[0010] Figure 4 This is an illustration of an example scene captured from one or more video cameras and sensors. Detailed Implementation
[0011] Embodiments of this disclosure are described herein. However, it should be understood that the disclosed embodiments are merely examples, and other embodiments may take various forms and alternative forms. The figures are not necessarily to scale; some features may be enlarged or minimized to show details of particular components. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching those skilled in the art to adopt the embodiments in various ways. As will be understood by those skilled in the art, various features illustrated and described with reference to any of the figures may be combined with features illustrated in one or more other figures to produce embodiments not explicitly illustrated or described. The combinations of illustrated features provide representative embodiments of typical applications. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of this disclosure may be desired.
[0012] According to an embodiment, the embodiment includes a framework for multimodal neurosymbolic scene understanding. This framework may also be referred to as a system. The framework may include a combination of hardware and software. From the hardware perspective, data (“modalities”) from various sensor devices flows to software components via wireless protocols. From there, initial software processes combine and transform these sensor modalities to provide predictive context for further downstream software processes, such as machine learning models, artificial intelligence frameworks, and web applications for user localization and visualization. These components of the system collectively enable scene understanding, environmental event detection, and inference paradigms, where sub-events are detected and classified at a lower level, more abstract events are inferred at a higher level, and information at both levels is made available to the operator or end user, even though the probability of events spans arbitrary time periods. Because these software processes fuse multiple sensor modalities together, may include neural networks (NNs) as event prediction models, and may include symbolic knowledge representation and reasoning (KRR) frameworks as temporal inference engines (e.g., spatiotemporal inference engines), the system can be said to perform multimodal neurosymbolic inference for scene understanding.
[0013] Figure 1 A schematic diagram of a monitoring facility or setup 1 is shown. Monitoring facility 1 includes a monitoring module arrangement 2 and an evaluation device 3. The monitoring module arrangement 2 includes multiple monitoring modules 4. The monitoring module arrangement 2 is arranged on the ceiling of the monitoring area 5. The monitoring module arrangement 2 is configured for visual, image-based, and / or video-based monitoring of the monitoring area 5.
[0014] In each case, the monitoring module 4 includes multiple cameras 6. Specifically, in one embodiment, the monitoring module 4 may include at least three cameras 6. The cameras 6 may be configured as color cameras, and particularly as compact cameras, such as smartphone cameras. The cameras 6 may have a viewing direction 7, a viewing angle, and a field of view 8. The cameras 6 of the monitoring module 4 are arranged with similarly aligned viewing directions 7. In particular, the cameras 6 are arranged such that in each case, the cameras 6 have overlapping fields of view 8 on a pair-by-pair basis. The monitoring cameras 6 may be arranged in a fixed position within the monitoring module 4 and / or arranged at fixed camera intervals between each other.
[0015] In one embodiment, the monitoring modules 4 can be coupled to each other mechanically and via a data communication connection. In another embodiment, a wireless connection can also be utilized. In one embodiment, the monitoring module arrangement 2 can be obtained through the coupling of the monitoring modules 4. One monitoring module 4 of the monitoring module arrangement 2 is configured as a collection transmission module 10. The collection transmission module 10 has a data interface 11. The data interface can specifically form a communication interface. Monitoring data from all monitoring modules 4 is supplied to the data interface 11. The monitoring data includes image data recorded by the camera 6. The data interface 11 is configured to jointly supply all image data to the evaluation device 3. For this purpose, the data interface 11 can be coupled to the evaluation unit 3, particularly via a data communication connection. The monitoring modules can communicate via a wireless data connection (e.g., Wi-Fi, LTE, cellular, etc.).
[0016] By utilizing surveillance facility 1, moving objects 9 can be detected and / or tracked within surveillance area 5. For this purpose, surveillance module 4 supplies surveillance data to evaluation device 3. Surveillance data may include camera data and other data acquired from various sensors monitoring the environment. Such sensors may include hardware sensor devices comprising any one or a combination of: ecological sensors (temperature, pressure, humidity, etc.), visual sensors (surveillance cameras), depth sensors, thermal imagers, location metadata (geospatial time series), wireless signal receivers (WiFi, Bluetooth, UWB, etc.), and acoustic sensors (vibration, audio), or any other sensors configured to collect information. Camera data may include images of surveillance area 5 monitored using camera 6. Evaluation device 3 may, for example, evaluate and / or monitor surveillance area 5 in a stereoscopic manner.
[0017] Figure 2 This is an overview system diagram of a wireless system 200 according to an embodiment of the present invention. In one embodiment, the wireless system 200 may include a wireless unit 201 for generating and transmitting Channel State Information (CSI) data or any wireless signals and data. In a monitoring scenario, the wireless unit 201 may communicate with the mobile device (e.g., a cellular phone, wearable device, tablet computer) of employee 215 or customer 207. For example, employee 215's mobile device may send wireless signal 219 to wireless unit 201. Upon receiving a wireless packet, system unit 201 obtains the associated CSI value or any other data associated with the packet reception. Furthermore, the wireless packet may contain identifiable information about the device ID, such as a MAC address used to identify employee 215. Therefore, system 200 and wireless unit 201 may determine various hotspots without utilizing data exchanged from employee 215's device.
[0018] While WiFi can be used as a wireless communication technology, any other type of wireless technology can also be utilized. For example, Bluetooth can be used if the system can obtain CSI from a wireless chipset. As shown in wireless units 201 and 203, the system unit can be able to include a WiFi chipset with up to three antennas attached. Wireless unit 201 may include a camera that monitors various people walking around the POI. In another example, wireless unit 203 may not include a camera and may only communicate with mobile devices.
[0019] System 200 can cover various aisles (in other environments), such as 209, 211, 213, and 214. Aisles can be defined as walking paths between shelves 205 or store walls. Data collected between the various aisles 209, 211, 213, and 214 can be used to generate heatmaps and monitor store traffic. The system can analyze data from all aisles and use this data to identify traffic in other areas of the store. For example, data collected from the mobile devices of various customers 207 can identify areas of the store receiving high traffic. This data can be used to place certain products. By utilizing this data, store managers can determine the location of high-traffic properties relative to low-traffic properties.
[0020] CSI data can be transmitted in packets found in wireless signals. In one example, wireless signal 221 can be generated by customer 207 and their associated mobile devices. System 200 can use various information found in wireless signal 221 to determine whether customer 207 is an employee or other characteristics. Customer 207 can also communicate with wireless unit 203 via signal 222. Furthermore, packet data found in wireless signal 221 can communicate with either wireless unit 201 or unit 203. Packet data in wireless signals 221, 219, and 217 can be used to provide information related to motion prediction and traffic data related to employee and customer mobile devices, etc.
[0021] While the wireless transceiver 201 can transmit CSI data, it can also utilize other sensors, devices, sensor streams, and software. These hardware sensor devices include any one or a combination of the following: ecological sensors (temperature, pressure, humidity, etc.), visual sensors (surveillance cameras), depth sensors, thermal imagers, location metadata (geospatial time series), wireless signal receivers (WiFi, Bluetooth, UWB, etc.), and acoustic sensors (vibration, audio), or any other sensors configured to collect information.
[0022] The various embodiments described can be based on a distributed messaging and application platform that facilitates communication between hardware sensor devices and software services. This embodiment can interface with hardware devices via a network interface card (NIC) or other similar hardware. These hardware sensor devices include any one or a combination of: ecological sensors (temperature, pressure, humidity, etc.), visual sensors (surveillance cameras), depth sensors, thermal imagers, location metadata (geospatial time series), wireless signal receivers (WiFi, Bluetooth, UWB, etc.), and acoustic sensors (vibration, audio), or any other sensors configured to collect information. Signals from these devices can flow across platforms as time-series data, video streams, and audio segments. The platform can interface with software services via application programming interfaces (APIs), enabling these software services to consume sensor data and transform it into data understandable across multiple platforms. Some software services can transform sensor data into metadata, which can then be provided to other software services as an auxiliary "view" or information of the sensor data. Building Information Modeling (BIM) software components illustrate this operation, taking user location information as input and providing contextualized geospatial information as output; this includes the user's proximity to objects of interest in the scene, which is crucial for spatiotemporal analysis performed by symbolic reasoning services (as described in more detail below). Other software services can consume both the raw and transformed data to make final predictions about scene events or generate environmental control commands.
[0023] Any communication platform providing such streaming facilities can be used in various embodiments. The system can allow manipulation of the resulting sensor data streams, predictive modeling based on those sensor data streams, visualization of actionable information, and spatially and temporally robust classification and disambiguation of scene events. In one embodiment, the "Security and Safety Things (SAST) platform" can be used as the communication platform underlying the system. In addition to the aforementioned utilities, the SAST platform can also be a mobile application ecosystem (Android), along with APIs that interface these mobile applications with sensor devices and software services. Other communication platforms can also be used for the same purpose, including but not limited to RTSP, XMPP, and MQTT.
[0024] A subset of software services in the system can be responsible for consuming and utilizing metadata about the sensors, raw sensor data, and state information about the overall system. After such raw sensor data is collected, it can be preprocessed to filter out noise. Additionally, these services can transform the sensor data to (i) generate machine learning features that predict scene events and / or (ii) generate control commands, alarms, or notifications that will directly affect the state of the environment.
[0025] Predictive models can utilize one or more sensor modalities as input, such as video frames and audio segments. The initial components of the predictive model (e.g., an "encoder") can perform a single-peak signal transformation on each modal input, producing as many intermediate features as were present at the start of the input modality. These features are state matrices composed of numerical values, each representing a functional mapping from the observed feature representations. In summary, all feature representations of the input can be characterized as a statistical embedding space, which expresses high-level semantic concepts as statistical patterns or clusters. Figure 3A and Figure 3B A depiction of such a computational pipeline is shown.
[0026] The embedding space of a unimodal mapping can be statistically coordinated (i.e., subject to conditions) in order to align two modes or impose constraints on one mode on another.
[0027] Alternatively, feature matrices from modalities can be summed, concatenated, or used to find their outer products (or equivalents); the results of these operations are then subjected to further functional mapping—this time, to a joint embedding space. Figure 3B The computational pipeline for such a method is illustrated. Using the final component of the predictive model (i.e., the "decoder"), samples from these embedding spaces (coordinated features, joint features, etc.) are then paired with labels and used for downstream statistical training and inference, such as event classification or control.
[0028] Examples of sensing, prediction, and control techniques that can be utilized in the embodiments include occupancy estimation using depth-based sensors, object detection using depth sensors, indoor occupant thermal comfort using body shape information, HVAC control based on occupancy trajectories, coordination of thermostatic control loads based on local energy use and the power grid, and time-series monitoring / prediction of future indoor thermal environmental conditions. All of these techniques can be integrated into a neurosymbolic scene understanding system to enable scene representation based on categorized events or to achieve environmental alteration. Many such statistical models exist as software services within the system, where the properties of the inputs, outputs, and intermediate transformations are determined by the type of target event to be predicted.
[0029] To enable temporally robust scene understanding in the system, the system may include a semantic model comprising (1) a domain ontology of indoor scenes (“DoORS”) and (2) a scalable set of inference rules for predicting human activity. A server such as the Apache Jena Fuseki server may be utilized and run in the backend to maintain (1) and (2): receiving sensor-based data from various sensors (e.g., SAST Android cameras), including Building Information Modeling (BIM) information, appropriately instantiating the DoORS knowledge graph, and sending the results of predefined SPARQL queries to the frontend, where predicted activity is overlaid on a live video feed.
[0030] First, the system can build a dataset of actions performed within a context of interest. The system can analyze certain activities that are agnostic to a wide variety of contexts, such as airports, shopping malls, retail spaces, and dining environments. Activities of interest may include "eating," "working on a laptop," "picking up an object from a shelf," "inspecting items in a store," etc.
[0031] A central concept in one embodiment could be an event-scene concept, defined as a subtype of scene that focuses on events occurring within the same spatiotemporal window. For example, “a soda can be taken from the refrigerator” could be modeled as a scene that includes human-centric events such as (1) “facing the refrigerator,” (2) “opening the refrigerator door,” (3) “reaching out one’s arm,” and (4) “grabbing the soda can.” Clearly, these events are temporally related: (2), (3), and (4) occur sequentially, while (1) continues for the entire duration of the preceding sequence (facing the refrigerator is the condition for interacting with an item placed in the refrigerator). In this way, the system can jointly model scenes as meaningful sequences (or combinations) of individual atomic events.
[0032] Beyond representing event scenarios, a key to enabling human activity prediction is incorporating sensor-based observations into the ontology. Specifically, the key observation type for use cases is based on the concept of distance; given a set of furniture in the scene—whose corresponding locations are known prior according to the corresponding BIM model—and the real-time location of people in the scene, DoorOS can be used to infer human activity based on proximity. For example, a person standing next to a coffee machine, extending their arm, is (likely) making coffee, and certainly not washing dishes at a distant sink.
[0033] Distance observations typically involve at least two physical entities (defined in the scene ontology by the class features of interest) and a metric. Because OWL / RDF's expressive power is insufficient for defining n-ary relations, in DoorORS, the system can materialize "distance" relations. For example, the system can create a class "Person_CoffeeMachine_Distance" whose instances take a person and a coffee machine as participants (both provided with unique IDs), and whose metric is associated with a precise numerical value indicating meters. Materialization is a widely used approach to achieve a trade-off between the complexity of the domain and the relative expressive power of ontology languages. In DoorORS, evaluating who is closest to the coffee machine at a given time, or whether a person is closer to the coffee machine than to other known elements in the room, is transformed into an observation identifying the minimum distance between a given person and a furniture element or bounded object. Note that the shortest distance between a person and an environmental element is "0," meaning that the object's (transformed) 2D coordinates fall within the coordinates of the bounding box of the person being considered.
[0034] As explained above, distances between people and environmental elements (like furniture or objects) are observed, measured in meters, and occur at specific times. When multiple people and environmental elements are present in a scene, distances are always represented as pairs of observations. Naturally, the temporal property of observations is crucial for inferring about activities: observations are part of events, and scenes typically consist of sequences of events. In this context, a scene like "person x takes a coffee break" could include "making coffee," "drinking coffee," "washing a cup in the sink," and / or "putting a cup in the dishwasher," where each of these events will depend on person x's varying proximity to the "coffee machine," "table," "sink," and "dishwasher." Distances are centered on the person's relative position and typically change at each moment; in DoorS, events / activities are predicted based on the observed sequence of distances, as in the example above, or based on the duration of the observed distances.
[0035] The results demonstrate that by leveraging two sensing modalities (video and spatial environment knowledge), the system can build software services that provide scene understanding beyond basic person detection from video analytics. Thus, additional scene understanding is created by utilizing more sensors. By working directly on a system with such a setup, such as on the SAST camera platform, the system enables rapid prototyping and quick transfer of results to a variety of use cases. While one embodiment pertains to a smart building use case, the approach remains applicable to many other fields. Figure 3A and Figure 3B Two possible computational pipelines for the proposed method are shown.
[0036] Figure 3A This is the first embodiment of a computational pipeline configured to understand multimodal scenarios. Figure 3B This is an alternative implementation of a computational pipeline that utilizes sensor data fusion. For example... Figure 3A As shown, the system may include a computational pipeline for multimodal scene understanding. The system can receive information from multiple sensors. In the embodiment shown below, two sensors are utilized; however, multiple sensors may be used. In one embodiment, sensor 301 may acquire acoustic signals, while sensor 302 may acquire image data. Image data may include still images or video images. The sensors can be any type of sensor, such as a LiDAR sensor, radar sensor, camera, video camera, sonar, microphone, or any of the aforementioned sensors or hardware.
[0037] At boxes 305 and 307, the system may involve data preprocessing. Data preprocessing may include transforming the data into a uniform structure or class. Preprocessing may be performed via onboard processing or off-board processors. Data preprocessing can help facilitate system-related processing, machine learning, or fusion processes by updating certain data, data structures, or other data deemed ready for processing.
[0038] At boxes 309 and 311, the system can encode the data using an encoder and apply feature extraction. At box 317, the encoded data or feature extraction can be sent to the spatiotemporal inference engine. The encoder can be a network (FC, CNN, RNN, etc.) that takes input (e.g., various sensor data or preprocessed sensor data) and outputs feature maps / vectors / tensors. These feature vectors can store information representing the input, its features. By converting characters into one-hot vector representations, each character of the input can be fed as input into the ML model / encoder. At the last timestep of the encoder, the final hidden representations of all previous inputs are passed as input to the decoder.
[0039] At boxes 313 and 315, the system can utilize a machine learning model or decoder to decode the data. The decoder can be used to output metadata to the temporal inference engine 317. The decoder can be a network (typically with the same network structure as the encoder, but with the opposite orientation) that takes feature vectors from the encoder and provides the best-closest match to the actual input or expected output. The decoder model can be able to decode the state representation vector and provide the probability distribution for each character. A softmax function can be used to generate the probability distribution vector for each character. This, in turn, helps generate the complete literal translation. The metadata can be used to facilitate scene understanding in a multimodal scene by indicating information captured from several sensors, which together can contribute to indicating the scene.
[0040] The spatiotemporal inference engine 317 can be configured to capture relationships between multimodal sensors to help determine various scenes and scenarios. Therefore, the spatiotemporal inference engine 317 can leverage metadata to capture such relationships. The spatiotemporal inference engine 317 can then feed current events into the model and perform predictions, outputting a predicted set of events and likelihood probabilities. Thus, the spatiotemporal inference engine enables the interpretation of large datasets (e.g., raw data with timestamps) into meaningful concepts at different levels of abstraction. This can include abstracting individual time points into longitudinal time intervals, calculating trends and gradients from a series of outcome measurements, and detecting different types of patterns that might otherwise be hidden in the raw data. The spatiotemporal inference engine can optionally work with a domain ontology 319. The domain ontology 319 can be an ontology that contains categories, attributes, and relationships between concepts, data, and entities that materialize one, several, or all of a public domain. Thus, by defining a set of concepts and categories representing a subject, an ontology is a way of showing the attributes of a subject domain and how they relate to each other.
[0041] Next, the temporal inference engine 317 can output scene inference in box 321. Scene inference can identify activities, determine control commands, or classify various events picked up by sensors. An example of a scene could be “taking a soda can from the refrigerator,” which can be generalized from several human-centered events collected by various sensors. For example, the previous example “taking a soda can from the refrigerator” can be modeled as a scene that includes human-centered events such as (1) “facing the refrigerator,” (2) “opening the refrigerator door,” (3) “reaching out an arm,” and (4) “grabbing the soda can.” Clearly, these events are temporally related: (2), (3), and (4) occur sequentially, while (1) lasts for the entire duration of the preceding sequence (facing the refrigerator is a condition for interacting with an item placed in the refrigerator). In this way, the system can be able to jointly model scenes as meaningful sequences (or combinations) of individual atomic events. Thus, the system can analyze and parse different events in light of threshold time periods, compare and contrast them with other identified events, and determine scenes or sequences based on those events. Therefore, when something lasts for the entire duration, the system requirement might be that the camera and sensors use sensor data to identify the first event ("facing the refrigerator"), which must occur throughout the entire time period compared to other events (events 2-4). Furthermore, the system can analyze event sequences to identify a scene.
[0042] At box 323, the system can output visualizations and controls. For example, if the system identifies a specific type of scenario, it can generate environmental control commands. Such commands could include providing an alarm or starting data recording based on the identified scenario type. In another embodiment, an alarm could be output, recording could be started, and so on.
[0043] Figure 3B This is an alternative embodiment of the computational pipeline. The alternative embodiment may include, for example, a process that allows the fusion module 320 to obtain features from the feature extractor or decoder. The fusion module can then fuse all the data to generate a dataset to be fed into a single machine learning model / decoder.
[0044] Figure 4 This is an example involving scene understanding by multiple people. Figure 4 In this scenario, the scene could include multiple people (e.g., in an instance of the DoorS class "Customer"), one person walking across a table, and another person washing their hands in a sink. The system can correctly identify the person whose bounding box includes the sink (distance = "0.0") as "washing (an instance of the DoorS class "Activity")," and it can also infer that this type of washing activity is merged into the DoorS class "CustomerActivityNoPRoduct" (e.g., at the bottom) because no object (an instance of the DoorS class "Product") is detected. The inference process is initiated by a query that compares distance-based metrics between people and objects in the scene and triggers rule-based inference to predict the most likely activity (e.g., at the top right). Note that this example is generated from a demo of the system, in which it is shown that the system can classify a person "walking" across a table as irrelevant and can identify such activities in a scene by leveraging machine learning, without requiring knowledge-based reasoning.
[0045] The processes, methods, or algorithms disclosed herein are deliverable to / implemented by a processing device, controller, or computer, which may include any existing programmable electronic control unit or dedicated electronic control unit. Similarly, the processes, methods, or algorithms may be stored in various forms as data and instructions executable by a controller or computer, including but not limited to information permanently stored on non-writable storage media such as ROM devices and information reproducibly stored on writable storage media such as floppy disks, magnetic tapes, CDs, RAM devices, and other magnetic and optical media. The processes, methods, or algorithms may also be implemented in a software executable object. Alternatively, the processes, methods, or algorithms may be embodied, wholly or partially, using suitable hardware components—such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), state machines, controllers, or other hardware components or devices, or combinations of hardware, software, and firmware components.
[0046] While exemplary embodiments have been described above, they are not intended to describe all possible forms encompassed by the claims. The terms used in this specification are descriptive and not limiting, and it should be understood that various changes may be made without departing from the spirit and scope of this disclosure. As previously described, features of various embodiments may be combined to form further embodiments of the invention that may not be explicitly described or illustrated. While various embodiments may have been described as providing advantages over other embodiments or prior art implementations in one or more desired features, or being preferred over other embodiments or prior art implementations, those skilled in the art will recognize that one or more features or characteristics may be compromised depending on the specific application and implementation to achieve desired overall system properties. These properties may include, but are not limited to, cost, strength, durability, lifecycle cost, merchantability, appearance, packaging, size, suitability, weight, manufacturability, ease of assembly, etc. Accordingly, while any embodiment may be described as less desirable than other embodiments or prior art implementations in one or more features, these embodiments are not outside the scope of this disclosure and may be desirable for a particular application.
Claims
1. A system for image processing, comprising: The first sensor is configured to capture at least one or more images; The second sensor is configured to capture sound information; A processor that communicates with a first sensor and a second sensor, wherein the processor is programmed to: Receive the one or more images and the sound information; Preprocessing the one or more image and sound information, wherein the preprocessing includes converting the one or more images and sound information into a uniform structure or class; The encoder is used to extract one or more data features associated with image and sound information; Metadata is output to the spatiotemporal inference engine via the decoder, wherein the spatiotemporal inference engine is configured to capture the relationships between multimodal sensors to help determine various scenarios, thereby using the metadata to capture the relationship between the first sensor and the second sensor, wherein the spatiotemporal inference engine is configured to interpret large datasets into meaningful concepts at different levels of abstraction, including abstracting individual time points into longitudinal time intervals, calculating trends and gradients from a series of outcome measurements, and detecting different types of patterns, and wherein metadata is derived using the decoder and the one or more data features; The spatiotemporal inference engine and metadata are used to determine one or more scenes associated with the image and sound information; and Output control commands in response to one or more of the scenarios.
2. The system according to claim 1, wherein, The spatiotemporal inference engine communicates with the domain ontology database and uses the domain ontology database to determine the one or more scenarios.
3. The system according to claim 2, wherein, The domain ontology database includes information indicating one or more scenarios that utilize the metadata.
4. The system according to claim 2, wherein, The domain ontology database is stored on a remote server that communicates with the processor.
5. The system according to claim 1, wherein, The system includes a third sensor configured to capture temperature information, and the processor communicates with the third sensor and receives the temperature information and extracts one or more associated data features from the temperature information.
6. The system according to claim 1, wherein, The processor is further programmed to fuse one or more data features associated with the image and sound information before outputting metadata.
7. The system according to claim 1, wherein, The processor is further programmed to separately extract one or more data features associated with image and sound information to multiple decoders.
8. The system according to claim 1, wherein, The decoder is associated with a machine learning network.
9. A system for image processing, comprising: The first sensor is configured to capture a first set of information indicating the environment; The second sensor is configured to capture a second set of information indicating the environment; A processor that communicates with a first sensor and a second sensor, wherein the processor is programmed to: Receive the first and second information sets indicating the environment; Preprocessing the first and second information sets of the indication environment, wherein the preprocessing includes converting the first and second information sets into a uniform structure or class; The encoder is used to extract one or more data features associated with the first information set and the second information set; Metadata is output to the spatiotemporal inference engine via the decoder, wherein the spatiotemporal inference engine is configured to capture the relationships between multimodal sensors to help determine various scenarios, thereby using the metadata to capture the relationship between the first sensor and the second sensor, wherein the spatiotemporal inference engine is configured to interpret large datasets into meaningful concepts at different levels of abstraction, including abstracting individual time points into longitudinal time intervals, calculating trends and gradients from a series of outcome measurements, and detecting different types of patterns, and wherein metadata is derived using the decoder and one or more data features; Use a spatiotemporal inference engine and metadata to determine one or more scenarios associated with the first information set and the second information set; and Output control commands in response to one or more of the scenarios.
10. The system according to claim 9, wherein, The first and second information sets contain different types of data.
11. The system according to claim 9, wherein, The first sensor may include a temperature sensor, a pressure sensor, a vibration sensor, a humidity sensor, or a carbon dioxide sensor.
12. The system according to claim 9, wherein, The processor is further programmed to preprocess a first set of information and a second set of information indicating the environment before extracting the one or more data features using the encoder.
13. The system according to claim 9, wherein, The system includes a fusion module for fusing a fused dataset from a first information set and a second information set.
14. The system according to claim 13, wherein, The metadata was extracted from the fused dataset.
15. A system for image processing, comprising: The first sensor is configured to capture a first set of information indicating the environment; The second sensor is configured to capture a second set of information indicating the environment; A processor that communicates with a first sensor and a second sensor, wherein the processor is programmed to: Receive the first and second information sets indicating the environment; Preprocessing the first and second information sets of the indication environment, wherein the preprocessing includes converting the first and second information sets into a uniform structure or class; Extract one or more data features associated with a first information set and a second information set indicating the environment; Metadata is output to the spatiotemporal inference engine via the decoder, wherein the spatiotemporal inference engine is configured to capture the relationships between multimodal sensors to help determine various scenarios, thereby using the metadata to capture the relationship between the first sensor and the second sensor, wherein the spatiotemporal inference engine is configured to interpret large datasets into meaningful concepts at different levels of abstraction, including abstracting individual time points into longitudinal time intervals, calculating trends and gradients from a series of outcome measurements, and detecting different types of patterns, and wherein the metadata indicates one or more data features. Use metadata to identify one or more scenarios associated with the first information set and the second information set; and Output control commands in response to one or more of the scenarios.
16. The system according to claim 15, wherein, The system includes a decoder configured to utilize a machine learning network.
17. The system of claim 15, wherein, The first and second information sets contain different types of data.
18. The system according to claim 15, wherein, The first sensor may include a temperature sensor, a pressure sensor, a vibration sensor, a humidity sensor, or a carbon dioxide sensor.
19. The system according to claim 15, wherein, The system includes a fusion module for fusing a fused dataset from a first information set and a second information set.
20. The system according to claim 19, wherein, The fused dataset is sent to a machine learning model to output metadata associated with the fused dataset.
Citation Information
Patent Citations
Temporal fusion of multimodal data from multiple data acquisition systems to automatically recognize and classify an action
US20170220854A1
Classifying audio scene using synthetic image features
US20220044071A1