Method for integrating optical and electromagnetic sensor data

The transformer-based neural system integrates visual and radio frequency sensor data using self-attention and cross-modal techniques, addressing the fragmentation issue in existing methods to enhance scene understanding and decision-making in complex environments.

DE102024206255A1Pending Publication Date: 2026-01-08ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024206255
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-03
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing multimodal data processing techniques fail to effectively integrate visual and radio frequency sensor data into a unified analysis framework, neglecting the complementary strengths of these modalities and resulting in a fragmented understanding of the environment.

Method used

A transformer-based neural system employing self-attention mechanisms and cross-modal interaction techniques is used to preprocess and transform optical and electromagnetic sensor data, integrating and combining features from both modalities to generate a unified representation.

Benefits of technology

This approach enables a more accurate and holistic understanding of the scene, improving decision-making capabilities, especially in complex environments, and provides real-time processing capabilities essential for applications like autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for integrating optical and electromagnetic sensor data, comprising the following steps, which are performed by the transformer-based neural system (30), in particular a transformer-based neural network (30): - Receiving (101) optical and electromagnetic sensor data from one or more sources in an environment, - Preprocessing (102) the received optical and electromagnetic sensor data to generate representations of features suitable for input into the transformer-based neural system (30), - Transforming (103) the preprocessed optical and electromagnetic sensor data via the transformer-based neural system (30) which employs one or more self-attention mechanisms and / or one or more cross-modal interaction techniques to integrate and combine the generated representations of the features, - Generating (104) a unified representation based on transforming (103).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for integrating optical and electromagnetic sensor data. It also relates to a training method, computer programs, a node, and a system for this purpose. State of the art

[0002] Significant progress has been made in the field of semantic augmentation and scene understanding using neural networks. These approaches have demonstrated promising results in improving the understanding of visual data and enable various applications such as autonomous driving, surveillance systems, and augmented reality. However, a limitation of existing methods lies in their inability to effectively integrate camera information with wireless sensor data.

[0003] Existing multimodal data processing techniques often fail to effectively integrate visual and radio frequency sensor data into a unified analysis framework. While individual sensors may be capable of handling individual data types, such as visual or radio signals, these approaches frequently result in a fragmented understanding of the environment. Fusing these heterogeneous data sources remains a significant challenge due to the lack of robust algorithms capable of seamlessly combining and interpreting information from different modalities. Furthermore, existing solutions typically rely on processing individual modalities, neglecting the complementary strengths of visual and radio frequency sensors.

[0004] It is therefore an object of the invention to overcome the disadvantages of the prior art. In particular, an object of the present invention is to provide a method for processing data from different modalities that enables a more accurate and holistic understanding of the scene. Disclosure of the invention

[0005] According to aspects of the invention, a method with the features of claim 1, a training method with the features of claim 6, a node with the features of claim 7 or 8, a system with the features of claim 9, and a computer program with the features of claim 10 or 11 are provided. Further features and details of the invention are disclosed in the respective dependent claims, the description, and the drawings. Features and details described in the context of the method also correspond to the inventive computer program, the inventive node, and the inventive system, and vice versa.

[0006] According to one aspect of the invention, a method for integrating optical and electromagnetic sensor data is provided. The method comprises the following steps, which are performed by a transformer-based neural system, in particular a transformer-based neural network: - Receiving optical and electromagnetic sensor data from one or more sources in an environment, - Preprocessing the received optical and electromagnetic sensor data to generate representations of features suitable for input into the transformer-based neural system, - Transforming the preprocessed optical and electromagnetic sensor data via the transformer-based neural system, which employs one or more self-attention mechanisms and / or one or more cross-modal interaction techniques to integrate and combine the generated representations of the features, - Generating a unified representation based on transformation.

[0007] In the transformation step, features from both modalities, i.e., the optical and the electromagnetic modality, are preferably integrated and combined. The invention may provide a transformative approach for integrating optical and electromagnetic sensor data by employing a transformer-based neural system. Optical sensor data can include visual data. Electromagnetic sensor data can be provided as radio frequency sensor data. This has the advantage that, by leveraging the strengths of each modality, the inventive method enables a significantly improved level of scene understanding within an environment. The transformer-based neural network can be adapted to recognize patterns, anomalies, and correlations between different types of features, leading to improved decision-making capabilities.For example, according to the invention, the system can identify changes in electromagnetic signals that indicate specific objects or events, even when visual data is limited or unavailable. This can be advantageously used in applications such as the detection and tracking of objects in complex environments, where a combination of sensor modalities may be crucial for reliable performance.

[0008] The invention's self-awareness mechanisms and cross-modal interaction techniques enable the neural network to effectively integrate these additional modalities and create a unified representation that combines and / or considers the strengths of each modality. The multimodal transformer model can be particularly useful in applications such as autonomous vehicles or robotics, where accurate scene understanding and robustness are crucial.

[0009] Furthermore, the invention enables real-time processing capabilities that are crucial for applications such as autonomous navigation, where timely decision-making is essential for safe and efficient operation. In particular, the decentralized architecture can provide significant advantages in terms of reduced latency and bandwidth limitations, making it an advantageous solution for applications with high data transmission requirements. By integrating optical / visual and radio frequency sensor data, the system can provide a comprehensive understanding of environmental conditions, thereby enabling more accurate predictions and decision-making.A possible decentralized architecture of the transformer-based neural system allows for real-time monitoring and control of complex systems in an industrial environment, leading to increased efficiency, reduced downtime, and improved overall performance.

[0010] A transformer-based neural system, specifically a transformer-based neural network, can be understood as a type of system or architecture that uses a model to handle data sequences. Furthermore, a transformer can employ self-attention mechanisms. These mechanisms allow the transformer model to weigh the importance of individual parts of the input data, regardless of their sequence in a received data stream. This differs from other neural network models, such as RNNs or LSTMs, which processed data sequentially and often struggled with far-reaching dependencies. Transformers can consist of an encoder and a decoder. The encoder processes the entire input sequence simultaneously, using self-attention to compute a representation of the input.The decoder then uses the encoder's output, along with previous outputs, to predict the next element in the sequence. The transformer-based neural system can learn from different parts of the input data simultaneously, making it highly effective for tasks involving complex input structures such as multiple data types or modalities.

[0011] In the context of a transformer-based neural network, a modality refers to a type of data or input that the system or network processes. Common modalities include text, images, audio, and video. Transformers can be designed to handle one or more modalities and often employ different techniques to encode each type of input into a format suitable for processing by the neural network.

[0012] In the context of scene understanding, this capability can enable transformers to effectively process diverse sensory inputs and provide a robust framework for improving semantic understanding and data interoperability.

[0013] In another example, the distributed architecture of the invention can be further optimized for real-time processing by incorporating techniques such as asynchronous communication and decentralized decision-making. This allows the system to operate more efficiently in scenarios where latency is critical, such as autonomous high-speed vehicles or real-time monitoring applications.

[0014] Furthermore, the decentralized architecture proposed in another embodiment of the invention has the potential to revolutionize the way we process and analyze sensor data. By distributing the processing load across a network of nodes, each with its own local transformer model, the system can reduce latency and bandwidth limitations while maintaining the benefits of a unified representation of the environment.

[0015] It is also possible that the process includes the following additional step during production: - Generating a semantic map representing combined features from the received, preferably fused, optical and electromagnetic data, wherein the semantic map includes enhanced object detection, one or more classification outputs and / or at least one transformed cross-modal data representation with respect to the environment.

[0016] This feature enables improved object detection, classification outputs, and transformed cross-modal data representation with respect to the environment. The generated semantic map can advantageously provide valuable information about the environment, enabling more accurate scene understanding and decision-making capabilities. It is possible that the transformer-based neural network architecture of the invention can be further enhanced by incorporating additional modalities, such as lidar or inertial navigation unit data, to create a multimodal sensor fusion framework. This allows the system to leverage the strengths of each modality, resulting in an even more accurate and robust scene understanding. For example, the addition of lidar data could provide high-resolution 3D point clouds, enabling precise object detection and tracking.Meanwhile, inertial navigation unit data could be used to estimate the system's pose and speed, enabling more accurate navigation and localization.

[0017] It is also possible that the process includes at least one of the following additional steps during preprocessing: - Transforming optical sensor data into electromagnetic sensor data and vice versa, - Converting the electromagnetic sensor data into a format suitable for integration with the optical sensor data, - Adapting the optical and / or electromagnetic sensor data into a consistent input format for the transformer-based neural system.

[0018] The provided feature may have the advantage of transforming optical sensor data into electromagnetic sensor data and vice versa, as well as converting the electromagnetic sensor data into a format suitable for integration with the optical sensor data. This enables a unified representation of the environment by combining features from both modalities. The transformation and conversion steps can be performed during the preprocessing phase, allowing the transformer-based neural network to effectively integrate and combine features from both modalities. It is possible that the fusion of optical and electromagnetic data, by leveraging the complementary strengths of each modality, will enable a more comprehensive understanding of the environment.For example, the transformer-based neural network can be trained to recognize patterns in the electromagnetic spectrum that indicate specific objects or structures, even if these are partially obscured or obscured by environmental factors. This can provide significant improvements in object detection and classification accuracy, particularly in scenarios where visual data is limited or unavailable. Furthermore, the ability to transform visual data into channel impulse responses and vice versa can facilitate better communication and integration between different types of sensors and systems.

[0019] It is possible that the procedure includes the following further step: - Classifying and detecting objects based on the inclusion of electromagnetic sensor data obtained via radio sensing of the environment and optical data.

[0020] This integration advantageously enables improved object detection and classification by incorporating electromagnetic properties determined via radio sensing. Furthermore, the dual-modal approach ensures robustness in diverse environments, including poor visibility. Additionally, cross-modal data transformation facilitates better communication and integration between different types of sensors and systems.

[0021] It is possible that the transformer-based neural system can be applied to various domains by integrating visual and radio frequency sensor data. Examples include enhanced surveillance systems to provide a unified representation of the environment, enabling improved object detection and classification; real-time monitoring and tracking of moving objects in complex environments, including those with poor visibility, by leveraging the strengths of both modalities; and intelligent traffic management systems that use the fusion of visual and radio frequency sensor data to optimize traffic flow and reduce congestion.

[0022] Furthermore, it is possible for the optical sensor data to include at least one image and / or the electromagnetic sensor data to include at least one channel impulse response. This integration allows the system to leverage the strengths of both modalities, such as improved scene understanding, enhanced object detection and classification, and robustness in diverse environments.

[0023] The method according to the invention can be used for real-time processing capabilities, which are essential for applications such as autonomous vehicles where accurate detection and tracking of objects are crucial. By incorporating electromagnetic properties determined via radio sensing, the system can classify and detect environmental phenomena with higher accuracy, enabling more effective monitoring of natural disasters or industrial processes. The transformer-based neural network's ability to transform visual data into channel impulse responses and vice versa allows for improved communication and integration between different types of sensors and systems, making it ideal for smart city infrastructure management.

[0024] Another aspect of the invention is a method for training a transformer-based neural system, preferably a transformer-based neural network, comprising the following steps: - Providing a dataset that includes optical sensor data and electromagnetic sensor data, - Preprocessing the optical sensor data to normalize it and format it as inputs to the transformer-based neural system, - Preprocessing the electromagnetic sensor data to convert it into a format suitable for integration with the optical sensor data, - Setting up the transformer-based neural system with initial weights and customized configuration to handle multimodal inputs, - parallel feeding of the pre-processed optical and electromagnetic sensor data into the transformer-based neural system, - Transforming the optical and electromagnetic sensor data via the transformer-based neural system, which employs one or more self-attention mechanisms and / or cross-modal interaction techniques to integrate and combine the generated representations of the features, - Outputting a unified representation that combines features from both modalities, and - Training the transformer-based neural system using a loss function selected from a group that includes cross-entropy loss, mean squared error loss, contrastive loss, or combined loss.

[0025] The optical sensor data is preferably visual data, which is provided or generated, for example, by cameras. Accordingly, the training method according to the invention offers the same advantages as those described in detail with regard to the method according to the invention.

[0026] Another aspect of the invention is a local node for integrating optical and electromagnetic sensor data, comprising a local transformer-based neural system, in particular a transformer-based neural network. The local node is configured to perform steps for receiving and preprocessing optical and electromagnetic sensor data via the local transformer-based neural system and includes means for transmitting the preprocessed data to a central node. Accordingly, the local node according to the invention provides the same advantages as described in detail with respect to the method according to the invention. These include improved data fusion, improved detection and classification of objects, robustness in different environments, and cross-modal data transformation.Furthermore, the local node receives optical and electromagnetic sensor data, preprocesses it via the transformer-based neural network, and sends the preprocessed data to a central node, thus enabling decentralized processing and reducing latency and bandwidth limitations. This local node can also be configured to operate autonomously, allowing it to make decisions without relying on central processing. This autonomy enables real-time processing capabilities, making it suitable for applications such as autonomous navigation, where timely decision-making is critical. Additionally, this local node can be integrated with other nodes to form a distributed network, enabling more complex integration and semantic analysis tasks.This distributed architecture allows nodes to communicate and share information, enabling a more comprehensive understanding of the environment.

[0027] Another aspect of the invention is a central node for integrating optical and electromagnetic sensor data, comprising means for receiving the pre-processed data and a central transformer-based neural system, in particular a transformer-based neural network, wherein the central transformer-based neural system is designed to perform steps for transforming the pre-processed data and generating a unified representation.

[0028] Accordingly, the central node according to the invention offers the same advantages as described in detail with regard to the method according to the invention.

[0029] The central node can seamlessly integrate optical and electromagnetic sensor data, leveraging the strengths of both modalities to provide a unified representation of the environment, thereby improving scene understanding, object detection, and classification accuracy. In the distributed setup, nodes can be strategically placed to minimize latency and maximize coverage, enabling seamless data transmission and fusion. Furthermore, the lightweight transformers at each node can be customized to handle specific task-oriented operations such as object detection or scene understanding, further reducing the computational load on the central transformer model.

[0030] Another aspect of the invention is a system for integrating optical and electromagnetic sensor data, comprising at least one local node and one central node. Accordingly, the system offers the same advantages as those described in detail with respect to the method.

[0031] It is possible that the system, particularly the distributed system operating in a decentralized manner, offers the further advantage of reduced latency and bandwidth constraints, since each node preprocesses its input data locally, for example, using a smaller, lightweight transformer model, before sending processed features to a central transformer for more complex integration and semantic analysis tasks. This approach enables real-time processing capabilities, which are essential for applications such as autonomous navigation, where prompt decision-making is crucial.

[0032] Furthermore, this distributed architecture can be extended to include other types of sensors or modalities, such as lidar, radar, or even human input, enabling the integration of diverse data streams into a unified view. The fusion of heterogeneous data sources can provide unprecedented insights and decision-making capabilities, revolutionizing various domains such as autonomous vehicles, smart cities, and environmental monitoring.

[0033] Another aspect of the invention is a first computer program, in particular a first computer program product, comprising instructions which, when executed by a computer at a local node, cause the computer to process data by performing steps for receiving and preprocessing, and then transmitting the preprocessed data to a computer at a central node. Accordingly, the computer program according to the invention offers the same advantages as those described in detail with respect to the method according to the invention.

[0034] Another aspect of the invention is a second computer program, in particular a second computer program product, comprising instructions which, when executed by a computer of a central node, cause the computer to receive the preprocessed data from a computer of a local node by performing steps to transform the received data and generate a unified representation for processing the received data. Accordingly, the computer program according to the invention offers the same advantages as those described in detail with respect to the method according to the invention.

[0035] In a further aspect of the invention, a data processing device can be provided which is configured to execute the method according to the invention. The device can, for example, be a computer, in particular a computer with at least one local node and / or one central node, which executes the computer program according to the invention. The respective computer can comprise at least one processor that can be used to execute the respective computer program. Furthermore, a non-volatile data storage device can be provided in which the respective computer program is stored and from which the processor can read the respective computer program for its execution.

[0036] According to another aspect of the invention, a computer-readable storage medium can be provided, comprising the respective computer program according to the invention and / or instructions which, when executed by a respective computer, cause the respective computer to perform the steps of the method according to the invention. The storage medium can be designed as a data storage device, for example, a hard disk and / or non-volatile memory and / or a memory card and / or a solid-state drive. The storage medium can, for example, be integrated into the computer.

[0037] Furthermore, the method according to the invention can be implemented as a computer-implemented method. Alternatively or additionally, at least one of the disclosed method steps can be computer-implemented and / or automated.

[0038] Further advantages, features, and details of the invention are evident from the following description, in which embodiments of the invention are described in detail with reference to the drawings. In this context, the features mentioned in the claims and in the description can be essential to the invention, either individually or in any combination. The drawings show: Fig. 1: a method, a first and a second computer program, a system, a local node and a central node according to embodiments of the invention, Fig. 2: a schematic flowchart of a process according to embodiments of the invention, Fig. 3: a training method according to embodiments of the invention.

[0039] The core of the invention lies in the use of a transformer-based neural network designed to fuse and process data from both optical sensors and / or cameras and / or electromagnetic or high-frequency sensors.

[0040] The transformer-based neural system, or network, can employ advanced algorithms to merge visual and electromagnetic data into a cohesive model. This leads to improved scene understanding by leveraging the strengths of both sensory inputs, something unattainable with current single-modality systems. By incorporating electromagnetic properties acquired through radio sensing, the system can classify and detect objects with greater accuracy. This can be particularly useful in environments where visual data alone is ambiguous or insufficient.

[0041] The dual-modality approach can ensure reliable performance under diverse environmental conditions. In contrast to systems based solely on optical sensors, the present invention can also be used effectively in poor visibility conditions, such as fog or darkness, where high-frequency sensors can still capture critical data.

[0042] The invention can provide cross-modal data transformation by transforming visual data into channel impulse responses and vice versa. This capability can facilitate better communication and integration between different types of sensors and systems, thereby improving interoperability and functionality.

[0043] The neural transformer network can improve the semantic interpretation of the scene by providing not only a visual representation but also an electromagnetic context. This leads to a more comprehensive and detailed understanding of the environment.

[0044] Fig. Figure 1 shows a method 100, a first and a second computer program 50, 60, a system 30, a local node 10 and a central node 20 according to embodiments of the invention.

[0045] Fig. Figure 1 shows in particular an embodiment of a method 100 for integrating optical and electromagnetic sensor data. The method 100 comprises the following steps, which are performed by a transformer-based neural system 30: In step 101, optical and electromagnetic sensor data are received from one or more sources in an environment. In step 102, optical and electromagnetic sensor data are preprocessed to generate representations of features suitable for input into the transformer-based neural system 30. Then, in step 103, the optical and electromagnetic sensor data are transformed by the transformer-based neural system, which employs one or more self-attention mechanisms and / or one or more cross-modal interaction techniques, to integrate and combine features from both modalities.Step 104 generates a uniform representation based on the transformation in step 103.

[0046] Fig. Figure 2 shows a schematic flowchart of a process according to embodiments of the invention. Fig. Figure 2 illustrates in particular a flowchart of the method according to the invention with respect to a distributed network or system architecture.

[0047] The present invention uses a transformer-based neural network 30 or system 30 designed for integrating optical and electromagnetic sensor data. Optical sensor data can include visual data such as an image or a stream of images. Electromagnetic sensor data can include radio frequency data, for example, a channel impulse response based on radio sensing.

[0048] The transformer-based neural network 30 can comprise an input layer 31, a preprocessing module 32, a fusion transformer layer 33, and an output layer 34, as shown in Fig. 2 shown.

[0049] The input layer 31 can, for example, process two inputs for the transformer-based neural network 30, or system 30. Network 30 can process two primary types of input. One of these inputs can be optical sensor data or visual data. Optical or visual data can include one or more images, for example, captured by sensors such as one or more different cameras positioned at various locations in the environment, providing views and associated data at different angles and resolutions. The other input can include electromagnetic sensor data or radio frequency data. The radio frequency data can include one or more channel impulse responses from transmitters and receivers in the environment, providing electromagnetic properties of a scene.Furthermore, the inputs can include images and high-frequency data samples from a simultaneous snapshot.

[0050] A preprocessing module 32 can normalize the optical or visual data and format it into a consistent input structure for the transformer-based neural network 30. Furthermore, the module 32 can convert electromagnetic sensor data, such as raw radio channel impulse response data, into a format suitable for integration with the optical sensor data, for example, feature extraction and normalization. Another processing step performed by the module 32 could include, for example, line-of-sight path detection. The processing module 32 may include at least one submodule, such as an image processing submodule and / or a radio data processing submodule (in Fig. 2 not shown).

[0051] The fusion transformer layer 33 can employ self-attention mechanisms that allow the network 30 to weigh and prioritize information from both modalities or data types, for example, optical or visual data and electromagnetic data, such as radio frequency data. These two data types can be processed simultaneously for optimal data fusion. Furthermore, the transformer 33 can enable cross-modal interaction between the visual and radio frequency data, thereby enhancing the network 30's ability to learn from both data types or both modalities, such as visual and radio frequency data.

[0052] Output layer 34 can generate a unified representation that combines features from both visual and radio frequency data, thereby enhancing the semantic understanding of the scene within the environment. Furthermore, output layer 34 can provide transformed data outputs, where visual data can be represented as channel impulse responses and vice versa, enabling cross-modal analysis. An output can include one or more integrated semantic maps of the environment, including enhanced object detection and / or classification outputs and / or transformed cross-modal data representations.

[0053] In another embodiment, the transformer-based neural system 30 or network 30 can be implemented as a distributed architecture, as in Fig. 2 shown.

[0054] Given the potential spatial distribution of data sources, such as multiple cameras and high-frequency sensors, within the system according to the invention, the system 30 can benefit from a distributed approach. In such an approach, the transformer-based system 30, or transformer model, can be designed to operate in a decentralized manner, with data from various sources being processed locally at or near the source before being integrated into a central transformer-based neural network. The system 30 can further comprise a central node 20 and at least one local node 10, each node 10, 20 being a transformer-based neural network. This makes it possible to reduce latency and bandwidth limitations that can arise when large volumes of raw data are transmitted over a communication network.

[0055] The at least one local node 10 can receive optical and / or electromagnetic data from one or more cameras and one or more sensors 201. The at least one local node 10 can locally preprocess this received input data using its (local) preprocessing module 32L before feeding this data into the local transformer-based neural network 33L for further processing.

[0056] It is possible that the transformer-based neural network 33L of local node 10 can comprise a smaller, lightweight transformer model. The at least one local transformer 33L can handle initial feature extraction and preliminary data fusion tasks. The preprocessed features, which are significantly more compact than the raw data, can then be sent to a central node 20.

[0057] The central node 20 can include a central transformer model 33C, which is capable of performing the more complex integration and semantic analysis tasks by processing the transmitted data 202. Then, based on transforming the transmitted data in the central transformer model 33C, the central transformer-based network of the central node 20 can generate a unified representation 203.

[0058] This multi-stage processing approach can not only improve the efficiency and scalability of the system, but also support real-time processing functions that are essential for applications such as autonomous navigation and dynamic environmental monitoring.

[0059] Fig. Figure 3 shows a training method 300 according to embodiments of the invention. Fig.3 in particular represents a method 300 for training a transformer-based neural network 30 according to embodiments of the invention, comprising the following steps: Step 301 provides a dataset that includes optical sensor data captured by cameras and electromagnetic sensor data from different environments and conditions.

[0060] In the next step, 302, the optical sensor data is preprocessed to normalize it and format it as input for the transformer-based neural network. Then, in step 303, the electromagnetic sensor data is preprocessed to convert it into a format suitable for integration with the optical sensor data. In step 304, the transformer-based neural system is set up with initial weights and a customized configuration to handle multimodal input. In the next step, 305, the preprocessed optical and electromagnetic sensor data are fed into the transformer-based neural network.In step 306, the optical and electromagnetic sensor data are transformed via the transformer-based neural system, which employs one or more self-awareness mechanisms and / or one or more cross-modal interaction techniques, to integrate and combine features from both modalities. Then, in the next step 307, a unified representation combining features from both modalities is output. In step 308, the transformer-based neural network is trained using a loss function selected from a group that includes cross-entropy loss, mean squared error loss, contrastive loss, or combined loss.

[0061] In another embodiment, the process flow of a training pipeline may include further or different training steps, such as processing batches of input data, each batch comprising synchronized pairs of visual and radio frequency data.

[0062] Furthermore, a step can be performed to calculate loss based on the difference between the network output and the actual data labels during training (loss calculation and backward propagation). Suitable loss functions for this multimodal system may include the following: - Cross-entropy loss: Ideal for classification tasks within the network, especially useful when categorizing objects detected in semantic maps. - Mean Squared Error Loss (MSE Loss): Effective for regression tasks, for example estimating the parameters of channel impulse responses or adjusting spatial coordinates during image transformations. - Contrastive loss: Useful in scenarios where the goal is to

[0063] Learning embeddings from both modalities, which are closer together for similar pairs and further apart for dissimilar pairs, improves the model's ability to distinguish between different types of sensor data. - Combined loss: A combination of different loss functions could be used to handle different aspects of training simultaneously, for example combining MSE for continuous data regression and cross-entropy for classification tasks.

[0064] Furthermore, it is possible to periodically validate the transformer model using a separate dataset to monitor its performance and generalizability.

[0065] In another embodiment, the training of the transformer-based neural network can include various fine-tuning and optimization steps by refining network parameters based on validation results to optimize performance for specific applications or conditions.

[0066] The preceding explanation of the embodiments describes the present invention in the context of examples. Naturally, individual features of the embodiments can be combined with one another as desired, insofar as technically feasible, without departing from the scope of protection of the present invention.

Claims

[1] Method (100) for integrating optical and electromagnetic sensor data, comprising the following steps, which are performed by a transformer-based neural system (30), in particular a transformer-based neural network (30): - Receiving (101) optical and electromagnetic sensor data from one or more sources in an environment, - Preprocessing (102) the received optical and electromagnetic sensor data to generate representations of features suitable for input into the transformer-based neural system (30), - Transforming (103) the preprocessed optical and electromagnetic sensor data via the transformer-based neural system (30) which employs one or more self-attention mechanisms and / or cross-modal interaction techniques to integrate and combine the generated representations of the features, - Generating (104) a unified representation based on transforming (103). [2] Method (100) according to claim 1, characterized by , that the process (100) during the generation (104) includes the following further step: - Generating a semantic map representing combined features from the received optical and electromagnetic data, wherein the semantic map includes enhanced object detection, one or more classification outputs and / or at least one transformed cross-modal data representation with respect to the environment. [3] Method (100) according to any one of the preceding claims, characterized by , that the process (100) during the processing (102) includes at least one of the following further steps: - Transforming optical sensor data into electromagnetic sensor data and vice versa, - Converting the electromagnetic sensor data into a format suitable for integration with the optical sensor data, - Adapting the optical and / or electromagnetic sensor data into a consistent input format for the transformer-based neural system (30). [4] Method (100) according to any one of the preceding claims, characterized by , that the procedure (100) includes the following further step: - Classifying and detecting objects based on the inclusion of electromagnetic sensor data obtained via radio sensing of the environment and optical data. [5] Method (100) according to any one of the preceding claims, characterized by that the optical sensor data includes at least one image and / or the electromagnetic sensor data includes at least one channel impulse response. [6] Method (300) for training a transformer-based neural system (30), in particular a transformer-based neural network (30), comprising the following steps: - Providing (301) a data set that includes optical sensor data and electromagnetic sensor data, - Preprocessing (302) the optical sensor data to normalize it and format it as inputs to the transformer-based neural system (30), - Preprocessing (303) the electromagnetic sensor data to convert it into a format suitable for integration with the optical sensor data, - Setting up (304) the transformer-based neural system (30) with initial weights and a customized configuration to handle multimodal inputs, - Feeding (305) the pre-processed optical and electromagnetic sensor data into the transformer-based neural system (30), - Transforming (306) the optical and electromagnetic sensor data via the transformer-based neural system (30) which employs one or more self-attention mechanisms and / or one or more cross-modal interaction techniques to integrate and combine the generated representations of the features, - Output (307) a unified representation that combines the generated representations of the features and - Training (308) the transformer-based neural system (30) using a loss function selected from a group that includes a cross-entropy loss, a mean squared error loss, a contrastive loss, or a combined loss. [7] Local node (10) for integrating optical and electromagnetic sensor data, comprising a local transformer-based neural system (30), in particular a transformer-based neural network (30), wherein the logical node (10) is configured to perform steps for receiving (101) optical and electromagnetic sensor data and preprocessing (102) the optical and electromagnetic sensor data via the local transformer-based neural system (30), and comprising means for sending the preprocessed data to a central node (20), in particular according to the respective steps (101, 102) of the method or methods (100, 300) according to one of the preceding claims. [8] Central node (20) for integrating optical and electromagnetic sensor data, comprising means for receiving the preprocessed data and a central transformer-based neural system (30), in particular a transformer-based neural network (30), wherein the central transformer-based neural system (30) is designed to perform steps for transforming (103) the preprocessed data and generating (104) a unified representation, in particular according to the respective steps (103, 104) of the method or methods (100, 300) according to one of the preceding claims. [9] System (30) for integrating optical and electromagnetic sensor data, comprising at least one local node (10) according to claim 7 and one central node (20) according to claim 8. [10] Computer program (50), comprising instructions which, when the program (50) is executed by a computer (11) of a local node (10), cause the computer (11) to process data by performing steps for receiving (101) and preprocessing (102) and to send the preprocessed data to a computer (21) of a central node (20), in particular according to the respective steps (101, 102) of the method or methods (100, 300) according to one of the preceding claims. [11] Computer program (60) comprising instructions which, when the program is executed by a computer (21) of a central node (20), cause the computer (21) to receive the pre-processed data from a computer (11) of a local node (10) by performing steps for transforming (103) and generating (104) in order to process the received data, in particular according to the respective steps (103, 104) of the method or methods (100, 300) according to any of the preceding claims.