Road environment detection method of multi-modal data fusion mechanism based on federal learning
Through the combination of generator and federated learning, the semantic fusion of visible light images and infrared images is achieved, solving the problems of data privacy protection and detection accuracy, and improving the detection capabilities of the autonomous driving system.
Patent Information
- Application Number
- CN202410326130.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-07-18
AI Technical Summary
In intelligent driving, it is difficult for the prior art to effectively integrate visible light images and infrared images while protecting data privacy to achieve accurate road environment detection.
Generator is used to perform semantic fusion of visible light images and infrared images, and the road environment detection model is trained through federated learning. High-quality fusion images are generated using the generative adversarial network, and weighted average of model parameters is performed on the client to achieve data privacy protection and accuracy of target detection.
It improves the accuracy of road environment detection and improves the accuracy of target detection while protecting data privacy. It is suitable for autonomous driving applications in diverse scenarios.
Smart Images

Figure CN120339976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a road environment detection method based on a multi-modal data fusion mechanism of federated learning, and belongs to the technical field of vehicle autonomous driving. Background Art
[0002] With the development of intelligent vehicles, autonomous driving has received increasing attention. By obtaining visible light image data through an in-vehicle camera and analyzing the image using an object detection model, road environment information can be obtained, providing effective information for the planning and control systems of autonomous driving vehicles, thereby improving the safety and ride comfort of the vehicles. However, in intelligent driving, relying solely on the visible light image data provided by the in-vehicle visible light camera is unreliable, and the visible light cameras used in intelligent driving systems are very sensitive to light conditions. Changes in these conditions can significantly affect the quality of the image, thus affecting the accuracy of object detection. Under low light conditions such as at night or in tunnels, the images captured by the visible light camera may become very dark, and details are difficult to identify, making object detection and recognition more difficult. When the vehicle is under strong light and reflection conditions, strong sunlight or other light sources (such as the headlights of oncoming vehicles) may cause the image to be overexposed, affecting the visibility of objects. Or when the light suddenly changes in scenarios such as entering or exiting a tunnel, it will also affect the camera detection. To cope with these scenarios, intelligent driving systems usually need to integrate data collected by other sensors to improve the detection accuracy and reliability in different environments. Integrating infrared thermal imaging technology into the object detection system of intelligent driving can bring various enhancements and advantages. Infrared thermal imaging does not depend on environmental light, so it can work effectively at night or under low light conditions. At the same time, in harsh weather conditions such as fog, smoke, and dust, infrared thermal imaging can penetrate these obstacles and provide clear images, enabling better identification and classification of some objects that are partially occluded or have a similar color to the background. Infrared thermal imaging also has better performance in distinguishing living bodies (such as pedestrians, animals) from non-living objects compared to visible light images. Since the body temperatures of humans and animals are usually higher than the surrounding environment, infrared thermal imaging can effectively identify these important objects, improving the detection accuracy and timeliness of pedestrians and animals in complex urban or rural environments. In terms of remote monitoring capabilities, infrared thermal imaging can detect heat sources at relatively long distances, enhancing the perception of distant objects. For a vehicle traveling at high speed, it can identify distant obstacles earlier and increase the reaction time.
[0003] Therefore, it is very important for intelligent vehicle driving to fuse visible light images and infrared images through multi-modal data fusion technology and explore the potential value of data. However, the image data information collected by intelligent vehicles may include user data such as vehicle geographical location and road environment outside the vehicle. Once leaked, it will violate personal privacy. Moreover, autonomous driving models often need to collect a large amount of data collected by vehicles to train a model with good performance, and vehicle data processors should adhere to principles such as in-vehicle processing, default non-collection, and desensitization processing during vehicle data processing activities. Under the current increasingly strict requirements for data privacy protection, how to effectively fuse multi-modal data while protecting data privacy and achieve intelligent road environment detection is a very important issue. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a road environment detection method based on a multi-modal data fusion mechanism of federated learning. First, the semantic fusion of visible light images and infrared images is performed by a generator, which can ensure the accuracy of subsequent detection results. Second, based on the object detection training method of federated learning, the accuracy of object detection can be improved while protecting data privacy.
[0005] To achieve the above purpose, the present invention is implemented by the following technical solutions:
[0006] The present invention discloses a road environment detection method based on a multi-modal data fusion mechanism of federated learning, including the following steps:
[0007] Obtain multi-modal data to be detected, where the multi-modal data includes visible light environment images and corresponding infrared environment images;
[0008] Input the multi-modal data to be detected into a trained generator to obtain a fused image;
[0009] Input the fused image into a trained road environment detection model to obtain corresponding environment detection results;
[0010] Among them, the road environment detection model is trained using federated learning, and the training method is as follows:
[0011] Based on the server, the initial model parameters of the road environment detection model are sent to each client respectively;
[0012] Based on any client, according to the initial model parameters, use the local image training set to train and update the road environment detection model to obtain local model parameters and send them to the server;
[0013] Based on the server, according to the local model parameters of all clients, the weighted average method is used to obtain the global model parameters and send them to each client for iterative update training until the final global model parameters are obtained;
[0014] Based on any client, according to the final global model parameters, a trained road environment detection model is obtained.
[0015] Furthermore, the image features of the fused image include the detail features of the visible light environment image and the thermal features of the infrared environment image.
[0016] Furthermore, the training method of the generator is as follows:
[0017] Construct an adversarial generation network, which includes a generator and a discriminator;
[0018] Obtain a multi-modal training set, which includes visible light environment training images, corresponding infrared environment training images, and real fused training images;
[0019] According to the multi-modal training set, perform iterative training on the adversarial generation network until the preset iterative termination condition is met, and obtain a trained generator;
[0020] Among them, each iterative training includes the following steps:
[0021] Based on the generator with the original parameters or updated parameters in the previous iteration, according to the visible light environment training image and the corresponding infrared environment training image, obtain the generated fused training image;
[0022] Based on the discriminator with the original parameters or updated parameters in the previous iteration, evaluate the authenticity of the generated fused training image according to the real fused training image;
[0023] Calculate the generator loss function and the discriminator loss function respectively;
[0024] By minimizing the generator loss function and the discriminator loss function, update the parameters of the generator and the discriminator respectively, and obtain the generator and the discriminator with the current iteration parameters updated.
[0025] Furthermore, the expression of the generator loss function is as follows:
[0026] L G =H(1,D(G(z)))
[0027] Among them, L Grepresents the generator loss function; H represents the cross-entropy function; D represents the generator network; G represents the discriminator network; z represents the input generated fused training image; D(G(z)) represents the judgment probability of the generated fused training image, where 1 represents that the data is absolutely real and 0 represents that the data is absolutely false; H(1, D(G(z)))) represents the distance between the judgment result and 1.
[0028] Further, the expression of the discriminator loss function is as follows:
[0029] L D = H(1, D(x)) + H(0, D(G(z)))
[0030] where, L D represents; the discriminator loss function represents; H represents the cross-entropy function; D represents the generator network; G represents the discriminator network; z represents the input generated fused training image; x represents the input real fused training image; H(1, D(x))) represents the distance between the real fused training image and 1; H(0, D(G(z))) represents the distance between the generated fused training image and 0.
[0031] Further, the road environment detection model includes the YOLOV5 model.
[0032] Further, the local image training set includes multiple local fused training images; the local fused training images are generated by a trained generator according to the local visible light environment training images and the corresponding infrared environment training images.
[0033] Further, the update expression of the local model parameters is as follows:
[0034]
[0035] where, represents the local model parameters of client n at the (t + 1)-th iteration; represents the local model parameters of client n at the t-th iteration; η represents the learning rate; represents the gradient of the loss function L calculated based on the image training set of local client n; D n represents the image training set of local client n.
[0036] Further, the expression of the global model parameters is as follows:
[0037]
[0038] where, Θ (t+1) represents the global model parameters at the (t + 1)-th iteration; N represents the number of clients; Denote the local model parameters of client n at the (t + 1)-th iteration.
[0039] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0040] For the road environment detection method based on the multi-modal data fusion mechanism of federated learning of the present invention, first, a generator is used to perform semantic fusion of visible light images and infrared images. The semantic information contained in the obtained fused image can provide more information for subsequent model detection, ensuring the accuracy of the detection results. Second, the object detection training method based on federated learning can solve the privacy and security problems of multi-modal fusion data during the model training process, and can improve the accuracy of object detection while protecting data privacy. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic flow chart of a road environment detection method based on a multi-modal data fusion mechanism of federated learning;
[0042] Figure 2 is a schematic diagram of semantic fusion of visible light images and infrared images based on a generative adversarial network. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and should not be used to limit the protection scope of the present invention.
[0044] This embodiment provides a road environment detection method based on a multi-modal data fusion mechanism of federated learning, including the following steps:
[0045] Obtain multi-modal data to be detected, where the multi-modal data includes visible light environment images and corresponding infrared environment images;
[0046] Input the multi-modal data to be detected into a trained generator to obtain a fused image;
[0047] Input the fused image into a trained road environment detection model to obtain corresponding environment detection results;
[0048] Among them, the road environment detection model is trained using federated learning, and the training method is as follows:
[0049] Based on the server, send the initial model parameters of the road environment detection model to each client respectively;
[0050] Based on any client, according to the initial model parameters, use the local image training set to train and update the road environment detection model, obtain local model parameters and send them to the server;
[0051] Based on the server, according to the local model parameters of all clients, the weighted average method is used to obtain the global model parameters and send them to each client for iterative update training until the final global model parameters are obtained;
[0052] Based on any client, according to the final global model parameters, a trained road environment detection model is obtained.
[0053] The technical concept of the present invention is as follows: First, the generator can perform semantic fusion of visible light images and infrared images, and the semantic information contained in the obtained fusion image can provide more information for subsequent model detection, ensuring the accuracy of the detection result. Second, based on the object detection training method of federated learning, it can solve the privacy and security problems of multi-modal fusion data in the model training process, and improve the accuracy of object detection while protecting data privacy.
[0054] First, train the generator based on the adversarial generative network.
[0055] This embodiment proposes a semantic fusion mechanism for visible light and infrared images. This mechanism can regard visible light images and infrared images as different semantic objects, and generate a fusion image containing semantic information through generative adversarial technology. The image features of the fusion image include the detail features of the visible light environment image and the thermal features of the infrared environment image, which can provide more feature information for subsequent model training and ensure the accuracy of model training.
[0056] A Generative Adversarial Network (GAN) consists of two competing neural networks: a Generator and a Discriminator. The task of the Generator is to accept random noise or input data and transform it into synthetic data similar to real data. It is typically composed of a deep neural network that can learn the distribution of the generated data. The goal of the Generator is to continuously improve the similarity between the generated data and the real data. The task of the Discriminator is to distinguish between the data generated by the Generator and the real data. It is also a deep neural network that accepts inputs from the Generator and real data and attempts to correctly classify them as "real" or "synthetic". The goal of the Discriminator is to distinguish between these two types of data as accurately as possible. The core idea of GAN is to train the model through the competition between the Generator and the Discriminator. The Generator attempts to generate increasingly realistic data to deceive the Discriminator, while the Discriminator attempts to become increasingly intelligent to better distinguish between real data and generated data. This competitive process leads to continuous improvement of the model. GAN uses two loss functions to train the Generator and the Discriminator. The loss function of the Generator encourages the generated data to be more similar to the real data, while the loss function of the Discriminator encourages it to correctly classify the data. The loss functions of the Generator and the Discriminator are adversarial to each other.
[0057] As Figure 2 shown, first, the visible light environment training images and the infrared environment training images are concatenated at the channel level and then fed into the Generator. The Generator learns how to combine the two types of images. It improves its fusion method by attempting to deceive the Discriminator. The Generator tries to create images that contain both the detailed features of the visible light images and the thermal features of the infrared images. The finally output image is the fused image. The Discriminator tries to distinguish between the generated fused image and the real fused image. This adversarial training helps the Generator better learn the fusion technique. After multiple iterations, the Generator is able to create high-quality fused images for subsequent object detection model training. This method uses the two main components of the adversarial generation network, the Discriminator and the Generator. After training, the Generator can create images that fuse visible light and infrared features. These images are visually rich and informative and are suitable for further model training.
[0058] Specifically, the training method of the Generator is as follows:
[0059] Construct an adversarial generation network, which includes a Generator and a Discriminator;
[0060] Obtain a multi-modal training set, which includes visible light environment training images, corresponding infrared environment training images, and real fused training images;
[0061] Iteratively train the adversarial generation network according to the multi-modal training set until the preset iteration termination condition is met, and obtain the trained generator;
[0062] Among them, each iterative training includes the following steps:
[0063] Based on the generator updated with the original parameters or the parameters updated in the previous iteration, obtain the generated fused training images according to the visible light environment training images and the corresponding infrared environment training images;
[0064] Based on the discriminator updated with the original parameters or the parameters updated in the previous iteration, evaluate the authenticity of the generated fused training images according to the real fused training images;
[0065] Calculate the generator loss function and the discriminator loss function respectively;
[0066] By minimizing the generator loss function and the discriminator loss function, update the parameters of the generator and the discriminator respectively to obtain the generator and the discriminator with the current iteration parameters updated.
[0067] Furthermore, the expression of the generator loss function is as follows:
[0068] L G =H(1, D(G(z)))
[0069] Among them, L G represents the generator loss function; H represents the cross-entropy function; D represents the generator network; G represents the discriminator network; z represents the input generated fused training image; D(G(z)) represents the judgment probability of the generated fused training image, 1 represents that the data is absolutely real, 0 represents that the data is absolutely false; H(1, D(G(z)))) represents the distance between the judgment result and 1. The standard for the generator to achieve good results is to make the evaluation of the generated data in the discriminator approach 1, that is, to make the discriminator judge the generated data as real data.
[0070] Cross-Entropy is a concept commonly used in information theory and machine learning, usually used to measure the difference or uncertainty between two probability distributions. In machine learning, cross-entropy is often used as a loss function, especially in classification problems, to measure the difference between the model's prediction and the actual label.
[0071] Furthermore, the expression of the discriminator loss function is as follows:
[0072] L D =H(1, D(x)) + H(0, D(G(z)))
[0073] Among them, L DRepresentation; discriminator loss function representation; H represents the cross-entropy function; D represents the generator network; G represents the discriminator network; z represents the input generated fused training image; x represents the input real fused training image; H(1, D(x)) represents the distance between the real fused training image and 1; H(0, D(G(z))) represents the distance between the generated fused training image and 0. Obviously, to achieve good performance, it is necessary to make the evaluation of real data in the discriminator approach 1, while the evaluation of generated data approaches 0.
[0074] Second, federated learning is used to train the road environment detection model.
[0075] The road environment detection model in this embodiment uses the YOLOV5 model. The YOLOv5 model is an advanced deep neural network architecture dedicated to real-time object detection. It adopts a single-shot multi-scale prediction method, which can simultaneously predict multiple bounding boxes and class probabilities in a single network forward pass, achieving high-precision and high-speed detection performance. The federated learning in this embodiment is a distributed machine learning method that allows multiple devices or servers to collaboratively train a model while keeping the data local, thereby improving privacy protection and reducing data transmission requirements.
[0076] Training the YOLOV5 model under the federated learning framework has significant advantages in protecting data privacy, enhancing model generalization ability, and improving data processing efficiency compared to traditional centralized learning methods. Federated learning allows the model to be trained locally, sharing only model parameters instead of raw data. Since the data is processed locally and only model updates need to be transmitted, the amount of data transmission and related costs can be significantly reduced. Since the model is independently trained in different vehicles and environments, the application ability and generalization of the model in diverse scenarios can be improved. Compared with centralized training, federated learning reduces the dependence on a central server and the demand for large-scale computing resources. In contrast, traditional centralized learning methods face various challenges. First, centralized data processing involves uploading all training data to a single server or data center, which not only increases the risk of unauthorized access or leakage of data but also may cause significant resource consumption due to data transmission and storage. Second, centralized training relies on a single data source or limited data samples, which may lead to insufficient model generalization ability, especially in diverse and non-uniform real-world scenarios. In addition, centralized learning faces challenges in computational efficiency and scalability when dealing with large-scale, heterogeneous data sets.
[0077] Federated learning is a method that conducts model training on multiple devices simultaneously and aggregates the updated model parameters. This method can effectively protect users' data privacy. In federated learning, the original data remains on local devices and does not need to be transmitted to a central server. Only the updated model parameters are sent and aggregated. This helps prevent the leakage or abuse of users' personal data. In addition, federated learning is also applicable to distributed data processing. In many scenarios, data is distributed across different geographical locations or organizations. Traditional centralized learning methods require data to be centrally stored in one place, which may incur significant data transmission and storage costs. Federated learning allows model training to be carried out while the data remains dispersed, reducing data transmission and storage requirements.
[0078] This embodiment combines machine learning and federated learning to construct a multi-vehicle collaborative intelligent road environment perception technology based on federated learning, and designs a machine learning training framework to enable collaboration among participating parties without directly exchanging user data. On the premise of protecting user data privacy, the data information of the road environment is processed and trained using a machine learning model to achieve intelligent road environment prediction and provide effective support information for autonomous driving vehicles.
[0079] As Figure 1 shown, in this federated learning architecture, there are n independent clients (cars), each of which holds a visible light and infrared training image dataset. It should be noted that the local image training set includes multiple local fused training images; the local fused training images are generated by a trained generator based on the local visible light environment training images and the corresponding infrared environment training images. The clients are connected to a central server, which is responsible for coordinating the model training process of these clients to achieve cross-node parameter optimization and aggregation. This method allows each client to independently process its multimodal data (visible light and infrared images) and perform effective model training and parameter update through the federated learning protocol without sharing the original data.
[0080] The training process of the model is divided into the following stages:
[0081] (1) The central server sends the initial model parameters to all participating clients and ensures that each client has the same model when starting training.
[0082] (2) Each client preprocesses its local multimodal data (visible light environment training images and the corresponding infrared environment training images) and generates new fused training images through semantic fusion. Then, the locally generated fused training images after semantic fusion are used to train the model to update the local model parameters. Finally, the trained local model parameters are uploaded to the central server.
[0083] (3) The central server collects all the uploaded local model parameters, uses weighted averaging to synthesize the global model parameters, and sends the updated global model parameters back to each client.
[0084] (4) The client receives the updated global model parameters from the server and synchronizes its local road environment detection model with these global model parameters. The client will test the synchronized road environment detection model on the local validation set to evaluate its performance.
[0085] Specifically, each client n uses its local image training set D n to update the model parameters. The update expression of the local model parameters is as follows:
[0086]
[0087] where represents the local model parameters of client n at the (t + 1)-th iteration; represents the local model parameters of client n at the t-th iteration; η represents the learning rate; represents the gradient of the loss function L calculated based on the image training set of local client n; D n represents the image training set of local client n.
[0088] The central server receives the updated model parameters from all clients and aggregates them. The expression of the global model parameters is as follows:
[0089]
[0090] where Θ (t+1) represents the global model parameters at the (t + 1)-th iteration; N represents the number of clients; represents the local model parameters of client n at the (t + 1)-th iteration.
[0091] In summary, this embodiment first proposes a semantic fusion mechanism for visible light and infrared images, formulates the image fusion problem as the fusion of different semantic objects, and uses generative adversarial technology to fuse different semantic objects. The fused image retains more semantic information, is visually rich and informative, is suitable for further model training, provides more information for subsequent model training, and ensures the accuracy of model training. Secondly, to solve the privacy and security problems of multi-modal fusion data in the model training process, a target detection method based on the federated learning framework is proposed. Federated learning can achieve data availability without visibility, providing an effective solution for in-vehicle data processing, improving the accuracy of target detection while protecting data privacy.
[0092] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0093] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0094] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0096] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A road environment detection method based on a multi-modal data fusion mechanism of federated learning, characterized in that It includes the following steps: Obtain the multi-modal data to be detected, where the multi-modal data includes a visible light environment image and a corresponding infrared environment image; Input the multi-modal data to be detected into the trained generator to obtain a fused image; Input the fused image into the trained road environment detection model to obtain the corresponding environment detection result; Among them, the road environment detection model is trained using federated learning, and the training method is as follows: Based on the server, send the initial model parameters of the road environment detection model to each client respectively; Based on any client, according to the initial model parameters, use the local image training set to train and update the road environment detection model to obtain the local model parameters and send them to the server; Based on the server, according to the local model parameters of all clients, use the weighted average method to obtain the global model parameters and send them to each client for iterative update training until the final global model parameters are obtained; Based on any client, according to the final global model parameters, obtain the trained road environment detection model.
2. The road environment detection method of the multi-modal data fusion mechanism based on federated learning according to claim 1, characterized in that, The image features of the fused image include the detailed features of the visible light environment image and the thermal features of the infrared environment image.
3. The road environment detection method based on the multi-modal data fusion mechanism of federated learning according to claim 1, characterized in that, The training method of the generator is as follows: Construct an adversarial generation network, which includes a generator and a discriminator; Obtain a multi-modal training set, which includes visible light environment training images, corresponding infrared environment training images, and real fused training images; According to the multi-modal training set, perform iterative training on the adversarial generation network until the preset iterative termination condition is met to obtain the trained generator; Among them, each iterative training includes the following steps: Based on the generator with the original parameters or the parameters updated in the previous iteration, according to the visible light environment training image and the corresponding infrared environment training image, obtain the generated fused training image; Based on the discriminator with the original parameters or the parameters updated in the previous iteration, evaluate the authenticity of the generated fused training image according to the real fused training image; Calculate the generator loss function and the discriminator loss function respectively; By minimizing the generator loss function and the discriminator loss function, update the parameters of the generator and the discriminator respectively to obtain the generator and the discriminator with the parameters updated in the current iteration.
4. The road environment detection method of the multi-modal data fusion mechanism based on federated learning according to claim 3, characterized in that, The expression of the generator loss function is as follows: L G = H(1, D(G(z))) Among them, L G represents the generator loss function; H represents the cross-entropy function; D represents the generator network; G represents the discriminator network; z represents the input generated fused training image; D(G(z)) represents the judgment probability of the generated fused training image, 1 represents that the data is absolutely real, and 0 represents that the data is absolutely fake; H(1, D(G(z)))) represents the distance between the judgment result and 1.
5. The road environment detection method of the multi-modal data fusion mechanism based on federated learning according to claim 3, characterized in that, The expression of the discriminator loss function is as follows: L D = H(1, D(x)) + H(0, D(G(z))) Among them, L D denotes; the discriminator loss function denotes; H denotes the cross-entropy function; D denotes the generator network; G denotes the discriminator network; z denotes the input generated fused training image; x denotes the input real fused training image; H(1, D(x)) represents the distance between the real fused training image and 1; H(0, D(G(z))) represents the distance between the generated fused training image and 0.
6. The road environment detection method of the multimodal data fusion mechanism based on federated learning according to claim 1, characterized in that, The road environment detection model includes the YOLOV5 model.
7. The road environment detection method based on the multi-modal data fusion mechanism of federated learning according to claim 1, characterized in that, The local image training set includes multiple local fused training images; the local fused training images are generated by the trained generator according to the local visible light environment training images and the corresponding infrared environment training images.
8. The road environment detection method based on the multi-modal data fusion mechanism of federated learning according to claim 1, characterized in that The update expression of the local model parameters is as follows: Among them, represents the local model parameters of client n at the (t + 1)-th iteration; represents the local model parameters of client n at the t-th iteration; η represents the learning rate; represents the gradient of the loss function L calculated based on the image training set of local client n; D n represents the image training set of local client n.
9. The road environment detection method based on the multi-modal data fusion mechanism of federated learning according to claim 1, characterized in that, The expression of the global model parameters is as follows: Among them, Θ (t+1) represents the global model parameters at the (t + 1)-th iteration; N represents the number of clients; represents the local model parameters of client n at the (t + 1)-th iteration.