A multi-dimensional target detection method based on deep learning multimodal fusion technology
By fusing multimodal features from sensor preprocessing and deep learning models, the problems of data processing and interpretability in multimodal data detection are solved, improving detection accuracy and robustness, and enhancing the model's adaptability and user trust.
Patent Information
- Application Number
- CN202510829206.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing multimodal data detection methods suffer from problems such as complex data preprocessing, difficulty in selecting learning algorithms, limited feature association extraction, need for optimization of feature fusion strategies, limited accuracy of dynamic weight adjustment mechanisms, complex implementation of real-time feedback closed-loop control systems, and limitations of interpretability mechanisms, resulting in limited detection accuracy and robustness.
Raw data is collected and preprocessed by deploying sensors. Image and sound data are processed using correction and noise reduction algorithms. Features are extracted by combining convolutional neural networks. Multimodal feature fusion is performed using a deep learning model. The model is optimized by a dynamic weight adjustment and real-time feedback closed-loop control system. An interpretability mechanism is introduced for visualization explanation.
It improves the accuracy and robustness of multimodal data detection, enhances the stability and adaptability of the model, provides reliable detection results and increases user trust, and achieves better human-computer interaction and decision support.
Smart Images

Figure CN120339645B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a multi-dimensional target detection method based on deep learning multimodal fusion technology. Background Technology
[0002] With the development of technology, multimodal data is becoming increasingly prevalent in various fields, such as medical imaging diagnosis, intelligent monitoring, and autonomous driving. In these applications, effectively fusing multimodal data and detecting targets is a significant challenge. Traditional target detection methods often only consider single-modal data, lacking comprehensive utilization of multimodal data, resulting in limited detection accuracy and robustness. To address this issue, researchers have proposed a multi-dimensional target detection method based on deep learning multimodal fusion technology.
[0003] Existing technologies suffer from drawbacks such as complex data preprocessing, difficulty in selecting learning algorithms, limited feature association extraction, need for optimization of feature fusion strategies, limited accuracy of dynamic weight adjustment mechanisms, complex implementation of real-time feedback closed-loop control systems, and limitations in interpretability mechanisms. Addressing these issues requires in-depth research and improvements in data processing, algorithm selection, feature fusion, weight adjustment, real-time feedback, and interpretability to further enhance the performance and application value of multi-dimensional target detection methods based on deep learning multimodal fusion technology. Summary of the Invention
[0004] This invention proposes a multi-dimensional target detection method based on deep learning multimodal fusion technology, which solves the problems of difficult data feature extraction and insufficient model interpretability in related technologies.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a multi-dimensional target detection method based on deep learning multimodal fusion technology, comprising the following steps:
[0006] Step S1: First, collect raw data through deployed sensors; preprocess the raw data to obtain consistent raw data;
[0007] Step S2: The original data with consistency is processed by correction algorithm and noise reduction algorithm respectively to obtain preprocessed image data and preprocessed sound data respectively;
[0008] Step S3: Input the preprocessed image data into a convolutional neural network for processing to obtain image features; process the preprocessed sound data using spectrograms, Mel spectrum, and power spectral density to obtain sound features;
[0009] Step S4: Input the image features and sound features into the deep learning model and use the attention mechanism to perform weighted fusion to obtain the fused features. Input the fused features into the classifier for processing to obtain the target detection results.
[0010] Step S5: Use a dynamic weight adjustment mechanism to collect the effectiveness information and feedback information of image features and sound features in real scenes, update the weights of the target detection results in the deep learning model based on the collected effectiveness information and feedback information, and obtain the final detection results;
[0011] Among them, a real-time feedback closed-loop control system is used to evaluate the performance of the deep learning model based on the collected validity information, feedback information and final detection results;
[0012] Among these, an interpretability mechanism is used to provide a visual explanation of the final detection results.
[0013] Furthermore, the specific process for obtaining preprocessed image data and preprocessed audio data is as follows:
[0014] First, sensors are deployed to collect raw data, including image and sound data.
[0015] The brightness, white balance, and color balance of the image data are adjusted to obtain processed image data. Then, the lens distortion in the processed image data is removed by the correction algorithm, the geometry of the processed image data is restored, and the geometry of the processed image data is adjusted to a uniform size to obtain preprocessed image data.
[0016] Noise reduction algorithms are used to remove background noise and background sounds from audio data to obtain processed audio data. Then, speech enhancement algorithms and audio segmentation are used to process the processed audio data to obtain preprocessed audio data.
[0017] Furthermore, the preprocessed image data is input into a convolutional neural network and processed sequentially through convolutional layers, pooling layers, and fully connected layers to extract local features, texture information, and high-level semantic features from the preprocessed image data. Image features are then constructed based on these local features, texture information, and high-level semantic features.
[0018] The preprocessed sound data is processed using spectrograms, Mel spectrograms, and power spectral density to obtain sound features.
[0019] Furthermore, the specific process for obtaining target detection results is as follows:
[0020] A multilayer perceptron, a deep learning model, is used to train and optimize image and sound features, learning the correlation information between them. By extracting cross-modal feature maps, the correlation information between image and sound features is aligned in the feature layer. An attention mechanism is then used in the feature layer to weight and fuse the correlation information between image and sound features to obtain fused features. The fused features are then input into a classifier for processing to obtain the target detection results.
[0021] Furthermore, the specific process for obtaining the final test results is as follows:
[0022] The effectiveness and feedback information of image features and sound features in real-world scenarios are collected through a dynamic weight adjustment mechanism. The effectiveness and feedback information are obtained by monitoring and recording the performance of deep learning models in real-world scenarios.
[0023] The weights of the target detection results in the deep learning model are updated based on the collected validity information and feedback information to obtain the final detection result.
[0024] Furthermore, the specific processing procedure of the real-time feedback closed-loop control system is as follows:
[0025] By collecting validity information and feedback information in real time and monitoring the final detection results output by the deep learning model through a real-time feedback closed-loop control system, the performance of the deep learning model is evaluated, and errors and defects of the deep learning model are promptly identified and corrected. Based on the real-time monitoring of the final detection results output by the deep learning model and the evaluation of the deep learning model's performance, the parameters of the deep learning model are updated using the gradient descent method to optimize the detection accuracy and performance of the deep learning model.
[0026] Furthermore, the specific processing procedure of the interpretability mechanism is as follows:
[0027] The final detection results are presented to the user using an interpretable mechanism, while heatmaps are used to show the regions that the deep learning model focuses on in image features, and spectrograms are used to demonstrate the deep learning model's analysis and focus on sound features.
[0028] Compared with existing technologies, the present invention has the following advantages:
[0029] (1) The present invention collects raw data by arranging sensors and preprocesses it to eliminate noise, distortion and inconsistency between image data and sound data, thereby improving the quality and comparability of image data and sound data. The consistency between image data and sound data helps to improve the stability and accuracy of deep learning models and provide more reliable results.
[0030] (2) This invention trains and optimizes image features and sound features by using a multilayer perceptron of a deep learning model. The deep learning model can learn the correlation and complementarity between different modal features and apply them to the feature extraction process. In the semantic correlation between image features and sound features, the deep learning model can learn the relationship between image features and sound features. By extracting cross-modal feature mapping, more accurate target detection can be achieved. This method of cross-modal feature extraction and complementary information fusion can make full use of the advantages of different modal features and improve the robustness and generalization ability of the deep learning model.
[0031] (3) In order to improve users’ trust and understanding of deep learning models, this invention introduces an interpretability mechanism. By using attention mechanism and visual explanation, deep learning models can provide visual explanations of decision basis and reasons, enabling users to better understand the decision-making process and reasons of deep learning models. Through the decision basis and important features of the interpretability mechanism, users can intuitively understand the reasoning process of deep learning models and obtain explanations of the decisions of deep learning models, which helps to increase users’ trust in deep learning models and provide better human-computer interaction and decision support. Attached Figure Description
[0032] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation Example 1
[0033] like Figure 1 As shown, the present invention provides a technical solution: a multi-dimensional target detection method based on deep learning multimodal fusion technology, comprising the following steps:
[0034] Step S1: First, collect raw data through deployed sensors; preprocess the raw data to obtain consistent raw data;
[0035] Step S2: The original data with consistency is processed by correction algorithm and noise reduction algorithm respectively to obtain preprocessed image data and preprocessed sound data respectively;
[0036] Step S3: Input the preprocessed image data into a convolutional neural network for processing to obtain image features; process the preprocessed sound data using spectrograms, Mel spectrum, and power spectral density to obtain sound features;
[0037] Step S4: Input the image features and sound features into the deep learning model and use the attention mechanism to perform weighted fusion to obtain the fused features. Input the fused features into the classifier for processing to obtain the target detection results.
[0038] Step S5: Use a dynamic weight adjustment mechanism to collect the effectiveness information and feedback information of image features and sound features in real scenes, update the weights of the target detection results in the deep learning model based on the collected effectiveness information and feedback information, and obtain the final detection results;
[0039] Among them, a real-time feedback closed-loop control system is used to evaluate the performance of the deep learning model based on the collected validity information, feedback information and final detection results;
[0040] Among these, an interpretability mechanism is used to provide a visual explanation of the final detection results.
[0041] The specific process for obtaining preprocessed image data and preprocessed audio data is as follows:
[0042] First, sensors are deployed, including cameras, microphones, accelerometers, etc., to collect raw data, including image data and sound data; the cameras include wide-angle or fisheye lenses.
[0043] The brightness, white balance, and color balance of the image data are adjusted to obtain processed image data. Then, the lens distortion in the processed image data is removed by the correction algorithm, the geometry of the processed image data is restored, and the geometry of the processed image data is adjusted to a uniform size to obtain preprocessed image data.
[0044] Noise reduction algorithms are used to remove background noise and background sounds from audio data to obtain processed audio data, thereby improving the quality of the audio data. The processed audio data is then further processed using speech enhancement algorithms and audio segmentation to obtain preprocessed audio data.
[0045] The preprocessed image data is input into a convolutional neural network (CNN), and then processed through convolutional layers, pooling layers and fully connected layers to extract local features, texture information and high-level semantic features from the preprocessed image data. Image features are then constructed based on the local features, texture information and high-level semantic features.
[0046] The preprocessed audio data is processed using spectrograms, mel spectrograms, and power spectral density to obtain audio features.
[0047] The specific process for obtaining target detection results is as follows:
[0048] Deep learning models are used to identify the intrinsic relationships between image and sound features. The multilayer perceptron of the deep learning model is used to train and optimize the image and sound features, learn the relationship information between image and sound features, and align the relationship information between image and sound features in the feature layer by extracting cross-modal feature maps, so as to improve the overall performance of the deep learning model.
[0049] At the feature layer, an attention mechanism is used to weight and fuse the correlation information between image features and sound features to obtain fused features. The fused features are then input into a classifier or regressor for processing to obtain the target detection results.
[0050] The specific process for obtaining the final test results is as follows:
[0051] Through a dynamic weight adjustment mechanism, the influence weights of the deep learning model on the final decision are automatically adjusted based on the effectiveness of image and sound features in real-world scenarios. This ensures that the deep learning model can make accurate and reliable detection decisions under different conditions. In implementing this dynamic weight adjustment, the effectiveness and feedback information of image and sound features in real-world scenarios are first collected. This is achieved by monitoring and recording the performance of the deep learning model in real-world scenarios. For image features, the adaptability of the deep learning model to different lighting conditions and background interference is statistically analyzed to evaluate the model's performance under these conditions. For sound feature detection accuracy, the detection accuracy of the deep learning model in noisy environments is evaluated.
[0052] By updating the weights of the target detection results in the deep learning model using validity information and feedback information, the final detection result is obtained. This makes the deep learning model more accurate and reliable in real-world scenarios. Through a dynamic weight adjustment mechanism, the adaptability and generalization ability of the deep learning model are improved, thus better adapting to different detection tasks and scenario requirements.
[0053] The specific processing procedure of the real-time feedback closed-loop control system is as follows:
[0054] A real-time feedback closed-loop control system is employed to track the detection accuracy and performance of the deep learning model in real-world scenarios. This system collects validity and feedback information in real time, monitors the final detection results output by the deep learning model, evaluates its performance, and promptly identifies and corrects errors and defects. Based on the real-time monitoring of the final detection results and the evaluation of the deep learning model's performance, gradient descent is used to update the model's parameters to optimize its detection accuracy and performance. The frequency of parameter updates can be adjusted according to actual needs to ensure the deep learning model can adapt to environmental changes and achieve better detection results.
[0055] The specific processing procedure of the interpretability mechanism is as follows:
[0056] Interpretability mechanisms enable deep learning models to provide visual explanations of their decision-making rationale and reasons. To enhance user trust and understanding of deep learning models, interpretability mechanisms present the final detection results to users in a visual manner. Simultaneously, the decision-making rationale and reasons of the deep learning model are presented to users in a visual way. For image features, heatmaps are used to show the regions that the deep learning model focuses on within the image features, explaining the model's decision-making rationale and the location of its attention. For sound features, spectrograms are used to demonstrate the deep learning model's analysis and focus on sound features. Through the decision-making rationale and key features presented by interpretability mechanisms, users can intuitively understand the reasoning process of deep learning models, which helps increase user trust in deep learning models and provides better human-computer interaction and decision support. Example 2
[0057] Based on Example 1:
[0058] In step S2, the white balance deviation of the original image data is within -50, lens distortion is removed, and preprocessed image data is obtained, with the distortion rate controlled at 0.01%.
[0059] In step S3, a convolutional neural network is used to extract features from the preprocessed image data to obtain image features. The convolutional layers in the convolutional neural network are used to extract local features from the preprocessed image data, and the number of filters in the convolutional layers is 32. The pooling kernel size in the pooling layers is 2. The feature extraction accuracy of the convolutional neural network is 85%.
[0060] In step S4, the association information between image features and sound features is learned through a deep learning model. The association information between image features and sound features is aligned in the feature layer by extracting cross-modal feature maps, with an alignment accuracy requirement of -0.1.
[0061] An attention mechanism is used to weight and fuse image features and sound features to obtain the fused features. The weight coefficient of the attention mechanism is 0.2.
[0062] The fused features are then input into a classifier or regressor for final decision-making, requiring an accuracy of over 90% for classification or regression. Example 3
[0063] Based on Example 1:
[0064] In step S2, the white balance deviation of the original image data is within +10, lens distortion is removed, and preprocessed image data is obtained, with the distortion rate controlled at 0.03%.
[0065] In step S3, a convolutional neural network is used to extract features from the preprocessed image data to obtain image features. The convolutional layers in the convolutional neural network are used to extract local features from the preprocessed image data, and the number of filters in the convolutional layers is 50. The pooling kernel size in the pooling layers is 3. The feature extraction accuracy of the convolutional neural network is 90%.
[0066] In step S4, the association information between image features and sound features is learned through a deep learning model. The association information between image features and sound features is aligned in the feature layer by extracting cross-modal feature maps, with an alignment accuracy requirement of ±0.1.
[0067] An attention mechanism is used to weight and fuse image features and sound features to obtain the fused features. The weight coefficient of the attention mechanism is 0.5.
[0068] The fused features are then input into a classifier or regressor for final decision-making, requiring an accuracy of over 90% for classification or regression. Example 4
[0069] Based on Example 1:
[0070] In step S2, the white balance deviation of the original image data is within +50, lens distortion is removed, and preprocessed image data is obtained, with the distortion rate controlled at 0.05%.
[0071] In step S3, a convolutional neural network is used to extract features from the preprocessed image data to obtain image features. The convolutional layers in the convolutional neural network are used to extract local features from the preprocessed image data, and the number of filters in the convolutional layers is 64. The pooling kernel size in the pooling layers is 4. The feature extraction accuracy of the convolutional neural network is 95%.
[0072] In step S4, the association information between image features and sound features is learned through a deep learning model. The association information between image features and sound features is aligned in the feature layer by extracting cross-modal feature maps, with an alignment accuracy requirement of ±0.1.
[0073] An attention mechanism is used to weight and fuse image features and sound features to obtain the fused features. The weight coefficient of the attention mechanism is 0.8.
[0074] The fused features are then input into a classifier or regressor for final decision-making, requiring an accuracy of over 90% for classification or regression.
[0075] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-dimensional target detection method based on deep learning multimodal fusion technology, characterized in that, Includes the following steps: Step S1: First, collect raw data through deployed sensors; preprocess the raw data to obtain consistent raw data; Step S2: The original data with consistency is processed by correction algorithm and noise reduction algorithm respectively to obtain preprocessed image data and preprocessed sound data respectively; Step S3: Input the preprocessed image data into a convolutional neural network for processing to obtain image features; process the preprocessed sound data using spectrograms, Mel spectrum, and power spectral density to obtain sound features; Step S4: Input the image features and sound features into the deep learning model and use the attention mechanism to perform weighted fusion to obtain the fused features. Input the fused features into the classifier for processing to obtain the target detection results. Step S5: Use a dynamic weight adjustment mechanism to collect the effectiveness information and feedback information of image features and sound features in real scenes, update the weights of the target detection results in the deep learning model based on the collected effectiveness information and feedback information, and obtain the final detection results; Among them, a real-time feedback closed-loop control system is used to evaluate the performance of the deep learning model based on the collected validity information, feedback information and final detection results; The specific processing procedure of the real-time feedback closed-loop control system is as follows: The system collects validity information and feedback information in real time and monitors the final detection results output by the deep learning model through a real-time feedback closed-loop control system. It evaluates the performance of the deep learning model and promptly identifies and corrects errors and defects in the deep learning model. Based on the real-time monitoring of the final detection results output by the deep learning model and the evaluation of the deep learning model's performance, the gradient descent method is used to update the parameters of the deep learning model to optimize the detection accuracy and performance of the deep learning model. Among these, an interpretability mechanism is used to provide a visual explanation of the final detection results; The specific processing procedure of the interpretability mechanism is as follows: The final detection results are presented to the user using an interpretable mechanism, while heatmaps are used to show the regions that the deep learning model focuses on in image features, and spectrograms are used to demonstrate the deep learning model's analysis and focus on sound features.
2. The target multi-dimensional detection method based on deep learning multimodal fusion technology according to claim 1, characterized in that: The specific process for obtaining preprocessed image data and preprocessed audio data is as follows: First, sensors are deployed to collect raw data, including image and sound data. The brightness, white balance, and color balance of the image data are adjusted to obtain processed image data. Then, the lens distortion in the processed image data is removed by the correction algorithm, the geometry of the processed image data is restored, and the geometry of the processed image data is adjusted to a uniform size to obtain preprocessed image data. Noise reduction algorithms are used to remove background noise and background sounds from audio data to obtain processed audio data. Then, speech enhancement algorithms and audio segmentation are used to process the processed audio data to obtain preprocessed audio data.
3. The target multi-dimensional detection method based on deep learning multimodal fusion technology according to claim 2, characterized in that: The preprocessed image data is input into a convolutional neural network and processed sequentially through convolutional layers, pooling layers, and fully connected layers to extract local features, texture information, and high-level semantic features from the preprocessed image data. Image features are then constructed based on these local features, texture information, and high-level semantic features. The preprocessed sound data is processed using spectrograms, Mel spectrograms, and power spectral density to obtain sound features.
4. The target multi-dimensional detection method based on deep learning multimodal fusion technology according to claim 3, characterized in that: The specific process for obtaining target detection results is as follows: The multilayer perceptron, a deep learning model, is used to train and optimize image and sound features, learn the correlation information between image and sound features, and align the correlation information between image and sound features in the feature layer by extracting cross-modal feature maps. In the feature layer, an attention mechanism is used to weight and fuse the correlation information between image features and sound features to obtain fused features. The fused features are then input into the classifier for processing to obtain the target detection result.
5. The target multi-dimensional detection method based on deep learning multimodal fusion technology according to claim 4, characterized in that: The specific process for obtaining the final test results is as follows: The effectiveness and feedback information of image features and sound features in real-world scenarios are collected through a dynamic weight adjustment mechanism. The effectiveness and feedback information are obtained by monitoring and recording the performance of deep learning models in real-world scenarios. The weights of the target detection results in the deep learning model are updated based on the collected validity information and feedback information to obtain the final detection result.
Citation Information
Patent Citations
Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device
CN119027775A
Multi-modal face recognition method and system based on deep learning
CN119942623A
High-resolution remote sensing image accurate classification system and method based on deep learning multi-modal fusion
CN120047721A