Target multi-dimensional detection method based on deep learning multi-modal fusion technology
Through multimodal fusion technology of sensor preprocessing and deep learning models, the problems existing in multimodal data detection are solved, and more stable, accurate and interpretable object detection is achieved, enhancing the robustness and adaptability of the model.
Patent Information
- Application Number
- CN202510829206.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
In the prior art, multimodal data detection methods have problems such as complex data preprocessing, difficulty in selecting learning algorithms, limited feature association extraction, unoptimized feature fusion strategy, limited accuracy of dynamic weight adjustment mechanism, complex real-time feedback closed-loop control system, and limitations of interpretability mechanism, resulting in insufficient detection accuracy and robustness.
By arranging sensors to collect raw data and preprocess them, image and sound data are processed using correction and noise reduction algorithms, features are extracted using convolutional neural networks, feature fusion is combined with multi-layer perceptrons and attention mechanisms of deep learning models, dynamic weight adjustment and real-time feedback closed-loop control system are adopted, and interpretability mechanisms are introduced for visual interpretation.
It improves the stability and accuracy of deep learning models, enhances the utilization of multimodal data, improves the robustness and generalization capabilities of detection, and improves the user's trust and understanding of the model.
Smart Images

Figure CN120339645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a target multi-dimensional detection method based on deep learning multi-modal fusion technology. Background Art
[0002] With the development of technology, multi-modal data has become increasingly common in various fields, such as medical image diagnosis, intelligent monitoring, autonomous driving, etc. In these applications, how to effectively fuse multi-modal data and detect targets is an important challenge. Traditional target detection methods often only consider single-modal data and lack the comprehensive utilization of multi-modal data, resulting in limited detection accuracy and robustness. To solve this problem, researchers have proposed a target multi-dimensional detection method based on deep learning multi-modal fusion technology.
[0003] In the prior art, there are disadvantages such as complex data preprocessing, difficult selection of learning algorithms, limited extraction of feature associations, suboptimal feature fusion strategies, limited accuracy of dynamic weight adjustment mechanisms, complex implementation of real-time feedback closed-loop control systems, and limitations of interpretability mechanisms. Solving these problems requires in-depth research and improvement in aspects such as data processing, algorithm selection, feature fusion, weight adjustment, real-time feedback, and interpretability to further enhance the performance and application value of the target multi-dimensional detection method based on deep learning multi-modal fusion technology. Summary of the Invention
[0004] The present invention proposes a target multi-dimensional detection method based on deep learning multi-modal fusion technology, which solves the problems of difficult data feature extraction and insufficient model interpretability in related technologies.
[0005] To achieve the above object, the present invention provides the following technical solution: A target multi-dimensional detection method based on deep learning multi-modal fusion technology, including the following steps: Step S1: First, collect raw data through arranged sensors; preprocess the raw data to obtain consistent raw data; Step S2: Process the consistent raw data respectively using a calibration algorithm and a noise reduction algorithm to obtain preprocessed image data and preprocessed sound data; Step S3: Input the preprocessed image data into a convolutional neural network for processing to obtain image features; process the preprocessed sound data through spectrogram, Mel spectrogram, and power spectral density to obtain sound features; Step S4: Input the image features and sound features into a deep learning model and use an attention mechanism for weighted fusion to obtain fused features, and input the fused features into a classifier for processing to obtain target detection results; Step S5: Adopt a dynamic weight adjustment mechanism to collect the effectiveness information and feedback information of image features and sound features in the real scene, and update the weights of the target detection results in the deep learning model according to the collected effectiveness information and feedback information to obtain the final detection result; Among them, a real-time feedback closed-loop control system is used to evaluate the performance of the deep learning model according to the collected effectiveness information, feedback information and the final detection result; Among them, the final detection result is visually explained through an interpretability mechanism.
[0006] Furthermore, the specific process of obtaining the preprocessed image data and preprocessed sound data is as follows: First, sensors are arranged to collect the original data through the arranged sensors. The original data includes image data and sound data; Adjust the brightness, white balance and color balance of the image data to obtain the processed image data. Then, use a calibration algorithm to remove the lens distortion in the processed image data, restore the geometric shape of the processed image data, and adjust the geometric shape size of the processed image data to a unified size to obtain the preprocessed image data; Use a noise reduction algorithm to remove the background noise and noise in the sound data to obtain the processed sound data, and process the processed sound data through a voice enhancement algorithm and audio segmentation to obtain the preprocessed sound data.
[0007] Furthermore, input the preprocessed image data into a convolutional neural network, and perform operations through a convolutional layer, a pooling layer and a fully connected layer in sequence to extract the local features, texture information and high-level semantic features of the preprocessed image data, and construct image features based on the local features, texture information and high-level semantic features; Process the preprocessed sound data using spectrogram, Mel spectrogram and power spectral density to obtain sound features.
[0008] Furthermore, the specific process of obtaining the target detection result is as follows: Use the multi-layer perceptron of the deep learning model to train and optimize the image features and sound features, learn the correlation information between the image features and sound features, and align the correlation information between the image features and sound features in the feature layer by extracting cross-modal feature maps; then use an attention mechanism in the feature layer to perform weighted fusion on the correlation information between the image features and sound features to obtain the fused features, and input the fused features into a classifier for processing to obtain the target detection result.
[0009] Furthermore, the specific process of obtaining the final detection result is as follows: Collect the effectiveness information and feedback information of image features and sound features in real scenarios through a dynamic weight adjustment mechanism, and obtain the effectiveness information and feedback information by monitoring and recording the performance of the deep learning model in real scenarios; Update the weights of the object detection results in the deep learning model according to the collected effectiveness information and feedback information to obtain the final detection results.
[0010] Furthermore, the specific processing process of the real-time feedback closed-loop control system is as follows: Collect the effectiveness information and feedback information in real time through the real-time feedback closed-loop control system and monitor the final detection results output by the deep learning model, evaluate the performance of the deep learning model, and timely discover and correct the errors and defects of the deep learning model. Based on the real-time monitoring of the final detection results output by the deep learning model and the evaluation of the performance of the deep learning model, use the gradient descent method to update the parameters of the deep learning model to optimize the detection accuracy and performance of the deep learning model.
[0011] Furthermore, the specific processing process of the interpretability mechanism is as follows: Use the interpretability mechanism to present the final detection results to the user. At the same time, use a heat map to show the areas of interest of the deep learning model in image features, and use a spectrogram to display the analysis and attention of the deep learning model to sound features.
[0012] Compared with the existing technologies, the present invention has the following beneficial effects: (1) The present invention collects raw data through the arranged sensors and performs preprocessing to eliminate the noise, distortion and inconsistency between image data and sound data, improve the quality and comparability of image data and sound data. The consistency guarantee of image data and sound data helps to improve the stability and accuracy of the deep learning model and provide more reliable results.
[0013] (2) The present invention trains and optimizes image features and sound features through the multi-layer perceptron of the deep learning model. The deep learning model can learn the associations and complementarities between different modal features and apply them to the feature extraction process. In the semantic associations between image features and sound features, the deep learning model can learn the relationships between image features and sound features. By extracting cross-modal feature mappings, more accurate object detection can be achieved. This method of cross-modal feature extraction and complementary information fusion can make full use of the advantages of different modal features and improve the robustness and generalization ability of the deep learning model.
[0014] (3) Through the interpretability mechanism, in order to improve users' trust and understanding of deep learning models, the present invention introduces an interpretability mechanism. By using the attention mechanism and visual explanations, the deep learning model can provide visual explanations of the decision-making basis and reasons, enabling users to better understand the decision-making process and reasons of the deep learning model. Through the decision-making basis and important features of the interpretability mechanism, users can intuitively understand the reasoning process of the deep learning model and obtain explanations for the decisions of the deep learning model, which helps to increase users' trust in the deep learning model and provide better human-computer interaction and decision support. Description of the Drawings
[0015] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments Embodiment 1
[0016] As Figure 1 shown, the present invention provides a technical solution: a target multi-dimensional detection method based on deep learning multi-modal fusion technology, including the following steps: Step S1: First, collect raw data through the arranged sensors; preprocess the raw data to obtain consistent raw data; Step S2: Process the consistent raw data respectively using a calibration algorithm and a noise reduction algorithm to obtain preprocessed image data and preprocessed sound data; Step S3: Input the preprocessed image data into a convolutional neural network for processing to obtain image features; process the preprocessed sound data through spectrogram, Mel spectrogram, and power spectral density to obtain sound features; Step S4: Input the image features and sound features into a deep learning model and perform weighted fusion using the attention mechanism to obtain fused features, and input the fused features into a classifier for processing to obtain a target detection result; Step S5: Adopt a dynamic weight adjustment mechanism to collect the effectiveness information and feedback information of the image features and sound features in the real scene, and update the weights of the target detection results in the deep learning model according to the collected effectiveness information and feedback information to obtain the final detection result; Among them, a real-time feedback closed-loop control system is used to evaluate the performance of the deep learning model according to the collected effectiveness information and feedback information and the final detection result; Among them, the final detection result is visually explained through the interpretability mechanism.
[0017] Among them, the specific process of obtaining the preprocessed image data and preprocessed sound data is as follows: First, sensors are arranged. The sensors include cameras, microphones, accelerometers, etc., to collect raw data, which includes image data and sound data; the camera includes a wide-angle or fish-eye lens; Adjust the brightness, white balance, and color balance of the image data to obtain processed image data. Then, use a correction algorithm to remove lens distortion in the processed image data, restore the geometric shape of the processed image data, and adjust the geometric shape size of the processed image data to a unified size to obtain preprocessed image data; Use a noise reduction algorithm to remove background noise and noise in the sound data to obtain processed sound data, so as to improve the quality of the sound data; process the processed sound data through a voice enhancement algorithm and audio segmentation to obtain preprocessed sound data.
[0018] Among them, input the preprocessed image data into a convolutional neural network (CNN), and perform operations through a convolutional layer, a pooling layer, and a fully connected layer in sequence to extract local features, texture information, and high-level semantic features of the preprocessed image data, and construct image features based on the local features, texture information, and high-level semantic features; Use spectrogram, mel spectrogram, and power spectral density to process the preprocessed sound data to obtain sound features.
[0019] Among them, the specific process of obtaining the object detection result is as follows: Use a deep learning model to identify the internal relationship between the image features and the sound features, use the multi-layer perceptron of the deep learning model to train and optimize the image features and the sound features, learn the association information between the image features and the sound features, and align the association information between the image features and the sound features in the feature layer by extracting cross-modal feature maps to improve the comprehensive performance of the deep learning model; In the feature layer, use an attention mechanism to perform weighted fusion on the association information between the image features and the sound features to obtain the fused features, and input the fused features into a classifier or a regressor for processing to obtain the object detection result.
[0020] Among them, the specific process of obtaining the final detection result is as follows: Through the dynamic weight adjustment mechanism, according to the effectiveness of image features and sound features in real scenarios, automatically adjust the influence weights of the deep learning model on the final decision, ensuring that the deep learning model can make accurate and reliable detection decisions in different situations. In the process of implementing dynamic weight adjustment, first collect the effectiveness information and feedback information of image features and sound features in real scenarios. Obtain the effectiveness information and feedback information by monitoring and recording the performance of the deep learning model in real scenarios. For image features, statistically analyze the adaptability of the deep learning model to factors such as different lighting conditions and background interference, and evaluate the deep learning model under factors such as different lighting conditions and background interference; for the detection accuracy of sound features, evaluate the detection accuracy of the deep learning model in a noisy environment. Update the weights of the object detection results in the deep learning model through the effectiveness information and feedback information to obtain the final detection result, and make the final detection result of the deep learning model more accurate and reliable in real scenarios. Through the dynamic weight adjustment mechanism, improve the adaptability and generalization ability of the deep learning model, so as to better adapt to different detection tasks and scenario requirements.
[0021] Among them, the specific processing process of the real-time feedback closed-loop control system is as follows: Adopt a real-time feedback closed-loop control system to track the detection accuracy and performance of the deep learning model in real scenarios. Collect the effectiveness information and feedback information in real time through the real-time feedback closed-loop control system and monitor the final detection result output by the deep learning model, evaluate the performance of the deep learning model, and timely discover and correct the errors and defects of the deep learning model. Based on the real-time monitoring of the final detection result output by the deep learning model and the evaluation of the performance of the deep learning model, use the gradient descent method to update the parameters of the deep learning model to optimize the detection accuracy and performance of the deep learning model. The frequency of updating the parameters of the deep learning model can be adjusted according to actual needs to ensure that the deep learning model can adapt to environmental changes in a timely manner and achieve better detection results.
[0022] Among them, the specific processing process of the interpretability mechanism is as follows: An interpretability mechanism enables deep learning models to provide visual explanations for decision-making bases and reasons. To improve users' trust and understanding of deep learning models, the interpretability mechanism is used to present the final detection results to users; at the same time, the decision-making bases and reasons of deep learning models are presented to users in a visual way. For image features, heatmaps are used to show the areas that deep learning models focus on in image features to explain the decision-making bases and the positions where attention is concentrated. For sound features, spectrograms are used to display the analysis and attention of deep learning models on sound features. Through the decision-making bases and important features of the interpretability mechanism, users can intuitively understand the reasoning process of deep learning models, which helps to increase users' trust in deep learning models and provides better human-computer interaction and decision-making support. Embodiment 2
[0023] Based on Embodiment 1: In step S2, the white balance deviation value of the image data in the original data is -50, and lens distortion is removed to obtain preprocessed image data, with the distortion rate controlled at 0.01%; In step S3, a convolutional neural network is used to extract features from the preprocessed image data to obtain image features; in the convolutional neural network, the convolutional layer is used to extract local features from the preprocessed image data, and the number of filters in the convolutional layer is 32; the pooling kernel size of the pooling layer is 2; the accuracy of feature extraction by the convolutional neural network is 85%; In step S4, the deep learning model learns the association information of image features and sound features, and aligns the association information of image features and sound features in the feature layer through extracting cross-modal feature mapping, with the alignment accuracy requirement being -0.1; The attention mechanism is used to perform weighted fusion on image features and sound features to obtain the fused features, and the weight coefficient of the attention mechanism is 0.2; The fused features are input into a classifier or a regressor for final decision-making, and the accuracy of classification or regression is required to reach more than 90%. Embodiment 3
[0024] Based on Embodiment 1: In step S2, the white balance deviation value of the image data in the original data is +10, and lens distortion is removed to obtain preprocessed image data, with the distortion rate controlled at 0.03%; In step S3, a convolutional neural network is used to extract features from the preprocessed image data to obtain image features; in the convolutional neural network, the convolutional layer is used to extract local features from the preprocessed image data, and the number of filters in the convolutional layer is 50; the pooling kernel size of the pooling layer is 3; the accuracy of feature extraction by the convolutional neural network is 90%; In step S4, the correlation information of the image features and the sound features is learned through a deep learning model, and the correlation information of the image features and the sound features is aligned in the feature layer by extracting the cross-modal feature mapping, and the alignment accuracy requirement is within ±0.1; The attention mechanism is used to perform weighted fusion on the image features and the sound features to obtain the fused features, and the weight coefficient of the attention mechanism is 0.5; The fused features are input into a classifier or a regressor for final decision-making, and the accuracy rate of classification or regression is required to reach more than 90%. Embodiment 4
[0025] Based on Embodiment 1: In step S2, the white balance deviation value of the image data in the original data is +50, and the lens distortion is removed to obtain the preprocessed image data, and the distortion rate is controlled within 0.05%; In step S3, a convolutional neural network is used to extract features from the preprocessed image data to obtain image features; in the convolutional neural network, the convolutional layer is used to extract local features from the preprocessed image data, the number of filters in the convolutional layer is 64; the pooling kernel size of the pooling layer is 4; the accuracy rate of feature extraction by the convolutional neural network is 95%; In step S4, the correlation information of the image features and the sound features is learned through a deep learning model, and the correlation information of the image features and the sound features is aligned in the feature layer by extracting the cross-modal feature mapping, and the alignment accuracy requirement is within ±0.1; The attention mechanism is used to perform weighted fusion on the image features and the sound features to obtain the fused features, and the weight coefficient of the attention mechanism is 0.8; The fused features are input into a classifier or a regressor for final decision-making, and the accuracy rate of classification or regression is required to reach more than 90%.
[0026] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A target multi-dimensional detection method based on deep learning multi-modal fusion technology, characterized in that Including the following steps: Step S1: First, collect the original data through the arranged sensors; preprocess the original data to obtain consistent original data; Step S2: Process the consistent original data using a calibration algorithm and a noise reduction algorithm respectively to obtain preprocessed image data and preprocessed sound data; Step S3: Input the preprocessed image data into a convolutional neural network for processing to obtain image features; process the preprocessed sound data through spectrogram, Mel spectrogram, and power spectral density to obtain sound features; Step S4: Input the image features and sound features into a deep learning model and use an attention mechanism for weighted fusion to obtain the fused features, and input the fused features into a classifier for processing to obtain the target detection result; Step S5: Use a dynamic weight adjustment mechanism to collect the effectiveness information and feedback information of the image features and sound features in the real scene, and update the weights of the target detection result in the deep learning model according to the collected effectiveness information and feedback information to obtain the final detection result; Among them, a real-time feedback closed-loop control system is used to evaluate the performance of the deep learning model according to the collected effectiveness information, feedback information, and the final detection result; Among them, the final detection result is visually explained through an interpretability mechanism.
2. The target multi-dimensional detection method based on the deep learning multi-modal fusion technology according to claim 1, characterized in that: The specific process of obtaining the preprocessed image data and preprocessed sound data is as follows: First, arrange the sensors, and collect the original data through the arranged sensors. The original data includes image data and sound data; Adjust the brightness, white balance, and color balance of the image data to obtain the processed image data, and then remove the lens distortion in the processed image data through a calibration algorithm, restore the geometric shape of the processed image data, and adjust the geometric shape size of the processed image data to a unified size to obtain the preprocessed image data; Use a noise reduction algorithm to remove the background noise and noise in the sound data to obtain the processed sound data, and process the processed sound data through a voice enhancement algorithm and audio segmentation to obtain the preprocessed sound data.
3. The multi-dimensional target detection method based on the deep learning multi-modal fusion technology according to claim 2, characterized in that: Input the preprocessed image data into a convolutional neural network, and perform operations through a convolutional layer, a pooling layer, and a fully connected layer in sequence to extract the local features, texture information, and high-level semantic features of the preprocessed image data, and construct image features based on the local features, texture information, and high-level semantic features; Use spectrogram, Mel spectrogram, and power spectral density to process the preprocessed sound data to obtain sound features.
4. The object multi-dimensional detection method based on the deep learning multi-modal fusion technology according to claim 3, characterized in that: The specific process of obtaining the target detection result is as follows: Use the multi-layer perceptron of the deep learning model to train and optimize the image features and sound features, learn the correlation information between the image features and sound features, and align the correlation information between the image features and sound features in the feature layer by extracting cross-modal feature maps; In the feature layer, use an attention mechanism to perform weighted fusion on the correlation information between the image features and sound features to obtain the fused features, and input the fused features into a classifier for processing to obtain the target detection result.
5. The target multi-dimensional detection method based on the deep learning multi-modal fusion technology according to claim 4, characterized in that: The specific process of obtaining the final detection result is as follows: Collect the validity information and feedback information of image features and sound features in the real scene through the dynamic weight adjustment mechanism, and obtain the validity information and feedback information by monitoring and recording the performance of the deep learning model in the real scene; Update the weight of the target detection result in the deep learning model according to the collected validity information and feedback information to obtain the final detection result.
6. The target multi-dimensional detection method based on the deep learning multi-modal fusion technology according to claim 5, characterized in that: The specific processing process of the real-time feedback closed-loop control system is as follows: Collect the validity information and feedback information in real time through the real-time feedback closed-loop control system and monitor the final detection result output by the deep learning model, evaluate the performance of the deep learning model, and timely discover and correct the errors and defects of the deep learning model. Based on the real-time monitoring of the final detection result output by the deep learning model and the evaluation of the performance of the deep learning model, use the gradient descent method to update the parameters of the deep learning model to optimize the detection accuracy and performance of the deep learning model.
7. The target multi-dimensional detection method based on the deep learning multi-modal fusion technology according to claim 6, characterized in that: The specific processing process of the interpretability mechanism is as follows: Use the interpretability mechanism to present the final detection result to the user, and at the same time use the heat map to show the area of interest of the deep learning model in the image features, and use the spectrogram to display the analysis and attention of the deep learning model to the sound features.
Citation Information
Patent Citations
Distributed sludge detection method based on deep learning
CN118351486A
Image recognition monitoring method based on unmanned aerial vehicle remote sensing technology
CN118823574A
Image processing analysis system based on deep learning
CN118840646A
Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device
CN119027775A
A method and system for automatic classification of video content recognition
CN119741636A