Battlefield emergency rescue unmanned aerial vehicle target detection method and equipment based on machine vision

Through the convolutional neural network optimized by multi-agent collaborative evolution, the problems of illumination changes, occlusion and multi-target coexistence in traditional battlefield target detection in complex environments are solved, and high-precision battlefield emergency rescue target detection is achieved.

CN120612631APending Publication Date: 2025-09-09THE 960TH HOSPITAL OF THE CHINESE PEOPLES LIBERATION ARMY JOINT LOGISTICS SUPPORT FORCE +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510779901.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional battlefield target detection methods have difficulty handling illumination changes, occlusions, multi-scale targets, and multi-target coexistence in complex environments, resulting in insufficient detection accuracy and adaptability, and unable to meet battlefield rescue needs.

Method used

A convolutional neural network based on multi-agent co-evolutionary optimization is adopted to improve the robustness and accuracy of feature extraction and target detection through multi-scale adaptive normalization, deformable co-pooling, co-evolutionary optimization algorithm and online incremental learning.

Benefits of technology

It effectively handles multi-scale targets and multi-target coexistence in complex battlefield environments, improves the accuracy and adaptability of target detection, and ensures real-time and rapid rescue efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612631A_ABST
    Figure CN120612631A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and data processing, in particular to a battlefield emergency rescue unmanned aerial vehicle target detection method and device based on machine vision. A convolutional neural network based on multi-agent co-evolution optimization is constructed to serve as a feature extraction and target detection model, the model is trained through collected images or video streams, training is completed when the number of iterations of the model is set, and then new images or video streams are input into the trained model to obtain a final detection result. And performing real-time analysis on the final detection to generate a corresponding action instruction. According to the method, the convolutional neural network based on multi-agent co-evolution optimization serves as an unmanned aerial vehicle target feature extraction and detection model, the unmanned aerial vehicle target features can be extracted and analyzed more accurately, then the rescue efficiency is improved, and the method can also be used for emergency rescue of disasters such as fire disasters, earthquakes and floods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and data processing technology, and in particular to a method and device for detecting battlefield emergency rescue drone targets based on machine vision. Background Art

[0002] The present invention relates to the field of artificial intelligence and data processing technology, and in particular to a method and device for detecting battlefield emergency rescue drone targets based on machine vision.

[0003] With the increasing complexity of modern warfare environments, traditional battlefield rescue methods and technologies face increasing challenges. Battlefield environments are often highly dynamic, characterized by complex backgrounds, variable lighting, sudden weather changes, and partial occlusion of targets. These factors place high demands on target detection technology, especially in rescue missions, where target detection accuracy directly impacts the timeliness and effectiveness of rescue efforts. Currently, traditional battlefield target detection relies heavily on manual operations or relatively simple image processing methods, but these methods often struggle to identify dynamic targets in complex environments, resolve targets of varying scales, or handle multiple interference factors. Existing target detection technologies often fall short of the target detection accuracy required for battlefield rescue missions, particularly at night, during sandstorms, or under complex lighting conditions. Furthermore, traditional target detection methods are limited in their ability to handle the coexistence and identification of multiple targets, leading to missed or false detections in practical applications.

[0004] The objective problems existing in the existing technologies are as follows: the normalization methods in the existing technologies are mostly standard normalization or simple illumination compensation, which cannot effectively adapt to the complex illumination changes, dynamic backgrounds and occlusions in battlefield environments. The quality of battlefield images is often greatly affected, affecting the accuracy of target detection. When processing images with targets of different scales, traditional convolutional neural networks often find it difficult to take into account the characteristics of small and large targets, resulting in some small targets being missed. The existing technologies have relatively simple multi-scale feature extraction mechanisms for targets and lack effective feature fusion methods. Traditional pooling operations are prone to losing spatial information when the target rotates or deforms, resulting in inaccurate representation of target features, especially when facing local occlusion and target deformation, the detection accuracy is significantly reduced. Existing technologies usually rely on a single convolutional network for target detection, which has difficulty handling the coexistence of multiple targets in complex environments. A single network is prone to falling into local optimality and ignoring the dynamic relationship between multiple targets. Traditional models often rely on batch processing and are difficult to update and adjust in real time in a dynamic battlefield environment. Sudden changes in the battlefield environment require stronger adaptability and real-time performance, which are often not effectively addressed in existing technologies.

[0005] Therefore, the present invention proposes a battlefield emergency rescue UAV target detection method and device based on machine vision to solve the above problems. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention develops a battlefield emergency rescue drone target detection method based on machine vision. The present invention uses a convolutional neural network based on multi-agent collaborative evolution optimization as a drone target feature extraction and detection model, which can more accurately extract and analyze drone target features, thereby improving rescue efficiency.

[0007] On the one hand, the technical solution to the technical problem of the present invention is a battlefield emergency rescue drone target detection method based on machine vision, comprising the following steps: S1. Collect image data through the camera on the drone and manually annotate the collected image data; S2. Construct a convolutional neural network based on multi-agent collaborative evolution optimization as a feature extraction and target detection model. Input the labeled image into the feature extraction and target detection model for training until the preset number of iterations is reached and the training is stopped. The trained feature extraction and target detection model is obtained. Then, new images or video streams are input into the trained model for processing to obtain the final target detection results. The feature extraction and target detection model includes input layer, preprocessing layer, feature extraction layer, pooling layer and decision fusion layer; S3. Analyze the final target detection results to obtain the target type, location, and threat level information, and generate corresponding action instructions based on mission requirements.

[0008] S1 is specifically as follows: the collected image data includes images or video streams. For video streams, they are decomposed into images frame by frame, and then the collected images and images obtained by decomposing the video streams are manually labeled. The labeled target categories include wounded people, rescue equipment, obstacles, threatening targets and interference background.

[0009] Furthermore, the image annotated in step S1 is input into the feature extraction and target detection model. The input layer of the model receives the image, and then the image is processed by the preprocessing layer for multi-scale adaptive normalization. The image normalization operation is as follows: The preprocessing layer performs multi-scale adaptive normalization on the image, combines contrast enhancement of local areas of the image with global illumination compensation of the image, dynamically standardizes the image, dynamically eliminates shadow interference through local mean, uses global standard deviation to constrain intensity distribution, and then combines illumination compensation factor to suppress overexposed areas, thereby reducing image noise and finally obtaining a normalized image.

[0010] Furthermore, a collaborative convolution module is used in the feature extraction layer of the feature extraction and target detection models, and a deformable collaborative pooling layer is used in the pooling layer, as follows: Feature extraction layer: A collaborative convolution module is used to perform cross-scale feature fusion through a multi-agent feature interaction mechanism. The agents are convolutional neural network submodules with different receptive fields. Each agent extracts features corresponding to a specific receptive field. Attention weighting and dynamic fusion coefficients are then combined to achieve multi-target coexistence. Pooling layer: A deformable collaborative pooling layer is used to perform adaptive feature sampling through deformation field interaction between intelligent agents, and partially occluded targets are represented by target deformation that enables the pooling window to adapt.

[0011] Furthermore, a collaborative evolutionary optimization algorithm is used in the decision fusion layer of the feature extraction and target detection model to establish a multi-objective fitness function in the parameter space and calculate the multi-objective fitness function. The calculation formula is as follows: in, Indicates the The fitness value of an agent, Indicates the The classification accuracy of each agent, represents the first fitness function hyperparameter, represents the second fitness function hyperparameter, Represents the current agent parameters and other agent parameters The mean cosine similarity of Sliding variance representing the parameter update amplitude; The elite retention strategy is adopted in the evolution process. The top 30% individuals are retained in each generation of evolution, and new individuals are generated through crossover mutation. The mutation probability , represents the maximum evolutionary generation, Represents the current evolution generation.

[0012] Furthermore, the loss function is calculated for the feature extraction and target detection model: Using multimodal contrast loss function ,distinguish between classes with high similarity by enhancing intra-class compactness and inter-class distinguishability; Constructing a cross-agent comparison constraint loss function Eliminate conflicting features between different agents, use the L2 norm to constrain the global consistency of feature maps of different agents, and make the multi-view features of the agents complement each other; Combined with the weight decay regularization term and cross entropy loss Constructing the total loss function .

[0013] Furthermore, a progressive pruning strategy based on contribution evaluation is adopted to gradually reduce the computational complexity of the model while ensuring the performance of the feature extraction and target detection models. The contribution evaluation specifically adopts the exponential moving average calculation function. ,The gradient feature product term is also used to optimize the sparsity of the UAV feature map; Among them, the pruning threshold in the progressive pruning strategy By dynamically adjusting the setting, when the agent contribution evaluation score is less than , freeze the agent parameters and start the compensation mechanism to obtain the compensation feature map.

[0014] Furthermore, the feature extraction and target detection models use online incremental learning to adapt to dynamic changes in the battlefield environment, specifically using a keyframe memory library to adapt to environmental changes; The output of each intelligent agent is fused to ensure the reliable output of other parts when some intelligent agents fail.

[0015] Furthermore, after multiple iterations of the feature extraction and target detection model, training is stopped until a preset number of iterations is reached, resulting in a trained feature extraction and target detection model. The new image is then input into the trained model for processing. The new image undergoes multi-scale adaptive normalization, dynamically adjusts illumination and contrast, and then performs forward propagation of the convolutional neural network. Multiple agents collaborate to extract multi-scale features, then adaptively adjust the feature sampling positions. Through a pruning mechanism, only agents with qualified contribution scores are activated. Subsequently, the outputs of each agent are fused through evidence theory to calculate the composite reliability. Finally, the category with the maximum reliability is taken as the final target detection result. If an environmental change is detected, online incremental learning is triggered and the feature extraction and target detection models are updated using the current frame.

[0016] On the other hand, the present invention also provides a battlefield emergency rescue drone target detection device based on machine vision, including a device for executing processing instructions for each step in a battlefield emergency rescue drone target detection method based on machine vision; Acquire real-time images or video streams through visible light cameras, infrared cameras, multispectral sensors, and infrared thermal imaging sensors with high resolution and high frame rate; High-speed storage device for storing collected image data; Embedded computing platform to support factual reasoning of deep learning models; A drone platform equipped with cameras and sensors for collecting data and executing action commands.

[0017] The effects provided in the summary of the invention are only the effects of the embodiments, rather than all the effects of the invention. The above technical solution has the following advantages or beneficial effects: (1) The present invention adopts a multi-scale adaptive normalization method, combines local area contrast enhancement with global illumination compensation, and dynamically normalizes the input image, solving the problems of illumination variation, shadow interference, and overexposed areas, improving image quality and enhancing the robustness of target detection. It can still effectively process image data in complex battlefield environments, such as at night, in sandstorms, and under overexposure conditions; (2) The present invention adopts a multi-agent collaborative convolution module, which enables each agent to process features of different scales through convolution kernels with different receptive fields. By combining the spatial attention mechanism and dynamic weight coefficient, the problem of missed detection when small and large targets appear at the same time is solved, and the detection capability of multi-scale targets is enhanced; (3) The deformable collaborative pooling layer is used to enable the pooling window to be adaptively adjusted, thereby improving the ability to represent deformable targets and maintaining good target recognition capabilities when facing partial occlusion; (4) Through the collaborative evolution optimization algorithm and the adoption of a multi-objective fitness function, the necessary differences between intelligent agents are ensured, effectively avoiding the homogenization problem. Through the cosine similarity penalty term, the feature extraction of intelligent agents is diversified, which improves the adaptability of the model in complex environments. (5) The multimodal contrast loss function is used to enhance the intra-class compactness and inter-class distinguishability, reducing the misclassification problem caused by inter-class similarity. In addition, the cross-agent contrast constraint loss function ensures the feature consistency of multi-agent outputs, further improving the accuracy of target detection; (6) Using a progressive pruning strategy and online incremental learning mechanism, the computational load of the model is gradually optimized by evaluating the contribution of the intelligent agent, thus streamlining the model while ensuring performance. At the same time, the online incremental learning module ensures that the model can quickly adapt to environmental changes. (7) Through the multi-agent decision fusion module, the outputs of multiple agents are integrated using evidence theory. When some agents fail due to interference or failure, reliable outputs can still be maintained, reducing misjudgments and improving the robustness of the system.

[0018] In summary, the present invention can effectively process target images of different scales, quickly adapt to environmental changes, and handle the problem of multi-target coexistence in complex environments, thereby improving the accuracy and precision of target detection and maintaining reliable and stable output. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0020] Figure 1 Schematic diagram of the method of the present invention.

[0021] Figure 2 The figures are comparison graphs of the performance surfaces of the traditional normalization method and the multi-scale adaptive normalization method in the present invention at three angles, wherein Figure (a-1), Figure (a-2) and Figure (a-3) are the performance surfaces of the traditional normalization method at three angles, and Figure (b-1), Figure (b-2) and Figure (b-3) are the performance surfaces of the multi-scale adaptive normalization method at three angles in the present invention.

[0022] Figure 3 This is a comparison chart of the anti-occlusion performance of the traditional pooling method and the deformable pooling method in this invention.

[0023] Figure 4 This is a comparison diagram of the feature space distribution of traditional loss and the multimodal contrast loss in the present invention. DETAILED DESCRIPTION

[0024] In order to clearly illustrate the technical features of this solution, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.

[0025] Example 1 A battlefield emergency rescue drone target detection method based on machine vision includes the following steps: S1. Collect image data through the camera on the drone and manually annotate the collected image data; S2. Construct a convolutional neural network based on multi-agent collaborative evolution optimization as a feature extraction and target detection model. Input the labeled image into the feature extraction and target detection model for training until the preset number of iterations is reached and the training is stopped. The trained feature extraction and target detection model is obtained. Then, new images or video streams are input into the trained model for processing to obtain the final target detection results. The feature extraction and target detection model includes input layer, preprocessing layer, feature extraction layer, pooling layer and decision fusion layer; S3. Analyze the final target detection results to obtain the target type, location, and threat level information, and generate corresponding action instructions based on mission requirements.

[0026] In a specific implementation manner, S1 is specifically as follows: The collected image data includes images or video streams. For video streams, they are decomposed into images frame by frame, and then the collected images and images obtained by decomposing the video streams are manually annotated. The annotated target categories include wounded people, rescue equipment, obstacles, threatening targets and interference background.

[0027] In a specific embodiment, the image annotated in step S1 is input into the feature extraction and target detection model. The input layer of the model receives the image, and then the image is processed by the preprocessing layer for multi-scale adaptive normalization. The image normalization operation is as follows: The preprocessing layer performs multi-scale adaptive normalization on the image, combines contrast enhancement of local areas of the image with global illumination compensation of the image, dynamically standardizes the image, dynamically eliminates shadow interference through local mean, uses global standard deviation to constrain intensity distribution, and then combines illumination compensation factor to suppress overexposed areas, thereby reducing image noise and finally obtaining a normalized image.

[0028] The calculation formula of the multi-scale adaptive normalization method is as follows: , in, Represents the pixel points of the input image; Represents pixel points Normalized pixel value; Represents pixel points The original pixel value of Indicated by pixel The size of the center is The window local mean of Indicates the size of the window, set ,For local occlusion and shadow interference, dynamic calculation of the local mean can eliminate the ,local intensity distortion caused by occlusion or uneven illumination; Represents the standard deviation of the global pixel values ​​of the input image. It can cope with drastic changes in global illumination, such as the alternation between bright explosions and night mode, constrain the overall intensity distribution, and prevent normalization failure caused by intensity differences across scenes. Indicates the parameter to prevent division by zero, set ; Indicates the illumination compensation factor, set ,In overexposed areas, such as areas where sunlight directly shines on metal equipment, the ,illumination compensation factor can suppress the pixel value drift in the overexposed ,area; Represents the mean of the global pixel values ​​of the original image.

[0029] In a specific embodiment, a collaborative convolution module is used in the feature extraction layer of the feature extraction and target detection model, and a deformable collaborative pooling layer is used in the pooling layer, as follows: Feature extraction layer: A collaborative convolution module is used to perform cross-scale feature fusion through a multi-agent feature interaction mechanism. The agents are convolutional neural network submodules with different receptive fields. Each agent extracts features corresponding to a specific receptive field. Attention weighting and dynamic fusion coefficients are then combined to achieve multi-target coexistence. The calculation formula in the feature extraction layer is as follows: , in, represents fusion features; Represents the normalized image of the input; Indicates the number of agents, set ; Indicates the The convolution kernel of each agent, The sizes of the convolution kernels of the four agents are 3×3, 5×5, 7×7, and 9×9 respectively. The multi-size convolution kernels capture features of different scales to cover multi-scale targets; represents the spatial attention weight of the corresponding agent, , represents a 1×1 convolution operation, Represents the Sigmoid activation function; represents the dynamic weight coefficient, , represents the Softmax function, represents global average pooling, represents a fully connected network.

[0030] Pooling layer: A deformable collaborative pooling layer is used to perform adaptive feature sampling through deformation field interaction between intelligent agents, and partially occluded targets are represented by target deformation that enables the pooling window to adapt.

[0031] The calculation formula of the pooling layer is as follows: , in, Deformation pooling output, characterizing the feature response after adaptive deformation, and coupling deformation pooling with contextual dilated convolution to occlude targets; represents a 3×3 sampling grid, is the lateral offset, is the longitudinal offset, and ; represents the learnable weight; Indicates the lateral deformation offset, Indicates the longitudinal deformation offset, The offset of the deformable pooling is generated by a preset deformation field prediction neural network to cope with target deformation. The deformation offset is predicted by context features to make the pooling window adaptive to the target shape; The calculation formula for the offset of deformable pooling is as follows: , in, Representing contextual features, it extracts them through convolution with a dilation rate of 2, resolves local occlusions, expands the receptive field to capture the context around the occluded area, and assists in predicting more reasonable deformation offsets. represents the hyperbolic tangent function; for The convolution operation limits the deformation offset to a reasonable physical range to avoid excessive distortion of features.

[0032] In a specific embodiment, a co-evolutionary optimization algorithm is used in the decision fusion layer of the feature extraction and target detection model to establish a multi-objective fitness function in the parameter space and calculate the multi-objective fitness function. The calculation formula is as follows: , in, Indicates the The fitness value of an agent, Indicates the The classification accuracy of each agent, Represents the first fitness function hyperparameter, set , Represents the second fitness function hyperparameter, set , Represents the current agent parameters and other agent parameters The mean cosine similarity of is used to prevent the homogenization of intelligent agents. The sliding variance of the parameter update amplitude is used to improve the model's anti-interference ability; The elite retention strategy is adopted in the evolution process. The top 30% individuals are retained in each generation of evolution, and new individuals are generated through crossover mutation. The mutation probability , represents the maximum evolutionary generation, Represents the current evolution generation.

[0033] In a specific implementation, the loss function is calculated for the feature extraction and target detection model: Using multimodal contrast loss function ,distinguish between classes with high similarity by enhancing intra-class compactness and inter-class distinguishability; Multimodal Contrastive Loss Function The calculation formula is as follows: , in, Represents a set of similar positive samples, including data augmentation versions such as color enhancement and rotation enhancement. represents the same type of positive samples, represents the negative sample set, represents negative samples, Indicates the temperature coefficient, set , Indicates cosine similarity calculation, represents the anchor sample feature, represents the positive sample feature, Represents negative sample features; Constructing a cross-agent comparison constraint loss function Eliminate conflicting features between different agents, use the L2 norm to constrain the global consistency of feature maps of different agents, and make the multi-view features of the agents complement each other; Cross-agent comparison constraint loss function The calculation formula is as follows: , in, Indicates the The feature map of an agent, Indicates the The feature map of an agent, represents the L2 norm, represents global average pooling; Combined with the weight decay regularization term and cross entropy loss Constructing the total loss function .

[0034] Total loss function The calculation formula is as follows: .

[0035] In the specific implementation, a progressive pruning strategy based on contribution evaluation is adopted to gradually reduce the computational complexity of the model while ensuring the performance of the feature extraction and target detection models. The contribution evaluation is specifically performed using the exponential moving average calculation function. ,The gradient feature product term is also used to optimize the sparsity of the UAV feature map; Among them, the pruning threshold in the progressive pruning strategy By dynamically adjusting the setting, when the agent contribution evaluation score is less than , freeze the agent parameters and start the compensation mechanism to obtain the compensation feature map.

[0036] The calculation formula of the progressive pruning strategy based on contribution evaluation is as follows: , in, represents the agent contribution score, Indicates the The agent contribution score of the iteration, represents the gradient feature product term, Indicates the The first iteration The feature map of an agent, Represents the exponential moving average calculation function; The calculation formula of the exponential moving average calculation function is as follows: , in, Indicates the exponential moving average calculation function attenuation factor, set , Represents the input of the exponential moving average calculation function, Indicates the attenuation factor, set , Indicates the result of the exponential moving average calculation function of the previous iteration; The pruning threshold is used to cope with sudden changes in the target's frequency. The dynamic threshold retains agents with contributions above the average level, which can avoid over-pruning caused by a fixed threshold. The pruning threshold calculation formula is as follows: , in, represents the mean of all agent contribution scores, represents the standard deviation of contribution scores; When the agent's contribution score is less than When , freeze the agent parameters and start the compensation mechanism to obtain the compensation feature map , the calculation formula is as follows: , Among them, the compensation feature map Integrates the features of the active agent output, represents an unfrozen agent, represents the learnable compensation weight.

[0037] In a specific embodiment, the feature extraction and target detection model uses online incremental learning to adapt to dynamic changes in the battlefield environment, specifically using a keyframe memory library to adapt to environmental changes; The calculation formula for online incremental learning is as follows: , , in, Indicates updating the memory bank. Indicates the entropy threshold, set , represents the keyframe memory library, which contains 1000 high entropy samples. Indicates the The predicted probability distribution of samples The entropy of , Represents the KL divergence weight coefficient, set , Indicates updating the i-th sample in the memory bank, Indicates updating the samples in the memory library. represents the Kullback-Leibler divergence, which measures the difference between the probability distributions predicted by the old and new models. Characterize the output of the new model The distribution is as close as possible to the output of the old model , which can avoid losing old knowledge due to over-adaptation to new data, Represents the adaptive loss function. In the incremental learning phase, as the loss function of incremental learning, it dynamically balances the adaptation to the new environment and the retention of old knowledge, enabling the model to achieve rapid and stable iteration in the complex and dynamic environment of the battlefield. The output of each intelligent agent is fused to ensure the reliable output of other parts when some intelligent agents fail.

[0038] The calculation formula for decision fusion is as follows: , , in, Indicates the Agent pairs The basic probability distribution of For the target categories, Indicates the Agent pairs The evidence value, Indicates the Agent pairs The evidence value, Indicates the target categories, Indicates the total number of categories, Representation category The composite reliability of Represents the basic probability distribution synthesis of A agents, Indicates that all agents are not assigned to The probability product of The final category is .

[0039] In a specific embodiment, the feature extraction and target detection model is trained after multiple iterations until a preset number of iterations is reached and training is stopped. The trained feature extraction and target detection model is obtained, and then a new image is input into the trained model for processing. The new image is subjected to multi-scale adaptive normalization, and the illumination and contrast are dynamically adjusted. Then, the forward propagation of the convolutional neural network is performed, and multiple agents collaborate to extract multi-scale features. The feature sampling position is then adaptively adjusted, and a pruning mechanism is used to activate only agents with contribution scores that meet the standards. Then, the outputs of each agent are fused through evidence theory to calculate the synthetic reliability. Finally, the maximum reliability category is taken as the final target detection result. If an environmental change is detected, online incremental learning is triggered and the feature extraction and target detection models are updated using the current frame.

[0040] Example 2 A machine vision-based battlefield emergency rescue drone target detection device, including a device for executing processing instructions for each step in a machine vision-based battlefield emergency rescue drone target detection method; Acquire real-time images or video streams through visible light cameras, infrared cameras, multispectral sensors, and infrared thermal imaging sensors with high resolution and high frame rate; High-speed storage device for storing collected image data; Embedded computing platform to support factual reasoning of deep learning models; A drone platform equipped with cameras and sensors for collecting data and executing action commands.

[0041] Example 3 like Figure 2 As shown in the figure, the detection performance under the coupling of different illumination intensities and shadow coverage is visualized by three-dimensional surface, and the environmental adaptability of the multi-scale adaptive normalization method in the present invention is compared with that of the traditional normalization method. Figure 2It can be seen that the performance surface of the traditional normalization method and the multi-scale adaptive normalization performance surface are compared from three perspectives. The traditional method shows significant performance collapse in areas with drastic lighting changes and high shadow coverage. However, the present invention maintains stable and high precision under complex lighting interactions through local mean dynamic compensation and global illumination constraint mechanism, indicating the effect of dynamic normalization processing on suppressing intensity distortion in battlefield environments.

[0042] Example 4 like Figure 3 As shown in the figure, the ability of different pooling methods to maintain feature integrity is evaluated by increasing the occlusion ratio test, and the spatial adaptation characteristics of traditional pooling (maximum pooling and average pooling) and deformable collaborative pooling are compared. Figure 3 It can be seen that the traditional method experiences a step-by-step loss of feature information when the occlusion rate increases. However, the present invention couples deformation field prediction with contextual dilated convolution to make the pooling window adaptively fit the target contour, significantly alleviating the feature misalignment problem caused by target deformation and local occlusion, demonstrating the strong characterization capability of the deformation sampling mechanism for dynamic targets on the battlefield.

[0043] Example 5 like Figure 4 As shown in the figure, we use the feature space dimensionality reduction to visualize the impact of different loss functions on the inter-class boundary and compare the fine-grained discrimination ability of traditional loss and multimodal contrast loss. Figure 4 It can be seen that traditional loss methods will cause large-scale overlap of inter-class features, while the present invention forms a separation manifold with clear boundaries through tight constraints within the class and cross-agent consistency optimization. This shows that the contrast loss mechanism in the present invention has the ability to accurately identify similar targets on the battlefield and can solve the problem of false detection caused by inter-class confusion.

[0044] Although the above describes the specific implementation methods of the invention in conjunction with the accompanying drawings, it does not limit the scope of protection of the invention. Based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.

Claims

1. A battlefield emergency rescue drone target detection method based on machine vision, characterized by: The following steps are involved: S1. Collect image data through the camera on the drone and manually annotate the collected image data; S2. Construct a convolutional neural network based on multi-agent collaborative evolution optimization as a feature extraction and target detection model. Input the labeled image into the feature extraction and target detection model for training until the preset number of iterations is reached and the training is stopped. The trained feature extraction and target detection model is obtained. Then, new images or video streams are input into the trained model for processing to obtain the final target detection results. The feature extraction and target detection model includes input layer, preprocessing layer, feature extraction layer, pooling layer and decision fusion layer; S3. Analyze the final target detection results to obtain the target type, location, and threat level information, and generate corresponding action instructions based on mission requirements.

2. The machine vision-based battlefield emergency rescue drone target detection method according to claim 1, characterized in that: S1 is as follows: The collected image data includes images or video streams. For video streams, they are decomposed into images frame by frame, and then the collected images and images obtained by decomposing the video streams are manually annotated. The annotated target categories include wounded people, rescue equipment, obstacles, threatening targets and interference background.

3. The machine vision-based battlefield emergency rescue drone target detection method according to claim 2, characterized in that: The image annotated in step S1 is input into the feature extraction and target detection model. The input layer of the model accepts the image, and then the image is processed by the preprocessing layer for multi-scale adaptive normalization. The image normalization operation is as follows: The preprocessing layer performs multi-scale adaptive normalization on the image, combines contrast enhancement of local areas of the image with global illumination compensation of the image, dynamically standardizes the image, dynamically eliminates shadow interference through local mean, uses global standard deviation to constrain intensity distribution, and then combines illumination compensation factor to suppress overexposed areas, thereby reducing image noise and finally obtaining a normalized image.

4. The method for detecting battlefield emergency rescue drone targets based on machine vision according to claim 3, wherein The feature extraction layer of the feature extraction and target detection model uses a collaborative convolution module, and the pooling layer uses a deformable collaborative pooling layer, as follows: Feature extraction layer: A collaborative convolution module is used to perform cross-scale feature fusion through a multi-agent feature interaction mechanism. The agents are convolutional neural network submodules with different receptive fields. Each agent extracts features corresponding to a specific receptive field. Attention weighting and dynamic fusion coefficients are then combined to achieve multi-target coexistence. Pooling layer: A deformable collaborative pooling layer is used to perform adaptive feature sampling through deformation field interaction between intelligent agents, and partially occluded targets are represented by target deformation that enables the pooling window to adapt.

5. The method for detecting battlefield emergency rescue drone targets based on machine vision according to claim 4 is characterized in that The decision fusion layer of the feature extraction and target detection model uses a co-evolutionary optimization algorithm to establish a multi-objective fitness function in the parameter space and calculate the multi-objective fitness function. The calculation formula is as follows: , in, Indicates the The fitness value of an agent, Indicates the The classification accuracy of each agent, represents the first fitness function hyperparameter, represents the second fitness function hyperparameter, Represents the current agent parameters and other agent parameters The mean cosine similarity of Sliding variance representing the parameter update amplitude; The elite retention strategy is adopted in the evolution process. The top 30% individuals are retained in each generation of evolution, and new individuals are generated through crossover mutation. The mutation probability , represents the maximum evolutionary generation, Represents the current evolution generation.

6. The machine vision-based battlefield emergency rescue drone target detection method according to claim 5, characterized in that: Calculate the loss function for the feature extraction and target detection model: Using multimodal contrast loss function ,distinguish between classes with high similarity by enhancing intra-class compactness and inter-class distinguishability; Constructing a cross-agent comparison constraint loss function Eliminate conflicting features between different agents, use the L2 norm to constrain the global consistency of feature maps of different agents, and make the multi-view features of the agents complement each other; Combined with the weight decay regularization term and cross entropy loss Constructing the total loss function .

7. The machine vision-based battlefield emergency rescue drone target detection method according to claim 6 is characterized by: A progressive pruning strategy based on contribution evaluation is adopted to gradually reduce the computational complexity of the model while ensuring the performance of the feature extraction and target detection models. The contribution evaluation specifically adopts the exponential moving average calculation function. ,The gradient feature product term is also used to optimize the sparsity of the UAV feature map; Among them, the pruning threshold in the progressive pruning strategy By dynamically adjusting the setting, when the agent contribution evaluation score is less than , freeze the agent parameters and start the compensation mechanism to obtain the compensation feature map.

8. The machine vision-based battlefield emergency rescue drone target detection method according to claim 7, characterized in that: The feature extraction and target detection models use online incremental learning to adapt to dynamic changes in the battlefield environment, specifically using a keyframe memory library to adapt to environmental changes; The output of each intelligent agent is fused to ensure the reliable output of other parts when some intelligent agents fail.

9. The machine vision-based battlefield emergency rescue drone target detection method according to claim 8, characterized in that: After multiple iterations of the feature extraction and target detection model, training is stopped until the preset number of iterations is reached. The trained feature extraction and target detection model is then obtained. The new image is then input into the trained model for processing. The new image undergoes multi-scale adaptive normalization, and the illumination and contrast are dynamically adjusted. The convolutional neural network then undergoes forward propagation. Multiple agents collaborate to extract multi-scale features, then adaptively adjust the feature sampling positions. Through a pruning mechanism, only agents with qualified contribution scores are activated. Next, the outputs of each agent are fused through evidence theory to calculate the composite reliability. Finally, the category with the maximum reliability is taken as the final target detection result. If an environmental change is detected, online incremental learning is triggered and the feature extraction and target detection models are updated using the current frame.

10. A machine vision-based battlefield emergency rescue drone target detection device, characterized by: A device for executing processing instructions for each step of a method for detecting a target of a battlefield emergency rescue drone based on machine vision as described in any one of claims 1 to 9; Acquire real-time images or video streams through visible light cameras, infrared cameras, multispectral sensors, and infrared thermal imaging sensors with high resolution and high frame rate; High-speed storage device for storing collected image data; Embedded computing platform to support factual reasoning of deep learning models; A drone platform equipped with cameras and sensors for collecting data and executing action commands.

Citation Information

Cited By

  • Equipment fault diagnosis method and device based on AI large model, equipment and medium

    CN121211288A