A camouflage target detection system and detection method for UAV scenarios

Through the boundary enhancement module and perception decoder of GBNet, the problem of vague boundary boundaries of camouflage in complex environments is solved, and higher recognition accuracy and segmentation accuracy are achieved, adapting to changes in complex scenarios.

CN118644795BActive Publication Date: 2025-06-17SHANDONG WEIRAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410972777.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2025-06-17
Estimated Expiration
2044-07-19

AI Technical Summary

Technical Problem

In complex environments, the boundaries of the camouflage target are blurred, making it difficult to accurately identify and locate the camouflage object, and the prior art performs poorly when dealing with boundary-related details.

Method used

A gated boundary awareness network GBNet is proposed, and a boundary enhancement module and boundary awareness decoder are introduced. Background information is selectively filtered through gated convolution blocks and injected boundary enhancement features to improve the accuracy of boundary recognition.

Benefits of technology

It improves the identification accuracy and segmentation accuracy of camouflage target boundaries, enhances the ability to understand camouflage target boundaries in complex scenarios, and effectively responds to challenges such as lighting changes and perspective distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118644795B_ABST
    Figure CN118644795B_ABST
Patent Text Reader

Abstract

The present invention provides a camouflage object detection model and detection method for drone scenarios, belonging to the technical field of drone image recognition and segmentation based on computer vision. The model of the present invention includes a boundary enhancement module and a boundary-aware decoder. The boundary-aware module selectively filters out redundant background information through gated convolutional blocks, providing effective guidance for accurately generating boundaries. At the same time, the boundary-aware decoder further enriches the representational ability of the decoder by injecting boundary enhancement features. The introduction of the boundary enhancement module aims to address the problem of blurred boundaries of camouflage targets and selectively filter out redundant background information through gated convolutional blocks. The design of the boundary-aware decoder further emphasizes the key role of boundary features in the process of identifying camouflage targets. By cleverly injecting boundary enhancement features, the decoder's ability to understand the boundaries of camouflage targets in complex scenarios is strengthened, thus more accurately completing the target segmentation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision-based UAV image recognition and segmentation, and particularly relates to a camouflage object target detection system and detection method for UAV scenarios. Background Art

[0002] Today, with the continuous development of UAV technology, UAVs are increasingly widely used in military, agricultural, environmental monitoring and other fields. However, in complex environments, identifying and locating camouflaged targets remains a daunting task. The boundaries of camouflaged targets are usually blurred, which makes it quite challenging to accurately identify the camouflaged objects blended into the complex scene. Therefore, it is necessary to overcome the blurred boundaries to ensure the precise positioning of camouflaged objects, especially in a changing environment.

[0003] With the rapid development of computer vision, significant progress has been made in camouflaged target detection methods. Currently, these methods can be broadly classified into three main types. The first type involves the meticulous design of network modules or architectures to deeply explore the distinguishable features of camouflaged targets. The second type adopts a multi-task joint learning method to extract features from multiple dimensions, including image depth, frequency domain, and the boundaries of camouflaged targets. The third type adopts a bionic method to simulate the mechanism in the predation process to improve the performance of camouflaged target detection.

[0004] Despite the considerable achievements of these methods, a daunting challenge still remains: the boundaries of camouflaged targets are usually inherently blurred, making it difficult to accurately identify the camouflaged objects blended into the complex scene. In addition, it is worth noting that although boundary clues are introduced, due to the insufficient consideration of the complexity of the edge background by models such as BGNet, irrelevant information cannot be effectively filtered out, resulting in poor performance when dealing with boundary-related details. Summary of the Invention

[0005] In view of the above problems, the present invention proposes an innovative gated boundary-aware network GBNet to address various challenges in the field of camouflaged target detection. Specifically, GBNet introduces a novel boundary enhancement module that selectively filters out redundant background information through gated convolutional blocks, providing effective guidance for accurately generating boundaries. At the same time, a novel boundary-aware decoder is also designed to further enrich the representational ability of the decoder by injecting boundary-enhanced features.

[0006] The first aspect of the present invention proposes a camouflage object detection model for UAV scenarios, characterized in that: for a given input image H*W*C, its spatial resolution is H*W, with C channels, and each pixel is assigned a class label, where 0 represents the background and 1 represents the camouflage object; the camouflage object detection model aims to predict a pixel-by-pixel label map and includes a backbone network, a boundary enhancement module BEM, and a boundary-aware decoder connected in sequence;

[0007] The backbone network uses a pyramid vision transformer as an encoder to extract features at different stages and send them to the boundary enhancement module;

[0008] The boundary enhancement module uses a gated convolutional module GCB to generate high-quality boundary cues and integrates them into the boundary-aware decoder, thereby enriching the representation ability of the decoder features;

[0009] The boundary-aware decoder includes a boundary injection module BIB and a context fusion module CAB; the boundary injection module integrates boundary information to enhance the representation of target features, and the context fusion module performs comprehensive fine-grained feature capture, and then obtains the final output.

[0010] Preferably, the backbone network is used to extract features in four stages, generating features at different levels, expressed as:

[0011]

[0012] where represents the feature information extracted from the layer of the backbone network, F represents the set of

[0013] Preferably, the boundary enhancement module passes the feature representation through a channel reduction module, and each module includes three 3x3 convolutional layers to extract multi-level boundary features; the channel size is set to 64, expressed as:

[0014]

[0015] where, represents the boundary feature of the layer, represents the set of boundary features;

[0016] Subsequently, , , is upsampled to The size is used for subsequent splicing operations; then, adjacent features , are combined and fed into the gated convolutional module to comprehensively sample boundary information; then, the same operation is repeated, and the features obtained through the gated convolutional module are concatenated to obtain ; finally, , are added together, and after performing 1x1 convolution, a mask with a channel size of 1 is obtained, representing the boundary , formalized as:

[0017] .

[0018] Preferably, the gated convolutional module GCB is specifically:

[0019] First, along the first dimension, the low-level vector tensor from and the high-level vector tensor are concatenated to form a temporary feature. Subsequently, global average pooling is performed on the temporary feature to obtain a globally aggregated feature; after that, feature transformation is performed through two 1x1 convolutional layers; subsequently, a gating mechanism is formed by the weights generated by the Sigmoid function; finally, the generated attention weights are multiplied by the original , and the result is added to, and through a 1x1 convolutional layer, the final weighted feature is obtained.

[0020] Preferably, the boundary injection module BIB is specifically:

[0021] The boundary injection module BIB accepts two inputs, namely the boundary enhanced feature and the corresponding feature from the backbone network; after the first BIB, the remaining input comes from the output of the previous BIB. Since the previous one already contains boundary information, the subsequent input has the same characteristics; then, element-wise multiplication is performed on these two inputs, and then a skip connection operation is performed for addition; subsequently, 3x3 convolution is applied to obtain a feature map annotated with boundary information ;

[0022] Subsequently, the feature map goes through three 1x1 convolutions, aiming to reduce the channel dimension to relieve the computational burden, and is divided into three branches, namely , and ; for branch , after 1x1 convolution, then 3x3 convolution, and finally the result is added to the original feature to obtain ;

[0023] For the branch , the over-feature map is subjected to average pooling, then one-dimensional convolution is performed for global context fusion. Subsequently, a Sigmoid activation function is applied to generate a weight mask, and then this mask is used to adaptively adjust the feature maps of branches and , emphasizing the information at different scales in the context of camouflaged object detection. The resulting adaptive feature map is denoted as ;

[0024] For the branch , the features are directly multiplied by the attention weights obtained from branch to obtain ;

[0025] Finally, and are concatenated, and then a 3x3 convolution operation is performed to obtain the fused output feature .

[0026] Preferably, the context fusion module CAB is specifically:

[0027] CAB aggregates information from adjacent regions to process the previously obtained fused output feature ; the low-level features and the high-level features undergo a series of convolution and pooling operations to obtain local and global information respectively. Subsequently, by introducing a Sigmoid activation function, the outputs of the left and right channels perform dynamic interaction. The weight range generated by the Sigmoid activation function is between 0 and 1, which is used to adjust the contributions of the low-level and high-level features; finally, through convolution and batch normalization, the adjusted features are integrated to produce a more expressive and targeted representation.

[0028] The second aspect of the present invention provides a method for detecting camouflaged object targets in a drone scenario, including the following processes:

[0029] Obtain an image by shooting with a drone;

[0030] Input the image into the camouflaged object target detection model for drone scenarios as described in the first aspect for target detection;

[0031] Output the target image after segmentation processing.

[0032] In the third aspect of the present invention, a camouflage target detection device for drone scenarios is provided. The device includes at least one processor and at least one memory, and the processor is coupled to the memory; the memory stores a computer executable program of the camouflage target detection model for drone scenarios as described in the first aspect; when the processor executes the computer executable program stored in the memory, the processor executes a camouflage target detection method for drone scenarios.

[0033] In the fourth aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program or instruction of the camouflage target detection model for drone scenarios as described in the first aspect. When the program or instruction is executed by a processor, the processor executes a camouflage target detection method for drone scenarios.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] The technical solution of the present invention has significant benefits and effects in the field of segmenting camouflage targets in complex drone scenarios. Traditionally, the similarity between drone camouflage targets and the background is high, making it difficult to accurately segment. However, the model of the present invention overcomes this problem. The innovation of the present invention lies in improving the segmentation accuracy, being able to more precisely distinguish objects with relatively high similarity, and ensuring a clear separation between the target boundary and the background. This not only helps to improve the accuracy of target detection but also effectively addresses challenges in complex scenarios, such as lighting changes and perspective distortion.

[0036] In the design of GBNet, the introduction of the boundary enhancement module aims to address the problem of the blurred boundaries of camouflage targets, and through the gated convolution block, it cleverly realizes the selective filtering of redundant background information. This not only helps to improve the sensitivity of the network to the boundaries of camouflage targets but also effectively reduces the influence of environmental noise, thereby enhancing the overall performance.

[0037] The design of the boundary-aware decoder further emphasizes the key role of boundary features in the process of identifying camouflage targets. By cleverly injecting boundary enhancement features, it strengthens the decoder's ability to understand the boundaries of camouflage targets in complex scenarios, thus more accurately completing the target segmentation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following description is only one embodiment of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1This is the overall structural block diagram of the camouflage object detection model GBNet of the present invention.

[0040] Figure 2 This is the schematic diagram of the network structure of the gated convolution module GCB of the present invention.

[0041] Figure 3 This is the schematic diagram of the network structure of the boundary injection module BIB of the present invention.

[0042] Figure 4 This is the schematic diagram of the network structure of the context aggregation module CAB of the present invention.

[0043] Figure 5 This is an example diagram of the model detection result in Embodiment 1 of the present invention.

[0044] Figure 6 This is a schematic diagram of the simple structure of the camouflage object detection device in Embodiment 2. Detailed implementation manners

[0045] Embodiment 1:

[0046] The present invention will be further described below in conjunction with specific embodiments.

[0047] Existing research in the field of biological vision shows that when an observer detects a camouflage object pattern that is closely similar to the background, the observer first generates potential region proposals. Subsequently, the observer's attention shifts from these potential regions to the surrounding environment. Through comparison with the surrounding environment, some unreasonable regions are excluded separately. During this comparison process, the overall outline of the target is outlined through supplementary information of detailed clues. Finally, the camouflage object can be accurately identified. Therefore, the present invention has conducted in-depth research on the problem of unclear recognition of the camouflage target boundary in the images taken by drones.

[0048] The present invention designs the GBNet of the gated boundary perception network. Specifically, GBNet includes two parts: the boundary enhancement module and the boundary perception decoder. The boundary perception module selectively filters out redundant background information through the gated convolution block, providing effective guidance for accurately generating the boundary. At the same time, the boundary perception decoder further enriches the representation ability of the decoder by injecting boundary enhancement features.

[0049] In the design of GBNet, the introduction of the boundary enhancement module aims to address the problem of the blurred boundary of the camouflage target, and selectively filters out redundant background information through the gated convolution block. This not only helps to improve the sensitivity of the network to the boundary of the camouflage target, but also effectively reduces the influence of environmental noise, thereby enhancing the overall performance.

[0050] The design of the boundary-aware decoder further emphasizes the crucial role of boundary features in the process of identifying camouflaged targets. By skillfully injecting boundary-enhanced features, the decoder's ability to understand the boundaries of camouflaged targets in complex scenes is strengthened, thus more accurately completing the target segmentation task.

[0051] As Figure 1 shown, this is the specific structure of the model of the present invention. First, features of four stages are extracted through the backbone network, and then the features of the four stages are sent to the boundary enhancement module for boundary information extraction. Subsequently, the extracted boundary information is injected into the boundary-aware decoder. The boundary-aware decoder consists of two parts, the boundary injection module and the context fusion module. The boundary injection module integrates boundary information to enhance the representation of target features, and the up-down fusion module conducts comprehensive fine-grained feature capture, and then obtains the final output.

[0052] Specifically, for a given input image H*W*C with a spatial resolution of H*W and C channels, each pixel is assigned a class label, where 0 represents the background and 1 represents the camouflaged object; the camouflaged object detection model aims to predict a pixel-by-pixel label map and includes a backbone network, a boundary enhancement module BEM, and a boundary-aware decoder connected in sequence; the backbone network, using the Pyramid Vision Transformer as the encoder, is used to extract features at different stages and send them to the boundary enhancement module; the boundary enhancement module uses the gated convolutional module GCB to generate high-quality boundary clues and integrates them into the boundary-aware decoder, thereby enriching the representation ability of the decoder features; the boundary-aware decoder includes a boundary injection module BIB and a context fusion module CAB; the boundary injection module integrates boundary information to enhance the representation of target features, and the up-down fusion module conducts comprehensive fine-grained feature capture, and then obtains the final output.

[0053] For the backbone network, the Pyramid Vision Transformer is used as the encoder to generate features at different levels, expressed as:

[0054]

[0055] where represents the feature information extracted from the th layer of the backbone network, and F represents the set of .

[0056] (I) Boundary Enhancement Module BEM

[0057] Using the boundary enhancement module, screen and select from features at different scales to learn high-quality boundary information. Subsequently, these features are passed to the boundary-aware decoder module, and under the guidance of the boundary mask generated by the boundary enhancement module BEM, the learning process is promoted. The entire model is trained in an end-to-end manner.

[0058] The boundary enhancement module aims to generate high-quality boundary cues using gated convolution blocks and integrate them into the boundary-aware decoder, thus enriching the representation ability of the decoder features. Specifically, the feature map is passed through the channel reduction module, and each module consists of three 3x3 convolutional layers to extract multi-level boundary features. The channel size is set to 64, denoted as:

[0059]

[0060] where represents the boundary feature of the th layer, and represents the set of boundary features.

[0061] Subsequently, , , are upsampled to the size of for subsequent concatenation operations. Then, adjacent features , are combined together and fed into the gated convolution module to comprehensively sample boundary information. Then, the same operation is repeated, and the features obtained through the gated convolution module are concatenated to get . Then, , are added together, and after a 1x1 convolution, a mask with a channel size of 1 is obtained, representing the boundary , which can be formalized as:

[0062] 。

[0063] As Figure 2 shows, the gated convolution module is the key to the boundary information enhancement module. Specifically, first, the vector tensors from (low level) and (high level) are concatenated along the first dimension to form temporary features. Subsequently, global average pooling is performed on the temporary features to obtain globally aggregated features. After that, feature transformation is performed through two 1x1 convolutional layers. Subsequently, the weights generated by the Sigmoid function constitute a gating mechanism. Finally, the generated attention weights are multiplied by the original (low level), the result is added to (high level), and through a 1x1 convolutional layer, the final weighted features are obtained.

[0064] (II) Boundary Injection Module BIB

[0065] As shown Figure 3 in the figure, it is a schematic diagram of the network structure of the Boundary Injection Module (BIB). The boundary-enhanced features obtained from the Boundary Enhancement Module (BEM) can be used as priors to improve the image representation ability of the features generated by the encoder. Specifically, the Boundary Information Injection Module (BIB) accepts two inputs: the boundary-enhanced features and the corresponding features from the backbone network . Note that after the first BIB, the remaining inputs come from the output of the previous BIB. Since the previous one already contains the boundary information, the subsequent inputs have the same characteristics. Then, an element-wise multiplication is performed on these two inputs, followed by a skip connection operation for addition. Subsequently, a 3x3 convolution is applied to obtain a feature map annotated with boundary information , which can be formulated as:

[0066]

[0067] Subsequently, the feature map goes through three 1x1 convolutions, aiming to reduce the channel dimension to alleviate the computational burden, and is divided into three branches, which are , and . This division aims to effectively model different scales or contexts, thereby comprehensively capturing image information.

[0068] For branch , it goes through a 1x1 convolution, then a 3x3 convolution, and finally the result is added to the original feature to obtain . This design is particularly important in scenarios where the boundaries of camouflaged objects are not obvious. The strategy aims to enhance the representation of features related to fuzzy boundaries to improve the model's ability to handle these challenging situations. The 1x1 convolution helps capture channel information, while the subsequent 3x3 convolution is used for spatial relationship modeling. By fusing the modified features with the original features, the network pays more attention to refining the details related to the unclear object boundaries, thereby improving the segmentation performance, which can be formulated as:

[0069]

[0070] For branch , the feature map is averaged pooled, then a one-dimensional convolution is performed for global context fusion. Subsequently, a Sigmoid activation function is applied to generate a weight mask. Then this mask is used to adaptively adjust the feature maps of branches and , emphasizing information at different scales in the context of camouflaged object detection. The resulting adaptive feature map is denoted as .

[0071] For the branch , the features are directly multiplied by the attention weights obtained from the branch to obtain . This mechanism allows the information in the branch to be modulated by the attention weights. This approach helps the model to more flexibly learn the associations between different features and adjust the contributions of the branch according to the attention weights. This correlation helps the network to better adapt to the complexity of the scenarios in the camouflaged object detection task.

[0072] Finally and are concatenated, followed by a 3x3 convolution operation to obtain the fused output features . The formal formula is as follows:

[0073]

[0074] This design aims to make full use of the information of different branches and scales extracted through the convolution operation to achieve comprehensive feature integration through the convolution operation.

[0075] (III) Context Aggregation Module CAB

[0076] As Figure 4 shown, it is a schematic diagram of the network structure of the context aggregation module CAB; CAB aggregates information from adjacent regions to process the previously obtained fused features . This operation enhances the model's perception of the context relationships within the feature map. By considering the interactions and dependencies between adjacent elements in the feature map, the context aggregation block promotes better context understanding, enabling the network to capture spatial dependencies and enhance its ability to identify complex patterns, especially in the context of camouflaged object detection. This helps to more effectively extract features and improve the overall performance of the model.

[0077] First the low-level features and high-level features

[0078] The core idea behind this design is that through the learned attention weights, this module can adaptively emphasize or weaken the importance of low-level and high-level features. This dynamic weight adjustment helps enhance the model's perception of different scales and contexts in the image, endowing the model with adaptability and flexibility. Finally, through convolution and batch normalization, the adjusted features are integrated to produce a more expressive and targeted representation, thereby improving the accuracy and robustness of camouflaged object detection.

[0079] (4) Explanation of experimental results

[0080] Table 1 Comparison of model recognition results

[0081]

[0082] Table 1 clearly shows the comparison results between the GBNet model and other methods, revealing its significant advantages in the field of object recognition. Compared with traditional methods, GBNet can not only more accurately locate the target position in complex backgrounds, but also effectively distinguish the boundary between the target and the background, improving the accuracy and stability of segmentation. This advantage is attributed to the fact that the GBNet model adopts an advanced deep learning architecture, which can learn and extract features from a large amount of complex data, thus better coping with changing environmental conditions and visual noise.

[0083] Meanwhile, as Figure 5 shown, it can be seen that the GBNet model demonstrates excellent capabilities in complex backgrounds and can accurately identify the position of the target. Traditional object recognition models often perform poorly in the face of complex backgrounds because background noise and the visual similarity of the target often cause misjudgments or inaccurate positioning. However, through its advanced algorithms and deep learning techniques, the GBNet model effectively overcomes these challenges and ensures high accuracy in complex environments.

[0084] Example 2:

[0085] As Figure 6As shown in the figure, the present invention also provides a device for detecting camouflage object targets in a drone scenario. The device includes at least one processor and at least one memory, and also includes a communication interface and an internal bus. A computer execution program of the camouflage object target detection model for the drone scenario as described in Embodiment 1 is stored in the memory. When the processor executes the computer execution program stored in the memory, the processor can execute a method for detecting camouflage object targets in a drone scenario. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the buses in the drawings of this application do not limit to only one bus or one type of bus. The memory may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disc, etc.

[0086] The device can be provided as a terminal, a server or other forms of devices.

[0087] Figure 6 It is a block diagram of a device shown for exemplary purposes. The device may include one or more of the following components: a processing component, a memory, a power component, a multimedia component, an audio component, an input / output (I / O) interface, a sensor component, and a communication component. The processing component generally controls the overall operation of the electronic device, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component may include one or more processors to execute instructions to complete all or part of the steps of the above method. In addition, the processing component may include one or more modules to facilitate the interaction between the processing component and other components. For example, the processing component may include a multimedia module to facilitate the interaction between the multimedia component and the processing component.

[0088] The memory is configured to store various types of data to support the operation of the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0089] The power supply component provides power to various components of the electronic device. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device. The multimedia component includes a screen that provides an output interface between the electronic device and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component includes a front camera and / or a rear camera. When the electronic device is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0090] The audio component is configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive external audio signals when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory or transmitted via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals. The I / O interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a start button, and a lock button.

[0091] The sensor assembly includes one or more sensors for providing status assessment of various aspects of the electronic device. For example, the sensor assembly can detect the on / off state of the electronic device, the relative positioning of components, such as the display and keypad of the electronic device, the sensor assembly can also detect a change in the position of the electronic device or a component of the electronic device, the presence or absence of user contact with the electronic device, the orientation or acceleration / deceleration of the electronic device, and the temperature change of the electronic device. The sensor assembly can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0092] The communication component is configured to facilitate communication between the electronic device and other devices in a wired or wireless manner. The electronic device can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0093] In an exemplary embodiment, the electronic device can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.

[0094] Embodiment 3:

[0095] The present invention also provides a computer-readable storage medium storing a computer program or instruction of the camouflage object detection model for drone scenarios as described in Embodiment 1, and when the program or instruction is executed by a processor, it can cause the processor to execute a camouflage object detection method for drone scenarios.

[0096] Specifically, a system, apparatus or device equipped with a readable storage medium can be provided. On this readable storage medium, software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system, apparatus or device is made to read and execute the instructions stored in the readable storage medium. In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments. Therefore, the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present invention.

[0097] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tape, etc. The storage medium can be any available medium accessible by a general or special-purpose computer.

[0098] It should be understood that the above processor can be a central processing unit (English: Central Processing Unit, abbreviated: CPU), and can also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated: ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0099] It should be understood that the storage medium is coupled to the processor, so that the processor can read information from the storage medium and can write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, abbreviated: ASIC). Of course, the processor and the storage medium can also exist as discrete components in a terminal or a server.

[0100] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0101] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider). In some embodiments, by using the status information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0102] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

[0103] Although the specific implementation manners of the present invention have been described above, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A camouflage target detection system for drone scenes, characterized by: For a given input image H*W*C, whose spatial resolution is H*W, has C channels, and each pixel is assigned a category label, where 0 represents background and 1 represents camouflaged object; the camouflaged object detection system aims to predict a pixel-by-pixel label mapping and includes a backbone network, a boundary enhancement module BEM and a boundary-aware decoder connected in sequence; The backbone network uses a pyramid visual transformer as an encoder to extract features at different stages and send them to the boundary enhancement module; The backbone network is used to extract features of four stages and generate features of different levels, which are expressed as: in Indicates that the The feature information extracted by the layer, F represents A collection of; The boundary enhancement module uses the gated convolution module GCB to generate high-quality boundary cues and integrates them into the boundary-aware decoder to enrich the representation capability of the decoder features; The boundary-aware decoder includes a boundary injection module BIB and a context fusion module CAB; the boundary injection module integrates boundary information to enhance the representation of target features, and the context fusion module performs comprehensive and fine-grained feature capture and then obtains the final output; The boundary enhancement module represents the feature Passed through channel reduction modules, each module includes three 3x3 convolutional layers to extract multi-level boundary features; the channel size is set to 64, expressed as: in, Indicates The boundary characteristics of the layer, Represents a collection of boundary features; Afterwards, , , Upsample to The size of is used for subsequent splicing operations; then, the adjacent features , Combined together and fed into the gated convolution module to fully sample the boundary information; then, repeat the same operation and connect the features obtained by the gated convolution module to obtain ; Finally, , Adding, after a 1x1 convolution, we get a mask with a channel size of 1, indicating the boundary , formalized as: The gated convolution module GCB is specifically: First, connect the Low-level vector tensor of and high-level vector tensors , forming temporary features, and then performing global average pooling on the temporary features to obtain global aggregated features; after that, feature conversion is performed through two 1x1 convolutional layers; then, the weights generated by the Sigmoid function form a gating mechanism; finally, the generated attention weights are combined with the original Multiply and add the result And pass through a 1x1 convolution layer to get the final weighted features; The boundary injection module BIB is specifically: The boundary injection module BIB accepts boundary enhancement features and the corresponding features from the backbone network Two inputs; after the first BIB, the remaining input comes from the output of the previous BIB, and since the previous one already contains boundary information, the subsequent input has the same characteristics; then, element-wise multiplication is performed on these two inputs, followed by a skip connection operation for addition; 3x3 convolution is then applied to obtain a feature map annotated with boundary information ; Then, the feature map After three 1x1 convolutions, the channel dimension is reduced to reduce the computational burden and is divided into three branches, namely , and ; For branches , after 1x1 convolution, then 3x3 convolution, and finally adding the result to the original feature, we get ; For branches , the feature map is average pooled, and then a one-dimensional convolution is performed for global context fusion. Subsequently, a Sigmoid activation function is applied to generate a weight mask, and then this mask is used to adaptively adjust the branch and The feature map emphasizes the information of different scales in the context of camouflaged object detection, and the obtained adaptive feature map is recorded as ; For branches , features directly from the branch Multiply the obtained attention weights and get ; at last, and Connect and then perform 3x3 convolution operation to obtain the fused output features ; The context fusion module CAB is specifically: CAB aggregates information from neighboring regions to process the previously obtained fused output features. ; Low-level features and high-level features After convolution and pooling operations, local and global information are obtained respectively. Then, by introducing the Sigmoid activation function, the outputs of the left and right channels interact dynamically. The weight generated by the Sigmoid activation function ranges from 0 to 1 and is used to adjust the contribution of low-level and high-level features. Finally, the adjusted features are integrated through convolution and batch normalization.

2. A method for detecting camouflaged objects in drone scenes, characterized in that: The process includes: Acquire images through drone photography; Inputting the image into the camouflage target detection system for drone scenes as claimed in claim 1 to perform target detection; Output the target image after segmentation processing.

3. A camouflage target detection device for drone scenes, characterized by: The device includes at least one processor and at least one memory, the processor and the memory are coupled; the memory stores a computer execution program of the camouflage target detection system for drone scenes according to claim 1; when the processor executes the computer execution program stored in the memory, the processor executes a camouflage target detection method for drone scenes.

4. A computer-readable storage medium, wherein the computer-readable storage medium stores a computer program or instruction of the camouflage target detection system for drone scenes as claimed in claim 1, wherein when the program or instruction is executed by a processor, the processor executes a camouflage target detection method for drone scenes.