Air-ground cooperative target detection method and device based on cross-domain cross-adaptation network, equipment and medium
Through the cross-domain cross-adaptation network method, bird's-eye view feature maps of different domains are obtained and fused, which solves the problem of information positioning and fusion in air-ground collaborative target detection, and improves the accuracy of target detection.
Patent Information
- Application Number
- CN202510414255.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-25
AI Technical Summary
The existing air-ground collaborative target detection technology is difficult to locate information at the three-dimensional spatial level and interact and fuse the observed same target, resulting in inaccurate target detection results.
Through cross-domain cross-adaptation network, bird's-eye view feature maps under different domains are obtained, multi-level feature information is extracted, and the weight information of each level feature map in each domain is determined based on the cascading results. The feature map is fused based on the weight information, and finally a fusion feature map is generated for object detection.
Effectively align and accurately integrate information from different perspectives to achieve accurate positioning of the target and improve the accuracy of target detection.
Smart Images

Figure CN120374937A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of collaborative perception, and more specifically, to an air-ground collaborative target detection method, device, equipment and medium based on a cross-domain cross-adaptation network. Background Art
[0002] Air-ground collaborative target detection refers to achieving target detection, recognition, positioning, etc. by integrating the observation information of the air perspective and the ground perspective at the same moment. Currently, most air-ground collaborative target detection architectures need to perform two-dimensional target detection based on air and ground observation means respectively, and do not perform information positioning in the three-dimensional space layer and interactively fuse the observed same target.
[0003] For the existing two-dimensional target detection technology, usually YOLO (You Only Look Once) detection heads are respectively deployed at both the air and ground ends to observe the target from two perspectives simultaneously. However, due to reasons such as being unable to obtain the positioning of the target in the same coordinate system, it is difficult to distinguish whether the targets observed by multiple agents are the same target. Summary of the Invention
[0004] (I) Technical Problems to be Solved
[0005] The present invention provides an air-ground collaborative target detection method, device, equipment and medium based on a cross-domain cross-adaptation network, which is used to at least partially solve one of the above technical problems.
[0006] (II) Technical Solutions
[0007] According to the first aspect of the present invention, there is provided an air-ground collaborative target detection method based on a cross-domain cross-adaptation network, including: obtaining bird's-eye view feature maps in different domains; respectively extracting multi-level feature information from the bird's-eye view feature maps of each domain to obtain multi-level feature maps corresponding to the bird's-eye view feature maps of this domain; where the scale information corresponding to different-level feature maps is different; respectively cascading the feature maps of different domains, and determining the weight information corresponding to each-level feature map in each domain according to the cascading result; fusing the bird's-eye view feature maps of different domains based on the weight information to obtain a fused feature map; performing target detection based on the fused feature map to obtain a target detection result.
[0008] According to an embodiment of the present invention, respectively cascading the feature maps of different domains and determining the weight information corresponding to each-level feature map in each domain according to the cascading result includes: performing the following operations on the multi-level feature maps in each domain: respectively cascading each-level feature map with the bird's-eye view feature map of another domain to obtain a cascaded feature map at this level; performing correlation analysis on the cascaded feature map to obtain the weight information corresponding to each-level feature map.
[0009] According to an embodiment of the present invention, fusing the bird's-eye view feature maps in different domains based on weight information to obtain a fused feature map, including: respectively fusing the multi-level feature maps in each domain based on the weight information corresponding to each level of feature maps in each domain to obtain an enhanced feature map for that domain; fusing the enhanced feature maps of different domains to obtain a fused feature map.
[0010] According to an embodiment of the present invention, fusing the enhanced feature maps of different domains to obtain a fused feature map, including: connecting the enhanced feature maps of different domains to obtain an enhanced cascaded feature map; generating an attention feature map corresponding to each domain based on the enhanced cascaded feature map; fusing the attention feature maps of different domains to obtain a fused feature map.
[0011] According to an embodiment of the present invention, generating an attention feature map corresponding to each domain based on the enhanced cascaded feature map, including: performing a linear transformation on the enhanced cascaded feature map to generate a first vector and a second vector; respectively determining the attention feature map corresponding to the current domain based on the first vector, the second vector, and the enhanced feature map of the current domain.
[0012] According to an embodiment of the present invention, fusing the attention feature maps of different domains to obtain a fused feature map, including: fusing the attention feature maps of different domains through learnable parameters to obtain a fused feature map; wherein, the learnable parameters are used to adjust the weights of the attention feature maps of different domains in the fused feature map.
[0013] According to an embodiment of the present invention, obtaining the bird's-eye view feature maps in different domains, including: receiving a first optical image captured by a first device and a second optical image captured by a second device, wherein there is a perspective difference between the images captured by the first device and the second device; respectively extracting depth information from the first optical image and the second optical image to generate corresponding three-dimensional features; respectively constructing bird's-eye view top views in different domains based on the three-dimensional features.
[0014] According to a second aspect of the present invention, there is provided an air-ground collaborative target detection device for a cross-domain cross-adaptation network, including: an acquisition module for acquiring bird's-eye view feature maps in different domains; an extraction module for respectively extracting multi-level feature information from the bird's-eye view feature maps of each domain to obtain the multi-level feature maps corresponding to the bird's-eye view feature maps of that domain; wherein, the scale information corresponding to different levels of feature maps is different; a cascading module for respectively cascading the feature maps of different domains and determining the weight information corresponding to each level of feature maps in each domain according to the cascading result; a fusion module for fusing the bird's-eye view feature maps in different domains based on the weight information to obtain a fused feature map; and a detection module for performing target detection based on the fused feature map to obtain a target detection result.
[0015] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0016] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0017] (III) Advantageous effects
[0018] The method, device and electronic device for air-ground collaborative target detection based on a cross-domain cross-adaptation network provided by the present invention at least include the following advantageous effects:
[0019] After converting the information of multiple domains into the BEV space, aligning and enhancing it, using the BEV space attention, and adjusting the scale of the feature maps of different domains through the learned weight values under different domains, the fusion effect of the fused feature map is improved, the unknown offset is effectively corrected, and the fused feature map contains more accurate feature information. By effectively aligning and accurately fusing the information from different perspectives, the accurate positioning of the target is achieved, and the accuracy of target detection is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0021] Figure 1 Schematically shows a flowchart of a method for air-ground collaborative target detection based on a cross-domain cross-adaptation network according to an embodiment of the present invention;
[0022] Figure 2 Schematically shows a schematic diagram of a method for air-ground collaborative target detection based on a cross-domain cross-adaptation network according to an embodiment of the present disclosure;
[0023] Figure 3 Schematically shows a flowchart of cascading the feature maps of different domains respectively and determining the corresponding weight information of each layer of feature maps under each domain according to the cascading result according to an embodiment of the present disclosure;
[0024] Figure 4 Schematically shows a flowchart of fusing the bird's-eye view feature maps of different domains based on the weight information to obtain a fused feature map according to an embodiment of the present disclosure;
[0025] Figure 5 Schematically shows a structural block diagram of a device for air-ground collaborative target detection based on a cross-domain cross-adaptation network according to an embodiment of the present invention;
[0026] Figure 6 A block diagram of an electronic device schematically showing an air-ground collaborative target detection method based on a cross-domain cross-adaptation network according to an embodiment of the present invention. Detailed implementation manners
[0027] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0028] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the feature, step, operation and / or component, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0029] In the present invention, unless otherwise clearly defined and limited, the terms "installed", "connected", "connected", "fixed", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection, an electrical connection or can communicate with each other; it may be a direct connection, or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0030] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "length", "circumferential", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the subsystem or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0031] Throughout the drawings, the same elements are denoted by the same or similar reference numerals. When it may cause confusion in the understanding of the present invention, the conventional structures or configurations will be omitted. And the shapes, sizes, and positional relationships of the components in the drawings do not reflect the actual sizes, proportions, and actual positional relationships. In addition, any reference signs located between parentheses should not be construed as limiting.
[0032] Similarly, to streamline the present invention and assist in understanding one or more of the various disclosed aspects, in the above description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. Descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0033] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0034] The embodiments of the present disclosure provide a ground-air collaborative target detection method, device, and electronic device based on a cross-domain cross-adaptation network. Before introducing the technical solutions provided by the embodiments of the present disclosure, the related technologies involved in the present disclosure will be described first.
[0035] Based on the deficiencies of existing two-dimensional target detection technologies, those skilled in the art proposed a feature map that can unify different field-of-view information to the same level - the Bird's Eye View (BEV) feature map. The bird's eye view is a perspective of viewing an object or scene from above, just like a bird looking down at the ground from the air.
[0036] The BEV feature map has now been widely used in collaborative perception tasks among multiple intelligent agents. By converting the feature information of optical images into the BEV feature space, the information of different fields of view can be presented at the unified feature level. However, most current collaborative perception research works based on the BEV feature map are carried out on the basis that all intelligent agents are homogeneous, that is, all intelligent agents have the same sensor model and share the same detection model. But with the development of unmanned platform technologies such as unmanned aerial vehicles and unmanned vehicles, the ground-air collaborative target detection task with the aerial perspective as the main observation perspective plays an increasingly important role in practical applications.
[0037] Taking an unmanned vehicle as the ground perspective and an unmanned aerial vehicle as the aerial perspective as an example, the inventors analyzed and found that: Different from the single perspective of multiple vehicles collaborating for observation, when an unmanned aerial vehicle enters the collaborative network, the multi-agent in the collaborative observation task evolves into a heterogeneous form.
[0038] Among multi-agents in the heterogeneous form, there are problems such as too large an information gap between different perspectives and relatively low sensitivity to high-value information in images. This results in that when generating the BEV feature map of each domain, the BEV feature map of this domain contains more specific sensitive information of this domain, while ignoring the important information contained in other domains, inevitably leading to an inter-domain information difference (domain gap). The inter-domain information difference will hinder the fusion of perceptual information by the agent and at the same time reduce the collaborative detection ability of the agent, resulting in inaccurate target detection results.
[0039] Based on the above findings, the inventors provide a ground-air collaborative target detection method based on a cross-domain cross-adaptation network in an embodiment of the present invention, including: obtaining bird's-eye view feature maps in different domains; respectively extracting multi-level feature information from the bird's-eye view feature maps of each domain to obtain the multi-level feature maps corresponding to the bird's-eye view feature maps of this domain; wherein, the scale information corresponding to different-level feature maps is different; respectively cascading the feature maps of different domains, and determining the weight information corresponding to each level of feature map in each domain according to the cascading result; fusing the bird's-eye view feature maps in different domains based on the weight information to obtain a fused feature map; performing target detection based on the fused feature map to obtain a target detection result.
[0040] Figure 1 Schematically shows a flowchart of a ground-air collaborative target detection method based on a cross-domain cross-adaptation network according to an embodiment of the present invention. Figure 2 Schematically shows a schematic diagram of a ground-air collaborative target detection method based on a cross-domain cross-adaptation network according to an embodiment of the present disclosure.
[0041] As Figure 1 、 Figure 2 shown, the signal photon point cloud data filtering method based on linear feature analysis in this embodiment includes operations S110 to S150.
[0042] In operation S110, obtain bird's-eye view feature maps in different domains.
[0043] In some embodiments, the bird's-eye view feature map is obtained by converting optical images taken by different devices. After obtaining optical images from different perspectives, the optical images are input into a BEV conversion network to obtain bird's-eye view feature maps in different domains.
[0044] See Figure 2, taking the first device as a vehicle and the second device as a drone at different flight altitudes as an example, obtain a ground perspective image (i.e., the first optical image) captured by the first device and an aerial perspective image (i.e., the second optical image) captured by the second device, and input the ground perspective image and the aerial gum image into the BEV conversion network. The BEV conversion network converts the ground perspective image and the aerial perspective image to obtain a vehicle perspective BEV image and a drone perspective BEV image.
[0045] In the specific implementation process, the BEV conversion network can perform the BEV perspective conversion based on the LSS (Lift, Splat, Shoot) method. Operation S110 may include: receiving the first optical image captured by the first device and the second optical image captured by the second device. Among them, there is a perspective difference between the images captured by the first device and the second device; respectively extract depth information from the first optical image and the second optical image to generate corresponding three-dimensional features; respectively construct bird's-eye view top views in different domains based on the three-dimensional features.
[0046] In operation S120, respectively extract multi-level feature information from the bird's-eye view feature maps of each domain to obtain the multi-level feature maps corresponding to the bird's-eye view feature maps of that domain; among them, the scale information corresponding to different-level feature maps is different.
[0047] In some embodiments, after obtaining the BEV feature maps of each domain, it is necessary to perform adaptive processing on the BEV feature maps of different domains to align the BEV feature maps of different domains. The present disclosure proposes to extract multi-scale feature information from the BEV feature maps of each domain to obtain the multi-level feature maps corresponding to the BEV feature maps of each domain, so as to use this multi-scale information to adapt to the feature information of other domains.
[0048] In the specific implementation process, respectively extract multi-scale network features from the BEV feature maps of each domain to obtain the corresponding multi-level feature maps of the BEV feature maps of that domain. Optionally, through a Feature Pyramid Networks (FPN) structure, sample the BEV feature map f i to extract multi-level feature information and obtain the m-level feature map f of the BEV feature map of that domain i m , where m is a positive integer greater than 1. In this embodiment, m = 4. And, after obtaining the m-level feature map f of the BEV feature map of that domain i m scale these different-level feature maps to the same size to maintain the spatial information consistency of each level.
[0049] In operation S130, the feature maps of different domains are concatenated respectively, and the weight information corresponding to the feature maps of each level under each domain is determined according to the concatenation result.
[0050] In some embodiments, after obtaining the multi-level feature maps of each domain, the feature maps of each level under the current domain are concatenated with the feature maps of other domains respectively, and the correlation analysis is performed on the concatenated features to determine the weight information of the feature maps of each level under each domain.
[0051] In operation S140, the bird's-eye view feature maps under different domains are fused based on the weight information to obtain a fused feature map.
[0052] In some embodiments, based on the weight information corresponding to the multi-level feature maps under each domain, the multi-level feature maps under this domain are fused to obtain the enhanced feature map f of this domain i sum The enhanced feature maps of different domains are fused to obtain the fused feature map f p See Figure 2 Operation S120 to operation S140 are executed in the cross-domain fusion module.
[0053] In operation S150, object detection is performed based on the fused feature map to obtain an object detection result.
[0054] In some embodiments, the fused feature map f p is formed by fusing multi-source data. The BEV feature maps of different domains capture different characteristics of the scene respectively, realizing cross-domain information complementarity, thus effectively enhancing the feature expression ability in the fused feature map and helping to improve the accuracy of object detection. For example, the features under the fused feature map can be detected through an object detection framework (CenterNet) that abandons anchor points to obtain the object detection result.
[0055] Figure 3 Schematically shows a flowchart of concatenating the feature maps of different domains respectively and determining the weight information corresponding to the feature maps of each level under each domain according to the concatenation result according to an embodiment of the present disclosure.
[0056] As Figure 3 shown, the concatenating the feature maps of different domains respectively and determining the weight information corresponding to the feature maps of each level under each domain according to the concatenation result of this embodiment includes operations S310 to S320, and operations S310 to S320 are performed on the multi-level feature maps under each domain.
[0057] In operation S310, each level of feature map is concatenated with the bird's-eye view feature map of another domain respectively to obtain the concatenated feature map at this level.
[0058] In some embodiments, the feature maps of each level in the current domain are concatenated with the bird's-eye view feature maps of another domain to obtain the concatenated feature maps at this level in this domain :
[0059]
[0060] where i represents the domain where the feature map of this level is located, m represents the level of the feature map of this level, veh represents the ground domain corresponding to the vehicle, and uav represents the air domain corresponding to the drone.
[0061] In operation S320, correlation analysis is performed on the concatenated feature maps to obtain the weight information corresponding to the feature maps of this level.
[0062] In some embodiments, correlation analysis is performed on the obtained concatenated feature maps:
[0063]
[0064] where includes the autocorrelation analysis results of the bird's-eye view feature map of the current domain and the feature map of the m-th level in the current domain, and the cross-correlation analysis results between the feature map of the m-th level in the current domain and the bird's-eye view feature maps of other domains, and T is the matrix transpose symbol.
[0065] After obtaining the correlation analysis results the weight information of the feature maps of each level in the current domain is determined based on the correlation analysis results :
[0066]
[0067] where represents the weight information calculated for each concatenated feature map through the correlation results.
[0068] Figure 4 Schematically shows a flowchart of fusing the bird's-eye view feature maps in different domains based on the weight information to obtain the fused feature map according to an embodiment of the present disclosure.
[0069] As Figure 4 shown, fusing the bird's-eye view feature maps in different domains based on the weight information in this embodiment to obtain the fused feature map includes operations S410 to S420.
[0070] In operation S410, the multi-level feature maps of this domain are fused respectively based on the weight information corresponding to the feature maps of each level in each domain to obtain the enhanced feature map of this domain.
[0071] In some embodiments, based on the weight information corresponding to each hierarchical feature map, multi-level feature maps in the same domain are fused to obtain an enhanced feature map of the domain. For example, based on the weight information corresponding to each hierarchical feature map, multi-level feature maps in the ground domain and the air domain are fused to obtain an enhanced feature map of the ground domain. and an enhanced feature map of the air domain. .
[0072] The enhanced feature map obtained by adaptively fusing multi-level feature maps in each domain through weight information retains both details and semantics, and optimizes the utilization efficiency of features through weight information adjustment, which can effectively improve the expression ability of features and their performance in subsequent tasks.
[0073] In operation S420, enhanced feature maps of different domains are fused to obtain a fused feature map.
[0074] In some embodiments, by adjusting the weights between enhanced feature maps of different domains, the attention degree to different domains is improved to obtain a fused feature map. This fused feature map will pay more attention to the BEV features with high information volume, making it play a more important role in the final fusion result and effectively enhancing the effect of feature fusion.
[0075] In the specific implementation process, operation S420 may further include operations S421 to S423.
[0076] In operation S421, enhanced feature maps of different domains are connected to obtain an enhanced cascaded feature map.
[0077] In some embodiments, enhanced feature maps of different domains are connected along the channel dimension to obtain an aerial cascaded feature map (i.e., an enhanced cascaded feature map).
[0078] In operation S422, an attention feature map corresponding to each domain is generated based on the enhanced cascaded feature map.
[0079] In some embodiments, a linear transformation is performed on the enhanced cascaded feature map to generate a first vector K cat and a second vector V. cat . And based on the first vector, the second vector, and the enhanced feature map of the current domain, an attention feature map of the domain is generated.
[0080] In the specific implementation process, the first vector K cat is a key vector generated by splicing the features of two domains, which is used to calculate the correlation with other queries (Q) in the attention mechanism to determine the importance of different positions. The second vector V catIt is the value vector generated after splicing the features of the two domains, containing the fused cross-domain information, and finally forming an enhanced feature representation through weighted aggregation of the attention weights.
[0081] The first vector and the second vector enable the attention mechanism to utilize the complementary information of different domains simultaneously. For example, the BEV features of a vehicle may contain details of the ground road (such as lane lines, obstacles), and the BEV features of a drone may cover a larger global scene (such as terrain, building layout). The first vector and the second vector encode the BEV features of different domains in the same way, providing a global-local joint representation for subsequent attention.
[0082] And, respectively based on the enhanced feature map of the ground domain Generate the ground domain query vector Q veh Based on the enhanced feature map of the ground domain Generate the ground domain query vector Q uav .
[0083] For the query vectors of each domain, calculate the attention between each domain's query vector and the first vector and the second vector respectively to obtain the attention feature maps f of different domains i . Among them, the attention feature map f i The expression is:
[0084]
[0085] Among them, T is the matrix transpose symbol, and d k Is the dimension size of the key vector.
[0086] Since the output of the cross-domain attention in the cross-domain fusion module is the weighted sum of V cat , and V cat Contains all the information from the two domains. By using the shared V cat In the cross-domain attention mechanism, the two domains can converge towards the direction of the central representation.
[0087] In operation S423, fuse the attention feature maps of different domains to obtain a fused feature map.
[0088] In some embodiments, through a settable learnable parameter λ, fuse the attention feature map f i Obtained after the attention mechanism to obtain a fused feature map f p , where the learnable parameter is used to adjust the weights of the attention feature maps of different domains in the fused feature map:
[0089]
[0090] After converting the information of multiple domains into the BEV space, aligning, and enhancing it, the fusion of BEV features needs to be performed. Since the observation perspectives and distances of each domain for the target are different, small perspectives on the ground at close range usually observe more self-feature information of the target, while the self-feature information of the target obtained from large perspectives in the air at long distances is relatively less, which easily leads to inconsistent information content in the enhanced feature maps of different domains. To solve this problem, this application uses BEV space attention to adjust the proportion of feature maps in different domains through the learned weight values under different domains, thereby improving the fusion effect of the fused feature map and making the fused feature map contain more accurate feature information.
[0091] Based on the above air-ground collaborative target detection method based on the cross-domain cross-adaptation network, the present invention also provides an air-ground collaborative target detection device based on the cross-domain cross-adaptation network. The following will be combined with Figure 5 to describe this device in detail.
[0092] Figure 5 Schematically shows a structural block diagram of an air-ground collaborative target detection device based on a cross-domain cross-adaptation network according to an embodiment of the present invention.
[0093] As Figure 5 shown, the air-ground collaborative target detection device 500 based on the cross-domain cross-adaptation network of this embodiment includes an acquisition module 510, an extraction module 520, a cascading module 530, a fusion module 540, and a detection module 550.
[0094] The acquisition module 510 is used to acquire bird's-eye view feature maps in different domains. In one embodiment, the acquisition module 510 can be used to perform the operation S110 described above, which will not be elaborated here.
[0095] The extraction module 520 is used to extract multi-level feature information from the bird's-eye view feature maps of each domain respectively, and obtain the multi-level feature maps corresponding to the bird's-eye view feature maps of this domain; among them, the scale information corresponding to different-level feature maps is different. In one embodiment, the extraction module 520 can be used to perform the operation S120 described above, which will not be elaborated here.
[0096] The cascading module 530 is used to cascade the feature maps of different domains respectively, and determine the weight information corresponding to each level of feature map in each domain according to the cascading result. In one embodiment, the cascading module 530 can be used to perform the operation S130 described above, which will not be elaborated here.
[0097] The fusion module 540 is used to fuse the bird's-eye view feature maps in different domains based on the weight information to obtain a fused feature map. In one embodiment, the fusion module 540 can be used to perform the operation S140 described above, which will not be elaborated here.
[0098] The detection module 550 is used to perform object detection based on the fused feature map to obtain the object detection result. In one embodiment, the detection module 550 can be used to execute the operation S150 described above, which will not be elaborated here.
[0099] According to an embodiment of the present invention, any multiple modules among the acquisition module 510, the extraction module 520, the cascade module 530, the fusion module 540, and the detection module 550 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the acquisition module 510, the extraction module 520, the cascade module 530, the fusion module 540, and the detection module 550 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the acquisition module 510, the extraction module 520, the cascade module 530, the fusion module 540, and the detection module 550 can be at least partially implemented as a computer program module, which can execute the corresponding functions when the computer program module is run.
[0100] Figure 6 Schematically shows a block diagram of an electronic device for an air-ground collaborative object detection method based on a cross-domain cross-adaptation network according to an embodiment of the present invention.
[0101] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. The processor 601 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 can also include on-board memory for caching purposes. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0102] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via the bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the program can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in one or more memories.
[0103] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. The drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage portion 608 as needed.
[0104] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, a method according to the embodiments of the present disclosure is implemented.
[0105] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.
[0106] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all be included within the protection scope of the present invention.
Claims
1. An air-ground collaborative target detection method based on a cross-domain cross-adaptation network, characterized in that The method includes: Obtaining bird's-eye view feature maps in different domains; Respectively extracting multi-level feature information from the bird's-eye view feature maps of each domain to obtain multi-level feature maps corresponding to the bird's-eye view feature maps of that domain; wherein, the scale information corresponding to different-level feature maps is different; Cascading the feature maps of different domains respectively, and determining the weight information corresponding to each-level feature map in each domain according to the cascading result; Fusing the bird's-eye view feature maps in different domains based on the weight information to obtain a fused feature map; Performing object detection based on the fused feature map to obtain an object detection result.
2. The air-ground collaborative target detection method according to claim 1, characterized in that: The cascading the feature maps of different domains respectively, and determining the weight information corresponding to each-level feature map in each domain according to the cascading result includes: Performing the following operations on the multi-level feature maps in each domain: Cascading each-level feature map with the bird's-eye view feature map of another domain respectively to obtain a cascaded feature map at that level; Performing correlation analysis on the cascaded feature map to obtain the weight information corresponding to each-level feature map.
3. The air-ground collaborative target detection method according to claim 1, characterized in that: The fusing the bird's-eye view feature maps in different domains based on the weight information to obtain a fused feature map includes: Fusing the multi-level feature maps of each domain respectively based on the weight information corresponding to each-level feature map in that domain to obtain an enhanced feature map of that domain; Fusing the enhanced feature maps of different domains to obtain a fused feature map.
4. The air-ground collaborative target detection method according to claim 3 is characterized in that: The fusing the enhanced feature maps of different domains to obtain a fused feature map includes: Connecting the enhanced feature maps of different domains to obtain an enhanced cascaded feature map; Generating an attention feature map corresponding to each domain based on the enhanced cascaded feature map; Fusing the attention feature maps of different domains to obtain a fused feature map.
5. The air-ground collaborative target detection method according to claim 4, characterized in that: The generating an attention feature map corresponding to each domain based on the enhanced cascaded feature map includes: Performing a linear transformation on the enhanced cascaded feature map to generate a first vector and a second vector; Determining the attention feature map corresponding to the current domain respectively based on the first vector, the second vector and the enhanced feature map of the current domain.
6. The air-ground collaborative target detection method according to claim 4, characterized in that: The fusing the attention feature maps of different domains to obtain a fused feature map includes: Fusing the attention feature maps of different domains through learnable parameters to obtain a fused feature map; wherein, the learnable parameters are used to adjust the weights of the attention feature maps of different domains in the fused feature map.
7. The method for collaborative ground and air target detection according to claim 1, wherein The obtaining bird's-eye view feature maps in different domains includes: Receiving a first optical image captured by a first device and a second optical image captured by a second device, wherein there is a perspective difference between the images captured by the first device and the second device; Extracting depth information from the first optical image and the second optical image respectively to generate corresponding three-dimensional features; Constructing bird's-eye view top views in different domains respectively based on the three-dimensional features.
8. An air-ground collaborative target detection device based on a cross-domain cross-adaptation network, characterized in that The device includes: An obtaining module, configured to obtain bird's-eye view feature maps in different domains; An extracting module, configured to extract multi-level feature information from the bird's-eye view feature maps of each domain respectively to obtain multi-level feature maps corresponding to the bird's-eye view feature maps of that domain; wherein, the scale information corresponding to different-level feature maps is different; A cascading module, configured to cascade feature maps in different domains respectively, and determine the weight information corresponding to each layer of feature maps in each domain according to the cascading result; A fusion module, configured to fuse the bird's-eye view feature maps in different domains based on the weight information to obtain a fused feature map; and A detection module, configured to perform object detection based on the fused feature map to obtain an object detection result.
9. An electronic device, comprising: One or more processors; A memory, configured to store one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.