Image Processing Method, Storage Medium, and Electronic Device

Feature fusion of multi-scale feature maps through multi-branch network structure solves the problem of low image detection accuracy and achieves more efficient and accurate detection effects.

CN115100417BActive Publication Date: 2025-07-04ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210662381.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-07-04
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

In the prior art, the accuracy of image detection is low, especially in the detection method with low complexity, and detection efficiency and accuracy are difficult to balance.

Method used

The multi-branch network structure is used to feature fusion of multi-scale feature maps, and the multiple first feature maps are fused and detected through multiple branches to improve the accuracy of the detection results and reduce the amount of parameters during the fusion process.

Benefits of technology

The accuracy and efficiency of image detection are improved, and the problem of low detection accuracy in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100417B_ABST
    Figure CN115100417B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing method, a storage medium, and an electronic device. Among them, the method includes: obtaining a target image, where the target image includes a target object; performing multi-scale feature extraction on the target image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; and performing detection on at least one second feature map to obtain a detection result of the target object. The present invention solves the technical problem of low accuracy in detecting images in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular, to an image processing method, a storage medium, and an electronic device. Background Art

[0002] Currently, in computer vision tasks, deep learning technology is generally used to detect objects in images. The complexity of deep learning technology is relatively high, resulting in low detection efficiency. For detection methods with relatively low complexity, the detection accuracy is poor.

[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of the present application provide an image processing method, a storage medium, and an electronic device to at least solve the technical problem of low accuracy in detecting images in related technologies.

[0005] According to one aspect of the embodiments of the present application, an image processing method is provided, including: obtaining a target image, where the target image includes a target object; performing multi-scale feature extraction on the target image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and performing detection on at least one second feature map to obtain a detection result of the target object.

[0006] According to one aspect of the embodiments of the present application, an image processing method is provided, including: obtaining a target remote sensing image, where the target remote sensing image includes a target object; performing multi-scale feature extraction on the target remote sensing image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and performing detection on at least one second feature map to obtain a detection result of the target object.

[0007] According to one aspect of the embodiments of the present application, an image processing method is provided, including: obtaining an agricultural remote sensing image, where the agricultural remote sensing image includes crops; performing multi-scale feature extraction on the agricultural remote sensing image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and performing detection on at least one second feature map to obtain a detection result of the crops.

[0008] According to one aspect of the embodiments of the present application, there is provided an image processing method, including: obtaining a remote sensing image of a building, where the remote sensing image of the building includes a target building; performing multi-scale feature extraction on the remote sensing image of the building to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and performing detection on the at least one second feature map to obtain a detection result of the target building.

[0009] According to one aspect of the embodiments of the present application, there is provided an image processing method, including: a cloud server obtains a target image, where the target image includes a target object; the cloud server performs multi-scale feature extraction on the target image to obtain a plurality of first feature maps; the cloud server uses a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and the cloud server performs detection on the at least one second feature map to obtain a detection result of the target object.

[0010] In the embodiments of the present application, first, a target image is obtained, where the target image includes a target object; multi-scale feature extraction is performed on the target image to obtain a plurality of first feature maps; a multi-branch network structure is used to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to represent performing feature fusion on the plurality of first feature maps through a plurality of branches to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches, and detection is performed on the at least one second feature map to obtain a detection result of the target object, achieving an improvement in the accuracy of the detection result of the target object. It is easy to notice that feature fusion can be performed on the first feature map through a plurality of branches, thereby improving the accuracy of the obtained at least one second feature map, and the plurality of branches can also reduce the number of parameters in the fusion process, thereby improving the efficiency of fusion, and further solving the technical problem of low accuracy in image detection in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0012] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing the image processing method according to the embodiments of the present application;

[0013] Figure 2It is a flowchart of an image processing method according to Embodiment 1 of the present application;

[0014] Figure 3 It is a schematic diagram of a giraffe target detector according to an embodiment of the present application;

[0015] Figure 4 It is a schematic diagram of a detector according to an embodiment of the present application;

[0016] Figure 5 It is a schematic diagram of a multi-branch network structure according to an embodiment of the present application;

[0017] Figure 6 It is a schematic diagram of detection by a target detection model according to an embodiment of the present application;

[0018] Figure 7 It is a schematic diagram of target detection model training according to an embodiment of the present application;

[0019] Figure 8 It is a flowchart of an image processing method according to Embodiment 2 of the present application;

[0020] Figure 9 It is a flowchart of an image processing method according to Embodiment 3 of the present application;

[0021] Figure 10 It is a flowchart of an image processing method according to Embodiment 4 of the present application;

[0022] Figure 11 It is a flowchart of an image processing method according to Embodiment 5 of the present application;

[0023] Figure 12 It is a schematic diagram of an image processing apparatus according to Embodiment 6 of the present application;

[0024] Figure 13 It is a schematic diagram of an image processing apparatus according to Embodiment 7 of the present application;

[0025] Figure 14 It is a schematic diagram of an image processing apparatus according to Embodiment 8 of the present application;

[0026] Figure 15 It is a schematic diagram of an image processing apparatus according to Embodiment 9 of the present application;

[0027] Figure 16 It is a structural block diagram of a computer terminal according to an embodiment of the present application. Detailed implementation manners

[0028] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0030] First, some nouns or terms that appear during the description of the embodiments of the present application are applicable to the following explanations:

[0031] GFPN: Generalized-Feature Pyramid Networks, a generalized feature pyramid network structure, which acts on the detection feature fusion part and is used for feature fusion at different scales.

[0032] GFocalV2: Generalized Focal Loss V2, a special network structure used for the matching of targets and predicted values in object detection and for feature detection.

[0033] YOLOv1, YOLOv2, YOLOv3, YOLOv4, YOLOv5, YOLOX: A series of special picture-based detection methods.

[0034] Currently, picture-based object detection is a basic technology in machine vision tasks and is widely used in industries such as remote sensing, security, land, water conservancy, and retail. Object detection technology based on deep learning is the current mainstream method, but the computational complexity of this solution is often very high and it is difficult to meet practical applications.

[0035] The present application provides an image processing method that can improve the accuracy of detection results while improving the efficiency of image detection.

[0036] Embodiment 1

[0037] According to an embodiment of the present application, an embodiment of an image processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0038] The method embodiment provided by the first embodiment of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method according to an embodiment of the present application. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more (shown as 102a, 102b,..., 102n in the figure) processors (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.

[0039] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is used for processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image processing method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned image processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0041] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0042] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0043] It should be noted here that in some alternative embodiments, the above Figure 1 shown computer device (or mobile device) may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to illustrate the types of components that may exist in the above computer device (or mobile device).

[0044] Under the above operating environment, the present application provides an image processing method as Figure 2 shown. Figure 2 is a flowchart of the image processing method according to Embodiment 1 of the present application. As Figure 2 shown, the method includes:

[0045] Step S202, obtaining a target image.

[0046] Wherein, the target image contains a target object.

[0047] The above-mentioned target image may be an image containing a target object to be detected. Among them, the target image may be a remote sensing image obtained by a drone and / or a satellite, and the target image may also be an image captured by a photographing device.

[0048] The above-mentioned target object may be a specific object to be detected in the target image.

[0049] In an agricultural scenario, the target image may be an agricultural remote sensing image, and the target object may be a crop to be detected.

[0050] In a building scenario, the target image may be a building remote sensing image, and the target object may be a building to be detected.

[0051] In an alternative embodiment, in order to better process the target image, the obtained target image may be transmitted to a corresponding processing device for processing. For example, it may be directly transmitted to a user's computer terminal (such as a laptop, a personal computer, etc.) for processing, or transmitted to a cloud server for processing through the user's computer terminal. It should be noted that since processing the target image requires a large amount of computing resources, in the embodiments of the present application, the processing device is taken as an example of a cloud server for description.

[0052] For example, in order to facilitate the user to upload the target image, an interactive interface may be provided to the user. Among them, the interactive interface includes controls such as "Select Image", "Upload", and "Image Display". The user can click the "Select Image" button to determine the target image to be uploaded, and upload the target image to the cloud server for processing by clicking the "Upload" button. In addition, in order to facilitate the user to confirm whether the selected target image is the target image to be processed, the selected target image may be displayed in the "Image Display" area. After the user confirms that it is correct, data upload is performed by clicking the "Upload" button.

[0053] It should be noted that data interaction may be performed between the client and the cloud server through a specific interface. The client may pass the description page of the target object selected by the user into the interface function and use it as a parameter of the interface to achieve the purpose of uploading the description page of the target object to the cloud server.

[0054] Step S204: Perform multi-scale feature extraction on the target image to obtain a plurality of first feature maps.

[0055] In an alternative embodiment, multi-scale feature extraction may be performed on the target image to facilitate obtaining a plurality of first feature maps of different scales, and a plurality of first feature maps may be fused through a preset feature fusion strategy to improve the detection accuracy of the feature maps.

[0056] In another alternative embodiment, the feature extraction layer in the Giraffe Target Detector (GiraffeDet) can be used to perform multi-scale feature extraction on the target image to obtain multiple first feature maps, where the scales of the multiple first feature maps are different. Optionally, the feature extraction layer contains multiple scale layers, and the corresponding sizes of each scale layer are different. Multi-scale feature extraction can be performed on the target image according to the multiple scale layers to obtain multiple first feature maps.

[0057] Step S206: Use a multi-branch network structure to perform feature fusion on the multiple first feature maps to obtain at least one second feature map.

[0058] Among them, the multi-branch network structure is used to perform feature fusion on the multiple first feature maps through multiple branches.

[0059] The above multi-branch network structure may include two or more branches. The specific number of branches included can be set by the user or flexibly adjusted according to requirements.

[0060] In an alternative embodiment, the multi-branch network can be used to perform repeated feature fusion on the multiple first feature maps according to a preset feature fusion strategy to obtain multiple second feature maps, where the feature fusion strategy can be the feature fusion strategy used by GiraffeDet.

[0061] Figure 3 It is a schematic diagram of a giraffe target detector according to an embodiment of the present application. As Figure 3 shown, S1-S5 are multiple first feature maps with different scales extracted by the feature extraction layer. The multiple first feature maps can be fused repeatedly, and the fused feature maps can be connected in the form of log2 N to obtain the finally required feature map for detection. In the fusion part, the method shown in Figure 3 can be used to perform upsampling and downsampling on feature maps of different sizes. Among them, the upward arrow represents the downsampling process, and the upward arrow represents the upsampling process. Then, the first feature maps of different sizes can be stitched and input into the corresponding multi-branch network structure. The first feature maps of different sizes are fused through the multi-branch network structure to obtain the fused features, such as S5_0. Then, the first feature maps of different sizes and the fused features can be continuously fused repeatedly to stack the final, such as S5_N, where N represents the number of stacking times. Figure 3 In Nare connected in the form of

[0062] In an alternative embodiment, the framework for detecting a target object in a target image may be backbone (feature extraction network) - neck (fusion of high-resolution detailed feature maps and low-resolution semantic feature maps at a shallow layer) - head (detector). Currently, the general computational ratio of backbone:neck:head is 2-4:1:2. Among them, the computational ratio of the fusion part is relatively small, resulting in relatively low subsequent computational accuracy. To solve this problem, this application adjusts the computational ratio to 1:5:1, increasing the computational ratio of the fusion part, which can further improve the detection accuracy.

[0063] It should be noted that the above computational ratio refers to the distribution of the amount of computation. Assuming that a total of 100 GFLOPs of computation is required, then the computation of the backbone part will approximately account for 2 / 5 * 100 = 40 GFLOPS; the specific computation is the standard operation such as convolution in the neural network.

[0064] Step S208: Detect at least one second feature map to obtain a detection result of the target object.

[0065] The above detection result may be a target detection box, where the target detection box is used to label the target object in the target image; the above detection result may also include the category of the target object, where the category of the target object may be labeled next to the target detection box or at other specified places.

[0066] In an alternative embodiment, multiple first feature maps may be fused to obtain at least one second feature map. For second feature maps of different scales, different detectors may be used for detection. Different detectors have the same structure, that is, the length, width, and height corresponding to different detectors are the same, but the parameters in the detectors are different. Since the accuracies of second feature maps of different scales are different, therefore, by using detectors with different accuracies to detect second feature maps of different scales, the detection effect for the feature map of this scale can be further improved, thereby improving the detection accuracy for the feature map of this scale. The detection accuracy of the detector can be adjusted so that the detector has a relatively high detection accuracy for detecting a second feature map with a relatively high accuracy, thereby improving the detection accuracy, and the detection accuracy of the detector can be adjusted so that the detector has a relatively low detection accuracy for detecting a second feature map with a relatively low accuracy, thereby improving the detection accuracy. Through this setting, the detection accuracy can be made higher under the same amount of computation.

[0067] In another alternative embodiment, GFocalV2 Head (detection algorithm) is used as the basic detector structure.

[0068] Figure 4 It is a schematic diagram of a detector according to an embodiment of the present application. The detector shown on the left is the way of the detector to be used for the target, which uses the same detector for feature maps of different scales for detection; the detector shown on the right is the way of the detector used in the present application, which uses detectors with different parameters but the same structure for feature maps of different scales, thereby improving the detection accuracy.

[0069] In another alternative embodiment, a target image can be obtained, where the target image contains a target object; a multi-scale feature extraction can be performed on the target image by using a feature extraction layer in the target detection model to obtain a plurality of first feature maps; a feature fusion can be performed on the plurality of first feature maps by using a multi-branch network structure in the target detection model to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; a detector in the target detection model can be used to detect at least one second feature map to obtain a detection result of the target object.

[0070] Through the above steps, first, a target image is obtained, where the target image contains a target object; a multi-scale feature extraction is performed on the target image to obtain a plurality of first feature maps; a feature fusion is performed on the plurality of first feature maps by using a multi-branch network structure to obtain at least one second feature map, where the multi-branch network structure is used to represent performing feature fusion on the plurality of first feature maps through multiple branches to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches, and at least one second feature map is detected to obtain a detection result of the target object, achieving an improvement in the accuracy of the detection result of the target object. It is easy to notice that feature fusion can be performed on the first feature maps through multiple branches, thereby improving the accuracy of the obtained at least one second feature map, and the multiple branches can also reduce the number of parameters in the fusion process, thereby improving the fusion efficiency, and further solving the technical problem of low accuracy in image detection in the related art.

[0071] In the above embodiment of the present application, the multi-branch network structure includes: a first branch and a second branch, where the output of the first branch is connected to the output of the second branch.

[0072] The above first branch may include a 1×1 convolutional layer, and the above second branch may include N 1×1 and 3×3 convolutional blocks, where the N convolutional blocks can be connected in sequence. It should be noted that the 3×3 in the second branch can be replaced with other convolutional layers. For example, 3×3 can be replaced with 5×5, but not limited thereto.

[0073] In an alternative embodiment, the multi-branch network structure may include: a first convolutional layer, a first branch, a second branch, the output of the first convolutional layer is connected to the inputs of the first branch and the second branch, and the output of the first branch is connected to the output of the second branch.

[0074] The above-mentioned first convolutional layer may be 1×1.

[0075] Figure 5 It is a schematic diagram of a multi-branch network structure according to an embodiment of the present application. As Figure 5 shown, multiple first feature maps to be fused can be processed and then combined to obtain a combined feature map. The combined feature map can be respectively input into the first branch and the second branch for processing to obtain two output feature maps output by the two branches. The two output feature maps can be concatenated to obtain a second feature map.

[0076] In another alternative embodiment, the multi-branch network structure may further include multiple branches, which are not limited herein. Among them, the first branch may include multiple sub-branches, and the second branch may also include multiple sub-branches.

[0077] In the above embodiments of the present application, the multi-branch network structure is used to perform feature fusion on multiple first feature maps to obtain at least one second feature map, including: performing channel merging on the multiple first feature maps to obtain a combined feature map; using the first branch to perform convolutional processing on the combined feature map to obtain a first output feature; using the second branch to perform convolutional processing on the combined feature map to obtain a second output feature; performing channel merging on the first output feature and the second output feature to obtain at least one second feature map.

[0078] The above-mentioned first branch may include convolutional layers with different convolutional kernels. By using the convolutional layers corresponding to convolutional kernels of different sizes to process the combined feature map, the computational amount of convolutional operations can be reduced, thereby improving the fusion efficiency.

[0079] In an alternative embodiment, the first convolutional layer can be used to directly perform channel merging on multiple first feature maps to obtain the above-mentioned combined feature map. The first convolutional layer can also be used to perform convolutional operations on multiple first feature maps to obtain multiple third feature maps, so as to unify the sizes of the multiple first feature maps, facilitate subsequent merging processes. The channels of the multiple third feature maps can be merged to achieve the effect of concatenating the multiple third feature maps to obtain a combined feature map. The first branch can be used to perform convolutional processing on the combined feature map to obtain a first output feature. The second branch can be used to perform convolutional processing on the combined feature to obtain a second output feature. The channels of the first output feature and the second output feature can be merged to obtain a second feature map.

[0080] Further, the second feature map can be used as the first feature map, and a multi-branch network can be used to continue processing multiple first feature maps to obtain a new second feature map. By repeatedly performing this step multiple times, multiple second feature maps can be obtained.

[0081] In the above embodiments of the present application, the first branch includes: at least one convolutional block, and each convolutional block in the multiple convolutional blocks contains multiple sub-convolutional layers, and the convolutional kernels of the multiple sub-convolutional layers are different.

[0082] In an alternative embodiment, the number of convolutional blocks can be determined according to the number of input first feature maps. If the number of input first feature maps is N, then the number of convolutional blocks is N.

[0083] The sizes of the convolutional kernels of the above multiple sub-convolutional layers are different. Among them, the convolutional kernel corresponding to the convolutional layer in front of the convolutional block in the multiple sub-convolutional layers can be smaller than the convolutional kernel corresponding to the convolutional layer behind the convolutional block.

[0084] The above multiple sub-convolutional layers can be 1×1 and 3×3 respectively.

[0085] The multiple sub-convolutional layers included in the above multiple convolutional blocks can be the same. The multiple sub-convolutional layers included in the above multiple convolutional blocks can be different.

[0086] In the above embodiments of the present application, using the first branch to perform convolutional processing on the merged feature map to obtain a first output feature includes: using at least one convolutional block to perform convolutional processing on the merged feature map to obtain a first output feature.

[0087] The above convolutional block contains a first sub-convolutional layer and a second sub-convolutional layer. Among them, the first sub-convolutional layer can be a convolutional layer with a smaller convolutional kernel, and the second sub-convolutional layer can be a convolutional layer with a larger convolutional kernel. The first sub-convolutional layer can be 1×1, and the second sub-convolutional layer can be 3×3.

[0088] In an alternative embodiment, in the case of including one convolutional block, the merged feature can be convolved using the first sub-convolutional layer in the convolutional block to obtain a processing result, and the processing result can be input into the second sub-convolutional layer, and the second sub-convolutional layer can be used to perform a convolution operation on the processing result to obtain a first output feature.

[0089] In the above embodiments of the present application, using the second branch to perform convolutional processing on the merged feature map to obtain a second output feature includes: using a three-convolutional layer to perform convolutional processing on the merged feature map to obtain a second output feature.

[0090] The above third convolutional layer can be 1×1.

[0091] In the above embodiments of the present application, the method further includes: using a target detection model to detect a target image to obtain a detection result of a target object, where the target detection model is trained based on target sample images, and the target sample images are obtained by performing data augmentation on a plurality of sample images through sample detection frames.

[0092] In an alternative embodiment, the target image can be input into the target detection model, and the target detection model is used to detect the target image to obtain a detection result of the target object. Among them, the main framework of the target detection model can be backbone-neck-head. First, the backbone can be used to perform multi-scale feature extraction on the target image to obtain a plurality of first feature maps, and then the neck can use a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; finally, the head is used to detect at least one second feature map to obtain a detection result of the target object.

[0093] In an alternative embodiment, the target detection model can be a commonly used detection model. The sample images sampled during its training are different from the sample images commonly used. Specifically, data augmentation is performed on a plurality of sample images through sample detection frames to obtain target sample images.

[0094] In another alternative embodiment, the target detection model can include a feature extraction layer, a multi-branch network structure, and a detection layer. Among them, the feature extraction layer is used to perform multi-scale feature extraction on the target image to obtain a plurality of first feature maps, the multi-branch network structure is used to fuse the plurality of first feature maps to obtain at least one second feature map, and the detection layer is used to detect at least one second feature map to obtain a detection result of the target object. Among them, the multi-branch network structure can include a first branch and a second branch, and the output of the first branch is connected to the output of the second branch.

[0095] Figure 6 It is a schematic diagram of detection by a target detection model according to an embodiment of the present application. The target image to be detected can be input into the target detection model, and the target detection model can output a detection result of the target object in the target image, and the detection result is obtained by annotating the target object with a target detection frame.

[0096] In the above embodiments of the present application, the method further includes: obtaining a plurality of sample images and sample detection frames corresponding to the plurality of sample images, wherein the sample detection frames are used to label target objects in the sample images; determining a preset number of sample detection frames in the plurality of sample images as target detection frames; performing data augmentation on the target objects corresponding to the target detection frames to obtain target sample images; and training an initial detection model using the target sample images to obtain a target detection model.

[0097] In an alternative embodiment, the plurality of sample images may be mixed first to obtain a mixed image, and then a preset number of sample detection frames among the plurality of sample detection frames included in the mixed image are determined as target detection frames.

[0098] In another alternative embodiment, a plurality of sample images and sample detection frames corresponding to the plurality of sample images may be randomly selected from a training dataset. Data augmentation may be performed on the plurality of sample images according to the box level (region level) of the plurality of sample images. First, a global data augmentation may be performed on the plurality of sample images once to obtain the augmented plurality of sample images, and then a preset number of sample detection frames are randomly selected from the augmented plurality of sample images as target detection frames for local data augmentation. Data augmentation may be performed on the target objects corresponding to the target detection frames to obtain target sample images. Finally, the initial detection model is trained using the target sample images to obtain a target model.

[0099] In another alternative embodiment, a preset number of sample detection frames may be randomly selected from the plurality of sample images as target detection frames for local data augmentation. Data augmentation may be performed on the target objects corresponding to the target detection frames to obtain the augmented plurality of sample images. Then, a global data augmentation is performed on the augmented plurality of sample images once to obtain target sample images. Finally, the initial detection model is trained using the target sample images to obtain a target model.

[0100] The above global data augmentation may include color transformation, rotation, contrast enhancement, random erasing, scaling, cropping, etc. for the plurality of sample images. The above local data augmentation may include color transformation, rotation, contrast enhancement, random erasing, scaling, etc. for the target objects in the target detection frames.

[0101] In the above embodiments of the present application, determining a preset number of sample detection frames in the plurality of sample images as target detection frames includes: splicing the plurality of sample images to obtain an initial sample image; mixing the initial sample image and a preset sample image to obtain a mixed image; and determining a preset number of sample detection frames in the mixed image as target detection frames.

[0102] In an alternative embodiment, multiple sample images can be sequentially stitched together to obtain an initial sample image. Optionally, multiple sample images can be stitched onto a white-background image to fill the white-background image, thereby obtaining the initial sample image. If the multiple sample images do not fill the white-background image, a part of the sample images can be randomly selected from the training data to fill the white-background image until it is full. Optionally, the multiple sample images can also be randomly stitched. Other methods for stitching the multiple sample images are also possible and are not limited herein.

[0103] The above-mentioned preset quantity can be a pre-set quantity, or can also be determined according to the quantity of sample detection frames. For example, the preset quantity can be 30% of the sample detection frames.

[0104] In another alternative embodiment, a mosaic module (mosaic) can be used to sequentially stitch multiple sample images together to obtain an initial sample image, where the mosaic is used to enrich the detection background and detection objects in the target image, so as to enrich the data set. After obtaining the initial sample image, the initial sample image can be scaled and cropped to obtain a processed initial sample image. A mixup module can be used to mix the processed initial sample image and a preset sample image to obtain a mixed image, so as to further enhance the data and thus enrich the features in the sample image. A preset quantity of the sample detection frames among the multiple sample detection frames included in the mixed image can be determined as target detection frames.

[0105] Figure 7 It is a schematic diagram of training a target detection model according to an embodiment of the present application. First, multiple sample images containing sample detection frames are input into the mosaic module, and the mosaic module is used to stitch the multiple sample images together to obtain an initial sample image. Then, the mixup module is used to mix the initial sample image and a preset sample image to obtain a mixed image. Finally, the region-level module determines a preset quantity of the sample detection frames in the mixed image as target detection frames, and data augmentation is performed on the target objects corresponding to the target detection frames to obtain target sample images; the initial detection model is trained using the target sample images to obtain the target detection model.

[0106] In the above embodiments of the present application, the method further includes: outputting a detection result; receiving a first feedback result, where the first feedback result is used to modify the channels in the merged feature map according to the detection result; and updating the merged feature map based on the first feedback result.

[0107] In an alternative embodiment, the detection result can be output and displayed to the user's client. The user can modify the channels in the merged feature map according to the detection result to obtain a first feedback result, so as to update the merged feature map according to the first feedback result. Detection can be performed based on the updated merged feature map to obtain a detection result with higher accuracy.

[0108] The neural network structure proposed in this application can use a lightweight network structure (csp_darknet) as the backbone, a large proportion of csp_GFPN as the neck, and a scale-decouple (different scales) GFocalv2 as the head. At the same time, combined with a training method based on learning data augmentation, it finally realizes fast and high-precision detection of target objects in target images. Greatly reduce the resource usage when deploying the neural network structure.

[0109] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that the image processing method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.

[0111] Embodiment 2

[0112] According to an embodiment of the present application, an embodiment of an image processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0113] Figure 8 is a flowchart of an image processing method according to Embodiment 2 of the present application, asFigure 8 As shown, the method may include the following steps:

[0114] Step S802, obtaining a target remote sensing image.

[0115] Wherein, the target remote sensing image contains a target object.

[0116] Step S804, performing multi-scale feature extraction on the target remote sensing image to obtain a plurality of first feature maps.

[0117] Step S806, using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map.

[0118] Wherein, the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches.

[0119] Step S808, detecting at least one second feature map to obtain a detection result of the target object.

[0120] In the above embodiments of the present application, the multi-branch network structure includes: a first branch and a second branch, wherein the output of the first branch is connected to the output of the second branch.

[0121] In the above embodiments of the present application, using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map includes: performing channel merging on the plurality of first feature maps to obtain a merged feature map; using the first branch to perform convolution processing on the merged feature map to obtain a first output feature; using the second branch to perform convolution processing on the merged feature map to obtain a second output feature; performing channel merging on the first output feature and the second output feature to obtain at least one second feature map.

[0122] In the above embodiments of the present application, the first branch includes: at least one convolution block, and each convolution block in the plurality of convolution blocks contains a plurality of sub-convolution layers, and the convolution kernels of the plurality of sub-convolution layers are different.

[0123] In the above embodiments of the present application, the method further includes: using a target detection model to detect the target image to obtain a detection result of the target object, wherein the target detection model is trained based on a target sample image, and the target sample image is obtained by performing data augmentation on a plurality of sample images through a sample detection frame.

[0124] In the above embodiments of the present application, the method further includes: obtaining a plurality of sample images and sample detection frames corresponding to the plurality of sample images, where the sample detection frames are used to label target objects in the sample images; determining a preset number of sample detection frames in the plurality of sample images as target detection frames; performing data augmentation on the target objects corresponding to the target detection frames to obtain target sample images; and training an initial detection model using the target sample images to obtain a target detection model.

[0125] In the above embodiments of the present application, determining a preset number of sample detection frames in the plurality of sample images as target detection frames includes: splicing the plurality of sample images to obtain an initial sample image; mixing the initial sample image and a preset sample image to obtain a mixed image; and determining a preset number of sample detection frames in the mixed image as target detection frames.

[0126] In the above embodiments of the present application, the method further includes: outputting a detection result; receiving a first feedback result, where the first feedback result is used to modify channels in a merged feature map according to the detection result; and updating the merged feature map based on the first feedback result.

[0127] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0128] Embodiment 3

[0129] According to an embodiment of the present application, an embodiment of an image processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0130] Figure 9 is a flowchart of an image processing method according to Embodiment 3 of the present application. As Figure 9 shown, the method may include the following steps:

[0131] Step S902, obtaining an agricultural remote sensing image.

[0132] Among them, the agricultural remote sensing image contains crops.

[0133] Step S904, performing multi-scale feature extraction on the agricultural remote sensing image to obtain a plurality of first feature maps.

[0134] Step S906, using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map.

[0135] Among them, the multi-branch network structure is used to perform feature fusion on multiple first feature maps through multiple branches.

[0136] Step S908, detecting at least one second feature map to obtain the detection result of the crop.

[0137] In the above embodiments of the present application, the multi-branch network structure includes: a first branch and a second branch, wherein the output of the first branch is connected to the output of the second branch.

[0138] In the above embodiments of the present application, using the multi-branch network structure to perform feature fusion on multiple first feature maps to obtain at least one second feature map includes: performing channel merging on multiple first feature maps to obtain a merged feature map; using the first branch to perform convolution processing on the merged feature map to obtain a first output feature; using the second branch to perform convolution processing on the merged feature map to obtain a second output feature; performing channel merging on the first output feature and the second output feature to obtain at least one second feature map.

[0139] In the above embodiments of the present application, the first branch includes: at least one convolution block, and each convolution block in the multiple convolution blocks includes multiple sub-convolution layers, and the convolution kernels of the multiple sub-convolution layers are different.

[0140] In the above embodiments of the present application, the method further includes: using a target detection model to detect an agricultural remote sensing image to obtain the detection result of the crop, wherein the target detection model is trained based on target sample images, and the target sample images are obtained by performing data augmentation on multiple sample images through sample detection frames.

[0141] In the above embodiments of the present application, the method further includes: obtaining multiple sample images and sample detection frames corresponding to the multiple sample images, wherein the sample detection frames are used to label the crops in the sample images; determining a preset number of sample detection frames in the multiple sample images as target detection frames; performing data augmentation on the crops corresponding to the target detection frames to obtain target sample images; using the target sample images to train an initial detection model to obtain a target detection model.

[0142] In the above embodiments of the present application, determining a preset number of sample detection frames in the multiple sample images as target detection frames includes: splicing the multiple sample images to obtain an initial sample image; mixing the initial sample image and a preset sample image to obtain a mixed image; determining a preset number of sample detection frames in the mixed image as target detection frames.

[0143] In the above embodiments of the present application, the method further includes: outputting a detection result; receiving a first feedback result, wherein the first feedback result is used to modify the channels in the merged feature map according to the detection result; updating the merged feature map based on the first feedback result.

[0144] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0145] Embodiment 4

[0146] According to an embodiment of the present application, an embodiment of an image processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0147] Figure 10 is a flowchart of an image processing method according to Embodiment 4 of the present application, as Figure 10 shown, the method may include the following steps:

[0148] Step S1002, obtain a remote sensing image of a building.

[0149] Among them, the remote sensing image of the building contains the target building.

[0150] Step S1004, perform multi-scale feature extraction on the remote sensing image of the building to obtain a plurality of first feature maps.

[0151] Step S1006, use a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map.

[0152] Among them, the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches.

[0153] Step S1008, detect at least one second feature map to obtain a detection result of the target building.

[0154] In the above embodiments of the present application, the multi-branch network structure includes: a first branch and a second branch, where the output of the first branch is connected to the output of the second branch.

[0155] In the above embodiments of the present application, using a multi-branch network structure to perform feature fusion on a plurality of first feature maps to obtain at least one second feature map includes: performing channel merging on the plurality of first feature maps to obtain a merged feature map; performing convolution processing on the merged feature map using the first branch to obtain a first output feature; performing convolution processing on the merged feature map using the second branch to obtain a second output feature; performing channel merging on the first output feature and the second output feature to obtain at least one second feature map.

[0156] In the above embodiments of the present application, the first branch includes: at least one convolutional block, and each convolutional block among the multiple convolutional blocks includes multiple sub-convolutional layers, and the convolutional kernels of the multiple sub-convolutional layers are different.

[0157] In the above embodiments of the present application, the method further includes: using a target detection model to detect a remote sensing image of a building to obtain a detection result of the target building, where the target detection model is trained based on a target sample image, and the target sample image is obtained by performing data augmentation on multiple sample images through sample detection frames.

[0158] In the above embodiments of the present application, the method further includes: obtaining multiple sample images and sample detection frames corresponding to the multiple sample images, where the sample detection frames are used to label the target buildings in the sample images; determining a preset number of sample detection frames among the multiple sample images as target detection frames; performing data augmentation on the target buildings corresponding to the target detection frames to obtain a target sample image; and training an initial detection model using the target sample image to obtain a target detection model.

[0159] In the above embodiments of the present application, determining a preset number of sample detection frames among the multiple sample images as target detection frames includes: splicing the multiple sample images to obtain an initial sample image; mixing the initial sample image and a preset sample image to obtain a mixed image; and determining a preset number of sample detection frames in the mixed image as target detection frames.

[0160] In the above embodiments of the present application, the method further includes: outputting a detection result; receiving a first feedback result, where the first feedback result is used to modify the channels in the merged feature map according to the detection result; and updating the merged feature map based on the first feedback result.

[0161] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0162] Embodiment 5

[0163] According to an embodiment of the present application, an embodiment of an image processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0164] Figure 11 is a flowchart of an image processing method according to Embodiment 5 of the present application. As Figure 11 shown, the method may include the following steps:

[0165] Step S1102, the cloud server obtains the target image.

[0166] Among them, the target image contains the target object.

[0167] Step S1104, the cloud server performs multi-scale feature extraction on the target image to obtain multiple first feature maps.

[0168] Step S1106, the cloud server uses a multi-branch network structure to perform feature fusion on the multiple first feature maps to obtain at least one second feature map.

[0169] Among them, the multi-branch network structure is used to perform feature fusion on the multiple first feature maps through multiple branches.

[0170] Step S1108, the cloud server detects at least one second feature map to obtain the detection result of the target object.

[0171] In the above embodiments of the present application, the multi-branch network structure includes: a first branch and a second branch, where the output of the first branch is connected to the output of the second branch.

[0172] In the above embodiments of the present application, the cloud server uses a multi-branch network structure to perform feature fusion on the multiple first feature maps to obtain at least one second feature map, including: the cloud server performs channel merging on the multiple first feature maps to obtain a merged feature map; the cloud server uses the first branch to perform convolution processing on the merged feature map to obtain a first output feature; the cloud server uses the second branch to perform convolution processing on the merged feature map to obtain a second output feature; the cloud server performs channel merging on the first output feature and the second output feature to obtain at least one second feature map.

[0173] In the above embodiments of the present application, the first branch includes: at least one convolution block, and each convolution block in the multiple convolution blocks contains multiple sub-convolution layers, and the convolution kernels of the multiple sub-convolution layers are different.

[0174] In the above embodiments of the present application, the method further includes: the cloud server uses a target detection model to detect the target image to obtain the detection result of the target object, where the target detection model is trained based on target sample images, and the target sample images are obtained by performing data augmentation on multiple sample images through sample detection frames.

[0175] In the above embodiments of the present application, the method further includes: the cloud server obtains a plurality of sample images and sample detection frames corresponding to the plurality of sample images, where the sample detection frames are used to label target objects in the sample images; the cloud server determines a preset number of sample detection frames in the plurality of sample images as target detection frames; the cloud server performs data augmentation on the target objects corresponding to the target detection frames to obtain target sample images; the cloud server uses the target sample images to train an initial detection model to obtain a target detection model.

[0176] In the above embodiments of the present application, the cloud server determines a preset number of sample detection frames in the plurality of sample images as target detection frames, including: the cloud server splices the plurality of sample images to obtain an initial sample image; the cloud server mixes the initial sample image and a preset sample image to obtain a mixed image; the cloud server determines a preset number of sample detection frames in the mixed image as target detection frames.

[0177] In the above embodiments of the present application, the method further includes: the cloud server outputs a detection result; the cloud server receives a first feedback result, where the first feedback result is used to modify channels in the merged feature map according to the detection result; the cloud server updates the merged feature map based on the first feedback result.

[0178] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0179] Embodiment 6

[0180] According to an embodiment of the present application, there is also provided an image processing apparatus for implementing the above image processing method. Figure 12 It is a schematic diagram of an image processing apparatus according to Embodiment 6 of the present application. As Figure 12 shown, the apparatus 1200 includes: an acquisition module 1202, an extraction module 1204, a fusion module 1206, and a detection module 1208.

[0181] Among them, the acquisition module is used to acquire a target remote sensing image, where the agricultural remote sensing image contains a target object; the extraction module is used to perform multi-scale feature extraction on the target remote sensing image to obtain a plurality of first feature maps; the fusion module is used to perform feature fusion on the plurality of first feature maps by using a multi-branch network structure to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; the detection module is used to detect at least one second feature map to obtain a detection result of the target object.

[0182] It should be noted here that the above-mentioned acquisition module 1202, extraction module 1204, fusion module 1206, and detection module 1208 correspond to steps S202 to S208 in Embodiment 1. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned modules can run in the computer terminal 10 provided in Embodiment 1 as part of the device.

[0183] In the above-mentioned embodiments of the present application, the multi-branch network structure includes: a first branch and a second branch, wherein the output of the first branch is connected to the output of the second branch.

[0184] In the above-mentioned embodiments of the present application, the fusion module includes: a first processing unit and a merging unit.

[0185] Among them, the merging unit is used to perform channel merging on multiple first feature maps to obtain a merged feature map; the first processing unit is used to perform convolutional processing on the merged feature map using the first branch to obtain a first output feature; the first processing unit is also used to perform convolutional processing on the merged feature map using the second branch to obtain a second output feature; the merging unit is also used to perform channel merging on the first output feature and the second output feature to obtain at least one second feature map.

[0186] In the above-mentioned embodiments of the present application, the first branch includes: at least one convolutional block, and each convolutional block in the multiple convolutional blocks includes multiple sub-convolutional layers, and the convolutional kernels of the multiple sub-convolutional layers are different.

[0187] In the above-mentioned embodiments of the present application, the detection module is further used to detect the target image using a target detection model to obtain the detection result of the target object, wherein the target detection model is trained based on a target sample image, and the target sample image is obtained by performing data augmentation on multiple sample images through sample detection frames.

[0188] In the above-mentioned embodiments of the present application, the device further includes: an acquisition module, a determination module, an enhancement module, and a training module.

[0189] Among them, the acquisition module is used to acquire multiple sample images and sample detection frames corresponding to the multiple sample images, wherein the sample detection frames are used to label the target objects in the sample images; the determination module is used to determine a preset number of sample detection frames in the multiple sample images as target detection frames; the enhancement module is used to perform data augmentation on the target objects corresponding to the target detection frames to obtain a target sample image; the training module is used to train an initial detection model using the target sample image to obtain a target detection model.

[0190] In the above-mentioned embodiments of the present application, the determination module includes: a splicing unit, a mixing unit, and a determination unit.

[0191] Among them, the splicing unit is used to splice multiple sample images to obtain an initial sample image; the mixing unit is used to mix the initial sample image and a preset sample image to obtain a mixed image; the determination unit is used to determine a preset number of sample detection frames in the mixed image as target detection frames.

[0192] In the above embodiments of the present application, the device further includes: an output module, a receiving module, and an updating module.

[0193] Among them, the output module is used to output the detection result; the receiving module is used to receive a first feedback result, where the first feedback result is used to modify the channels in the merged feature map according to the detection result; the updating module is used to update the merged feature map based on the first feedback result.

[0194] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0195] Embodiment 7

[0196] According to an embodiment of the present application, there is also provided an image processing device for implementing the above image processing method. Figure 13 It is a schematic diagram of an image processing device according to Embodiment 7 of the present application. As Figure 13 shown, the device 1300 includes: an acquisition module 1302, an acquisition module 1302, an acquisition module 1302, an acquisition module 1302.

[0197] Among them, the acquisition module is used to acquire agricultural remote sensing images, where the agricultural remote sensing images contain crops; the extraction module is used to perform multi-scale feature extraction on the agricultural remote sensing images to obtain a plurality of first feature maps; the fusion module is used to use a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; the detection module is used to detect at least one second feature map to obtain the detection result of the crops.

[0198] It should be noted here that the above acquisition module 1302, extraction module 1304, fusion module 1306, and detection module 1308 correspond to steps S802 to S808 in Embodiment 2. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0199] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0200] Embodiment 8

[0201] According to an embodiment of the present application, there is also provided an image processing apparatus for implementing the above image processing method. Figure 14 It is a schematic diagram of an image processing apparatus according to Embodiment 8 of the present application, as Figure 14 shown. The apparatus 1400 includes: an acquisition module 1402, an extraction module 1404, a fusion module 1406, and a detection module 1408.

[0202] Among them, the acquisition module is used to acquire a remote sensing image of a building, where the remote sensing image of the building includes a target building; the extraction module is used to perform multi-scale feature extraction on the remote sensing image of the building to obtain a plurality of first feature maps; the fusion module is used to perform feature fusion on the plurality of first feature maps by using a multi-branch network structure to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; the detection module is used to detect at least one second feature map to obtain a detection result of the target building.

[0203] It should be noted here that the above acquisition module 1402, extraction module 1404, fusion module 1406, and detection module 1408 correspond to steps S902 to S908 in Embodiment 3. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the apparatus, can run in the computer terminal 10 provided in Embodiment 1.

[0204] In the above embodiments of the present application, the method further includes: using a target detection model to detect the remote sensing image of the building to obtain a detection result of the target building, where the target detection model is trained based on a target sample image, and the target sample image is obtained by performing data augmentation on a plurality of sample images through a sample detection frame.

[0205] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0206] Embodiment 9

[0207] According to an embodiment of the present application, there is also provided an image processing apparatus for implementing the above image processing method. Figure 15 It is a schematic diagram of an image processing apparatus according to Embodiment 9 of the present application, as Figure 15As shown, the device 1500 includes: an acquisition module 1502, an extraction module 1504, a fusion module 1506, and a detection module 1508.

[0208] Among them, the acquisition module is used to obtain a target remote sensing image through a cloud server, where the target remote sensing image contains a target object; the extraction module is used to perform multi-scale feature extraction on the target remote sensing image through the cloud server to obtain a plurality of first feature maps; the fusion module is used to perform feature fusion on the plurality of first feature maps through the cloud server using a multi-branch network structure to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; the detection module is used to detect at least one second feature map through the cloud server to obtain a detection result of the target object.

[0209] It should be noted here that the above acquisition module 1502, extraction module 1504, fusion module 1506, and detection module 1508 correspond to steps S1002 to S1008 in Embodiment 4. The functions of the four modules are the same as those of the corresponding steps in terms of the implemented examples and application scenarios, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0210] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.

[0211] Embodiment 10

[0212] An embodiment of the present invention can provide an electronic device, which includes: a memory and a processor. The processor is used to run a program stored in the memory. The electronic device can be a computer terminal, and the computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal can also be replaced with a mobile terminal or other terminal devices.

[0213] Optionally, in this embodiment, the above computer terminal can be located in at least one of multiple network devices in a computer network.

[0214] In this embodiment, the above computer terminal may execute the program code of the following steps in the image processing method: obtaining a target image, where the target image includes a target object; performing multi-scale feature extraction on the target image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and performing detection on at least one second feature map to obtain a detection result of the target object.

[0215] Optionally, Figure 16 is a structural block diagram of a computer terminal according to an embodiment of the present application. As Figure 16 shown, the computer terminal A may include: one or more (only one is shown in the figure) processors and a memory.

[0216] Among them, the memory may be used to store software programs and modules, such as the program instructions / modules corresponding to the image processing method and device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above image processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal A through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0217] The processor may call the information and application programs stored in the memory through a transmission device to execute the following steps: obtaining a target image, where the target image includes a target object; performing multi-scale feature extraction on the target image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and performing detection on at least one second feature map to obtain a detection result of the target object.

[0218] Optionally, the above processor may also execute the program code of the following steps: performing channel merging on the plurality of first feature maps to obtain a merged feature map; performing convolution processing on the merged feature map using a first branch to obtain a first output feature; performing convolution processing on the merged feature map using a second branch to obtain a second output feature; and performing channel merging on the first output feature and the second output feature to obtain at least one second feature map.

[0219] Optionally, the above-mentioned processor can also execute the program code of the following steps: detecting a target image by using a target detection model to obtain a detection result of a target object, where the target detection model is trained based on target sample images, and the target sample images are obtained by performing data augmentation on a plurality of sample images through sample detection frames.

[0220] Optionally, the above-mentioned processor can also execute the program code of the following steps: obtaining a plurality of sample images and sample detection frames corresponding to the plurality of sample images, where the sample detection frames are used to label target objects in the sample images; determining a preset number of sample detection frames in the plurality of sample images as target detection frames; performing data augmentation on the target objects corresponding to the target detection frames to obtain target sample images; and training an initial detection model by using the target sample images to obtain a target detection model.

[0221] Optionally, the above-mentioned processor can also execute the program code of the following steps: splicing a plurality of sample images to obtain an initial sample image; mixing the initial sample image and a preset sample image to obtain a mixed image; and determining a preset number of sample detection frames in the mixed image as target detection frames.

[0222] Optionally, the above-mentioned processor can also execute the program code of the following steps: outputting a detection result; receiving a first feedback result, where the first feedback result is used to modify channels in a merged feature map according to the detection result; and updating the merged feature map based on the first feedback result.

[0223] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: obtaining a target remote sensing image, where the target remote sensing image contains a target object; performing multi-scale feature extraction on the target remote sensing image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and detecting the at least one second feature map to obtain a detection result of the target object.

[0224] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: obtaining an agricultural remote sensing image, where the agricultural remote sensing image contains crops; performing multi-scale feature extraction on the agricultural remote sensing image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and detecting the at least one second feature map to obtain a detection result of the crops.

[0225] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: obtain a remote sensing image of a building, where the remote sensing image of the building includes a target building; perform multi-scale feature extraction on the remote sensing image of the building to obtain a plurality of first feature maps; use a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; perform detection on at least one second feature map to obtain a detection result of the target building.

[0226] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: the cloud server obtains a target image, where the target image includes a target object; the cloud server performs multi-scale feature extraction on the target image to obtain a plurality of first feature maps; the cloud server uses a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; the cloud server performs detection on at least one second feature map to obtain a detection result of the target object.

[0227] By adopting the embodiment of the present application, first, a target image is obtained, where the target image includes a target object; multi-scale feature extraction is performed on the target image to obtain a plurality of first feature maps; a multi-branch network structure is used to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to represent performing feature fusion on the plurality of first feature maps through multiple branches to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches, and detection is performed on at least one second feature map to obtain a detection result of the target object, thereby improving the accuracy of the detection result of the target object. It is easy to notice that feature fusion can be performed on the first feature maps through multiple branches, thereby improving the accuracy of the at least one second feature map obtained, and the multiple branches can also reduce the number of parameters in the fusion process, thereby improving the fusion efficiency, and further solving the technical problem of low accuracy in image detection in the related art.

[0228] Those of ordinary skill in the art can understand that Figure 15 The structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 15 It does not limit the structure of the above electronic device. For example, the computer terminal 10 may further include more Figure 15more or fewer components (such as network interfaces, display devices, etc.) shown in the figure, or having a configuration different from that Figure 15 shown in the figure.

[0229] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0230] Embodiment 11

[0231] An embodiment of the present invention further provides a storage medium. Optionally, in this embodiment, the above storage medium can be used to store the program code executed by the image processing method provided in the above Embodiment 1.

[0232] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0233] Optionally, in this embodiment, the storage medium is set to store the program code for executing the following steps: obtaining a target image, where the target image includes a target object; performing multi-scale feature extraction on the target image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; and performing detection on at least one second feature map to obtain the detection result of the target object.

[0234] Optionally, the above storage medium is further set to store the program code for executing the following steps: performing channel merging on the plurality of first feature maps to obtain a merged feature map; performing convolution processing on the merged feature map using a first branch to obtain a first output feature; performing convolution processing on the merged feature map using a second branch to obtain a second output feature; and performing channel merging on the first output feature and the second output feature to obtain at least one second feature map.

[0235] Optionally, the above storage medium is further set to store the program code for executing the following steps: using a target detection model to detect the target image to obtain the detection result of the target object, where the target detection model is trained based on target sample images, and the target sample images are obtained by performing data augmentation on a plurality of sample images through sample detection frames.

[0236] Optionally, the above storage medium is further configured to store program code for performing the following steps: obtaining a plurality of sample images and sample detection frames corresponding to the plurality of sample images, wherein the sample detection frames are used to label target objects in the sample images; determining a preset number of sample detection frames in the plurality of sample images as target detection frames; performing data augmentation on the target objects corresponding to the target detection frames to obtain target sample images; and training an initial detection model using the target sample images to obtain a target detection model.

[0237] Optionally, the above storage medium is further configured to store program code for performing the following steps: splicing a plurality of sample images to obtain an initial sample image; mixing the initial sample image and a preset sample image to obtain a mixed image; and determining a preset number of sample detection frames in the mixed image as target detection frames.

[0238] Optionally, the above storage medium is further configured to store program code for performing the following steps: outputting a detection result; receiving a first feedback result, wherein the first feedback result is used to modify channels in a merged feature map according to the detection result; and updating the merged feature map based on the first feedback result.

[0239] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a target remote sensing image, wherein the target remote sensing image includes a target object; performing multi-scale feature extraction on the target remote sensing image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, wherein the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and detecting the at least one second feature map to obtain a detection result of the target object.

[0240] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining an agricultural remote sensing image, wherein the agricultural remote sensing image includes crops; performing multi-scale feature extraction on the agricultural remote sensing image to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, wherein the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through a plurality of branches; and detecting the at least one second feature map to obtain a detection result of the crops.

[0241] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a remote sensing image of a building, where the remote sensing image of the building includes a target building; performing multi-scale feature extraction on the remote sensing image of the building to obtain a plurality of first feature maps; using a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; and performing detection on the at least one second feature map to obtain a detection result of the target building.

[0242] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: a cloud server obtains a target image, where the target image includes a target object; the cloud server performs multi-scale feature extraction on the target image to obtain a plurality of first feature maps; the cloud server uses a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches; and the cloud server performs detection on the at least one second feature map to obtain a detection result of the target object.

[0243] By adopting the embodiment of the present application, first, a target image is obtained, where the target image includes a target object; multi-scale feature extraction is performed on the target image to obtain a plurality of first feature maps; a multi-branch network structure is used to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to represent performing feature fusion on the plurality of first feature maps through multiple branches to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches, and detection is performed on the at least one second feature map to obtain a detection result of the target object, thereby achieving an improvement in the accuracy of the detection result of the target object. It is easy to notice that feature fusion can be performed on the first feature maps through multiple branches, thereby improving the accuracy of the at least one second feature map obtained, and the multiple branches can also reduce the number of parameters in the fusion process, thereby improving the fusion efficiency, and further solving the technical problem of low accuracy in image detection in the related art.

[0244] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0245] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0246] In several embodiments provided by this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0247] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0248] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0249] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs and other various media that can store program codes.

[0250] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An image processing method, characterized in that, Including: Obtain a target image, where the target image includes a target object; Perform multi-scale feature extraction on the target image to obtain multiple first feature maps; Use a multi-branch network structure to perform feature fusion on the multiple first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the multiple first feature maps through multiple branches. The multi-branch network structure includes a first branch and a second branch. The convolutional block of the first branch contains multiple sub-convolutional layers, and the convolutional kernels of the multiple sub-convolutional layers have different sizes; Perform detection on the at least one second feature map to obtain a detection result of the target object; Wherein, the multi-branch network structure further includes: a first convolutional layer, the output of the first convolutional layer is connected to the inputs of the first branch and the second branch, and the at least one second feature map is obtained by splicing the output feature maps of the first branch and the second branch. The output feature map is obtained by the first branch and the second branch processing the merged feature map in parallel, and the merged feature map is obtained by the first convolutional layer performing channel merging on the multiple first feature maps.

2. The method according to claim 1, characterized in that, The multi-branch network structure includes: a first branch and a second branch, wherein the output of the first branch is connected to the output of the second branch.

3. The method according to claim 2, characterized in that, Using a multi-branch network structure to perform feature fusion on the multiple first feature maps to obtain at least one second feature map includes: Perform channel merging on the multiple first feature maps to obtain a merged feature map; Use the first branch to perform convolutional processing on the merged feature map to obtain a first output feature; Use the second branch to perform convolutional processing on the merged feature map to obtain a second output feature; Perform channel merging on the first output feature and the second output feature to obtain the at least one second feature map.

4. The method according to claim 3, wherein The first branch includes: at least one convolutional block, and each convolutional block in the multiple convolutional blocks contains multiple sub-convolutional layers, and the convolutional kernels of the multiple sub-convolutional layers are different.

5. The method according to claim 1, characterized in that, The method further includes: Use a target detection model to perform detection on the target image to obtain the detection result of the target object, where the target detection model is trained based on target sample images, and the target sample images are obtained by performing data augmentation on multiple sample images through sample detection frames.

6. The method according to claim 5, wherein The method further includes: Obtain the multiple sample images and the sample detection frames corresponding to the multiple sample images, where the sample detection frames are used to label the target objects in the sample images; Determine a preset number of sample detection frames in the multiple sample images as target detection frames; Perform data augmentation on the target objects corresponding to the target detection frames to obtain target sample images; Use the target sample images to train an initial detection model to obtain the target detection model.

7. The method according to claim 6, wherein Determining a preset number of sample detection frames in the multiple sample images as target detection frames includes: Splice the multiple sample images to obtain an initial sample image; Mix the initial sample image and a preset sample image to obtain a mixed image; Determine the preset number of the sample detection frames in the mixed image as the target detection frames.

8. The method according to claim 3, characterized in that The method further includes: Output the detection result; Receive a first feedback result, where the first feedback result is used to modify the channels in the merged feature map according to the detection result; Update the merged feature map based on the first feedback result.

9. An image processing method, characterized in that, It includes: Obtain a target remote sensing image, where the target remote sensing image contains a target object; Perform multi-scale feature extraction on the target remote sensing image to obtain a plurality of first feature maps; Use a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches, the multi-branch network structure includes a first branch and a second branch, and the convolutional block of the first branch contains a plurality of sub-convolutional layers, and the convolutional kernel sizes of the plurality of sub-convolutional layers are different; Detect the at least one second feature map to obtain a detection result of the target object; Wherein, the multi-branch network structure further includes: a first convolutional layer, the output of the first convolutional layer is connected to the inputs of the first branch and the second branch, the at least one second feature map is obtained by splicing the output feature maps of the first branch and the second branch, the output feature map is obtained by the first branch and the second branch processing the merged feature map in parallel, and the merged feature map is obtained by the first convolutional layer performing channel merging on the plurality of first feature maps.

10. An image processing method, characterized in that, It includes: Obtain a building remote sensing image, where the building remote sensing image contains a target building; Perform multi-scale feature extraction on the building remote sensing image to obtain a plurality of first feature maps; Use a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches, the multi-branch network structure includes a first branch and a second branch, and the convolutional block of the first branch contains a plurality of sub-convolutional layers, and the convolutional kernel sizes of the plurality of sub-convolutional layers are different; Detect the at least one second feature map to obtain a detection result of the target building; Wherein, the multi-branch network structure further includes: a first convolutional layer, the output of the first convolutional layer is connected to the inputs of the first branch and the second branch, the at least one second feature map is obtained by splicing the output feature maps of the first branch and the second branch, the output feature map is obtained by the first branch and the second branch processing the merged feature map in parallel, and the merged feature map is obtained by the first convolutional layer performing channel merging on the plurality of first feature maps.

11. The method according to claim 10, wherein The method further includes: Use a target detection model to detect the building remote sensing image to obtain the detection result of the target building, where the target detection model is trained based on a target sample image, and the target sample image is obtained by performing data augmentation on a plurality of sample images through sample detection frames.

12. An image processing method, characterized in that, It includes: The cloud server obtains a target image, where the target image contains a target object; The cloud server performs multi-scale feature extraction on the target image to obtain a plurality of first feature maps; The cloud server uses a multi-branch network structure to perform feature fusion on the plurality of first feature maps to obtain at least one second feature map, where the multi-branch network structure is used to perform feature fusion on the plurality of first feature maps through multiple branches, the multi-branch network structure includes a first branch and a second branch, and the convolutional block of the first branch contains a plurality of sub-convolutional layers with different convolutional kernel sizes; The cloud server detects the at least one second feature map to obtain a detection result of the target object; Among them, the multi-branch network structure further includes: a first convolutional layer, the output of the first convolutional layer is connected to the inputs of the first branch and the second branch, the at least one second feature map is obtained by splicing the output feature maps of the first branch and the second branch, the output feature maps are obtained by the first branch and the second branch processing the merged feature map in parallel, and the merged feature map is obtained by the first convolutional layer performing channel merging on the plurality of first feature maps.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where, when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 12.

14. An electronic device, characterized in that, Including: A memory and a processor, the processor is used to run the program stored in the memory, where, when the program runs, it executes the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Target detection method and system based on feedback mechanism

    CN111739062A

  • Wheat ear detection method based on Officient Det

    CN113158865A