Target detection method, device, storage medium and electronic device
By fusing the feature maps of multi-eye perspectives, the problem of difficulty in detecting blur or obstructing targets in the prior art is solved, and a higher target detection accuracy is achieved.
Patent Information
- Application Number
- CN202011326014.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-11-23
AI Technical Summary
Existing target detection methods are difficult to effectively detect fuzzy or obstructed targets, resulting in low detection accuracy.
By obtaining views of at least two targets generated based on the same scene, obtaining the feature map of each view, selecting one feature map as the reference map, and converting other feature maps into the pixel coordinate system corresponding to the reference map through grid processing, obtaining a projected feature map, and then fusing the reference map with the projected feature map, obtaining the total feature map, and detecting the target based on the total feature map.
Through the fusion of multi-eye viewing feature maps, feature information is enhanced, blurred or obstructed targets can be effectively detected, and the target detection accuracy is improved.
Smart Images

Figure CN114612875B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a target detection method, device, storage medium and electronic equipment. Background Art
[0002] Object detection is a technology that is involved in many scenarios. It is to detect the category and precise location of objects from images. For example, in intelligent driving scenarios, in order to ensure the normal driving of the vehicle, it is necessary to accurately detect motor vehicles, pedestrians, non-motor vehicles and other targets within a certain distance in front, so as to make intelligent decisions.
[0003] Currently, popular target detection methods include FasterRCNN, SSD and YOLO, but these methods are usually based on monocular images for target detection. This detection method cannot effectively detect blurred or occluded targets, and the detection accuracy is not high. Summary of the invention
[0004] In view of this, the present invention provides a target detection method, device, storage medium and electronic device, which are used to effectively detect blurred or obscured targets and improve target detection accuracy.
[0005] Specifically, the present invention is achieved through the following technical solutions:
[0006] According to a first aspect of the present invention, there is provided a target detection method, the method comprising:
[0007] Obtain at least two views of the target generated based on the same scene;
[0008] Obtaining a feature map of each of the views;
[0009] Selecting one of the feature maps as a reference map, and converting the other feature maps into a pixel coordinate system corresponding to the reference map through gridding processing, to obtain each projection feature map;
[0010] The reference image is merged with all the projection feature images to obtain a total feature image;
[0011] The target is detected based on the overall feature map.
[0012] In a possible implementation, the step of converting other feature maps into pixel coordinate systems corresponding to the reference map through gridding to obtain respective projection feature maps includes:
[0013] For any of the other feature maps, according to the internal and external parameters of the shooting device corresponding to the feature map, the pixel points in the feature map are projected to a predefined grid virtual space in the world coordinate system, and the feature values of the pixel points in the feature map are assigned to the corresponding grid vertices in the virtual space; the center of the virtual space is the coordinate origin of the world coordinate system;
[0014] According to the internal and external parameters of the shooting device corresponding to the reference image, the mesh vertices in the virtual space are projected to the pixel coordinate system corresponding to the reference image, and the values of the mesh vertices in the virtual space are assigned to the corresponding projection points to obtain the projection feature map corresponding to the feature map.
[0015] In a possible implementation, projecting the pixel points in the feature map to a predefined grid virtual space in a world coordinate system according to the internal and external parameters of the shooting device corresponding to the feature map includes:
[0016] Sampling is performed in the feature map using a linear interpolation method to obtain sampling points;
[0017] According to the internal and external parameters of the shooting device corresponding to the feature map, the sampling points are projected into the virtual space, and the feature values of the sampling points are assigned to the corresponding mesh vertices in the virtual space.
[0018] In a possible implementation, fusing the reference image with all the projection feature images to obtain a total feature image includes:
[0019] The reference image and the values of corresponding pixels in all the projected feature images are added to obtain a total feature image.
[0020] According to a second aspect of the present invention, there is provided an object detection device, the device comprising a module for executing the object detection method in the first aspect or any possible implementation manner of the first aspect.
[0021] According to a third aspect of the present invention, there is provided a storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the program implements the steps of the target detection method in the first aspect or any possible implementation of the first aspect.
[0022] According to a fourth aspect of the present invention, there is provided an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the target detection method in the first aspect or any possible implementation of the first aspect are implemented.
[0023] In a possible implementation, the electronic device is a train approach protection warning device or a photographing device.
[0024] The technical solution provided by the present invention brings at least the following beneficial effects:
[0025] The technical solution provided by the present invention first obtains at least two views of the target generated based on the same scene, then obtains the feature map of each view, and then selects one of the feature maps as the reference map, and converts the other feature maps into the pixel coordinate system corresponding to the reference map through gridding processing, so as to obtain each projection feature map, and then fuses the reference map with all the projection feature maps to obtain the total feature map, and detects the target based on the total feature map. That is to say, the present invention enhances the features by fusing the feature maps of multiple perspectives, so that blurred or obstructed targets can be effectively detected, thereby improving the target detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A schematic diagram of a flow chart of a target detection method provided by an embodiment of the present invention;
[0027] Figure 2 A schematic diagram of a flow chart of another target detection method provided by an embodiment of the present invention;
[0028] Figure 3 A schematic diagram of the structure of a target detection device provided by an embodiment of the present invention;
[0029] Figure 4 A schematic diagram of the structure of an image processing module in a target detection device provided by an embodiment of the present invention;
[0030] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0032] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0033] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0034] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0035] See also Figure 1 The embodiment of the present invention provides a target detection method, which can be applied to electronic devices with image processing functions, such as unmanned vehicles, automatically flying drones, mobile phones with multiple cameras, and the method may include the following steps:
[0036] S101, acquiring at least two target views generated based on the same scene;
[0037] Among them, the views of at least two targets are at least two pictures taken at the same time by at least two shooting devices with different viewing angles (such as visual sensors), for example, two pictures taken at the same time by two shooting devices with different viewing angles (such as binocular cameras).
[0038] S102, obtaining a feature map of each view;
[0039] In some embodiments, obtaining a feature map of each view in step S102 includes:
[0040] For each of the views, the view is input into a trained convolutional neural network to obtain a feature map of the view output by the convolutional neural network.
[0041] In an embodiment of the present invention, the convolutional neural network may be, for example, a Darknet53 network.
[0042] For example, the view size captured by the camera is (960, 512, 3). After the convolutional neural network, it is downsampled by 4 times to obtain a feature map of size (240, 128, 256).
[0043] S103, selecting one of the feature maps as a reference map, and converting the other feature maps into a pixel coordinate system corresponding to the reference map through gridding processing, to obtain respective projection feature maps;
[0044] In some embodiments, in step S103, other feature maps are converted into the pixel coordinate system corresponding to the reference map by gridding processing to obtain various projection feature maps, including:
[0045] For any of the other feature maps, according to the internal and external parameters of the shooting device corresponding to the feature map, the pixel points in the feature map are projected to a predefined grid virtual space in the world coordinate system, and the feature values of the pixel points in the feature map are assigned to the corresponding grid vertices in the virtual space; the center of the virtual space is the coordinate origin of the world coordinate system;
[0046] According to the internal and external parameters of the shooting device corresponding to the reference image, the mesh vertices in the virtual space are projected to the pixel coordinate system corresponding to the reference image, and the values of the mesh vertices in the virtual space are assigned to the corresponding projection points to obtain the projection feature map corresponding to the feature map.
[0047] In the embodiment of the present invention, the virtual space G represents the area to be detected, and the virtual space G includes the area where the target appears. According to the appearance of the target, for example, the length a=10 meters, the width b=10 meters, and the height c=10 meters of the virtual space G can be defined. The virtual space G is in the world coordinate system, and the coordinate origin of the world coordinate system can be defined as the center of the virtual space G.
[0048] In some embodiments, the interval between each grid in the virtual space G can be defined as 50 mm, and the value of each grid vertex can be initialized to 0.
[0049] In some embodiments, the step of projecting the pixel points in the feature map to a predefined gridded virtual space in a world coordinate system according to the internal and external parameters of the shooting device corresponding to the feature map includes:
[0050] Sampling is performed in the feature map using a linear interpolation method to obtain sampling points;
[0051] According to the internal and external parameters of the shooting device corresponding to the feature map, the sampling points are projected into the virtual space, and the feature values of the sampling points are assigned to the corresponding mesh vertices in the virtual space.
[0052] S104, fusing the reference image with all the projection feature images to obtain a total feature image;
[0053] In some embodiments, in step S104, the reference image is merged with all the projection feature images to obtain a total feature image, including:
[0054] The reference image and the values of corresponding pixels in all the projected feature images are added to obtain a total feature image.
[0055] For example, the base image and the values of corresponding pixels in all projected feature images may be directly added to obtain a total feature image, or the base image and the values of corresponding pixels in all projected feature images may be weighted added to obtain a total feature image, which is not limited in this embodiment of the present invention.
[0056] Of course, other methods may also be used to fuse the reference image with all the projection feature images to obtain the total feature image, which is not limited in the embodiment of the present invention.
[0057] S105: Detect the target based on the total feature map.
[0058] In some embodiments, the target can be detected on the overall feature map through RPN (Region Proposal Networks).
[0059] The following takes binocular detection as an example to illustrate the process of the target detection method provided by the embodiment of the present invention. Figure 2 shown.
[0060] S201, obtaining two target views P1 and P2 generated based on the same scene;
[0061] Among them, P1 is the image captured by the first camera, and P2 is the image captured by the second camera.
[0062] S202, obtaining the feature graph F1 of P1 and the feature graph F2 of P2 respectively through the Darknet53 network;
[0063] S203, selecting F1 as the reference image, projecting the pixel points in F2 to a predefined grid virtual space G in the world coordinate system according to the internal and external parameters of the second camera corresponding to F2, and assigning the feature values of the pixel points in F2 to the corresponding grid vertices in the virtual space G (i.e., the projection points of the pixel points in F2), and the center of the virtual space G is the coordinate origin of the world coordinate system;
[0064] S204, projecting the mesh vertices in the virtual space G to the pixel coordinate system corresponding to F1 according to the internal and external parameters of the first camera corresponding to F1, and assigning the values of the mesh vertices in the virtual space G to the corresponding projection points, to obtain a projection feature map f21 corresponding to F2;
[0065] S205, adding the values of corresponding pixels in F1 and f21 to obtain a total feature map F;
[0066] S206: Detect the target on the total feature map F through RPN.
[0067] Based on the same inventive concept, see Figure 3An embodiment of the present invention provides a target detection device, including: a view acquisition module 11, a feature map acquisition module 12, an image processing module 13, an image fusion module 14 and a target detection module 15.
[0068] A view acquisition module 11 is configured to acquire at least two target views generated based on the same scene;
[0069] A feature map acquisition module 12 is configured to acquire a feature map of each of the views;
[0070] The image processing module 13 is configured to select one of the feature maps as a reference map, and transform the other feature maps into the pixel coordinate system corresponding to the reference map through gridding processing to obtain various projection feature maps;
[0071] An image fusion module 14 is configured to fuse the reference image with all the projection feature images to obtain a total feature image;
[0072] The target detection module 15 is configured to detect the target based on the overall feature map.
[0073] In some embodiments, Figure 4 As shown, the image processing module 13 includes:
[0074] The first image processing submodule 131 is configured to project the pixel points in any of the other feature maps to a predefined grid virtual space in a world coordinate system according to the internal and external parameters of the shooting device corresponding to the feature map, and assign the feature values of the pixel points in the feature map to the corresponding grid vertices in the virtual space; the center of the virtual space is the coordinate origin of the world coordinate system;
[0075] The second image processing submodule 132 is configured to project the mesh vertices in the virtual space to the pixel coordinate system corresponding to the reference image according to the internal and external parameters of the shooting device corresponding to the reference image, and assign the values of the mesh vertices in the virtual space to the corresponding projection points to obtain a projection feature map corresponding to the feature map.
[0076] In some embodiments, the first image processing submodule 131 is configured to:
[0077] Sampling is performed in the feature map using a linear interpolation method to obtain sampling points;
[0078] According to the internal and external parameters of the shooting device corresponding to the feature map, the sampling points are projected into the virtual space, and the feature values of the sampling points are assigned to the corresponding mesh vertices in the virtual space.
[0079] In some embodiments, the feature map acquisition module 12 is configured to:
[0080] For each of the views, the view is input into a trained convolutional neural network to obtain a feature map of the view output by the convolutional neural network.
[0081] In some embodiments, the image fusion module 14 is configured to:
[0082] The reference image and the values of corresponding pixels in all the projected feature images are added to obtain a total feature image.
[0083] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0084] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0085] Based on the same inventive concept, an embodiment of the present invention further provides a storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of the target detection method in any possible implementation manner described above are implemented.
[0086] Alternatively, the storage medium may be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0087] Based on the same inventive concept, an embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the target detection method in any possible implementation manner described above.
[0088] Based on the same inventive concept, see Figure 5 An embodiment of the present invention further provides an electronic device, including a memory 71 (e.g., a non-volatile memory), a processor 72, and a computer program stored in the memory 71 and executable on the processor 72. When the processor 72 executes the program, the steps of the target detection method in any possible implementation method described above are implemented, which may be equivalent to the target detection device described above. Of course, the processor may also be used to process other data or operations.
[0089] The electronic device may be a photographing device, such as a camera, a gimbal with a camera, or a drone. The drone may include at least a binocular vision sensor. The electronic device may be a train approach protection warning device, which may be mounted on a train. In addition, the electronic device may also be, for example, an unmanned car, a mobile terminal with multiple cameras, such as a mobile phone with dual cameras, and the like.
[0090] like Figure 5 As shown, the electronic device may also generally include: a memory 73, a network interface 74, and an internal bus 75. In addition to these components, other hardware may also be included, which will not be described in detail.
[0091] It should be pointed out that the above-mentioned target detection device can be implemented by software. As a device in a logical sense, it is formed by the processor 72 of the electronic device in which it is located reading the computer program instructions stored in the non-volatile memory into the memory 73 for execution.
[0092] The embodiments of the subject matter and functional operations described in this specification may be implemented in the following: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules in computer program instructions encoded on a tangible non-temporary program carrier to be executed by a data processing device or to control the operation of the data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information and transmit it to a suitable receiver device for execution by a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0093] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform corresponding functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits), and the apparatus can also be implemented as special purpose logic circuits.
[0094] Computers suitable for executing computer programs include, for example, general and / or special microprocessors, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory and / or a random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, the computer will also include one or more large-capacity storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to this large-capacity storage device to receive data from it or to transmit data to it, or both. However, the computer does not necessarily have such a device. In addition, the computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0095] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0096] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of the specific embodiments of specific inventions. Certain features described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may work in certain combinations as described above and even initially claim protection, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of a sub-combination.
[0097] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or requiring that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
[0098] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.
[0099] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A target detection method, characterized in that: The method comprises: Obtain at least two views of the target generated based on the same scene; Obtaining a feature map of each of the views; Selecting one of the feature maps as a reference map, and converting the other feature maps into a pixel coordinate system corresponding to the reference map through gridding processing, to obtain each projection feature map; The reference image is merged with all the projection feature images to obtain a total feature image; Detecting the target based on the overall feature map; The other feature maps are converted into the pixel coordinate system corresponding to the reference map by gridding processing to obtain each projection feature map, including: For any of the other feature maps, according to the internal and external parameters of the shooting device corresponding to the feature map, the pixel points in the feature map are projected to a predefined grid virtual space in the world coordinate system, and the feature values of the pixel points in the feature map are assigned to the corresponding grid vertices in the virtual space; the center of the virtual space is the coordinate origin of the world coordinate system; According to the internal and external parameters of the shooting device corresponding to the reference image, the mesh vertices in the virtual space are projected to the pixel coordinate system corresponding to the reference image, and the values of the mesh vertices in the virtual space are assigned to the corresponding projection points to obtain the projection feature map corresponding to the feature map.
2. The method according to claim 1, characterized in that The step of projecting the pixel points in the feature map to a predefined grid virtual space in a world coordinate system according to the internal and external parameters of the shooting device corresponding to the feature map includes: Sampling is performed in the feature map using a linear interpolation method to obtain sampling points; According to the internal and external parameters of the shooting device corresponding to the feature map, the sampling points are projected into the virtual space, and the feature values of the sampling points are assigned to the corresponding mesh vertices in the virtual space.
3. The method according to claim 1, characterized in that The step of fusing the reference image with all the projection feature images to obtain a total feature image comprises: The reference image and the values of corresponding pixels in all the projected feature images are added to obtain a total feature image.
4. A target detection device, characterized in that: The device comprises: A view acquisition module, configured to acquire at least two views of a target generated based on the same scene; A feature map acquisition module, configured to acquire a feature map of each of the views; An image processing module is configured to select one of the feature maps as a reference map, and transform the other feature maps into a pixel coordinate system corresponding to the reference map through gridding processing to obtain respective projection feature maps; An image fusion module, configured to fuse the reference image with all the projection feature images to obtain a total feature image; A target detection module, configured to detect the target based on the overall feature map; The image processing module comprises: The first image processing submodule is configured to project the pixel points in any of the other feature maps to a predefined grid virtual space in a world coordinate system according to the internal and external parameters of the shooting device corresponding to the feature map, and assign the feature values of the pixel points in the feature map to the corresponding grid vertices in the virtual space; the center of the virtual space is the coordinate origin of the world coordinate system; The second image processing submodule is configured to project the mesh vertices in the virtual space to the pixel coordinate system corresponding to the reference image according to the internal and external parameters of the shooting device corresponding to the reference image, and assign the values of the mesh vertices in the virtual space to the corresponding projection points to obtain a projection feature map corresponding to the feature map.
5. The device according to claim 4, characterized in that The first image processing submodule is configured as follows: Sampling is performed in the feature map using a linear interpolation method to obtain sampling points; According to the internal and external parameters of the shooting device corresponding to the feature map, the sampling points are projected into the virtual space, and the feature values of the sampling points are assigned to the corresponding mesh vertices in the virtual space.
6. The device according to claim 4, characterized in that The image fusion module is configured as follows: The reference image and the values of corresponding pixels in all the projected feature images are added to obtain a total feature image.
7. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 3 are implemented.
9. The electronic device according to claim 8, characterized in that: The electronic device is a train approach protection warning device or a photographing device.
Citation Information
Patent Citations
Tracking system based on binocular camera shooting
CN101344965A