Drone Aerial Perspective Personnel Recognition Method, Medium and Electronic Device
Through the target yolo-v5 algorithm and super-resolution processing, combined with color, gradient direction histogram and contour features, the problem of discontinuity of identity information of drones at high altitude recognition is solved, and continuous tracking and trajectory continuity is achieved.
Patent Information
- Application Number
- CN202311722447.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-12-14
AI Technical Summary
When a drone flies at high altitude, it is difficult to continuously identify personnel identity information, resulting in breakpoints in tracking trajectory.
The target yolo-v5 algorithm is used to preprocess and segment the image, combined with super-resolution processing and feature extraction, and obtain color, gradient direction histogram and contour features to achieve continuous recognition and tracking of the target person.
When the drone cannot collect face images, the matching degree recognition of color, gradient direction histogram and contour features ensures the continuity of the tracking trajectory and avoids tracking breakpoints.
Smart Images

Figure CN117765276B_ABST
Abstract
Description
Background Art
[0002] Drones usually fly at a relatively high altitude and can only capture facial information of people at certain specific times. Most of the time, only personal information of people can be obtained. In the existing technology, facial information is mainly used for personnel identification when tracking people. However, since facial information of people cannot be collected at any time during the entire tracking process, it may lead to breakpoints in the tracking trajectory and affect the tracking result. Summary of the Invention
[0003] The technical problem to be solved by this application is: how to ensure that identity information can be recognized at all times when tracking people in the case of not being able to recognize the face, and to ensure that there are no breakpoints in the tracking trajectory.
[0004] According to the first aspect of this application, a method for identifying people from the aerial perspective of a drone is provided, including:
[0005] S100, obtaining a plurality of target images; wherein, the target image is an image obtained by preprocessing the original image collected by the drone.
[0006] S200, according to each target image, obtaining a first set of target image blocks YT=(YT1, YT2,..., YTi,..., YTj); i = 1, 2,..., j; wherein, j is the number of the first target image blocks included in all target images; the first target image block is an image block including the target person obtained by splitting the target image; each first target image block contains a target task; YTi is the i-th first target image block; the first target image block is obtained by identifying the target image according to the target yolo-v5 algorithm; the target yolo-v5 algorithm sets a preset detection scale, so that any two people in the target image are in different preset grids of the target yolo-v5 algorithm.
[0007] S300, according to YT, obtaining a second set of target image blocks XT=(XT1, XT2,..., XTi,..., XTj); wherein, XTi is the second target image block obtained by performing super-resolution processing on YTi; the resolution of the second target image block is greater than that of the first target image block.
[0008] S400, according to XT, obtaining a set of target feature vectors XL=(XL1, XL2,..., XLi,..., XLj); wherein, XLi is the target feature vector corresponding to XTi; XLi=(XLCi, XLTi, XLKi, XLRi); wherein, XLCi is the color feature corresponding to XTi; XLTi is the histogram of oriented gradients feature corresponding to XTi; XLKi is the contour feature corresponding to XTi; XLRi is the face feature corresponding to XTi.
[0009] For S500, if XLKi is empty, obtain the matching degree Pi between XLi and the feature vector CL to be processed; Pi = (XLi·CL) / (|XLi|×|CL|).
[0010] For S600, when Pi is greater than the preset matching degree threshold, determine the target person corresponding to XLi as the person to be tracked.
[0011] For S700, obtain the trajectory of the person to be tracked based on the preset recognition algorithm.
[0012] According to the second aspect of the present application, there is provided a non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, and the at least one instruction or at least one program segment is loaded and executed by a processor to implement the above-mentioned method for identifying a person from an aerial perspective of a drone.
[0013] According to the third aspect of the present application, there is provided an electronic device including a processor and the above-mentioned non-transitory computer-readable storage medium.
[0014] The present application has at least the following beneficial effects:
[0015] This application first obtains a number of target images obtained by preprocessing the original images collected by the drone; secondly, according to the target YOLO-v5 algorithm, the target images are split to obtain a number of image blocks containing target persons, obtaining the first set of target image blocks. In order to take into account the detection of both large and small targets, the target YOLO-v5 algorithm sets a preset detection scale, so that any two persons in the target image are in different preset grids of the target YOLO-v5 algorithm, improving the accuracy of the detection and extraction of small targets in the target image by the target YOLO-v5. Furthermore, since the flight altitude of the drone is relatively high, the resolution of the first target image blocks of the image blocks containing target persons obtained thereby is relatively low, so the first target image blocks are subjected to super-resolution processing; a number of second target image blocks are obtained. Then, according to each second target image block, the target feature vector of each target person is obtained. Here, the target feature vector includes color features, histogram of oriented gradients (HOG) features, contour features, and face features. The color feature is the feature extracted from the color of the image block, which can reflect information such as the clothing color and overall skin color of the target person; the HOG feature describes the distribution of the gradient intensity and gradient direction in the local area of the image block, and can characterize the appearance and shape of the object in the image block; the contour feature is the feature extracted from the edge of the person, which can characterize the human contour; the face feature can characterize the facial features of the person; the above features can all be used to identify the identity information of the person, and the face feature is more accurate in recognition. However, if the drone does not collect a face image, that is, the face feature is empty, at this time, the matching degree between the target feature vector and the feature vector to be processed in the target feature vector is obtained; if it is greater than the preset matching degree threshold, it is determined that the target person corresponding to XLi is the person to be tracked; and it is tracked to generate a trajectory. When this application cannot collect a face image, it uses the color features, HOG features, and contour features of the target person to realize the continuous identification and tracking of the target person, which can ensure the identification of the identity information of the target person at all times and ensure that the tracking trajectory does not have a breakpoint. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a flowchart of a method for identifying personnel from an aerial perspective of a drone provided by an embodiment of the present application;
[0018] Figure 2 It is a schematic structural diagram of the target YOLO-v5 algorithm provided by an embodiment of the present application;
[0019] Figure 3 Schematic diagrams of the FPN structure and the PAN structure of the original Yolo-v5 algorithm provided by the embodiments of the present application;
[0020] Figure 4 Schematic diagrams of the FPN structure and the PAN structure of the target Yolo-v5 algorithm provided by the embodiments of the present application;
[0021] Figure 5 Schematic diagram of the convolutional attention module of the target Yolo-v5 algorithm provided by the embodiments of the present application. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.
[0023] As Figure 1 shown, a method for identifying people from an aerial perspective of a drone according to an embodiment of the present application includes the following steps:
[0024] S100, obtaining a plurality of target images; wherein, the target images are images obtained by preprocessing the original images collected by the drone.
[0025] Specifically, step S100 includes:
[0026] S110, collecting a plurality of original images.
[0027] Here, a plurality of original images are collected by using the on-board camera of the drone, and the plurality of original images are transmitted to the terminal or the server side by using a wireless communication protocol.
[0028] S120, performing denoising processing and sharpening processing on each of the plurality of original images to obtain a plurality of target images.
[0029] Here, the terminal or the server side is used to preprocess each of the plurality of original images.
[0030] First, Gaussian filtering denoising processing is performed. A Gaussian filtering template is constructed and smoothly moved within the collected remote sensing image. The distribution parameter of the Gaussian function represents the width of the Gaussian function and determines the denoising degree of the remote sensing image, which can remove the influence of isolated points on the image.
[0031] Secondly, Laplace filter sharpening processing: mainly using the second-order derivative of the image to enhance the edge of the image, first operate the image processed by Gaussian filtering with a preset convolution template, and then attach it to the original image to obtain the sharpened result. The image processed by Gaussian filtering can achieve noise removal, but it will cause a certain degree of image blur. Then the purpose of edge sharpening can be achieved through Laplace filtering.
[0032] S200, according to each target image, obtain the first target image block set YT = (YT1, YT2, ..., YTi, ..., YTj); i = 1, 2, ..., j; wherein j is the number of first target image blocks contained in all target images; the first target image block is an image block including a target person obtained by segmenting the target image; each first target image block contains a target task; YTi is the i-th first target image block; the first target image block is obtained by identifying the target image according to the target yolo-v5 algorithm; the target yolo-v5 algorithm sets a preset detection scale so that any two persons in the target image are in different preset grids of the target yolo-v5 algorithm.
[0033] Specifically, the pre-processed image is subjected to character recognition. Figure 2 As shown, this application uses the target YOLO-V5 algorithm to detect pedestrians in the preprocessed image, and uses anchor frame recognition and cropping to obtain multiple local image blocks. Since the target person in the remote sensing image taken by the drone at high altitude is small, it is difficult to identify. Therefore, the following improvements are made to the original YOLO-V5 network structure:
[0034] The multi-scale fusion idea is embedded in the target Yolo-v5 algorithm. The target Yolo-v5 algorithm uses FPN and PAN structures for multi-scale fusion, such as Figure 3 The original yolo-v5 uses three different scales. In order to take into account the detection of large and small pedestrian targets, this application adds a preset detection scale. As an example, the preset detection scale can be 152*152 to ensure that each small pedestrian target falls in each grid, avoiding the previous situation where multiple small pedestrian targets fall in the same grid at the same time. Figure 4As shown, the Feature Pyramid Network (FPN) mainly consists of three parts: bottom-up, lateral connection, and top-down. The feature maps of four scales are transmitted bottom-up. After passing through the convolutional layer, the feature maps gradually become smaller, but the information they contain becomes richer. First, in the top-down process, the shallow feature maps are upsampled to the same size as the feature maps of the next lower layer. Then, the lateral connection reduces the dimension in a top-down manner to obtain the feature maps, and they are fused with the feature maps obtained by upsampling in the top-down process. Finally, the fused feature maps are processed using a 3×3 convolution to improve the non-linear expression ability of the network, thereby enhancing the performance of object detection.
[0035] In addition, the above-mentioned YOLO-v5 algorithm also sets up a convolutional attention module, a channel attention module, and a spatial attention module.
[0036] Specifically, most attention mechanisms usually extract feature maps according to the channel and spatial dimensions respectively, without learning the attention weights from both the channel and the spatial simultaneously, making it difficult to learn rich features and ensure the recognition performance of the network. Therefore, based on the original YOLO-v5 algorithm, a convolutional attention module is introduced to obtain the target YOLO-v5 algorithm, which also includes a channel attention module and a spatial attention module. The schematic diagram of the convolutional attention module is as Figure 5 shown. Specifically, the convolutional attention module is mainly applied in the Backbone stage of the network of the YOLO-v5 algorithm, changing the original CPL to CPL+CBAM. The advantage of CBAM is that it considers both spatial attention and channel attention simultaneously, can learn more rich features, and improve the performance of the network in recognizing pedestrians.
[0037] S300, according to YT, obtain the second target image block set XT=(XT1, XT2,..., XTi,..., XTj); where XTi is the second target image block obtained by performing super-resolution processing on YTi; the resolution of the second target image block is greater than that of the first target image block.
[0038] Specifically, the above-mentioned super-resolution processing is performed according to the bilinear interpolation method. Here, a super-resolution algorithm is used to clarify the first target image block with blurred details, aiming to reconstruct the first target image block from low resolution to high resolution. The specific super-resolution reconstruction method uses the simple and effective bilinear interpolation method, mainly by linearly extending the interpolation function of two variables and performing linear interpolation once in both the horizontal and vertical directions to achieve the reconstruction of the low-resolution first target image block. The reconstructed second target image block is clearer than before, laying a foundation for subsequent feature extraction and vectorization.
[0039] S400. According to XT, obtain the target feature vector set XL = (XL1, XL2,..., XLi,..., XLj); where XLi is the target feature vector corresponding to XTi; XLi = (XLCi, XLTi, XLKi, XLRi); where XLCi is the color feature corresponding to XTi; XLTi is the histogram of oriented gradients feature corresponding to XTi; XLKi is the contour feature corresponding to XTi; XLRi is the face feature corresponding to XTi.
[0040] Specifically, the color feature is the feature extracted from the color of the image block, which can reflect information such as the clothing color and overall skin color of the target person; the acquisition method of the color feature is as follows: convert the second target image block from the RGB color space to the HSV color space. Among them, the difference in the saturation component S is relatively obvious and can be used as the color feature for marking. S conforms to the following characteristics:
[0041] S = 1 - (3 / (R + G + B)) * [min(R, G, B)];
[0042] Among them, R, G, and B are the pixel intensity values of the three color channels in the RGB color space, and min(R, G, B) is the minimum intensity value of the pixel point in the three channels.
[0043] The histogram of oriented gradients feature describes the distribution of the gradient intensity and gradient direction in the local area of the image block, and can characterize the appearance and shape of the object in the image block;
[0044] The histogram of oriented gradients feature XLTi corresponding to XLi is obtained through the following steps:
[0045] S401. Normalize XTi to obtain XTGi;
[0046] S402. Calculate the horizontal gradient and vertical gradient of each preset pixel point of the image corresponding to XTGi;
[0047] S403. Divide the horizontal gradient and vertical gradient of each preset pixel point into several intervals respectively to obtain the gradient histogram of each preset pixel point;
[0048] S404. Obtain the distribution of the gradient amplitude in each interval of the gradient histogram of each preset pixel point;
[0049] S405. Generate several multi-dimensional feature vectors according to the distribution of the gradient amplitude in each interval of the gradient histogram of each preset pixel point;
[0050] S406. Concatenate the above multi-dimensional feature vectors and perform normalization processing to obtain the feature descriptor;
[0051] S407. Concatenate the feature descriptors corresponding to each preset pixel point to obtain XTi.
[0052] The contour feature is the feature of the human body edge extracted, which can represent the human body contour. Here, in order to reflect the difference between the human body and the surrounding background, the canny operator (edge detection operator) can be used to detect the edge of the whole human body as the contour feature.
[0053] The face feature can represent the facial features of a person. All the above features can be used to identify the personal identity information.
[0054] Obtain the target feature vector according to the above color feature, histogram of oriented gradients feature, contour feature and face feature.
[0055] S500. If XLKi is empty, obtain the matching degree Pi between XLi and the to-be-processed feature vector CL; Pi = (XLi · CL) / (|XLi| × |CL|).
[0056] Specifically, if the face feature is empty, that is, the UAV fails to collect a face image, at this time, identity recognition and tracking are performed according to the target feature vector. Obtain the matching degree Pi between XLi and the to-be-processed feature vector CL. Here, the higher the matching degree, the higher the similarity between XLi and the to-be-processed feature vector CL.
[0057] S600. When Pi is greater than the preset matching degree threshold, determine the target person corresponding to XLi as the person to be tracked.
[0058] Specifically, associate the target person corresponding to XLi with the person corresponding to the to-be-processed feature vector, that is, determine the target person corresponding to XLi as the person to be tracked.
[0059] S700. Obtain the trajectory of the person to be tracked based on a preset recognition algorithm.
[0060] Specifically, the preset recognition algorithm can be a re-identification algorithm. Based on the re-identification algorithm, track the person to be tracked to obtain the trajectory of the person to be tracked. And perform map rendering and visual display on the trajectory of the person to be tracked.
[0061] In an exemplary embodiment of the present application, after step S400, the above method further includes:
[0062] S501. If XLKi is not empty, obtain the matching degree PRi between XLKi and the to-be-processed face feature CLR in the to-be-processed feature vector; PRi = (XLKi · CLR) / (|XLKi| × |CLR|).
[0063] Specifically, since the confidence level of face features in identity recognition is relatively high, in this embodiment, if face features exist, the face features are used to match the face features to be processed in the feature vector to be processed.
[0064] S502. If PRi is greater than the preset face matching degree threshold, determine the target person corresponding to XLi as the person to be tracked, and proceed to step S700.
[0065] An embodiment of the present application also provides a computer program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the methods according to various exemplary embodiments described above in this specification.
[0066] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0067] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described here can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the methods according to the embodiments of the present disclosure.
[0068] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.
[0069] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.
[0070] An electronic device according to this embodiment of the present application. The electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0071] The electronic device is presented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one of the above-mentioned processors, at least one of the above-mentioned memories, and a bus connecting different system components (including the memory and the processor).
[0072] Among them, the memory stores program code, and the program code can be executed by the processor, so that the processor executes the steps according to various exemplary embodiments of the present application described in the "Exemplary Method" section above of this specification.
[0073] The memory may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) and / or a cache memory, and may further include a read-only memory (ROM).
[0074] The memory may also include a program / utilities having a set (at least one) of program modules. Such program modules include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0075] The bus may represent one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus structures.
[0076] The electronic device may also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device, and / or communicate with any device that enables the electronic device to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface. Moreover, the electronic device may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter. As shown in the figure, the network adapter communicates with other modules of the electronic device through the bus. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0077] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions for causing a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0078] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium having a program product stored thereon that can implement the methods described above in this specification. In some possible embodiments, various aspects of the present application can also be implemented in the form of a program product, which includes program code that, when the program product runs on a terminal device, causes the terminal device to execute the steps according to various exemplary embodiments described in the "Exemplary Methods" section above of this specification.
[0079] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0080] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0081] The program code included on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0082] The program code for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0083] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0084] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0085] The above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for identifying personnel from an aerial perspective of an unmanned aerial vehicle, characterized in that, Including: S100, obtaining a number of target images; where the target images are images obtained by preprocessing the original images collected by the drone; S200, obtaining a first set of target image patches YT=(YT1, YT2,..., YTi,..., YTj) according to each target image; i = 1, 2,..., j; where j is the number of first target image patches included in all target images; the first target image patch is an image patch including the target person obtained by splitting the target image; each first target image patch contains a target task; YTi is the i-th first target image patch; the first target image patch is obtained by identifying the target image according to the target yolo-v5 algorithm; the target yolo-v5 algorithm sets a preset detection scale so that any two people in the target image are in different preset grids of the target yolo-v5 algorithm; S300, obtaining a second set of target image patches XT=(XT1, XT2,..., XTi,..., XTj) according to YT; where XTi is the second target image patch obtained by performing super-resolution processing on YTi; the resolution of the second target image patch is greater than that of the first target image patch; S400, obtaining a set of target feature vectors XL=(XL1, XL2,..., XLi,..., XLj) according to XT; where XLi is the target feature vector corresponding to XTi; XLi=(XLCi, XLTi, XLKi, XLRi); where XLCi is the color feature corresponding to XTi; XLTi is the histogram of oriented gradients feature corresponding to XTi; XLKi is the contour feature corresponding to XTi; XLRi is the face feature corresponding to XTi; S500, if XLKi is empty, obtaining the matching degree Pi between XLi and the feature vector CL to be processed; Pi=(XLi·CL) / (|XLi|×|CL|); S600, when Pi is greater than the preset matching degree threshold, determining the target person corresponding to XLi as the person to be tracked; S700, obtaining the trajectory of the person to be tracked based on the preset recognition algorithm.
2. The method for identifying personnel from an aerial perspective of a drone according to claim 1, wherein After step S400, the method further includes: S501, if XLKi is not empty, obtaining the matching degree PRi between XLKi and the face feature CLR to be processed in the feature vector to be processed; PRi=(XLKi·CLR) / (|XLKi|×|CLR|); S502, if PRi is greater than the preset face matching degree threshold, determining the target person corresponding to XLi as the person to be tracked and entering step S700.
3. The method for identifying personnel from an aerial perspective of a drone according to claim 1, wherein Step S100 includes: S110, collecting a number of original images; S120, performing denoising processing and sharpening processing on each of the number of original images to obtain a number of target images.
4. The method for identifying personnel from an aerial perspective of an unmanned aerial vehicle according to claim 1, wherein The target yolo-v5 algorithm also sets a convolutional attention module, a channel attention module, and a spatial attention module.
5. The method for identifying a person from an aerial perspective of a drone according to claim 1, wherein XLTi is obtained through the following steps: S401, normalizing XTi to obtain XTGi; S402. Calculate the horizontal gradient and vertical gradient of each preset pixel of the image corresponding to XTGi; S403. Divide the horizontal gradient and vertical gradient of each preset pixel into several intervals respectively to obtain the gradient histogram of each preset pixel; S404. Obtain the distribution of the gradient magnitudes within each interval of the gradient histogram of each preset pixel; S405. Generate several multi-dimensional feature vectors based on the distribution of the gradient magnitudes within each interval of the gradient histogram of each preset pixel; S406. Concatenate the multi-dimensional feature vectors and perform normalization processing to obtain a feature descriptor; S407. Concatenate the feature descriptors corresponding to each preset pixel to obtain XTi.
6. The method for identifying personnel from an aerial perspective of an unmanned aerial vehicle according to claim 1, wherein The super-resolution processing is performed according to the bilinear interpolation method.
7. A non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by a processor to implement the method as described in any one of claims 1-6.
8. An electronic device, characterized in that, Comprising a processor and the non-transitory computer-readable storage medium as described in claim 7.
Citation Information
Patent Citations
Long-time unmanned aerial vehicle tracking and positioning method, system and device based on stereoscopic vision
CN111563916A
Power equipment bolt detection method and system based on unmanned aerial vehicle inspection
CN115546666A