Target positioning and identification method, device and readable storage medium
By using a real-time visual perception system that enables multiple drones to work together, and by processing and fusing facial feature vectors with edge devices, the system solves the problems of high cost, low accuracy, and poor real-time performance in drone-based person recognition, and achieves efficient and accurate person target localization and recognition.
Patent Information
- Application Number
- CN202310266610.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-13
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-03-13
AI Technical Summary
Existing drone-based person recognition technology suffers from high costs, low positioning accuracy, and poor real-time performance, especially in large-scale person recognition in open environments where efficiency and accuracy are insufficient.
A real-time visual perception system employing multi-drone collaborative operation acquires multiple facial image information through a swarm of drones, processes and extracts features using edge devices, determines three-dimensional coordinate points through multi-view cross-search using machine vision technology, and fuses facial feature vectors to achieve identity recognition.
It improves the positioning and recognition accuracy of human targets, reduces hardware costs, ensures real-time performance and system load balance, and enhances the efficiency of UAVs in tracking human targets in open environments.
Smart Images

Figure CN116469142B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless networks, in particular to a target positioning and identification method, device and computer readable storage medium. BACKGROUND
[0002] In recent years, the character recognition technology has been widely applied to improve public safety. Through the analysis of facial features, this kind of technology can effectively identify the identity information of the character. The traditional character recognition scheme mainly relies on the images captured by fixed position cameras, and the field of view of these cameras is limited, the tracking efficiency of moving targets is low, and large-scale character recognition cannot be performed in open environment crowd scenes.
[0003] The character recognition technology can be applied in unmanned aerial vehicle monitoring, thereby improving the ability of character positioning and identification. However, since the unmanned aerial vehicles currently applying the character positioning and identification technology need to be supported by high positioning function hardware, the orientation of a single unmanned aerial vehicle and a character greatly differs, and the limited on-board computing resources of the unmanned aerial vehicle and the large computing power required for character positioning and identification make the unmanned aerial vehicle overloaded, delayed and other factors, so the current use of unmanned aerial vehicles for character positioning and identification often has the technical problems of high cost, low character target positioning accuracy and poor character recognition accuracy and real-time performance. SUMMARY
[0004] The main purpose of the present application is to provide a target positioning and identification method, device and computer readable storage medium, which aims to solve the technical problems of high cost, low character target positioning accuracy and poor character recognition accuracy and real-time performance of the unmanned aerial vehicles currently applying the character recognition technology.
[0005] To achieve the above purpose, the present application provides a target positioning and identification method, which is applied to a target positioning and identification system; the target positioning and identification system comprises a group of unmanned aerial vehicles and a group of edge devices; the group of unmanned aerial vehicles and the group of edge devices are in wireless communication connection;
[0006] The target positioning and identification method comprises the following steps:
[0007] The group of unmanned aerial vehicles acquires a plurality of facial image information of a character target;
[0008] The target edge device that completes the time earliest in the group of edge devices is determined, and each facial image information is dispatched to the target edge device;
[0009] The target edge device converts the facial image coordinates in the facial image information into a three-dimensional coordinate point set, and determines the facial space position corresponding to the facial image coordinates according to the three-dimensional coordinate point set; and
[0010] extracting, by the target edge device, a face feature vector in each of the face image information, and fusing each of the face feature vector to determine identity information of the person target.
[0011] Optionally, the step of acquiring, by the UAV group, the plurality of face image information of the person target comprises:
[0012] collecting, by the UAV group, a plurality of scene images within a visual angle range at every preset shooting interval;
[0013] determining and acquiring, based on a preset target detection model, a human target image in the scene image;
[0014] acquiring, based on a preset face detection model, face image information in the human target image; the face image information represents a face image with a face detection frame and face image coordinates.
[0015] Optionally, the edge device group comprises a dispatching edge device and a non-dispatching edge device.
[0016] The step of determining the target edge device with the earliest completion time in the edge device group comprises:
[0017] sending, by the dispatching edge device, a test file package to each of the non-dispatching edge devices to determine a transmission time delay of the non-dispatching edge device at present; and
[0018] acquiring a task queue length, a central processing unit performance parameter and an image processor performance parameter of the non-dispatching edge device at present;
[0019] inputting the transmission time delay, the task queue length, the central processing unit performance parameter and the image processor performance parameter into a preset multivariate linear regression model to obtain a predicted completion time of the non-dispatching edge device;
[0020] comparing the predicted completion time of each of the non-dispatching edge devices to determine the target edge device with the earliest completion time in each of the non-dispatching edge devices.
[0021] Optionally, the step of converting, by the target edge device, the face image coordinates in the face image information into a three-dimensional coordinate point set comprises:
[0022] inputting, by the target edge device, the face image coordinates in the face image information and a preset height space limitation parameter into a preset two-dimensional-three-dimensional conversion model to convert the face image coordinates to obtain a three-dimensional coordinate point set with a plurality of three-dimensional coordinate points.
[0023] Optionally, the step of determining the face space position corresponding to the face image coordinate according to the set of three-dimensional coordinate points comprises:
[0024] traversing the three-dimensional coordinate points in the set of three-dimensional coordinate points, and projecting the traversed three-dimensional coordinate points to each face image information;
[0025] judging whether the traversed three-dimensional coordinate points are projected in the face detection frame in all face image information;
[0026] if the traversed three-dimensional coordinate points are in the face detection frame in all face image information, the traversed three-dimensional coordinate points are determined as the face space position corresponding to the face image coordinate.
[0027] Optionally, the face image coordinate comprises a face feature point; and the step of fusing each face feature vector to determine the identity information of the target person comprises:
[0028] inputting the face feature point and the resolution of each face image information into a preset fusion weight model to obtain a feature fusion weight of each face image information;
[0029] multiplying the face feature vector and the corresponding feature fusion weight to obtain a weighted face feature vector;
[0030] adding each weighted face feature vector to determine the identity information of the target person.
[0031] Optionally, the step of adding each weighted face feature vector to determine the identity information of the target person comprises:
[0032] adding each weighted face feature vector to obtain a fused face feature vector, and inputting the fused face feature vector into a support vector machine to determine the corresponding person category;
[0033] if the person category is a preset concerned person, outputting the space position and face image of the target person.
[0034] Optionally, the target positioning and recognition system further comprises a cloud server; and the edge device is in wireless communication connection with the cloud server.
[0035] After the step of determining the target edge device with the earliest completion time in the group of edge devices, the method further comprises:
[0036] If the predicted completion time of the target edge device is greater than the preset shooting interval, each of the face image information is sent to the cloud server; the cloud server is configured to determine the face space position corresponding to the face image coordinate; and determine the identity information of the target person.
[0037] In addition, to achieve the above object, the present application also provides a target positioning and identification system, comprising:
[0038] A UAV group is configured to obtain multiple face image information of a target person.
[0039] An edge device group is configured to determine a target edge device with the earliest completion time in the edge device group and dispatch each of the face image information to the target edge device.
[0040] The face image coordinate in the face image information is converted into a three-dimensional coordinate point set, and the face space position corresponding to the face image coordinate is determined according to the three-dimensional coordinate point set; and the face feature vector in each of the face image information is extracted, and each of the face feature vectors is fused to determine the identity information of the target person.
[0041] In addition, to achieve the above object, the present application also provides a target positioning and identification device, comprising a processor, a storage unit, and a target positioning and identification program stored in the storage unit and executable by the processor, wherein when the target positioning and identification program is executed by the processor, the steps of the target positioning and identification method are implemented.
[0042] The present application also provides a computer readable storage medium, wherein the target positioning and identification program is stored on the computer readable storage medium, and when the target positioning and identification program is executed by the processor, the steps of the target positioning and identification method are implemented.
[0043] The target positioning and identification method in the technical scheme of the present application first acquires multiple face image information of a character target through a UAV group, which can not only perform large-scale character identification in a crowd scene in an open environment and effectively improve the tracking efficiency of the character target based on the flexibility of the UAV, but also can acquire face image information of each character target at different angles, so as to improve the accuracy of face identification, and the UAV only transmits face image information, without the need to transmit all scene images within the angle of view, thereby reducing the data traffic of transmission, improving the efficiency of information transmission and ensuring the real-time performance of image processing. Then, the target edge device with the earliest completion time in the edge device group is determined, and each face image information is scheduled to the target edge device, so as to learn the load state and processing capacity of each edge device in the edge device group in real time, and select the target edge device with the earliest completion time at present, and the face image information acquired based on the UAV group at this time is determined as the current character target positioning and identification task scheduling to the target edge device, so that the task can be processed in time and quickly, thereby improving the efficiency of character target positioning and identification and greatly reducing the time delay of character target positioning and identification. Finally, the target edge device converts the face image coordinates in the face image information into a three-dimensional coordinate point set, and determines the face space position corresponding to the face image coordinates according to the three-dimensional coordinate point set, so that the UAV can determine various possible three-dimensional coordinate points, that is, various possible face space coordinates based on the two-dimensional face image coordinates acquired by the ordinary camera, and then determine the actual face space coordinates of the character target based on the three-dimensional coordinate point set, the face image information at different angles of multiple UAVs and machine vision technology, so that the face space coordinates of the character target are accurate and the hardware cost of the UAV positioning is greatly reduced. In addition, the target edge device extracts face feature vectors in each face image information and fuses the face feature vectors to determine the identity information of the character target, so that the face feature vectors in the face image information with different face angles are fused to obtain a relatively complete face feature vector of the character target, realize comprehensive analysis of the face feature vector, and greatly improve the recognition accuracy of the identity information of the character target. Moreover, the way of character target positioning and fusion identification in the present application is performed in the target edge device with the earliest completion time, so that the load of the whole system is more balanced, and the efficiency and timeliness of character positioning and identification are more excellent and stable. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 The structure diagram of the hardware running environment of the target positioning and identification device involved in the embodiment scheme of the present application;
[0045] Figure 2 This is a flowchart illustrating the first embodiment of the target localization and identification method of the present invention;
[0046] Figure 3 This is a detailed flowchart of step S10 in an embodiment of the target localization and identification method of the present invention;
[0047] Figure 4 This is a detailed flowchart of step S20 in an embodiment of the target localization and identification method of the present invention;
[0048] Figure 5 This is a schematic diagram illustrating the application scenarios involved in the target localization and identification method of the present invention;
[0049] Figure 6 This is a system block diagram of the heterogeneous device dynamic task scheduling involved in the target localization and identification method of the present invention;
[0050] Figure 7 This is a flowchart of the multi-view cross-search algorithm involved in the target localization and identification method of the present invention;
[0051] Figure 8 This is a schematic diagram of the fusion weight network structure and training process involved in the target localization and recognition method of the present invention;
[0052] Figure 9 This is a schematic diagram of the scenario corresponding to the dynamic scheduling and cross-location involved in the target localization and identification method of the present invention;
[0053] Figure 10 This is a schematic diagram of a multi-view fusion face recognition scenario involved in the target localization and recognition method of the present invention.
[0054] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0055] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0056] This invention provides a target localization and identification device. The device includes multiple drones and multiple edge devices with data processing and computing capabilities.
[0057] like Figure 1 As shown, Figure 1 This is a schematic diagram of the hardware operating environment of the target positioning and identification device involved in the embodiments of the present invention.
[0058] like Figure 1As shown, the target positioning and identification device can include a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a storage unit 1005, and a communication bus 1002. The communication bus 1002 is used to realize the connection communication between the components. The user interface 1003 can include a display (Display), an input unit such as a control panel. The optional user interface 1003 can also include a standard wired interface, a wireless interface. The network interface 1004 can optionally include a standard wired interface, a wireless interface (such as a WIFI interface). The storage unit 1005 can be a high-speed RAM storage unit, or a stable storage unit (non-volatile memory) such as a magnetic disk storage unit. The storage unit 1005 can also be a storage device independent of the aforementioned processor 1001. The storage unit 1005, as a computer storage medium, can include a target positioning and identification program.
[0059] Those skilled in the art can understand that Figure 1 The hardware structure shown in the foregoing embodiments does not constitute a limitation on the device, and can include more or fewer components than those shown, or combine certain components, or different component arrangements.
[0060] Continuing to refer to Figure 1 , Figure 1 The storage unit 1005 in the foregoing embodiments, as a computer readable storage medium, can include an operating system, a user interface module, a network communication module, and a target positioning and identification program.
[0061] In the foregoing embodiments, Figure 1 In the foregoing embodiments, the network communication module is mainly used to connect the server and communicate data with the server; and the processor 1001 can call the target positioning and identification program stored in the storage unit 1005 and perform the following operations:
[0062] Obtain multiple face image information of the character target through the UAV group;
[0063] Determine a target edge device with the earliest completion time in the edge device group and schedule each face image information to the target edge device;
[0064] Convert a face image coordinate in the face image information into a three-dimensional coordinate point set through the target edge device, and determine a face space position corresponding to the face image coordinate according to the three-dimensional coordinate point set; and
[0065] Extract a face feature vector in each face image information through the target edge device, and fuse each face feature vector to determine identity information of the character target.
[0066] Further, the processor 1001 can call a target positioning and identification program stored in the memory 1005, and further perform the following operations:
[0067] Collect multiple scene images within a visual angle range through the drone group every preset shooting interval length;
[0068] Determine and obtain a human target image in the scene image based on a preset target detection model;
[0069] Obtain face image information in the human target image based on a preset face detection model; the face image information represents a face image with a face detection frame and face image coordinates.
[0070] Further, the processor 1001 can call a target positioning and identification program stored in the memory 1005, and further perform the following operations:
[0071] Send a test file package to each non-scheduled edge device through the scheduling edge device to determine the current transmission delay of the non-scheduled edge device; and
[0072] Obtain the current task queue length, central processor performance parameter and image processor performance parameter of the non-scheduled edge device;
[0073] Input the transmission delay, task queue length, central processor performance parameter and image processor performance parameter into a preset multivariate linear regression model to obtain a predicted completion time of the non-scheduled edge device;
[0074] Compare the predicted completion times of each non-scheduled edge device to determine a target edge device with the earliest completion time among the non-scheduled edge devices.
[0075] Further, the processor 1001 can call a target positioning and identification program stored in the memory 1005, and further perform the following operations:
[0076] Input the face image coordinates in the face image information and a preset height space limitation parameter into a preset two-dimensional-three-dimensional conversion model through the target edge device to convert the face image coordinates to obtain a three-dimensional coordinate point set with multiple three-dimensional coordinate points.
[0077] Further, the processor 1001 can call a target positioning and identification program stored in the memory 1005, and further perform the following operations:
[0078] Iterate through the three-dimensional coordinate points in the three-dimensional coordinate point set, and project the iterated three-dimensional coordinate points to each face image information;
[0079] determining whether the traversed three-dimensional coordinate point is projected in the face detection frame of the face image information;
[0080] If the traversed three-dimensional coordinate point is projected in the face detection frame of the face image information, the traversed three-dimensional coordinate point is determined as the face space position corresponding to the face image coordinate.
[0081] Further, the processor 1001 can call the target positioning and identification program stored in the memory 1005, and further perform the following operations:
[0082] The face feature vector of each face image information is multiplied by the corresponding feature fusion weight to obtain a weighted face feature vector.
[0083] The face feature vector of each face image information is multiplied by the corresponding feature fusion weight to obtain a weighted face feature vector.
[0084] The face feature vector of each face image information is multiplied by the corresponding feature fusion weight to obtain a weighted face feature vector.
[0085] Further, the processor 1001 can call the target positioning and identification program stored in the memory 1005, and further perform the following operations:
[0086] The face feature vector of each face image information is multiplied by the corresponding feature fusion weight to obtain a weighted face feature vector.
[0087] If the target edge device is predicted to complete the time greater than the preset shooting interval length, each face image information is sent to the cloud server; the cloud server is used to determine the face space position corresponding to the face image coordinate; and determine the identity information of the target object.
[0088] Further, the processor 1001 can call the target positioning and identification program stored in the memory 1005, and further perform the following operations:
[0089] If the target edge device is predicted to complete the time greater than the preset shooting interval length, each face image information is sent to the cloud server; the cloud server is used to determine the face space position corresponding to the face image coordinate; and determine the identity information of the target object.
[0090] In order to facilitate understanding of each of the following embodiments in the present application, a series of main challenges and corresponding technical solutions for positioning and identifying the target object by using the unmanned aerial vehicle in the present application are briefly described:
[0091] In order to perform large-scale person recognition in open environment crowd scenes, a UAV can be used to perform person recognition and target localization tracking. The human body recognition and tracking solution based on UAV benefits from its wide field of view and high mobility, and can be applied to various application scenarios such as military operations and security services. However, there are great challenges in using a UAV for human body recognition and tracking, mainly as follows:
[0092] 1) Poor accuracy and real-time performance of person recognition
[0093] Currently, the face recognition solution based on deep neural network has achieved high accuracy, but the prerequisite is that there are sufficient face pixels in the image. However, due to the high flight height of a single UAV and different angles between the UAV and the person, the size of the face in the picture is often small, and the deflection angle of the face towards the UAV camera is too large, which makes it difficult to recognize. At the same time, the face recognition technology based on deep neural network consumes a large amount of computing resources. The on-board computing resources of the UAV are limited, and deploying the face recognition technology based on deep neural network on the UAV will cause the UAV to be overloaded, which brings high delay to the system, and is not ideal for real-time recognition.
[0094] 2) Insufficient accuracy and working range of target localization
[0095] In order to track the target person in the crowd, most advanced positioning technologies require a depth camera or a laser radar on the UAV, which is about 10 times more expensive than a traditional camera. In addition, in outdoor and long-distance scenes, the positioning accuracy and point cloud density of the depth camera and the laser radar decrease sharply, resulting in poor positioning accuracy and small working area.
[0096] In order to solve the above technical problems and the challenges required to be addressed, the main technical solution of the present application is:
[0097] This invention designs a real-time visual perception system for multi-drone collaborative operation, namely the target localization and recognition system mentioned in this invention, which can accurately and quickly locate and identify human targets in crowded places. This invention uses a drone platform with onboard computing capabilities and deploys a lightweight neural network model on the drones to detect people in a scene from the perspectives of different drones. Simultaneously, through analysis at the drone end, effective facial data is extracted and offloaded to edge devices to reduce data transmission traffic. To achieve person localization, this invention utilizes images from different perspectives of multiple drones and uses machine vision technology to calculate the person's three-dimensional position by cross-searching from multiple views. To achieve high-precision recognition, this invention acquires facial data of people from different perspectives of multiple drones, uses neural network technology to evaluate their recognizability, and performs fusion recognition based on their recognizability. To achieve load balancing and end-to-end latency optimization in multi-UAV systems, this invention uses a lightweight machine learning algorithm to dynamically evaluate the estimated completion time of tasks for each edge device in the system. Based on the Earliest Finish Time (EFT), the optimal completion device (target edge device) is selected, and data is offloaded to the optimal completion device for execution. In addition, when edge devices require a long data processing time, data can be offloaded to a cloud server. This enables the scheduling and balancing of workloads between edge devices and between edge devices and cloud servers, minimizing processing latency, improving the efficiency of target localization and recognition, and ensuring the real-time performance of target localization and recognition.
[0098] Furthermore, to facilitate understanding of the application scenarios involved in this invention, please refer to... Figure 5 , Figure 5 This is a schematic diagram illustrating the application scenarios involved in the target localization and identification method of the present invention. For example... Figure 5 As shown, the drone swarm includes three drones: drone 1, drone 2, and drone 3. Each drone collects facial image data (information) within a certain area in the air. After determining the current optimal device from the edge computing device, the drones are guided by scheduling decision information to unload the collected facial image data to the current optimal device. In this way, the current optimal device processes the task of locating and recognizing the target person to determine the spatial location and identity information of the target person.
[0099] This invention provides a target localization and identification method.
[0100] Please refer to Figure 2 , Figure 2A flowchart of a target positioning and identification method according to a first embodiment of the present application; in the first embodiment of the present application, the target positioning and identification method is applied to a target positioning and identification system; the target positioning and identification system comprises a group of unmanned aerial vehicles and a group of edge devices; the group of unmanned aerial vehicles and the group of edge devices are in wireless communication connection;
[0101] The target positioning and identification method comprises the following steps:
[0102] In step S10, the group of unmanned aerial vehicles acquires multiple face image information of the person target;
[0103] In the present embodiment, the target positioning and identification system is a comprehensive system with person target positioning and identification functions, which is composed of multiple unmanned aerial vehicles (group of unmanned aerial vehicles) and multiple edge devices (group of edge devices). The group of unmanned aerial vehicles is used to take pictures of the scene within the visual angle range from the air, and to acquire face image information of the person target from the scene; the edge device refers to a computing device or a data processing device within a certain range of the group of unmanned aerial vehicles, which is mainly used for scheduling of positioning and identification tasks, and for processing of positioning and identification tasks of the person target. The group of unmanned aerial vehicles and the group of edge devices can communicate through wireless communication technologies such as wifi local area network, ZigBee, 4G, 5G, etc. In order to improve the communication efficiency between the two, the spatial distance between the group of unmanned aerial vehicles and the group of edge devices should not be too far, which can be adjusted according to actual needs.
[0104] As for the group of unmanned aerial vehicles, it is composed of multiple unmanned aerial vehicles. The number of unmanned aerial vehicles in the group of unmanned aerial vehicles can be set according to actual needs. The current unmanned aerial vehicles are all equipped with cameras. In the present embodiment, a common camera can be selected, i.e. a camera with basic video recording and shooting functions, which does not need to have a ranging function or a special positioning device such as a laser radar, thereby reducing the cost of the group of unmanned aerial vehicles.
[0105] During the flight of the group of unmanned aerial vehicles, each unmanned aerial vehicle can acquire face image information of the person target in the scene within its own visual angle range, so that the group of unmanned aerial vehicles acquires multiple visual angle face image information. In these face image information, not only the face image information of each person target in the scene, but also the face image information of the same person target from different visual angles.
[0106] Specifically, the group of unmanned aerial vehicles first acquires scene images, i.e. acquires and takes scene images including person targets and other objects within their respective visual angle ranges, and then identifies and marks the person targets from the scene images based on a lightweight neural network, and further extracts face image information from the scene images.
[0107] Please refer to Figure 3In an embodiment, the step S10 comprises:
[0108] At step S11, the UAV group collects multiple scene images within the visual angle range at a preset shooting interval.
[0109] The preset shooting interval, i.e., the sampling frequency of the UAV collecting scene images, can be set according to actual needs, such as 1s, i.e., every 1s, the UAV in the UAV group collects scene images within the respective visual angle range, so that the UAV group obtains multiple scene images.
[0110] At step S12, a human target image in the scene image is determined and obtained based on a preset target detection model.
[0111] At step S13, a face image information in the human target image is obtained based on a preset face detection model; the face image information represents a face image with a face detection frame and face image coordinates.
[0112] In this embodiment, the UAV has on-board computing capability, and the scene images captured by the UAV are processed on the on-board side by the target detection model and the face detection model. Moreover, the above image recognition models carried by the UAV are all lightweight neural networks, thereby reducing the load of the UAV and using more computing power for other work, preventing the UAV from causing losses due to flight abnormalities caused by excessive load.
[0113] Specifically, the lightweight target detection algorithm YoloX-tiny and the face detection algorithm RetinaFace are deployed on the on-board side of the UAV. Compared with the traditional Yolo algorithm, YoloX-tiny removes the Anchor mechanism, reducing the complexity and parameter amount of the model. At the same time, YoloX-tiny is adjusted and trained on the VisDrone UAV target detection dataset, and data enhancement methods such as random scaling, cropping, arrangement (Mosaic), and multi-image mixing (Mix Up) are used to strengthen the training of the network to better detect the person in the UAV visual angle. The above two algorithm models will extract the face image, face feature coordinates, and boundary box position (face detection frame) containing the face in the scene image.
[0114] After obtaining the scene image, the light-weight target detection model is used to identify and determine which are the human image from the scene image, and the human target image (human image) is extracted from the scene image. Then, the light-weight face detection model is used to obtain the face image information in the human target image, and the face detection frame is labeled on the face image during the identification of the human target image by the face detection model, and the image coordinates of the face image in the entire scene image are determined, so that the face image information with the face detection frame and the face image coordinates is obtained.
[0115] The embodiment uses a light-weight image recognition neural network, and the unmanned aerial vehicle only needs to perform the task of extracting the face image from the scene in parallel, thereby reducing the running load of the unmanned aerial vehicle, ensuring the normal operation of other necessary modules such as shooting and flight of the unmanned aerial vehicle, and thus the unmanned aerial vehicle does not need to carry heavy and complex on-board devices and cooling devices, thereby reducing the flight pressure and energy consumption. Furthermore, the unmanned aerial vehicle sends the extracted face image information, which is much smaller than the scene image, to the edge device, thereby ensuring the real-time performance of data exchange and reducing the time delay.
[0116] In step S20, the target edge device with the earliest completion time in the edge device group is determined, and each face image information is scheduled to the target edge device.
[0117] The edge device group needs to specify at least one edge device as a scheduling edge device for scheduling the current positioning and identification task. The scheduling edge device can be selected from any one or several edge devices in the edge device group, or can be selected according to the performance of the edge device, which is not limited here.
[0118] The scheduling edge device performs performance evaluation and prediction of the time required to complete the current task on other non-scheduling edge devices to determine the target edge device with the earliest completion time in the edge device group and schedule each face image information to the target edge device.
[0119] Specifically, after determining the target edge device, the scheduling edge device can send the related scheduling decision information to each unmanned aerial vehicle, thereby commanding the unmanned aerial vehicle group to unload all the face image information to the target edge device, and completing the current human target positioning and identification task through the target edge device with the earliest completion time. Thus, the processing efficiency of the human target positioning and identification task is improved, and real-time positioning and identification can be ensured.
[0120] In an embodiment, the edge device group includes a scheduling edge device and a non-scheduling edge device.
[0121] Please refer to Figure 4The step S20 of determining the target edge device with the earliest completion time in the edge device group comprises:
[0122] The step S21 of sending a test file package to each non-scheduled edge device by the scheduled edge device to determine the current transmission delay of the non-scheduled edge device; and
[0123] After the scheduled edge device is selected, the other edge devices are defined as non-scheduled edge devices.
[0124] The state collection and scheduling of the other non-scheduled edge devices are performed by the scheduled edge device. Specifically, the shooting interval of the scene image by the UAV group can be determined by sending a test file package to each non-scheduled edge device, receiving the state feedback data of the non-scheduled edge device to obtain the maximum transmission delay Delay j of the non-scheduled edge device at the current time.
[0125] The step S22 of obtaining the current task queue length, the central processor performance parameter and the image processor performance parameter of the non-scheduled edge device;
[0126] The state feedback data of the non-scheduled edge device is received to obtain the current task queue length Queue j , the CPU and GPU performance parameters CPU j , GPU j of the non-scheduled edge device.
[0127] The step S23 of inputting the transmission delay, the task queue length, the central processor performance parameter and the image processor performance parameter into a preset multivariate linear regression model to obtain the predicted completion time of the non-scheduled edge device;
[0128] The step S24 of comparing the predicted completion time of each non-scheduled edge device to determine the target edge device with the earliest completion time in each non-scheduled edge device.
[0129] The above maximum transmission delay Delay j , the current task queue length Queue j , the CPU and GPU performance parameters CPU j , GPU jThe input is input into a preset multivariable linear regression (MLR) model in the scheduling edge device, and the model outputs the earliest completion time of the task of each non-scheduling edge device. After the UAV completes a shooting, the scheduling edge device compares the predicted completion times of the non-scheduling edge devices, and arranges the subsequent positioning and recognition tasks after the shooting to the non-scheduling edge device with the shortest predicted completion time, that is, the target edge device corresponding to the earliest completion time. It should be noted that the scheduling edge device can also participate in the target positioning and recognition task, so that the scheduling edge device can also estimate its predicted completion time, so as to comprehensively compare all edge devices to determine the earliest completion time, that is, the current optimal edge device, and send the current task to the optimal device.
[0130] For the multivariable linear regression model in this embodiment, the corresponding mathematical expression can be:
[0131] EFT j =MLR(Delay j ,||Queue j ||,CPU j ,GPU j )=θ0+θ1Delay j +
[0132] θ2Queue j +θ3GPU j +θ4CPU j );
[0133] Wherein θ0, θ1, θ2, θ3, θ4 are updateable parameters of the multivariable linear regression; Delay j is the maximum transmission delay of the jth edge device; Queue j is the number of tasks queued in the jth edge device, that is, the task queue length; CPU j , GPU j respectively represent the CPU performance parameter and the GPU performance parameter of the jth edge device.
[0134] The multivariable linear regression model in this embodiment is a trainable and updateable model. In order to train the multivariable linear regression model, a trajectory pool can be maintained to store real historical records. Each trajectory is a pair of parameters (Delay j ,||Queue j ||,CPU j , GPU j ) and real completion time. The mean square error loss can be used to update the multivariable linear regression model in real time during system operation, so as to realize accurate prediction of the predicted completion time.
[0135] For the purpose of understanding the flow of this embodiment of the present application, please refer to Figure 6 , Figure 6 The system block diagram of the heterogeneous device dynamic task scheduling method involved in the target positioning and identification method of the present application. As shown in Figure 6 , the edge device group includes edge device 1, edge device j... edge device M, with M edge devices, the UAV group includes UAV 1, UAV i... UAV N, with N edge devices, in addition to the cloud device, which is not described here. Select any edge device as the scheduling edge device, collect the device state of other edge devices, and also collect the device state of the UAV, such as the shooting state and flight state of the UAV. Further, based on the device state information of the edge device, perform multivariate linear regression, and the trajectory pool stores the real historical records for the multivariate linear regression model. In the process of multivariate linear regression, the device selector selects the target edge device with the earliest completion time, and can generate a distribution table to number and sort the current tasks, such as task 1 corresponding to edge device 1.
[0136] In this embodiment of the present application, the target edge device with the earliest completion time at the current time is determined, which can greatly reduce the time delay of processing the target positioning and identification task, and balance the load between each edge device. In addition to improving the processing efficiency of the target positioning and identification task, it also ensures that each task can be carried out in an orderly manner, improving the stability and reliability of the operation of the edge device group.
[0137] Step S30, converting the face image coordinates in the face image information into a set of three-dimensional coordinate points by the target edge device, and determining the face space position corresponding to the face image coordinates according to the set of three-dimensional coordinate points; and
[0138] The face image information includes face image coordinates, which can be face feature coordinates. One of the face feature coordinates can be converted from two-dimensional coordinates to three-dimensional coordinates. The target edge device, as an execution device for processing the current human target positioning and recognition task, can first convert the two-dimensional face image coordinates to two-dimensional space coordinates of the human target through a preset image-space coordinate matrix, and then predict multiple possible three-dimensional coordinates of the human target, because the three-dimensional coordinates are only missing the depth coordinates (commonly represented by z-axis coordinates) compared with the two-dimensional coordinates. In this way, a set of three-dimensional coordinates can be obtained, and the real three-dimensional space coordinates of the human target are included in the set of three-dimensional coordinates. Then, to determine the spatial position of the human target, the multiple-view face images of the human target can be aligned to determine a unique three-dimensional coordinate, which is the face three-dimensional space coordinates, that is, the face space position, which is actually the spatial position of the human target.
[0139] In an embodiment, the step S30 of converting the face image coordinates in the face image information into a set of three-dimensional coordinates by the target edge device includes:
[0140] Step a: inputting the face image coordinates in the face image information and a preset height space limitation parameter into a preset two-dimensional-three-dimensional conversion model by the target edge device to convert the face image coordinates into a set of three-dimensional coordinates with multiple three-dimensional coordinates.
[0141] In this embodiment, the preset two-dimensional-three-dimensional conversion model is used to convert two-dimensional face image coordinates into three-dimensional face space coordinates.
[0142] For example, one of the drones D i is used to illustrate the definition of the three-dimensional coordinate position P W of a human in the real world as [x W , y W , z W ] T , and the coordinates (face image coordinates) of the human in the image taken by the drone D i are P i = [x i , y i ] T . According to the coordinate system conversion rule in machine vision, the mapping relationship between the three-dimensional point P W and the two-dimensional point P i in the image is:
[0143]
[0144] where C iD i represents the camera internal matrix of the UAV D i , which can establish the mapping relationship between the image pixel coordinate system and the camera coordinate system of the UAV D i i represents the rotation matrix and translation matrix of the UAV D i , which can establish the mapping relationship between the camera coordinate system and the world coordinate system of the UAV D i i i can be obtained through PNP visual positioning, and C i can be obtained through camera calibration technology. According to formula (1), the formula for obtaining three-dimensional coordinates from two-dimensional image coordinates can be derived:
[0145]
[0146] The formula (2) is the mathematical expression of the two-dimensional-three-dimensional conversion model. At this time, due to the lack of depth information, the three-dimensional space coordinates of the face cannot be directly determined from the two-dimensional coordinates of the face image, but a set of possible three-dimensional coordinate points, called epipolar lines, are obtained.
[0147] In order to reduce the processing and operation amount of data by the target edge device, the preset height space restriction parameters: the lowest height h min , the highest height h max , can be input at the same time as the face image coordinates P i , so as to narrow the range of the three-dimensional coordinate point set by using the height space restriction, and exclude some three-dimensional coordinate points that are not possible for the human target. The height space restriction parameter here is related to the spatial height of the human body (mainly the height), such as the height space restriction of 0-3m, or other height space restriction parameters can be set, which are not limited here.
[0148] In an embodiment, step S30, the step of determining the face space position corresponding to the face image coordinates according to the set of three-dimensional coordinate points, comprises:
[0149] Step b, traversing the three-dimensional coordinate points in the set of three-dimensional coordinate points, and projecting the traversed three-dimensional coordinate points into each face image information;
[0150] For this embodiment, it can be considered as a multi-view cross search algorithm. For one of the UAVs, after obtaining the set of three-dimensional coordinate points, the three-dimensional coordinate points in the set of three-dimensional coordinate points can be traversed from the first three-dimensional coordinate point, and the three-dimensional coordinate points can be projected into the view of all UAVs, that is, into all face images.
[0151] Step c, judging whether the traversed three-dimensional coordinate point is projected in the face detection frame of all face image information;
[0152] Step d, if the traversed three-dimensional coordinate point is in the face detection frame of all face image information, the traversed three-dimensional coordinate point is determined as the face space position corresponding to the face image coordinate.
[0153] Judging whether the three-dimensional coordinate point is projected in the face detection frame of all face images.
[0154] If the traversed three-dimensional coordinate point is projected in the face detection frame of all face image information, it means that the three-dimensional coordinate point is the three-dimensional space coordinate of the target, and at this time the face space position corresponding to the face image coordinate is determined, and the traversal can be stopped.
[0155] If the traversed three-dimensional coordinate point is not projected in the face detection frame of all face image information, it may be projected in part of the face detection frame or not projected in any face detection frame, and the next three-dimensional coordinate point needs to be traversed until the current traversed three-dimensional coordinate point is projected in the face detection frame of all face image information, so as to determine the space position of the target.
[0156] In order to further understand the multi-view cross search algorithm proposed in this embodiment, please refer to Figure 7 , Figure 7 The multi-view cross search algorithm flowchart involved in the target positioning and recognition method of the present application. The process of determining the face space position corresponding to the face image coordinate according to the three-dimensional coordinate point set can be:
[0157] All unmanned aerial vehicles perform face detection;
[0158] Select an arbitrary main-view unmanned aerial vehicle;
[0159] Obtain the epipolar line according to the detection result (face image coordinate);
[0160] Search the depth range according to the space limitation;
[0161] Traverse the three-dimensional coordinate points on the epipolar line;
[0162] Project to other unmanned aerial vehicle views;
[0163] Judge whether the projection of the point falls in the face detection frame of all unmanned aerial vehicles;
[0164] If yes, exit the traversal and use the current point as the position estimation point of the target; output the positioning result;
[0165] If not, continue to traverse the next point.
[0166] By the above embodiment of the present application, the different perspective images of multiple UAVs are used to calculate the three-dimensional position of a person by cross-searching from multiple views using machine vision technology, which realizes accurate positioning of a person in a three-dimensional real world using only two-dimensional images taken by a traditional camera, and aligns the facial sub-images of the same person from different UAV perspectives. This embodiment allows multiple UAVs to form a unified perception of a real scene.
[0167] In step S40, the target edge device extracts a face feature vector from each of the face image information and fuses each of the face feature vectors to determine the identity information of the target person.
[0168] After aligning the multi-perspective face images of the target person, the target edge device runs a multi-view fusion recognition algorithm module, which is provided with a fusion weight network (FWN). Based on image recognition technology (such as the ArcFace face recognition algorithm), the face feature vector (which can be a 1x512-dimensional vector) in each perspective face image is extracted, and in order to comprehensively utilize these face features at different angles and improve the accuracy of face recognition, the fusion weight network assigns a fusion weight to each face image to reflect its legibility. The legibility here is related to the perspective and resolution. The closer the perspective is to the front of the face, the higher the resolution, the better the legibility, and the higher the fusion weight.
[0169] Then, the fusion layer in the fusion weight network fuses the face feature vectors at different perspectives using the fusion weight values to generate a fused face feature vector. The fused face feature contains more information than the face feature extracted from a single image. Finally, the features are classified by a support vector machine algorithm (SVM). The support vector machine algorithm is trained on a database containing the target person's feature vector, and in the testing process, the fused face features are classified. The target person classified as the target person is considered to be the target person to be found, and the three-dimensional position and face image are obtained for further operations such as real-time outdoor personnel search, monitoring, and broadcasting of the face image.
[0170] In an embodiment, the face image coordinates include face feature coordinate points;
[0171] In step S40, the step of fusing each of the face feature vectors to determine the identity information of the target person includes:
[0172] In step e, the face feature coordinate points and resolution of each of the face image information are input into a preset fusion weight model to obtain the feature fusion weight of each of the face image information.
[0173] Step f, the face feature vector is multiplied by the corresponding feature fusion weight to obtain a weighted face feature vector;
[0174] Step g, adding each of the weighted face feature vectors to determine the identity information of the target person.
[0175] As Figure 8 shown, Figure 8 the fusion weight network structure and training process involved in the target positioning and identification method of the present application.
[0176] Please refer to Figure 8 the upper half of the figure, the input of the fusion weight network (weight fusion network) is the facial feature point and the resolution of the face image, and the output is a fusion weight for fusion. The neural network structure of the fusion weight network is a simple fully connected layer, which has 5 layers with dimensions of 180, 360, 180, 2 and 1 respectively. The above network dimension setting is derived from the angle range of the face orientation unmanned aerial vehicle, from 0 degrees to 360 degrees, so that the angle-related high-dimensional features can be extracted from the facial feature points and mapped into a 1-dimensional fusion weight, which reflects the easy recognition of the face angle.
[0177] After obtaining the feature fusion weight of each face image information, the original face feature vector is multiplied by the corresponding feature fusion weight to obtain a weighted face feature vector, and then each weighted face feature vector is added to obtain a fused face feature vector. Further, based on the fused face feature vector, the identity information of the target person is determined. Specifically, the fusion layer normalizes the fusion weight of each angle feature vector, multiplies it with the corresponding angle face feature vector, and adds all the multiplied face feature vectors to obtain a fused face feature vector. The fused face feature vector contains more details than the face feature vector extracted from a single image.
[0178] Classification is performed by support vector machine (SVM), the fused face feature is input into the support vector machine, and the support vector machine outputs the class of the feature, which is classified as a target person of a preset attention person, which is considered to be the target person to be found. The preset attention person or target person here can be set according to the actual application requirements. When the person class is identified as a preset attention person, the spatial position of the target person and its face image can be output to the designated system, so as to facilitate manual confirmation, tracking, search, search and other subsequent activities of the person.
[0179] That is, the step of adding each of the weighted face feature vectors to determine the identity information of the target person.
[0180] adding the weighted face feature vectors to obtain a fused face feature vector, and inputting the fused face feature vector into a support vector machine to determine a corresponding person category;
[0181] If the person category is a preset attention person, outputting the spatial position and face image of the person target.
[0182] Further, please refer to the lower half of Figure 8 For the fusion weight network, the network structure and the training process are shown in the figure. The fusion weight network inputs the facial feature points and the resolution of the face image, and outputs a fusion weight for fusion. A simple fully connected layer is used as the network structure of the fusion weight network to realize lightweight and fast calculation. In the training process of the fusion weight network, the ArcFace face recognition algorithm is used to extract features from the benchmark face image and the training set face image, and then the two features are compared to obtain the true weight of the training face image. The weight can represent the deviation of the training data from the benchmark face image (i.e. reflect its legibility). Then the weight output by the fusion weight network is calculated with the mean square loss function to update the fusion weight network in reverse.
[0183] In addition, due to the lightweight and fast characteristics of the fusion weight network, it is deployed on the processor (CPU) for inference, and the ArcFace face recognition algorithm is deployed on the graphics card (GPU) for inference, and the parallel work of the two is realized. At the same time, a thread pool is maintained, and a thread is created for each view of the face image, so as to realize multi-thread parallel processing, so as to further improve the efficiency of face image fusion recognition.
[0184] This embodiment of the present application uses face images from different unmanned aerial vehicle perspectives, fuses them together according to the weight reflecting legibility, and also through parallel computing, not only reduces the processing delay, but more importantly, compared with single-view face images, through multi-view face images and feature fusion of face images according to certain weight rules, the accuracy of person target recognition can be greatly improved.
[0185] In addition, in an embodiment, the target positioning and recognition system further comprises a cloud server; the edge device is in wireless communication connection with the cloud server;
[0186] After the step S20 of determining the target edge device with the earliest completion time in the edge device group, the method further comprises:
[0187] If the predicted completion time of the target edge device is greater than the preset photographing interval duration, each of the face image information is sent to the cloud server; the cloud server is configured to determine the face space position corresponding to the face image coordinate; and determine the identity information of the target person.
[0188] If the predicted completion time of the target edge device is greater than the preset photographing interval duration, that is, the earliest completion time is greater than the preset photographing interval duration, the predicted completion time of all the edge devices is greater than the preset photographing interval duration, which means that the current load of the edge device group is high. In order to process the face image data transmitted by the unmanned aerial vehicle in time, in this case, the unmanned aerial vehicle group can be guided to unload the face image information to the remote cloud server through the scheduling of the edge device, so that the cloud server is used as the target edge device in the above embodiments, and the same steps are performed to determine the space position and identity information of the target person.
[0189] The target positioning and identification method in the technical scheme of the present application first acquires multiple face image information of a character target through a UAV group, which can not only perform large-scale character identification in a crowd scene in an open environment and effectively improve the tracking efficiency of the character target based on the flexibility of the UAV, but also can acquire face image information of each character target at different angles, so as to improve the accuracy of face identification, and the UAV only transmits face image information, without the need to transmit all scene images within the angle of view, thereby reducing the data traffic of transmission, improving the efficiency of information transmission and ensuring the real-time performance of image processing. Then, the target edge device with the earliest completion time in the edge device group is determined, and each face image information is scheduled to the target edge device, so as to learn the load state and processing capacity of each edge device in the edge device group in real time, and select the target edge device with the earliest completion time at present, and the face image information acquired based on the UAV group at this time is determined as the current character target positioning and identification task scheduling to the target edge device, so that the task can be processed in time and quickly, thereby improving the efficiency of character target positioning and identification and greatly reducing the time delay of character target positioning and identification. Finally, the target edge device converts the face image coordinates in the face image information into a three-dimensional coordinate point set, and determines the face space position corresponding to the face image coordinates according to the three-dimensional coordinate point set, so that the UAV can determine various possible three-dimensional coordinate points, that is, various possible face space coordinates based on the two-dimensional face image coordinates acquired by the ordinary camera, and then determine the actual face space coordinates of the character target based on the three-dimensional coordinate point set, the face image information at different angles of multiple UAVs and machine vision technology, so that the face space coordinates of the character target are accurate and the hardware cost of the UAV positioning is greatly reduced. In addition, the target edge device extracts face feature vectors in each face image information and fuses the face feature vectors to determine the identity information of the character target. The face feature vectors in the face image information with different face angles are fused to obtain a relatively complete face feature vector of the character target, realize comprehensive analysis of the face feature vector, and greatly improve the recognition accuracy of the identity information of the character target. In addition, the way of character target positioning and fusion identification in the present application is performed in the target edge device with the earliest completion time, so that the load of the whole system is more balanced, and the efficiency and timeliness of character positioning and identification are more excellent and stable.
[0190] In order to further strengthen the understanding of each embodiment of the present application, the related description and supplement are made in combination with each actual application scene in the present application. Please refer to Figure 9 and Figure 10. Figure 9 The scene diagram corresponding to the dynamic scheduling and cross positioning involved in the target positioning and identification method of the present application; Figure 10 The scene diagram of multi-view fusion face recognition involved in the target positioning and identification method of the present application.
[0191] As Figure 9 shown, Figure 9 The lower half of the figure is called dynamic scheduling of heterogeneous devices (DTSH), which involves the process of selecting the current optimal edge device of the present application, that is, the process of determining the target edge device. The system includes edge devices and cloud devices (cloud servers). After the edge device obtains the device state, it is input into the multivariate linear regression to output the predicted completion time (predicted completion time). The target device is determined by the device selector. The target device here can be a target edge device or a cloud server.
[0192] Figure 9 The upper half of the figure is called multi-drone character cross positioning (MDPL). The drone 1, the drone 2, and the drone 3 all detect and obtain the faces of three character targets (character 1, character 2, and character 3). Through the target device, multi-view cross positioning is performed, and the face images of the same character target are aligned, thereby determining the spatial coordinates corresponding to the face image of character 1 (X1, Y1, Z1), the spatial coordinates corresponding to the face image of character 2 (X2, Y2, Z2), and the spatial coordinates corresponding to the face image of character 3 (X3, Y3, Z3). k k k K K K
[0193] As Figure 10 shown, Figure 10 Involving multi-view fusion face recognition (MVFI), in the multi-image feature fusion stage, taking the face image of character k as an example, including face image 1, face image 2, and face image 3, and inputting the facial feature coordinates and resolution in the face image information into the weighted fusion network (FWN) to obtain the fusion weight of each image. After combining the face features extracted by the face feature extraction network, the feature fusion is performed in the fusion layer of the fusion weight network to obtain the fusion feature. The fusion feature after fusion is compared with the face feature of the target character (character of interest), and finally the recognition result of the character target is output.
[0194] In addition, the present application also provides a computer readable storage medium. The computer readable storage medium of the present application stores a target positioning and identification program, wherein when the target positioning and identification program is executed by a processor, the steps of the target positioning and identification method as described above are implemented.
[0195] The method implemented when the target positioning and identifying program is executed can refer to the embodiments of the target positioning and identifying method of the present application, which will not be repeated here.
[0196] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0197] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0198] These computer program instructions can also be stored in a computer-readable storage unit that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable storage unit produce a product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0199] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0200] It is to be noticed that the singular word "a" or "an" should not be construed as meaning "one and only one" unless expressly so defined by the statements of the specification. The terms "comprising", "comprises" and "comprised of" should be interpreted as referring to the components, members, steps and / or elements of the object in which they are used. The reference signs in the claims should not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps other than those listed in a claim. The word "a" or "an" preceding the citation of a list of elements should not be construed as excluding the presence of additional such elements nor should it be construed as implying that the preceding list is exhaustive. The word "the" preceding the citation of a list of elements should not be construed as excluding the presence of one or more additional such elements nor should it be construed as implying that any such additional elements are an essential part of the preceding list. The word "substantially" does not exclude "completely" e.g. a composition which is "substantially non-toxic" can be a composition which is completely non-toxic. The word "comprise", and variations of the word such as "comprising" or "comprises", when used in a claim, will be understood to mean the inclusion of a stated integer or group of integers but not to the exclusion of any other integer or group of integers. The word "preferably" should not be understood as meaning "exclusively".
[0201] Although the preferred embodiments of the application have been described, those skilled in the art will be able to make additional modifications and variations to the described embodiments without departing from the inventive concepts disclosed in the specification. Accordingly, it is intended that the appended claims be construed to include all such modifications and variations as fall within the scope of the present application.
[0202] The above description is only preferred embodiments of the present application, and is not intended to limit the patent scope of the present application. Any equivalent structure variations made according to the disclosure and drawings of the present application, or direct / indirect application in other related technical fields, are included in the patent protection scope of the present application.
Claims
1. A method of target location and identification, characterized by, The target positioning and identification method is applied to a target positioning and identification system; the target positioning and identification system comprises a group of unmanned aerial vehicles and a group of edge devices; the group of unmanned aerial vehicles and the group of edge devices are in wireless communication connection; The target positioning and identification method comprises the following steps: obtaining multiple face image information of a character target by the group of unmanned aerial vehicles; determining a target edge device with the earliest completion time in the group of edge devices and scheduling each face image information to the target edge device; converting face image coordinates in the face image information into a set of three-dimensional coordinate points by the target edge device, and determining a face space position corresponding to the face image coordinates according to the set of three-dimensional coordinate points; and extracting face feature vectors in each face image information by the target edge device, and fusing each face feature vector to determine identity information of the character target; wherein the step of determining a face space position corresponding to the face image coordinates according to the set of three-dimensional coordinate points comprises: traversing three-dimensional coordinate points in the set of three-dimensional coordinate points, and projecting the traversed three-dimensional coordinate points into each face image information; determining whether the traversed three-dimensional coordinate points are projected in a face detection frame in all face image information; if the traversed three-dimensional coordinate points are in the face detection frame in all face image information, the traversed three-dimensional coordinate points are determined as the face space position corresponding to the face image coordinates.
2. The object locating and identifying method of claim 1, wherein, The step of obtaining multiple face image information of a character target by the group of unmanned aerial vehicles comprises: collecting multiple scene images within a visual angle range by the group of unmanned aerial vehicles every preset shooting interval; determining and obtaining a human target image in the scene image based on a preset target detection model; obtaining face image information in the human target image based on a preset face detection model; the face image information represents a face image with a face detection frame and face image coordinates.
3. The object locating and identifying method of claim 1, wherein, The group of edge devices comprises a scheduling edge device and a non-scheduling edge device; The step of determining a target edge device with the earliest completion time in the group of edge devices comprises: sending a test file package to each non-scheduling edge device by the scheduling edge device to determine a current transmission delay of the non-scheduling edge device; and obtaining a current task queue length, a central processing unit performance parameter and an image processor performance parameter of the non-scheduling edge device; inputting the transmission delay, the task queue length, the central processing unit performance parameter and the image processor performance parameter into a preset multivariate linear regression model to obtain a predicted completion time of the non-scheduling edge device; comparing the predicted completion times of each non-scheduling edge device to determine a target edge device with the earliest completion time in each non-scheduling edge device.
4. The object locating and identifying method of claim 1, wherein, The step of converting face image coordinates in the face image information into a set of three-dimensional coordinate points by the target edge device comprises: inputting, by the target edge device, the face image coordinates in the face image information and a preset height space limitation parameter into a preset two-dimensional-three-dimensional conversion model to convert the face image coordinates to obtain a three-dimensional coordinate point set having a plurality of three-dimensional coordinate points.
5. The object locating and identifying method of claim 1, wherein, The face image coordinates include face feature coordinate points; the step of fusing the face feature vectors to determine the identity information of the target person includes: inputting the face feature coordinate points and the resolution of each face image information into a preset fusion weight model to obtain a feature fusion weight of each face image information; multiplying the face feature vectors by the corresponding feature fusion weights to obtain weighted face feature vectors; adding the weighted face feature vectors to determine the identity information of the target person.
6. The object locating and identifying method of claim 5, wherein, The step of adding the weighted face feature vectors to determine the identity information of the target person includes: adding the weighted face feature vectors to obtain a fused face feature vector, and inputting the fused face feature vector into a support vector machine to determine a corresponding person category; if the person category is a preset concerned person, outputting the spatial position and face image of the target person.
7. The object locating and identifying method of claim 1, wherein, The target positioning and recognition system further includes a cloud server; the edge device is in wireless communication connection with the cloud server; After the step of determining the target edge device with the earliest completion time in the edge device group, the method further includes: if the predicted completion time of the target edge device is greater than a preset shooting interval, sending each face image information to the cloud server; the cloud server is configured to determine a face space position corresponding to the face image coordinates; and determine the identity information of the target person.
8. A target location and identification device, characterized by The target positioning and recognition device includes a processor, a storage unit, and a target positioning and recognition program stored on the storage unit and executable by the processor, wherein when the target positioning and recognition program is executed by the processor, the steps of the target positioning and recognition method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a target positioning and recognition program, wherein when the target positioning and recognition program is executed by the processor, the steps of the target positioning and recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Face fusion recognition method and device, electronic equipment and storage medium
CN108229330A
Detection and identification method and device of underwater robot and computer storage medium
CN112862865A
Method for identifying consignee based on face alignment and face identification in unmanned aerial vehicle delivery
CN114220157A