View angle selection method and device for multiple cameras placed in surrounding mode
By using Kalman filters and historical head trajectories in a multi-camera system to smoothly process and match the three-dimensional coordinates of the head, dynamically selecting the camera viewing angle, solving the problem of wrong viewing angle selection in multi-camera object detection scenarios, improving the recognition accuracy of target correspondence and correctness of viewing angle selection.
Patent Information
- Application Number
- CN202510093831.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
In multi-camera object detection and tracking scenarios, how to choose the most suitable viewing angle for object detection and tracking is an important research topic. In the case of similar detection targets, the prior art often leads to errors in identifying the corresponding relationship of detection targets between different cameras, further resulting in errors in viewing angle selection.
By obtaining the three-dimensional coordinates of the human head and the human body orientation from the current frame images collected by each camera, combining the Kalman filter and the historical human head trajectory, smoothing and matching, dynamically selecting the camera perspective of the optimal human body orientation.
The accuracy of identification of the corresponding relationships of different cameras detect targets is improved, and the error recognition caused by excessive similarity of detection targets is avoided, which enhances the accuracy of viewing angle selection.
Smart Images

Figure CN120047667A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual image processing, and particularly to a method and device for selecting perspectives of multiple cameras placed in a surrounding manner. Background Art
[0002] With the continuous development of computer vision technology, multi-camera systems have been widely used in many fields, including but not limited to intelligent monitoring, autonomous driving, virtual reality, augmented reality, 3D reconstruction, and target tracking. Especially in the target detection and tracking scenarios of multi-person multi-camera remote conferences, the collaborative work of multiple cameras can provide a comprehensive real-time perspective of each person.
[0003] However, in the above scenarios, how to select the most suitable perspective for target detection and tracking has always been an important research topic. Existing technologies usually extract features of the detection targets in each camera, perform feature comparison, and further calibrate the cameras in advance to determine the corresponding relationships of the detection targets recognized by different cameras, and select the camera corresponding to the appropriate perspective for each detection target. In the case of similar detection targets, the above methods often lead to incorrect recognition of the corresponding relationships of the detection targets between different cameras, further resulting in incorrect perspective selection.
[0004] Therefore, how to improve target recognition in the multi-camera target detection scenario with high similarity of detection targets is a technical problem that needs to be solved currently. Summary of the Invention
[0005] This application provides a method and device for selecting perspectives of multiple cameras placed in a surrounding manner to solve the technical problem that how to improve target recognition in the multi-camera target detection scenario with high similarity of detection targets is a technical problem that needs to be solved currently.
[0006] To solve the above technical problem, in a first aspect, an embodiment of this application provides a method for selecting perspectives of multiple cameras placed in a surrounding manner, including:
[0007] Obtain first detection data in the current frame image from the current frame images collected by each camera; wherein, the first detection data includes all first human head three-dimensional coordinates and human body orientations in the current frame image;
[0008] Match the first human head three-dimensional coordinates with a number of historical human head trajectories maintained by the corresponding camera;
[0009] Smooth the matched first human head three-dimensional coordinates through a Kalman filter and the historical human head trajectories to obtain corresponding second human head three-dimensional coordinates;
[0010] Determine the human body orientations corresponding to the detection targets in each of the current frame images according to the three-dimensional coordinates of the second human head, and select, according to the human body orientations, the camera with the smallest front view angle difference from all the cameras for the detection target.
[0011] Compared with the prior art, the embodiments of the present application have the following beneficial effects: By combining the Kalman filter with the first three-dimensional coordinate data obtained by real-time detection and the historical human head trajectories, the first three-dimensional coordinate data is corrected, improving the accuracy of the three-dimensional coordinates of the second human head under each camera, and further improving the recognition accuracy of the corresponding relationships of the detection targets under different cameras; When accurate three-dimensional coordinate data is obtained, the corresponding relationships of the detection targets under different cameras can be determined by the distances between the two three-dimensional coordinates, and further the human body orientations of the same detection target under different cameras can be determined. The above process does not require feature extraction of the pictures in the cameras, avoiding misidentifying the corresponding relationships of the detection targets under different cameras due to too high similarity of the detection targets, and further improving the correctness of the corresponding view selection of each detection target.
[0012] In some embodiments of the first aspect of the present application, obtaining the first detection data in the current frame image collected from each camera includes:
[0013] Input the current frame image into a first neural network to obtain second detection data; the second detection data includes all the two-dimensional coordinates of the first human heads in the current frame image and the human body orientations;
[0014] Input the current frame image into a second neural network to obtain third detection data; the third detection data includes the first depth map corresponding to the current frame image;
[0015] Obtain the three-dimensional coordinates of the first human head according to the two-dimensional coordinates of the first human head and the first depth map.
[0016] Compared with the prior art, the above embodiments have the following beneficial effects: By using the first neural network to identify the human heads in the current frame image, the two-dimensional pixel point coordinates of each target in the planar image are obtained. Further, by using the second neural network to estimate the depth map of the current frame image, the depth value reflects the relative position of the detection target in space. Therefore, the two-dimensional coordinates combined with the depth value can accurately obtain the three-dimensional coordinates of each human head, thereby improving the accuracy of the corresponding relationships of the human heads in the front frame images collected by different cameras subsequently.
[0017] In some embodiments of the first aspect of the present application, matching the three-dimensional coordinates of the first human head with a plurality of historical human head trajectories maintained by the corresponding camera includes:
[0018] The first detection data further includes the confidence levels corresponding to the three-dimensional coordinates of each of the first human heads;
[0019] According to the first threshold and the confidence level, all the three-dimensional head coordinates in the current frame image are divided into third three-dimensional head coordinates and fourth three-dimensional head coordinates; wherein, the confidence level of the third three-dimensional head coordinates is greater than the first threshold; the confidence level of the fourth three-dimensional head coordinates is less than or equal to the first threshold.
[0020] After matching each of the third three-dimensional head coordinates in the current frame image with each of the historical head trajectories, the remaining unmatched historical head trajectories are matched with the fourth three-dimensional head coordinates.
[0021] Compared with the prior art, the above embodiments have the following beneficial effects: In order to avoid the problem of incorrect matching between the three-dimensional head coordinates and the historical head trajectories caused by incorrect object recognition of the first neural network, the third three-dimensional head coordinates with higher confidence levels are preferentially matched with the historical head trajectories, so as to prevent the historical head trajectories corresponding to correctly recognized objects from being preempted by incorrectly recognized objects, thereby improving the accuracy of the matching between the three-dimensional head coordinates and the historical head trajectories.
[0022] In some embodiments of the first aspect of the present application, the step of, after matching each of the third three-dimensional head coordinates in the current frame image with each of the historical head trajectories, matching the remaining unmatched historical head trajectories with the fourth three-dimensional head coordinates includes:
[0023] Calculate and obtain a first cost matrix according to the first Euclidean distance between each of the third three-dimensional head coordinates and each of the historical head trajectories.
[0024] Optimize the first cost matrix through the Hungarian algorithm, and determine the historical head trajectories matched with each of the third three-dimensional head coordinates according to the optimized first cost matrix.
[0025] Calculate and obtain a second cost matrix according to the first Euclidean distance between the fourth three-dimensional head coordinates and the remaining each of the historical head trajectories.
[0026] Optimize the second cost matrix through the Hungarian algorithm, and determine the historical head trajectories matched with each of the fourth three-dimensional head coordinates according to the optimized second cost matrix.
[0027] Compared with the prior art, the above embodiments have the following beneficial effects: Since the closer the Euclidean distance between the three-dimensional coordinates of a human head and the historical human head trajectory is, the greater the probability that the three-dimensional coordinates of the human head match the historical human head trajectory. Therefore, a cost matrix is constructed based on the Euclidean distances between each three-dimensional coordinate of the human head and each historical human head trajectory, and the first cost matrix is optimized by the Hungarian algorithm to find a matching scheme that can minimize the sum of the Euclidean distances between each three-dimensional coordinate of the human head and the corresponding historical human head trajectory, thereby improving the accuracy of identifying the historical human head trajectories of each detection target in the current frame image.
[0028] In some embodiments of the first aspect of the present application, the matching of the first three-dimensional coordinates of the human head with a plurality of historical human head trajectories maintained by the corresponding camera further includes:
[0029] When there is no historical human head trajectory that matches the three-dimensional coordinates of the third human head, a historical human head trajectory corresponding to the three-dimensional coordinates of the third human head is established starting from the three-dimensional coordinates of the third human head;
[0030] When there is no historical human head trajectory that matches the three-dimensional coordinates of the fourth human head, the three-dimensional coordinates of the fourth human head are deleted;
[0031] When there is no first three-dimensional coordinate of the human head that matches the historical human head trajectory, the historical human head trajectory is deleted.
[0032] Compared with the prior art, the above embodiments have the following beneficial effects: Since the installation angles of each camera are different, it is possible that a certain detection target appears in the images captured by the camera only for a certain period of time. Therefore, when a detection target with a high confidence level appears for the first time, its historical human head trajectory is started to be maintained, and according to the matching results, the detection targets that fail to be successfully matched and have a low confidence level are deleted, eliminating the targets misidentified by the first neural network; at the same time, it is also possible that a certain detection target starts to disappear from the images captured by the camera for a certain period of time, resulting in only the historical human head trajectory of the target, but no corresponding first three-dimensional coordinates of the human head. Therefore, it is necessary to promptly eliminate the historical human head trajectory, thereby improving the correctness of the matching of historical human head trajectories between different cameras in the future.
[0033] In some embodiments of the first aspect of the present application, the smoothing the matched first three-dimensional coordinates of the human head through the Kalman filter and the historical human head trajectory to obtain the corresponding second three-dimensional coordinates of the human head includes:
[0034] Through the Kalman filter, the predicted coordinates corresponding to the first three-dimensional coordinates of the human head in the current frame image are obtained from the historical human head trajectory;
[0035] The predicted coordinates are corrected by combining the three-dimensional coordinates of the first human head through the Kalman filter, and the three-dimensional coordinates of the second human head are obtained.
[0036] Compared with the prior art, the above embodiments have the following beneficial effects: In order to further improve the accuracy of the three-dimensional coordinates of the first human head, the Kalman filter is used to maintain the historical human head trajectory of the detection target, smooth the data fluctuations caused by the detection errors of the neural network, and effectively improve the accuracy of subsequent data processing.
[0037] In some embodiments of the first aspect of the present application, the determining the human body orientation corresponding to each detection target in each of the current frame images according to the three-dimensional coordinates of the second human head includes:
[0038] A first transformation matrix between two corresponding point sets is obtained by minimizing the first matching distance between the corresponding point sets of any two of the current frame images; the points in the point set are the three-dimensional coordinates of the second human head corresponding to the current frame image;
[0039] Based on the point set of any current frame image, the points in other point sets are mapped into the reference through the first transformation matrix;
[0040] According to the second Euclidean distance between the points in the reference, several points belonging to the same detection target are determined, and the human body orientation corresponding to the points is used as the human body orientation of the detection target in the corresponding current frame image.
[0041] Compared with the prior art, the above embodiments have the following beneficial effects: When the three-dimensional coordinates of the second human head in any current frame image match the historical human head trajectory, the corresponding relationship between the three-dimensional coordinates of the second human head in the current frame images collected by different cameras is still unknown. Therefore, by matching the three-dimensional coordinates of the second human head in different current frame images and obtaining the first transformation matrix between different current frame images to determine the corresponding relationship between the three-dimensional coordinates of the second human head in the current frame images collected by different cameras, the human body orientation of the same human body under different camera perspectives can be accurately known, and then the camera perspective with the optimal human body orientation can be dynamically selected through the dynamically maintained historical human head trajectory, improving the accuracy of perspective selection.
[0042] In some embodiments of the first aspect of the present application, the obtaining a first transformation matrix between two corresponding point sets by minimizing the first matching distance between the corresponding point sets of any two of the current frame images includes:
[0043] According to the first transformation matrix, the second points closest to the first points in the first point set are determined from the second point set, and the third Euclidean distance between the first points and the second points is calculated and obtained;
[0044] Determine the first weight of the third Euclidean distance according to the depth value corresponding to the first point in the depth map, and obtain the first matching distance according to each of the third Euclidean distances and the corresponding first weight;
[0045] Adjust the first transformation matrix, and iteratively optimize the first matching distance until the first matching distance meets the second threshold.
[0046] Compared with the prior art, the above embodiments have the following beneficial effects: By combining the depth value with the Euclidean distance, the matching distance between two points is determined. The depth value reflects the reliability of the point to a certain extent. Therefore, the weight of the third Euclidean distance corresponding to the point is determined according to the depth value, and the first matching distance is calculated, which improves the accuracy of the evaluation of the matching degree. Therefore, when the first matching distance reaches the optimum, the corresponding first transformation matrix is more accurate when used to judge the corresponding relationship between the three-dimensional coordinates of the second human head in the current frame images collected by different cameras subsequently; In addition, since the first transformation matrix is obtained by iteratively optimizing the first matching distance, it is not necessary to calibrate the camera in advance to obtain an accurate first transformation matrix, which reduces the time cost of multi-camera view selection.
[0047] In some embodiments of the first aspect of the present application, the obtaining the first matching distance according to each of the third Euclidean distances and the corresponding first weight includes:
[0048] The first weight increases as the corresponding depth value decreases; the first point and the second point have a one-to-one matching relationship;
[0049] When there is no second point in the second point set that matches the first point, delete the first point from the first point set;
[0050] Multiply the third Euclidean distance by the first weight to obtain the second matching distance corresponding to the first point and the second point, and accumulate all the second matching distances to obtain the first matching distance.
[0051] Compared with the prior art, the above embodiments have the following beneficial effects: Since the smaller the depth value, the smaller the depth error, and the corresponding point is more reliable at this time, increasing the corresponding weight value can improve the accuracy of the evaluation of the matching relationship between the subsequent two point sets; In addition, by deleting redundant matching points during the matching process, the detection target of the multi-camera system is de-duplicated, and the accuracy of subsequent view selection is improved.
[0052] In a second aspect, an embodiment of the present application further provides a view selection device for a multi-camera placed in a surrounding manner, including: a first detection module, a first matching module, a Kalman filtering module, and a camera selection module;
[0053] Among them, the first detection module is used to obtain first detection data in the current frame image collected by each camera; among them, the first detection data includes all first human head three-dimensional coordinates and human body orientations in the current frame image.
[0054] The first matching module is used to match the first human head three-dimensional coordinates with a number of historical human head trajectories maintained by the corresponding camera.
[0055] The Kalman filtering module is used to perform smoothing processing on the matched first human head three-dimensional coordinates through a Kalman filter and the historical human head trajectories to obtain corresponding second human head three-dimensional coordinates.
[0056] The camera selection module is used to determine the human body orientations corresponding to each detection target in each of the current frame images according to the second human head three-dimensional coordinates, and select the camera with the smallest front view angle difference from all the cameras according to the human body orientations. Description of the Drawings
[0057] Figure 1 It is a schematic flowchart of a method for selecting a viewing angle of a multi-camera placed in a surrounding manner provided in some embodiments of the present application.
[0058] Figure 2 It is a schematic diagram of a camera deployment provided in some embodiments of the present application.
[0059] Figure 3 It is a schematic diagram of human body orientation judgment provided in some embodiments of the present application.
[0060] Figure 4 It is a schematic diagram of the structure of a first neural network provided in some embodiments of the present application.
[0061] Figure 5 It is a schematic diagram of the detection result of a first neural network provided in some embodiments of the present application.
[0062] Figure 6 It is a schematic diagram of the structure of a second neural network provided in some embodiments of the present application.
[0063] Figure 7 It is a schematic flowchart of the maintenance and update process of historical human head trajectories provided in some embodiments of the present application.
[0064] Figure 8 It is a schematic diagram of the process of mapping between two point sets provided in some embodiments of the present application.
[0065] Figure 9 It is a schematic diagram of the structure of a viewing angle selection device for a multi-camera placed in a surrounding manner provided in some embodiments of the present application. Detailed implementation manners
[0066] In the prior art, usually, feature extraction is performed on detection targets in each camera, and feature comparison is carried out. Further, by pre-calibrating the cameras, the corresponding relationship of the detection targets recognized by different cameras is determined, and a camera corresponding to a suitable viewing angle is selected for each detection target. In the case of similar detection targets, the above method often causes incorrect recognition of the corresponding relationship of the detection targets between different cameras, further resulting in incorrect viewing angle selection.
[0067] To solve the above technical problems, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0068] Embodiment 1
[0069] Please refer to Figure 1 , which is a method for selecting a viewing angle of multiple cameras placed in a surrounding manner provided by an embodiment of the present application, including S10 to S40, specifically:
[0070] S10: Obtain first detection data in the current frame image from the current frame images collected by each camera; wherein, the first detection data includes all first human head three-dimensional coordinates and human body orientations in the current frame image.
[0071] Refer to Figure 2 , which is a schematic diagram of camera deployment provided by some embodiments of the present application, wherein the cameras are placed in a surrounding manner, preferably surrounded at equal intervals, and the deployment order of the cameras is recorded, and the deployment order is used as a constraint condition for subsequent matching of detection targets (i.e., detection targets such as human heads, human faces, and human bodies), improving the matching accuracy. At the same time, arranging the cameras in a surrounding manner can also ensure that at least one camera can capture a clear frontal human body target.
[0072] Refer to Figure 3 , which is a schematic diagram for judging the human body orientation provided by some embodiments of the present application, wherein the definition of the human body orientation is the angle between the front of the human body (or human head, human face) and the camera. In some embodiments of the present application, preferably, the human body orientation angle with the human body facing away from the camera is 0 degrees, and the human body orientation angle with the human body facing the camera is 180 degrees. The closer the human body orientation angle is to 180 degrees, the more frontal the current human body is relative to the camera.
[0073] Further, in some embodiments of the present application, obtaining the first detection data in the current frame image collected from each camera includes:
[0074] Input the current frame image into a first neural network to obtain second detection data; the second detection data includes all the two-dimensional coordinates of the first human heads in the current frame image and the human body orientation;
[0075] Input the current frame image into a second neural network to obtain third detection data; the third detection data includes the first depth map corresponding to the current frame image;
[0076] Obtain the three-dimensional coordinates of the first human head according to the two-dimensional coordinates of the first human head and the first depth map.
[0077] The first neural network is used to identify the human heads in the current frame image to obtain the two-dimensional pixel point coordinates of each target in the planar image. Further, the second neural network is used to estimate the depth map of the current frame image. The depth value reflects the relative position of the detected target in space. Therefore, the two-dimensional coordinates combined with the depth value can accurately obtain the three-dimensional coordinates of each human head, thereby improving the accuracy of the head correspondence relationship in the previous frame images collected by different cameras subsequently.
[0078] Reference Figure 4 , which is a schematic structural diagram of a first neural network provided in some embodiments of the present application. Preferably, in some embodiments of the present application, the structure of the first neural network uses YOLOv7-tiny as the backbone network of the first neural network for feature extraction, uses SPPCSPC (Spatial Pyramid Pooling and Spatial Pyramid Context) as the NECK module for feature fusion, and inputs the fused features into several detection heads HEAD.
[0079] When performing object detection on the human face and the human head, if the same detection head is used, it will lead to the problem that the human face and the human head share a detection box (anchor). However, one detection box can only be used to detect one target, resulting in a missed detection problem. Therefore, preferably, in some embodiments of the present application, the first neural network is provided with three detection heads, which are respectively used to detect and output the detection data of the human head, the human face, and the human body, thereby avoiding the problem of low detection recall rate caused by the occupation of the detection box.
[0080] Preferably, in some embodiments of the present application, after the camera captures the current frame image, the current frame image is preprocessed and compressed into image data of a preset size, such as a size of 640*640*3.
[0081] Preferably, in some embodiments of the present application, the second detection data includes, but is not limited to: the center point of the detection box corresponding to the human face and the human head (the center point of the detection box corresponding to the human head is the two-dimensional coordinates of the first human head), width and height, confidence level, class score, the center point of the detection box corresponding to the human body, width and height, confidence level, class score, and the orientation of the human body. Refer to Figure 5 It is a schematic diagram of the detection result of a first neural network provided in some embodiments of the present application.
[0082] Preferably, in some embodiments of the present application, after the second detection data is obtained, due to the problem that one detection target has multiple detection boxes, the following steps need to be performed on the detection boxes to eliminate redundant detection boxes and ensure that each detected target in the current frame image has only one detection box: S11: Sort the detection boxes from high to low according to the confidence level of the detection boxes; S12: Select the detection box with the highest confidence level as the first detection box and remove the first detection box from the detection box list; S13: Calculate the proportion of the overlapping area of the first detection box and each detection box in the detection box list (Intersection over Union, IOU), and remove the detection boxes with IOU greater than a given threshold from the detection box list; S14: Repeat S12 to S13 until the detection box list is empty.
[0083] Preferably, in some embodiments of the present application, the first neural network is trained and obtained through the following steps: S15: For the conference room application scenario with multiple cameras arranged in a ring, collect images of multiple people in various types of conference room scenarios, different human head postures, different human body postures, different backgrounds, different genders, different lighting conditions, and different distances; S16: Mark the head boxes, face boxes, body boxes, and human body orientation information in the data collected in S15; S17: Perform data augmentation processing such as cropping, flipping, affine transformation, and color transformation on the data to enhance the sample diversity and improve the generalization ability of the model; S18: Use the Wise Intersection over Union (Wise-IoU) as the loss function for the target box, use the Binary Cross Entropy (BCE) as the loss function for the target score, and use the Mean Squared Error (MSE) as the loss function for the human body orientation; S19: Use the pre-trained weights of the backbone part of the first neural network pre-trained with the open-source ImageNet dataset, and further use the dataset obtained in S15 to S17, with an initial learning rate of 0.1, and train for 300 epochs to obtain the first neural network.
[0084] Refer to Figure 6, A schematic diagram of the structure of a second neural network provided in some embodiments of the present application, including: an encoder (Encoder) and a decoder (Decoder); wherein the encoder is composed of several layers of convolutional neural network layers (Convolutional Neural Networks, CNN) connected in sequence, preferably using 5 layers of CNN; the decoder is composed of several layers of CNN, a feature concatenation layer (Cat), and an upsampling layer (UpSample), preferably composed of three operation modules. After each operation module connects the upsampling layer to the feature concatenation layer, the feature concatenation layer is composed of CNN; the features extracted by the second to fourth layers of CNN in the encoder are respectively input to the three feature concatenation layers in the decoder.
[0085] Preferably, in some embodiments of the present application, before the current frame image is input to the second neural network, it is preprocessed, and the current frame image is scaled to a size of 640*480*3, and a depth map with a size of 240*320 is output.
[0086] Preferably, in some embodiments of the present application, the second neural network is trained and obtained through the following steps: S110: For the conference room application scenario with multiple cameras arranged in a surround manner, use RGB-Depth cameras to collect multi-person images under various types of conference room scenarios, with different head postures, different backgrounds, different lighting conditions, and different distances; S111: Use the depth map collected by the RGB-Depth camera as a label; S112: Perform data augmentation processing such as cropping, flipping, affine transformation, and color transformation on the data to improve sample diversity and enhance the generalization ability of the model; S113: Adopt an encoding-decoding network structure similar to Unet to meet the performance requirements of the device side; S114: Use MSE as the loss function; S115: First, pre-train with open-source data, and then train with the data collected in S110 to S112.
[0087] Preferably, in some embodiments of the present application, after obtaining the two-dimensional coordinates of the first human head and the first depth map, the three-dimensional coordinates of the first human head can be obtained through the following steps:
[0088] Based on the depth values of each pixel point in the first depth map, obtain the depth value corresponding to the two-dimensional coordinates of the first human head. Here, it is preferably to use the smoothed depth values of all pixel points within the human head detection box corresponding to the two-dimensional coordinates of the first human head, and then obtain the three-dimensional coordinates of the first human head based on the following formula:
[0089]
[0090] Wherein, (X, Y, Z) are the three-dimensional coordinates of the first human head; (u, v) are the two-dimensional coordinates of the first human head; depth is the depth value; f x , f y , c xand c y are the internal parameters of the camera.
[0091] S20: Match the three-dimensional coordinates of the first human head with a number of historical human head trajectories maintained by the corresponding camera.
[0092] Further, in some embodiments of the present application, the matching of the three-dimensional coordinates of the first human head with a number of historical human head trajectories maintained by the corresponding camera includes:
[0093] The first detection data further includes the confidence levels corresponding to the three-dimensional coordinates of each of the first human heads;
[0094] According to a first threshold and the confidence levels, divide all the three-dimensional coordinates of the first human heads in the current frame image into the three-dimensional coordinates of the third human heads and the three-dimensional coordinates of the fourth human heads; wherein, the confidence levels of the three-dimensional coordinates of the third human heads are greater than the first threshold; the confidence levels of the three-dimensional coordinates of the fourth human heads are less than or equal to the first threshold;
[0095] After matching the three-dimensional coordinates of each of the third human heads in the current frame image with each of the historical human head trajectories, match the remaining unmatched historical human head trajectories with the three-dimensional coordinates of the fourth human heads.
[0096] To avoid the problem of incorrect matching between the three-dimensional coordinates of the human head and the historical human head trajectories caused by incorrect object recognition by the first neural network, preferentially match the three-dimensional coordinates of the third human heads with higher confidence levels with the historical human head trajectories, and avoid the historical human head trajectories corresponding to the correctly recognized objects being preempted by the incorrectly recognized objects, thereby improving the accuracy of the matching between the three-dimensional coordinates of the human head and the historical human head trajectories.
[0097] Further, in some embodiments of the present application, after matching the three-dimensional coordinates of each of the third human heads in the current frame image with each of the historical human head trajectories, the matching of the remaining unmatched historical human head trajectories with the three-dimensional coordinates of the fourth human heads includes:
[0098] Calculate and obtain a first cost matrix according to the first Euclidean distances between the three-dimensional coordinates of each of the third human heads and each of the historical human head trajectories;
[0099] Optimize the first cost matrix through the Hungarian algorithm, and determine the historical human head trajectories matched with the three-dimensional coordinates of each of the third human heads according to the optimized first cost matrix;
[0100] Calculate and obtain a second cost matrix according to the first Euclidean distances between the three-dimensional coordinates of the fourth human heads and the remaining each of the historical human head trajectories;
[0101] Optimize the second cost matrix through the Hungarian algorithm, and determine the historical head trajectories matching each of the three-dimensional coordinates of the fourth head according to the optimized second cost matrix.
[0102] Since the closer the Euclidean distance between the three-dimensional coordinates of the head and the historical head trajectory, the greater the probability that the three-dimensional coordinates of the head match the historical head trajectory. Therefore, construct a cost matrix based on the Euclidean distance between each three-dimensional coordinate of the head and each historical head trajectory, optimize the first cost matrix through the Hungarian algorithm, and find a matching scheme that can minimize the sum of the Euclidean distances between each three-dimensional coordinate of the head and the corresponding historical head trajectory, so as to improve the accuracy of identifying the historical head trajectories of each detection target in the current frame image.
[0103] Preferably, in some embodiments of the present application, the first cost matrix and the second cost matrix may be in the following form:
[0104]
[0105] where, d i,j is the first Euclidean distance between the i-th historical head trajectory and the j-th three-dimensional coordinate of the head. This first Euclidean distance can be obtained by calculating the average value of the Euclidean distances between all points in the historical head trajectory and the three-dimensional coordinate of the head, or by using the Euclidean distance between the point in the historical head trajectory and the three-dimensional coordinate of the head. The present application does not limit the calculation method of the first Euclidean distance.
[0106] For example, after obtaining the above first cost matrix or second cost matrix, optimize the above matrix through the Hungarian algorithm. When the first cost matrix is the following matrix:
[0107]
[0108] Through the Hungarian algorithm, it can be known that d 1,4 、d 2,2 、d 3,1 and d 4,3 are the optimal matching results, and the total cost at this time is 29.
[0109] Furthermore, in some embodiments of the present application, the step of matching the first three-dimensional coordinates of the head with several historical head trajectories maintained by the corresponding camera further includes:
[0110] When there is no historical head trajectory matching the three-dimensional coordinates of the third head, establish a historical head trajectory corresponding to the three-dimensional coordinates of the third head with the three-dimensional coordinates of the third head as the starting point;
[0111] When there is no historical head trajectory matching the three-dimensional coordinates of the fourth head, delete the three-dimensional coordinates of the fourth head;
[0112] When there is no first head three-dimensional coordinate that matches the historical head trajectory, delete the historical head trajectory.
[0113] Since the installation angles of each camera are different, it is possible that a certain detection target appears in the images captured by the camera only for a certain period of time. Therefore, when a detection target with a high confidence level appears for the first time, start maintaining its historical head trajectory, and according to the matching results, delete the detection targets that fail to be successfully matched and have a low confidence level, and eliminate the targets misidentified by the first neural network. At the same time, it is also possible that a certain detection target starts to disappear from the images captured by the camera for a certain period of time, resulting in only the historical head trajectory of the target, but no corresponding first head three-dimensional coordinate. Therefore, it is necessary to promptly eliminate the historical head trajectory, so as to improve the correctness of the subsequent matching of historical head trajectories between different cameras.
[0114] Further, in some embodiments of the present application, the smoothing process of the matched first head three-dimensional coordinate through the Kalman filter and the historical head trajectory to obtain the corresponding second head three-dimensional coordinate includes:
[0115] Through the Kalman filter, obtain the predicted coordinate corresponding to the first head three-dimensional coordinate in the current frame image of the historical head trajectory;
[0116] Through the Kalman filter combined with the first head three-dimensional coordinate, correct the predicted coordinate to obtain the second head three-dimensional coordinate.
[0117] In order to further improve the accuracy of the first head three-dimensional coordinate, the historical head trajectory of the detection target is maintained through the Kalman filter to smooth the data fluctuations caused by the detection errors of the neural network, effectively improving the accuracy of subsequent data processing.
[0118] Preferably, in some embodiments of the present application, refer to Figure 7 , which is a schematic diagram of a historical head trajectory maintenance and update process provided in some embodiments of the present application, including:
[0119] S21: According to the confidence level of the head detection frame corresponding to the first head three-dimensional coordinate, divide all the first head three-dimensional coordinates corresponding to the current frame image to obtain the third head three-dimensional coordinate and the fourth head three-dimensional coordinate.
[0120] S22: First, match the third head three-dimensional coordinate with the historical head trajectory. If the match is successful, use the third head three-dimensional coordinate for the state update of the Kalman filter to obtain a stable second head three-dimensional coordinate, where the state update of the Kalman filter is performed through the following formula:
[0121]
[0122] where x is the three-dimensional coordinates of the second human head; is the predicted three-dimensional coordinates of the second human head by the Kalman filter; K is the Kalman gain; z is the three-dimensional coordinates of the third human head; H is the measurement matrix.
[0123] S23: If the three-dimensional coordinates of the third human head do not match any historical human head trajectories successfully, establish a historical human head trajectory for the three-dimensional coordinates of the third human head and start to maintain and update the historical human head trajectory; repeatedly execute S22 to S23 until all the three-dimensional coordinates of the third human head have completed the matching operation.
[0124] S24: Match the remaining historical human head trajectories with the three-dimensional coordinates of the fourth human head. If the match is successful, use the three-dimensional coordinates of the fourth human head for the state update of the Kalman filter, and replace the three-dimensional coordinates of the third human head in S22 with the three-dimensional coordinates of the fourth human head to obtain the three-dimensional coordinates of the second human head.
[0125] S25: If the three-dimensional coordinates of the fourth human head do not match any historical human head trajectories successfully, delete the three-dimensional coordinates of the fourth human head.
[0126] S26: If there is a historical human head trajectory that does not match any of the three-dimensional coordinates of the first human head successfully, mark the historical human head trajectory as lost, and continuously monitor whether there is a three-dimensional coordinate of the first human head in the subsequent specified number of frames of images that can match it. If there is still none, delete the historical human head trajectory.
[0127] S30: Smooth the matched three-dimensional coordinates of the first human head through the Kalman filter and the historical human head trajectory to obtain the corresponding three-dimensional coordinates of the second human head.
[0128] S40: Determine the human body orientation corresponding to each detection target in each of the current frame images according to the three-dimensional coordinates of the second human head, and select the camera with the smallest front view angle difference from all the cameras according to the human body orientation.
[0129] Further, in some embodiments of the present application, the determining the human body orientation corresponding to each detection target in each of the current frame images according to the three-dimensional coordinates of the second human head includes:
[0130] Obtain the first transformation matrix between the corresponding two point sets by minimizing the first matching distance between any two corresponding point sets of the current frame images; the points in the point set are the three-dimensional coordinates of the second human head corresponding to the current frame images;
[0131] Taking the point set of any current frame image as a reference, map the points in other point sets to the reference through the first transformation matrix;
[0132] Determine a plurality of points belonging to the same detection target according to the second Euclidean distance between the points in the reference, and use the human body orientation corresponding to the point as the human body orientation of the detection target in the corresponding current frame image.
[0133] When the second three-dimensional head coordinates in any current frame image match the historical head trajectory, the corresponding relationship between the second three-dimensional head coordinates in the current frame images collected by different cameras is still unknown. Therefore, by matching the second three-dimensional head coordinates in different current frame images, the first transformation matrix between different current frame images is obtained. After determining the corresponding relationship between the second three-dimensional head coordinates in the current frame images collected by different cameras, the human body orientation of the same human body under different camera perspectives can be accurately known. Furthermore, by dynamically maintaining the historical head trajectory, the camera perspective with the optimal human body orientation can be dynamically selected to improve the accuracy of perspective selection.
[0134] Preferably, referring to Figure 8 , it is a schematic diagram of the process of mapping between two point sets provided in some embodiments of the present application. Among the head point sets collected by Camera 1 and Camera 2, the points with the same color represent the same human body target. However, as can be seen from the Figure 8 upper two subgraphs, the perspectives of different cameras are different. Therefore, it is necessary to convert the perspective of Camera 2 to the perspective of Camera 1. Therefore, by multiplying the second three-dimensional head coordinates in the point set of Camera 2 by the first transformation matrix, it is converted to the coordinate system of the corresponding perspective of Camera 1, thereby obtaining Figure 8 the mapping result graph in the lower left. As can be seen from the mapping result graph, if there is a matching relationship between two points, the distance between the two points has reached the closest. Therefore, an accurate first transformation matrix helps to improve the accuracy of obtaining the matching relationship between subsequent two points.
[0135] Further, in some embodiments of the present application, the obtaining of the first transformation matrix between corresponding two point sets by minimizing the first matching distance between any two corresponding point sets of the current frame images includes:
[0136] According to the first transformation matrix, determine the second point closest to each first point in the first point set from the second point set, and calculate the third Euclidean distance between the first point and the second point;
[0137] According to the depth value corresponding to the first point in the depth map, determine the first weight of the third Euclidean distance, and obtain the first matching distance according to each third Euclidean distance and the corresponding first weight;
[0138] Adjust the first transformation matrix and iteratively optimize the first matching distance until the first matching distance meets the second threshold.
[0139] By combining the depth value with the Euclidean distance, the matching distance between two points is determined. The depth value reflects the reliability of the point to a certain extent. Therefore, according to the depth value, the weight corresponding to the third Euclidean distance of the point is determined, and the first matching distance is calculated, which improves the evaluation accuracy of the matching degree. Therefore, when the first matching distance reaches the optimum, the corresponding first transformation matrix is more accurate when judging the correspondence between the three-dimensional coordinates of the second human head in the current frame images collected by different cameras subsequently; in addition, since the first transformation matrix is obtained by iteratively optimizing the first matching distance, there is no need to calibrate the camera in advance to obtain an accurate first transformation matrix, which reduces the time cost of multi-camera view selection.
[0140] Further, in some embodiments of the present application, the obtaining of the first matching distance according to each of the third Euclidean distances and the corresponding first weights includes:
[0141] The first weight increases as the corresponding depth value decreases; the first point and the second point have a one-to-one matching relationship;
[0142] When there is no second point in the second point set that matches the first point, the first point is deleted from the first point set;
[0143] Multiply the third Euclidean distance by the first weight to obtain the second matching distance corresponding to the first point and the second point, and accumulate all the second matching distances to obtain the first matching distance.
[0144] Since the smaller the depth value, the smaller the depth error, and the corresponding point is more reliable at this time, increasing the corresponding weight value can improve the evaluation accuracy of the matching relationship between the subsequent two point sets; in addition, by deleting redundant matching points during the matching process, the purpose of removing duplicates in the detection target of the multi-camera system is achieved, and the accuracy of subsequent view selection is improved.
[0145] Preferably, in some embodiments of the present application, since the installation positions of the cameras are fixed, a rough initial first transformation matrix can be determined through the rotation ranges between different cameras, and the first transformation matrix is continuously iteratively optimized through the first matching distance, so that there is no need to calibrate the cameras in advance.
[0146] Preferably, in some embodiments of the present application, when selecting the camera with the smallest frontal view difference from all cameras according to the human body orientation, since the human body orientation corresponding to each human head in each frame of image is known, referring to Figure 3 it can be known that by selecting the camera with the human body orientation angle closest to 180 degrees from each camera, the camera with the smallest frontal view difference from the detection target can be obtained.
[0147] In summary, it can be seen that a perspective selection method for surrounding placement of multiple cameras provided by an embodiment of the present application has the following beneficial effects: By combining the Kalman filter with the first three-dimensional coordinate data obtained by real-time detection and the historical head trajectory, the first three-dimensional coordinate data is corrected, improving the accuracy of the second head three-dimensional coordinates under each camera and further enhancing the recognition accuracy of the corresponding relationship of detection targets under different cameras; When accurate three-dimensional coordinate data is obtained, the corresponding relationship of each detection target under different cameras can be determined by the distance between two three-dimensional coordinates, and further the human body orientation of the same detection target under different cameras can be determined. The above process does not require feature extraction of the pictures in the cameras, avoiding misidentifying the corresponding relationship of detection targets under different cameras due to too high similarity of detection targets, and further improving the correctness of perspective selection for each detection target.
[0148] Embodiment 2
[0149] Reference Figure 9 , a perspective selection device for surrounding placement of multiple cameras provided by an embodiment of the present application, includes: a first detection module 11, a first matching module 12, a Kalman filter module 13, and a camera selection module 14.
[0150] Further, in some embodiments of the present application, the first detection module 11 is configured to obtain first detection data in the current frame image from the current frame images collected by each camera; wherein, the first detection data includes all first head three-dimensional coordinates and the human body orientation in the current frame image; the first matching module 12 is configured to match the first head three-dimensional coordinates with a plurality of historical head trajectories maintained by the corresponding camera; the Kalman filter module 13 is configured to perform smoothing processing on the matched first head three-dimensional coordinates through a Kalman filter and the historical head trajectory to obtain corresponding second head three-dimensional coordinates; the camera selection module 14 is configured to determine the human body orientation corresponding to each detection target in each of the current frame images according to the second head three-dimensional coordinates, and select the camera with the smallest front view difference from all the cameras according to the human body orientation.
[0151] Further, in some embodiments of the present application, obtaining the first detection data in the current frame image from the current frame images collected by each camera includes: inputting the current frame image into a first neural network to obtain second detection data; the second detection data includes all first head two-dimensional coordinates and the human body orientation in the current frame image; inputting the current frame image into a second neural network to obtain third detection data; the third detection data includes a first depth map corresponding to the current frame image; and obtaining the first head three-dimensional coordinates according to the first head two-dimensional coordinates and the first depth map.
[0152] Further, in some embodiments of the present application, the matching of the three-dimensional coordinates of the first human head with a plurality of historical human head trajectories maintained by the corresponding camera includes: the first detection data further includes the confidence corresponding to each of the three-dimensional coordinates of the first human head; according to a first threshold and the confidence, all the three-dimensional coordinates of the first human head in the current frame image are divided into the three-dimensional coordinates of the third human head and the three-dimensional coordinates of the fourth human head; wherein, the confidence of the three-dimensional coordinates of the third human head is greater than the first threshold; the confidence of the three-dimensional coordinates of the fourth human head is less than or equal to the first threshold; after matching each of the three-dimensional coordinates of the third human head in the current frame image with each of the historical human head trajectories, the remaining unmatched historical human head trajectories are matched with the three-dimensional coordinates of the fourth human head.
[0153] Further, in some embodiments of the present application, after matching each of the three-dimensional coordinates of the third human head in the current frame image with each of the historical human head trajectories, the remaining unmatched historical human head trajectories are matched with the three-dimensional coordinates of the fourth human head, including: calculating and obtaining a first cost matrix according to the first Euclidean distance between each of the three-dimensional coordinates of the third human head and each of the historical human head trajectories; optimizing the first cost matrix by the Hungarian algorithm, and determining the historical human head trajectories matched with each of the three-dimensional coordinates of the third human head according to the optimized first cost matrix; calculating and obtaining a second cost matrix according to the first Euclidean distance between the three-dimensional coordinates of the fourth human head and the remaining each of the historical human head trajectories; optimizing the second cost matrix by the Hungarian algorithm, and determining the historical human head trajectories matched with each of the three-dimensional coordinates of the fourth human head according to the optimized second cost matrix.
[0154] Further, in some embodiments of the present application, the matching of the three-dimensional coordinates of the first human head with a plurality of historical human head trajectories maintained by the corresponding camera further includes: when there is no historical human head trajectory matching the three-dimensional coordinates of the third human head, establishing a historical human head trajectory corresponding to the three-dimensional coordinates of the third human head with the three-dimensional coordinates of the third human head as the starting point; when there is no historical human head trajectory matching the three-dimensional coordinates of the fourth human head, deleting the three-dimensional coordinates of the fourth human head; when there is no three-dimensional coordinate of the first human head matching the historical human head trajectory, deleting the historical human head trajectory.
[0155] Further, in some embodiments of the present application, the smoothing process of the matched three-dimensional coordinates of the first human head by the Kalman filter and the historical human head trajectory to obtain the corresponding three-dimensional coordinates of the second human head includes: obtaining, by the Kalman filter, the predicted coordinates corresponding to the three-dimensional coordinates of the first human head in the current frame image of the historical human head trajectory; correcting the predicted coordinates by combining the Kalman filter with the three-dimensional coordinates of the first human head to obtain the three-dimensional coordinates of the second human head.
[0156] Further, in some embodiments of the present application, determining the human body orientation corresponding to each detection target in each of the current frame images according to the three-dimensional coordinates of the second human head includes: obtaining a first transformation matrix between two corresponding point sets by minimizing a first matching distance between any two corresponding point sets of the current frame images; the points in the point set correspond to the three-dimensional coordinates of the second human head in the current frame image; taking the point set of any current frame image as a reference, mapping the points in other point sets to the reference through the first transformation matrix; determining several points belonging to the same detection target according to the second Euclidean distance between the points in the reference, and taking the human body orientation corresponding to the points as the human body orientation of the detection target in the corresponding current frame image.
[0157] Further, in some embodiments of the present application, obtaining the first transformation matrix between two corresponding point sets by minimizing the first matching distance between any two corresponding point sets of the current frame images includes: determining, according to the first transformation matrix, a second point closest to each first point in the first point set from the second point set, and calculating and obtaining a third Euclidean distance between the first point and the second point; determining a first weight of the third Euclidean distance according to the depth value corresponding to the first point in the depth map, and obtaining the first matching distance according to each third Euclidean distance and the corresponding first weight; adjusting the first transformation matrix and iteratively optimizing the first matching distance until the first matching distance meets a second threshold.
[0158] Further, in some embodiments of the present application, obtaining the first matching distance according to each third Euclidean distance and the corresponding first weight includes: the first weight increases as the corresponding depth value decreases; the first point and the second point have a one-to-one matching relationship; when there is no second point in the second point set that matches the first point, deleting the first point from the first point set; multiplying the third Euclidean distance by the first weight to obtain a second matching distance corresponding to the first point and the second point, and accumulating all the second matching distances to obtain the first matching distance.
[0159] It can be understood that the above device item embodiments correspond to the method item embodiments of the present invention. A perspective selection device for surrounding and placing multiple cameras provided by the embodiments of the present invention can implement any method item embodiment of the present invention, that is, the perspective selection method for surrounding and placing multiple cameras provided in Embodiment 1.
[0160] In summary, it can be seen that a perspective selection device for surrounding placement of multiple cameras provided by an embodiment of the present application has the following beneficial effects: By combining the Kalman filter with the first three-dimensional coordinate data obtained by real-time detection and the historical human head trajectory, the first three-dimensional coordinate data is corrected, improving the accuracy of the second human head three-dimensional coordinates under each camera, and further improving the recognition accuracy of the corresponding relationship of detection targets under different cameras; When accurate three-dimensional coordinate data is obtained, the corresponding relationship of each detection target under different cameras can be determined by the distance between the two three-dimensional coordinates, and further determine the human body orientation of the same detection target under different cameras. The above process does not require feature extraction of the pictures in the cameras, avoiding misidentifying the corresponding relationship of detection targets under different cameras due to too high similarity of detection targets, and further improving the correctness of the corresponding perspective selection of each detection target.
[0161] Embodiment III
[0162] Based on the above embodiment of the perspective selection method for surrounding placement of multiple cameras, another embodiment of the present application provides a perspective selection terminal device for surrounding placement of multiple cameras. The perspective selection terminal device for surrounding placement of multiple cameras includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the perspective selection method for surrounding placement of multiple cameras according to any embodiment of the present application is implemented.
[0163] Exemplarily, in this embodiment, the computer program can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the perspective selection device for surrounding placement of multiple cameras.
[0164] The perspective selection device for surrounding placement of multiple cameras can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The perspective selection terminal device for surrounding placement of multiple cameras may include, but is not limited to, a processor and a memory.
[0165] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the perspective selection device for the multi-cameras placed in a surrounding manner, and connects various parts of the entire perspective selection device for the multi-cameras placed in a surrounding manner through various interfaces and lines. The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory, the processor realizes various functions of the perspective selection device for the multi-cameras placed in a surrounding manner. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0166] Embodiment 4
[0167] Based on the embodiments of the above-described perspective selection method for multi-cameras placed in a surrounding manner, another embodiment of the present application provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the perspective selection method for multi-cameras placed in a surrounding manner according to any embodiment of the present application.
[0168] In this embodiment, the above storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0169] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for selecting a viewing angle of a plurality of cameras placed around, characterized in that: include: Acquire first detection data in the current frame image from the current frame image captured by each camera; wherein the first detection data includes all first head three-dimensional coordinates and body orientations in the current frame image; Matching the first head three-dimensional coordinates with a number of historical head trajectories maintained by the corresponding camera; Smoothing the matched three-dimensional coordinates of the first head through a Kalman filter and the historical head trajectory to obtain the corresponding three-dimensional coordinates of the second head; According to the second three-dimensional coordinates of the human head, the human body orientation corresponding to each detection target in each current frame image is determined, and according to the human body orientation, a camera with the smallest front viewing angle difference with the detection target is selected from all the cameras.
2. A method for selecting a viewing angle of a plurality of cameras placed around as claimed in claim 1, characterized in that: The step of acquiring first detection data in the current frame image from the current frame image collected by each camera includes: Input the current frame image into the first neural network to obtain second detection data; the second detection data includes all first head two-dimensional coordinates and the human body orientation in the current frame image; Input the current frame image into a second neural network to obtain third detection data; the third detection data includes a first depth map corresponding to the current frame image; The three-dimensional coordinates of the first head are acquired according to the two-dimensional coordinates of the first head and the first depth map.
3. The method for selecting a viewing angle of a plurality of cameras placed around as claimed in claim 1, characterized in that: The matching of the first head three-dimensional coordinates with a plurality of historical head trajectories maintained by the corresponding camera includes: The first detection data also includes the confidence level corresponding to each of the first three-dimensional coordinates of the head; According to the first threshold and the confidence level, all the first three-dimensional coordinates of the head in the current frame image are divided into third three-dimensional coordinates of the head and fourth three-dimensional coordinates of the head; wherein the confidence level of the third three-dimensional coordinates of the head is greater than the first threshold; and the confidence level of the fourth three-dimensional coordinates of the head is less than or equal to the first threshold; After matching each of the third head three-dimensional coordinates in the current frame image with each of the historical head trajectories, matching the remaining unmatched historical head trajectories with the fourth head three-dimensional coordinates.
4. A method for selecting a viewing angle of a plurality of cameras placed around as claimed in claim 3, characterized in that: After matching each of the third head three-dimensional coordinates in the current frame image with each of the historical head trajectories, matching the remaining unmatched historical head trajectories with the fourth head three-dimensional coordinates includes: Calculate and obtain a first cost matrix according to the first Euclidean distance between each of the third head three-dimensional coordinates and each of the historical head trajectories; Optimizing the first cost matrix by using the Hungarian algorithm, and determining the historical head trajectory matching the three-dimensional coordinates of each third head according to the optimized first cost matrix; Calculate and obtain a second cost matrix based on the first Euclidean distance between the fourth head three-dimensional coordinate and the remaining historical head trajectories; The second cost matrix is optimized by the Hungarian algorithm, and the historical head trajectories matching the three-dimensional coordinates of each fourth head are determined according to the optimized second cost matrix.
5. The method for selecting a viewing angle of a plurality of cameras placed around as claimed in claim 4, characterized in that: The matching of the first head three-dimensional coordinates with a plurality of historical head trajectories maintained by the corresponding camera further includes: When there is no historical head trajectory matching the third head three-dimensional coordinates, taking the third head three-dimensional coordinates as a starting point, establishing a historical head trajectory corresponding to the third head three-dimensional coordinates; When there is no historical head trajectory matching the fourth head three-dimensional coordinate, deleting the fourth head three-dimensional coordinate; When there is no first head three-dimensional coordinate matching the historical head trajectory, the historical head trajectory is deleted.
6. The method for selecting a viewing angle of a plurality of cameras placed around as claimed in claim 1, characterized in that: The step of smoothing the matched first head three-dimensional coordinates through the Kalman filter and the historical head trajectory to obtain the corresponding second head three-dimensional coordinates includes: Obtaining, by means of the Kalman filter, predicted coordinates of the historical head trajectory corresponding to the first head three-dimensional coordinates in the current frame image; The predicted coordinates are corrected by combining the first three-dimensional coordinates of the head through the Kalman filter to obtain the second three-dimensional coordinates of the head.
7. The method for selecting a viewing angle of a plurality of cameras placed around as claimed in claim 1, characterized in that: The determining, according to the second three-dimensional coordinates of the human head, the human body orientation corresponding to each detection target in each current frame image includes: By minimizing the first matching distance between any two corresponding point sets of the current frame image, a first transformation matrix between the corresponding two point sets is obtained; the points in the point set are the three-dimensional coordinates of the second head in the corresponding current frame image; Taking any point set of the current frame image as a reference, mapping points in other point sets to the reference through the first transformation matrix; According to the second Euclidean distance between the points in the benchmark, a plurality of points belonging to the same detection target are determined, and the human body orientation corresponding to the point is used as the human body orientation of the detection target in the corresponding current frame image.
8. A method for selecting a viewing angle of a plurality of cameras placed around as claimed in claim 2 or 7, characterized in that: The step of obtaining a first transformation matrix between any two corresponding point sets by minimizing a first matching distance between any two corresponding point sets of the current frame images comprises: According to the first transformation matrix, determining from the second point set the second point closest to each first point in the first point set, and calculating and obtaining the third Euclidean distance between the first point and the second point; Determine a first weight of the third Euclidean distance according to a depth value corresponding to the first point in the depth map, and obtain the first matching distance according to each third Euclidean distance and the corresponding first weight; The first conversion matrix is adjusted, and the first matching distance is iteratively optimized until the first matching distance meets a second threshold.
9. The method for selecting a viewing angle of a plurality of cameras placed around as claimed in claim 8, characterized in that: The acquiring the first matching distance according to each of the third Euclidean distances and the corresponding first weight includes: The first weight increases as the corresponding depth value decreases; the first point and the second point are in a one-to-one matching relationship; When there is no second point in the second point set that matches the first point, deleting the first point from the first point set; The third Euclidean distance is multiplied by the first weight to obtain a second matching distance corresponding to the first point and the second point, and all the second matching distances are accumulated to obtain the first matching distance.
10. A viewing angle selection device for surrounding multiple cameras, characterized in that: include: A first detection module, a first matching module, a Kalman filter module and a camera selection module; The first detection module is used to obtain first detection data in the current frame image from the current frame image captured by each camera; wherein the first detection data includes all first head three-dimensional coordinates and body orientation in the current frame image; The first matching module is used to match the first three-dimensional coordinates of the head with a plurality of historical head trajectories maintained by the corresponding camera; The Kalman filter module is used to smooth the matched three-dimensional coordinates of the first head through the Kalman filter and the historical head trajectory to obtain the corresponding three-dimensional coordinates of the second head; The camera selection module is used to determine the human body orientation corresponding to each detection target in each current frame image according to the second three-dimensional coordinates of the head, and select the camera with the smallest front viewing angle difference with the detection target from all the cameras according to the human body orientation.