A method, system and device for sampling and classifying images

By setting up multiple cameras on the shooting path, automatically acquiring and clustering pedestrian images, the problem of inefficient manual acquisition and classification in the training of the cross-lens tracking algorithm model is solved, and efficient image sampling and classification are achieved.

CN114724186BActive Publication Date: 2025-07-22JINAN BOGUAN INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210389297.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-07-22
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

In the prior art, the model training of portrait cross-lens tracking algorithm requires a large number of manual acquisition and classification of images, which is inefficient and has a large workload.

Method used

By setting up multiple cameras on the shooting path, images containing pedestrians are automatically acquired, pedestrian features are clustered, pedestrian images are generated, and they are used as training sets to reduce manual intervention.

Benefits of technology

It improves the shooting and classification efficiency of portrait images, reduces workload, and realizes automated image sampling and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724186B_ABST
    Figure CN114724186B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for sampling and classifying images, which relates to the field of image acquisition. Since there are multiple first cameras on the shooting path, when there is a pedestrian in the shooting image of any one of the first cameras, the first camera with a pedestrian in its shooting image will be controlled to take a picture to obtain an image containing the pedestrian. Then, the images taken by all the first cameras in the current period are acquired, and these images are clustered according to the pedestrian features in these images. The images belonging to the same pedestrian are clustered into an image set to obtain the image sets corresponding to each pedestrian. Finally, the set of these image sets is determined as the training set of the cross-camera tracking algorithm model for portraits. There is no need for manual shooting of images of each pedestrian and classification of the taken images, which improves the shooting and classification efficiency of portrait images and reduces the workload.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image acquisition, and particularly to a method, system and device for sampling and classifying images. Background Art

[0002] When training a cross-camera tracking algorithm model for human figures, not only a large number of human figure image materials are required as a training set, but also the image materials of each human figure need to be composed of images taken from different shooting angles of the human figure, such as the front, back, left side and right side of the human figure, etc. The prior art first manually determines a road section as a sampling road section within the allowable shooting range, then installs cameras at various positions on this road section, so as to manually shoot images of each shooting angle of a pedestrian when the pedestrian passes through this road section, and finally manually classify the images taken so that the images of the same pedestrian are grouped into one category to obtain image materials of multiple pedestrians. However, the method of manual acquisition is inefficient, and a large amount of work is required for manually classifying the images taken. Summary of the Invention

[0003] The purpose of the present invention is to provide a method, system and device for sampling and classifying images, which do not require manual shooting of images of each pedestrian and classification of the taken images, improve the shooting and classification efficiency of human figure images, and reduce the workload.

[0004] To solve the above technical problems, the present invention provides a method for sampling and classifying images, including:

[0005] Obtaining an image containing a pedestrian captured by any first camera on a shooting path;

[0006] Obtaining pedestrian features in the images captured by all the first cameras in the current period;

[0007] Clustering each of the images based on the pedestrian features to obtain a set of pedestrian images corresponding to each of the pedestrians;

[0008] Taking the set containing all the sets of pedestrian images as a training set of a cross-camera tracking algorithm model for human figures.

[0009] Preferably, before obtaining an image containing a pedestrian captured by any first camera on a shooting path, it further includes:

[0010] Determining all available paths between a preset first location and a preset second location;

[0011] Obtaining the positions and orientation angles of each second camera in each of the available paths;

[0012] According to the positions and orientation angles of the second cameras, four second cameras that are respectively used to capture the front, back, left side, and right side of a pedestrian and are located on the same available path are used as a set of camera units for the available path;

[0013] Determine the number of camera units for each of the available paths;

[0014] Use the available paths with the number of camera units greater than the preset number of camera units as the shooting paths;

[0015] Wherein, the second camera serves as the first camera.

[0016] Preferably, after determining all available paths between the preset first location and the preset second location, it further includes:

[0017] Determine the number of straight-line segments in each of the available paths;

[0018] Judge whether the number of straight-line segments is greater than the preset number of straight-line segments;

[0019] If so, delete the available path;

[0020] If not, retain the available path;

[0021] Obtain the positions and orientation angles of each second camera in each of the available paths, including:

[0022] Obtain the positions and orientation angles of each second camera in each of the retained available paths.

[0023] Preferably, before using, according to the positions and orientation angles of the second cameras, four second cameras that are respectively used to capture the front, back, left side, and right side of a pedestrian and are located on the same available path as a set of camera units for the available path, it further includes:

[0024] Determine the included angle degree between the center line of the shooting range of the second camera and the center line of the road segment where the second camera is located in a preset coordinate system, where the vertical axis of the preset coordinate system is the center line of the road segment and the horizontal axis is the perpendicular line of the center line of the road segment;

[0025] Determine the second camera with the included angle degree between the first preset angle and the second preset angle as the second camera for capturing the front of the pedestrian;

[0026] Determine the second camera with the included angle degree between the second preset angle and the third preset angle as the second camera for capturing the back of the pedestrian;

[0027] When the position of the second camera is on the left side of the center line of the road section, the second camera with the included angle between the third preset angle and the fourth preset angle is determined as the second camera for photographing the left side of the pedestrian;

[0028] When the position of the second camera is on the right side of the center line of the road section, the second camera with the included angle between the fourth preset angle and the first preset angle is determined as the second camera for photographing the right side of the pedestrian.

[0029] Preferably, before using the available path with the number of camera groups greater than the preset number of camera groups as the shooting path, it further includes:

[0030] Respectively determine the sum of the number of images captured by all the second cameras within the first preset time in each available path;

[0031] Rank the sums according to the numerical magnitude relationship;

[0032] Judge whether there is an available path in all the available paths where the number of camera groups is greater than the preset number of camera groups and the rank of the sum is before the preset rank;

[0033] If so, use the available path where the number of camera groups is greater than the preset number of camera groups and the rank of the sum is before the preset rank as the shooting path;

[0034] If not, enter the step of using the available path with the number of camera groups greater than the preset number of camera groups as the shooting path.

[0035] Preferably, after obtaining the pedestrian image sets corresponding to each pedestrian, it further includes:

[0036] Respectively determine the number of front images, back images, left side images, and right side images of the pedestrian captured in each pedestrian image set;

[0037] Judge whether the quantity difference between any two of the front image, back image, left side image, and right side image of the pedestrian is less than the preset difference;

[0038] If all are less than the preset difference, it is determined that the missing shot rate of the pedestrian image set is qualified;

[0039] If there is a quantity difference not less than the preset difference, it is determined that the missing shot rate of the pedestrian image set is unqualified.

[0040] Preferably, after obtaining the pedestrian image sets corresponding to each pedestrian, it further includes:

[0041] Determine the total number of the first cameras;

[0042] Based on the total number, determine the theoretical number of images in each of the pedestrian image sets;

[0043] Determine the actual number of images in each of the pedestrian image sets;

[0044] Determine the ratio of the actual number of images to the theoretical number of images;

[0045] Determine the average value of all the ratios;

[0046] Judge whether the average value is within a preset numerical range;

[0047] If so, determine that the actual number of images meets the preset number requirement;

[0048] If not, determine that the actual number of images does not meet the preset number requirement.

[0049] Preferably, cluster each of the images based on the pedestrian features to obtain the pedestrian image sets corresponding to each of the pedestrians, including:

[0050] Determine any one of the first cameras as an anchor camera among all the first cameras;

[0051] Using a preset similarity relationship formula, respectively determine the feature similarities between the pedestrian features of each image captured by the anchor camera and the pedestrian features of the images captured by the other first cameras;

[0052] Cluster the images with the feature similarities greater than the first preset similarity into the pedestrian image sets;

[0053] Merge the pedestrian image sets corresponding to the pedestrians with the historical image sets to obtain the new historical image sets corresponding to the pedestrians;

[0054] Use the new historical image sets as the pedestrian image sets.

[0055] The present application further provides an image sampling and classification system, including:

[0056] An image acquisition unit, configured to acquire an image including a pedestrian captured by any one of the first cameras on a shooting path;

[0057] A feature acquisition unit, configured to acquire the pedestrian features in the images captured by all the first cameras in the current period;

[0058] A clustering unit, configured to cluster each of the images based on the pedestrian features to obtain the pedestrian image sets corresponding to each of the pedestrians;

[0059] A determination unit for using the set containing all the pedestrian image sets as the training set of the cross-camera pedestrian tracking algorithm model.

[0060] This application also provides an image sampling and classification device, including:

[0061] A memory for storing a computer program;

[0062] A processor for implementing the steps of the image sampling and classification method as described above when executing the computer program.

[0063] The present invention provides an image sampling and classification method, system and device. Since there are multiple first cameras on the shooting path, when there is a pedestrian in the shooting screen of any one of the first cameras, the first camera with a pedestrian in its shooting screen will be controlled to take pictures to obtain images containing the pedestrian. Then, the images taken by all the first cameras in the current period are acquired, and these images are clustered according to the pedestrian features in these images. The images belonging to the same pedestrian are clustered into an image set to obtain the image sets corresponding to each pedestrian. Finally, the set of these image sets is determined as the training set of the cross-camera pedestrian tracking algorithm model. It is not necessary to manually take pictures of the images of each pedestrian and classify the taken images, which improves the efficiency of shooting and classifying portrait images and reduces the workload. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the prior art and the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0065] Figure 1 It is a flowchart of an image sampling and classification method provided by the present invention;

[0066] Figure 2 It is a schematic diagram of the shooting angle of the second camera provided by the present invention;

[0067] Figure 3 It is a schematic structural diagram of an image sampling and classification system provided by the present invention;

[0068] Figure 4 It is a schematic structural diagram of an image sampling and classification device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The core of the present invention is to provide a method, system and device for sampling and classifying images, which do not require manual shooting of images of each pedestrian and classifying the captured images, improving the shooting and classification efficiency of portrait images and reducing the workload.

[0070] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] Please refer to Figure 1 , Figure 1 , which is a flowchart of a method for sampling and classifying images provided by the present invention. The method includes:

[0072] S11: Obtain an image containing a pedestrian captured by any first camera on the shooting path;

[0073] The shooting path refers to a path that a pedestrian walks through in reality. Due to the need to consider privacy issues, the shooting path is usually some paths in the designated areas where shooting can be carried out. For example, two locations can be determined in the area where shooting can be carried out, and then all paths between these two locations can be calculated using a path planning algorithm such as the A* algorithm. These paths can all be used as shooting paths, or some shooting-friendly paths can be selected from these paths as shooting paths. When identifying whether there is a pedestrian in the shooting screen of the first camera, a neural network model can be set in the processor of the backend system connected to each first camera. When there is a pedestrian in the shooting screens of these first cameras, the trajectory of the pedestrian is tracked according to a preset tracking algorithm. After determining that the pedestrian leaves the shooting screen of the first camera, the most suitable frame of image is collected from all the trajectories obtained by tracking the pedestrian, such as the frame of image with the clearest pedestrian features, so as to facilitate the subsequent acquisition of the image containing the pedestrian in the first camera. When the processor detects a pedestrian based on the shooting screen of the camera and controls the first camera to shoot, an image containing the pedestrian can also be obtained. The neural network model is pre-trained by multiple images containing pedestrians. The present application does not limit how to obtain the image containing the pedestrian.

[0074] S12: Obtain the pedestrian features in the images captured by all first cameras in the current cycle;

[0075] The current cycle can be any set time. In order to cluster all the images captured by all the first cameras during the current cycle, it is necessary to obtain the pedestrian features in all the images captured by all the first cameras. The pedestrian features can include, but are not limited to, the face features and body features of the pedestrians.

[0076] S13: Cluster each image based on the pedestrian features to obtain a set of images corresponding to each pedestrian.

[0077] When clustering, if the pedestrian features include face features and body features, take the images in one first camera as a reference, and then compare the features with all other images. Then, cluster the images with similar features into an image set. When performing feature comparison, all the images containing face features can be clustered as one clustering step, and all the images containing body features can be clustered as another clustering step, that is, cluster the face features and body features separately. For example, there are two images A and B in the first camera, and the other images are a, b, c, and d. Then, a, b, c, and d all need to be compared with A and B for features. If it is determined that a and b are similar to A, and c and d are similar to B, then cluster a, b, and A into one image set, and cluster c, d, and B into another image set. Considering that face features can better represent the information of a pedestrian, the face image set and body image set of the same pedestrian can be combined based on the face feature image set to obtain the final image set of the pedestrian.

[0078] S14: Use the set containing all the pedestrian image sets as the training set of the portrait cross-camera tracking algorithm model.

[0079] The present invention provides a method, system, and device for sampling and classifying images. Since there are multiple first cameras on the shooting path, when there is a pedestrian in the shooting screen of any first camera, the first camera with a pedestrian in its shooting screen will be controlled to take pictures to obtain images containing the pedestrian. Then, all the images captured by all the first cameras during the current cycle are obtained, and these images are clustered according to the pedestrian features in these images. The images belonging to the same pedestrian are clustered into an image set to obtain the image sets corresponding to each pedestrian. Finally, the set of these image sets is determined as the training set of the portrait cross-camera tracking algorithm model, eliminating the need for manual shooting of images of each pedestrian and classification of the captured images, improving the shooting and classification efficiency of portrait images and reducing the workload.

[0080] Based on the above embodiments:

[0081] As a preferred embodiment, before controlling the first camera with a pedestrian in its shooting screen to take pictures, it further includes:

[0082] Determine all available paths between a preset first location and a preset second location;

[0083] Obtain the positions and orientation angles of each second camera in each available path;

[0084] According to the positions and orientation angles of the second cameras, take the 4 second cameras that are respectively used to capture the front, back, left side, and right side of a pedestrian on the same available path as a set of camera units for the available path;

[0085] Determine the number of camera units for each available path;

[0086] Take the available paths with the number of camera units greater than the preset number of camera units as shooting paths;

[0087] Wherein, the second camera serves as the first camera.

[0088] Considering the actual situation, when determining the shooting path, there are some unavailable paths between the first location and the second location, and not all paths are normal paths. For example, there may be a road under construction. In order to more effectively collect pedestrian images, it is necessary to determine the available paths between the two locations. Since the image material of a pedestrian requires images in 4 directions: the front, back, left side, and right side of the pedestrian, in order to ensure that 4 different side images of the pedestrian can be obtained, it is necessary to determine 4 second cameras that shoot different sides on the available path as a set of camera units. That is, usually a set of camera units can capture 4 different side images of a pedestrian. Please refer to Figure 2 , Figure 2 is a schematic diagram of the shooting angle of the second camera provided by the present invention. When the path itself is the X-axis, it can be seen that the camera located at a certain angle to the right of the X-axis, that is, the camera in the blank area to the right of the X-axis, can be used as the camera for shooting the right side of the pedestrian. Secondly, considering the issue of image acquisition efficiency, when the number of pedestrians passing through each available path is the same, the more the number of camera units on an available path, the more pedestrian images can be captured on that available path. Taking the available paths with the number of camera units greater than the preset number as shooting paths can improve the efficiency of image acquisition.

[0089] It should also be noted that although the second camera is defined as the second camera for photographing a pedestrian at a certain angle based on the position and orientation angle of the second camera, in actual applications, each second camera will still take pictures when it detects a pedestrian by itself or when other devices detect a pedestrian, not limited to taking pictures only when the pedestrian is at a certain angle. That is, the second camera defined as photographing the right side of the pedestrian will not only take pictures when it sees the right side of the pedestrian. Defining each second camera is to ensure that a pedestrian can obtain images of the four sides of the pedestrian after passing through the shooting of a camera group.

[0090] As a preferred embodiment, after determining all available paths between the preset first location and the preset second location, it further includes:

[0091] Determine the number of straight-line segments in each available path;

[0092] Judge whether the number of straight-line segments is greater than the preset number of straight-line segments;

[0093] If so, delete the available path;

[0094] If not, retain the available path;

[0095] Obtain the positions and orientation angles of each second camera in each available path, including:

[0096] Obtain the positions and orientation angles of each second camera in each retained available path.

[0097] In order to reduce the complexity of the shooting path, in this application, considering that there are many available paths between two locations, but some of them are more winding paths. These paths will have problems such as narrow camera fields of view, low pedestrian flow, and winding roads that are not easy to determine the second camera. Therefore, before defining the second camera, first delete the more winding available paths among all the available paths, that is, leave the available paths with better fields of view and straighter roads, which can reduce the complexity of the shooting path and facilitate subsequent determination of the positions of each second camera and clustering.

[0098] As a preferred embodiment, before taking the four second cameras located on the same available path and respectively used to photograph the front, back, left side, and right side of the pedestrian as a set of camera groups for the available path, it further includes:

[0099] Determine the included angle degree between the center line of the shooting range of the second camera and the center line of the road segment where the second camera is located in the preset coordinate system. The vertical axis of the preset coordinate system is the center line of the road segment, and the horizontal axis is the perpendicular line of the center line of the road segment;

[0100] Determine the second camera with an included angle between the first preset angle and the second preset angle as the second camera for shooting the front of the pedestrian;

[0101] Determine the second camera with an included angle between the second preset angle and the third preset angle as the second camera for shooting the back of the pedestrian;

[0102] When the position of the second camera is on the left side of the road section center line, determine the fourth camera with an included angle between the fifth preset angle and the third preset angle as the second camera for shooting the left side of the pedestrian;

[0103] When the position of the second camera is on the right side of the road section center line, determine the second camera with an included angle between the fourth preset angle and the first preset angle as the second camera for shooting the right side of the pedestrian.

[0104] In order to define the shooting orientations of each second camera, in this application, a coordinate system will be preset in advance. Please follow Figure 2 , Figure 2 is a schematic diagram of the shooting angle of the second camera provided by the present invention. In the preset coordinate system, take the middle line of the straight road section as the Y-axis and the perpendicular line of the middle line as the X-axis. The directions of the coordinate axes and the included angle between the second camera and the road section center line can be set according to the actual application scenario. For example, when the straight road section is a one-way street facing due north, the forward direction of the straight road section, that is, the north, can be used as the positive direction of the Y-axis, and the direction of the perpendicular line to the east can be used as the positive direction of the X-axis. The included angle and the definition of the second camera can be set with different included angles according to the width of the straight road section. For example, the second camera with an included angle between 60 degrees and 120 degrees can be used as the camera for shooting the front of the pedestrian, the second camera with an included angle between 240 and 300 degrees can be used as the camera for shooting the back of the pedestrian, the second camera with an included angle between 120 and 240 degrees on the left side of the road can be used as the camera for shooting the left side of the pedestrian, and the second camera with an included angle between 300 and 60 degrees on the right side of the road can be used as the camera for shooting the right side of the pedestrian. According to the actual width of the straight road section, these included angle definitions can be changed accordingly. In addition, with different directions of the coordinate axes, the definition of the second camera at the same position is also different. If the positive direction of the Y-axis in the coordinate system corresponding to the above-mentioned one-way street facing due north is changed from due north to due south, the second camera that originally shoots the front of the pedestrian can become the second camera that shoots the back of the pedestrian.

[0105] As a preferred embodiment, before using the available path with the number of camera groups greater than the preset number of camera groups as the shooting path, it further includes:

[0106] Determine the sum of the numbers of the images captured by all the second cameras within a first preset time in each available path respectively;

[0107] Rank the sums of the numbers according to the magnitude relationship of the numerical values;

[0108] Determine whether there is an available path in all the available paths where the number of camera groups is greater than a preset number of camera groups and the rank of the sum of the numbers is before a preset rank;

[0109] If so, use the available path where the number of camera groups is greater than the preset number of camera groups and the rank of the sum of the numbers is before the preset rank as the shooting path;

[0110] If not, enter the step of using the available path where the number of camera groups is greater than the preset number of camera groups as the shooting path.

[0111] In order to improve the image acquisition efficiency, in this application, considering that there are multiple available paths and the pedestrian flow on each available path is different, it can be seen that the more the pedestrian flow of an available path is, the more images can be captured within a cycle. That is, the greater the sum of the numbers of the images captured by all the second cameras in an available path is, the greater the pedestrian flow on this available path is. However, considering that the more the number of camera groups is, the more accurately the clustering of the images can be reflected. This is because the shooting pictures of different cameras are different. The more cameras there are, the more types of shooting pictures need to be processed, and the higher the difficulty of clustering is. If the images can be accurately clustered under a higher difficulty, it means that the accuracy is higher. Therefore, the available path where the number of camera groups is greater than the preset number and the rank of the sum of the image numbers is before the preset rank, that is, the available path with a larger pedestrian flow, can be used as the shooting path, so as to acquire images more efficiently. If there is no such available path, it is still determined according to the number of camera groups.

[0112] As a preferred embodiment, after obtaining the pedestrian image sets corresponding to each pedestrian, it further includes:

[0113] Determine the numbers of the front images, back images, left side images and right side images of the pedestrians captured in each pedestrian image set respectively;

[0114] Determine whether the difference between the numbers of any two of the front image, back image, left side image and right side image of the pedestrian is less than a preset difference;

[0115] If they are all less than the preset difference, it is determined that the missing shot rate of the pedestrian image set is qualified;

[0116] If there is a difference not less than the preset difference, it is determined that the missing shot rate of the pedestrian image set is unqualified.

[0117] In order to determine whether the missed shot rate is qualified, in this application, considering that in the actual application scenario, when collecting pedestrian images, factors such as the environment where the first camera is located and the training degree of the neural network model may affect the collection of pedestrian images, resulting in the situation of missing the collection of pedestrian images. Therefore, it is necessary to analyze the images taken in the current cycle. Since the same number of first cameras for shooting different sides of pedestrians are usually set on the shooting path, theoretically, after each pedestrian passes through the shooting path, the number of images of the 4 sides of the pedestrian should be the same. However, due to the existence of the above-mentioned influences, the actual number of images taken by each first camera may have a difference from the theoretically taken number of images. The missed shot rate is determined according to this difference. When the difference is small, it indicates a low missed shot rate, and when the difference is large, it indicates a high missed shot rate. For example, if there are 100 frontal images of a certain pedestrian, 90 right-side images, 85 left-side images, and 110 back images, when the preset difference is 10, comparing the differences between any two side images one by one, it can be seen that since the difference between the left-side image and the back image is greater than the preset difference, it indicates that the missed shot rate is unqualified.

[0118] In addition, considering that in the actual situation, there may also be a situation where the pedestrian in the captured image is suddenly affected by external factors and is determined as another pedestrian after being affected by the external factors. For example, if the pedestrian in the captured image is originally pedestrian A, if pedestrian A is suddenly blocked by other objects in the captured image for a while and then appears in the captured image again, then pedestrian A may be determined as pedestrian B. At this time, the image of this pedestrian will also be captured, which is equivalent to obtaining a duplicate image of this pedestrian. In order to avoid this situation, when determining the number of frontal images, back images, left-side images, and right-side images of pedestrians captured in each pedestrian image set, it is necessary to remove such duplicate images from the pedestrian image set. Specifically, it can be screened manually or through software algorithms. This application does not limit the comparison.

[0119] As a preferred embodiment, after obtaining the pedestrian image set corresponding to each pedestrian, it further includes:

[0120] Determine the total number of the first cameras;

[0121] Based on the total number, determine the theoretical number of images in each pedestrian image set;

[0122] Determine the actual number of images in each pedestrian image set;

[0123] Determine the ratio of the actual number of images to the theoretical number of images;

[0124] Determine the average value of all ratios;

[0125] Judge whether the average value is within the preset numerical range;

[0126] If so, it is determined that the actual number of images meets the preset number requirement;

[0127] If not, it is determined that the actual number of images does not meet the preset number requirement.

[0128] In order to determine the accuracy of clustering pedestrian images, in this application, considering that in the actual application scenario, the accuracy of clustering pedestrian images is related to the clustering algorithm itself and the environment where each first camera is located. These factors may lead to missed detection situations. Therefore, it is necessary to analyze the images captured in the current cycle. For example, when there is only one shooting path, theoretically, the actual total number of images of a pedestrian captured by all first cameras after the pedestrian passes through the shooting path should be equal to the number of first cameras. However, due to the above-mentioned impacts in the actual application scenario, there may be a difference between the actual total number and the theoretical number of images. A percentage can be calculated between the actual capture quantity and the theoretical value in each pedestrian image set after clustering on the same day. The average value of the percentages of all sets is calculated. This value is the clustering accuracy evaluation for each day, and the accuracy of clustering pedestrian images can be judged based on this evaluation value. For example, if there are actually 10 pedestrians passing through the shooting path and there are 4 first cameras on the shooting path, a total of 40 pedestrian images will be captured. It can be known that theoretically, 10 pedestrian image sets should be obtained finally and each pedestrian image set should contain 4 images. However, if the actual number of pedestrian images in a pedestrian image set is less than 4, it means that the clustering accuracy of pedestrian images is not accurate enough. Based on this, it can also be known that if 10 pedestrian image sets are not obtained after clustering finally, it can also indicate that the clustering accuracy of pedestrian images is not accurate enough.

[0129] As a preferred embodiment, each photo is clustered based on pedestrian features to obtain a pedestrian photo set corresponding to each pedestrian, including:

[0130] Determine any one of the first cameras as an anchor camera among all the first cameras;

[0131] Using a preset similarity relationship formula, respectively determine the feature similarity between the pedestrian features of each photo captured by the anchor camera and the pedestrian features of the photos captured by other first cameras;

[0132] Cluster the photos with feature similarity greater than the first preset similarity into a pedestrian photo set;

[0133] Merge the pedestrian photo set corresponding to the pedestrian with the historical photo set of the pedestrian to obtain a new historical photo set corresponding to the pedestrian;

[0134] Use the new historical photo set as the pedestrian photo set.

[0135] In order to simply cluster all the images captured by the first cameras into pedestrian image sets corresponding to each pedestrian. Specifically, first define any one of the first cameras as the anchor camera, and then compare all the images captured by the anchor camera with all the images captured by other first cameras for feature comparison. The preset similarity relationship formula can be the cosine distance formula or a neural network model pre-trained with multiple pedestrian images. This application does not limit this. When the preset similarity relationship formula is the cosine distance formula, if the pedestrian feature in the image captured by the anchor camera is X, and the pedestrian feature in the image captured by other first cameras is Y, then X * Y = ‖X‖ * ‖Y‖ * cosθ, where cosθ is the feature similarity. It is necessary to compare the features of all the images captured by the anchor camera with all the images captured by all other first cameras one by one, that is, calculate the cosine distance, to obtain the distance between the images in each anchor camera and all other images. Preset a distance value. For each image in the anchor camera, determine whether the distance between it and all other images is greater than the preset distance, and then cluster the images with a distance greater than the preset distance and the images of the corresponding anchor camera into the image set of this pedestrian. For example, there are two images A and B in the anchor camera, and there are a total of 4 images in all the images of all other first cameras, namely a, b, c, and d. Determine the distances between A and a, b, c, d, and also determine the distances between B and a, b, c, d. If the distances between A and a and b are greater than the preset distance, and the distances between B and c and d are greater than the preset distance, then cluster A, a, and b as the image set of pedestrian A, and cluster B, c, and d as the image set of pedestrian B. If the distances between a and A and B are both greater than the preset distance, but the distance of A is greater than that of B, then a will be clustered with A and will not be clustered with B.

[0136] It should be noted that if the pedestrian feature includes a face feature and a body feature, the feature similarity can be calculated twice respectively to obtain a face image set and a body image set, and then compare whether the number of the same pedestrians in these two image sets is greater than the preset number. Specifically, since the face feature can express the information of the pedestrian more clearly than the body feature, the face feature can be used as the first priority comparison feature. When comparing the pedestrian images in these two image sets, first compare the images containing the face feature, and then compare the images containing the body feature. If the number of the same pedestrians is greater than the preset number, it means that the pedestrians recorded in the two sets are similar, and then merge these two sets. If it is less than the preset number, it means that the pedestrians recorded in these two sets are not similar, indicating that the ability to identify pedestrians or the accuracy of extracting pedestrian features is not high at this time. At this time, the similarity calculation of the body feature can be performed again, and then merged to ensure the accuracy of the merged set.

[0137] After clustering, calculate each image set obtained by clustering and each historical image set, calculate the first centroid feature vector of each image set obtained by clustering, and then compare it with the second centroid feature vector of the historical image set. Specifically, it can also be compared through the cosine distance formula. If the distance between the two centroid feature vectors is greater than the preset distance, then in the way of image set clustering as described above, cluster the image sets obtained by clustering corresponding to the historical image sets with a distance greater than the preset distance to obtain a new image set. This new image set contains the images of a pedestrian taken historically and the images taken this time, and use this new image set as the pedestrian image set.

[0138] To determine the accuracy of clustering, after clustering, the images in an image set are all images of the same pedestrian, that is, at this time, the pedestrian image sets corresponding to each pedestrian are obtained. However, considering the different similarity relationship formulas used, there will be different clustering effects. For example, if the preset similarity relationship formula used is the cosine distance relationship formula, since it cannot accurately define the preset distance when it is first put into use, it may occur that the images of two pedestrians are clustered into the image set of one pedestrian; when the anchor camera is a camera that captures the back of the pedestrian, since the images captured by this camera are usually the back of the pedestrian and the recognition rate is low, it is easy to cluster the images of the backs of other pedestrians into the image set of this pedestrian. Moreover, due to the low recognition rate of the back of the pedestrian, it is also difficult to cluster the front and side images of this pedestrian into the image set of this pedestrian. So at this time, it can be statistically determined whether the number of front, back, left side, and right side images in each image set is the same. If not, it means that the preset distance needs to be modified. If there is a large difference in the number of images, it means that the selected anchor camera is unreasonable and other cameras need to be selected as the new anchor camera.

[0139] In addition, considering that the current cycle is usually long, perhaps several days or even longer, in order to merge the image sets in a timely manner, a merging cycle can be set again. For example, when the current cycle is 240 hours, each 24 hours can be used as a merging cycle, and a merge of face features and body features and a merge of the image sets obtained by clustering and the historical image sets are performed within each merging cycle.

[0140] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of an image sampling and classification system provided by the present invention. The system includes:

[0141] An image acquisition unit 11, configured to acquire an image containing a pedestrian captured by any first camera on the shooting path;

[0142] A feature acquisition unit 12, configured to acquire the pedestrian features in the images captured by all first cameras within the current cycle;

[0143] The clustering unit 13 is configured to cluster each image based on pedestrian features to obtain a set of pedestrian images corresponding to each pedestrian;

[0144] The determination unit 14 is configured to use the set containing all the sets of pedestrian images as the training set of the portrait cross-camera tracking algorithm model.

[0145] For a detailed introduction to an image sampling and classification system provided in this application, please refer to the embodiments of the above-mentioned image sampling and classification method, which will not be elaborated herein.

[0146] Please refer to Figure 4 , Figure 4 FIG. is a schematic structural diagram of an image sampling and classification device provided by the present invention. The device includes:

[0147] A memory 21 for storing a computer program;

[0148] A processor 22 for implementing the steps of the above-mentioned photo sampling and classification method when executing the computer program.

[0149] For a detailed introduction to an image sampling and classification device provided in this application, please refer to the embodiments of the above-mentioned image sampling and classification method, which will not be elaborated herein.

[0150] In the present specification, the embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0151] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.

Claims

1. A method for sampling and classifying images, characterized in that, Including: Obtain an image containing a pedestrian captured by any one of the first cameras on the shooting path; Obtain the pedestrian features in the images captured by all the first cameras in the current cycle; Cluster the images based on the pedestrian features to obtain a set of pedestrian images corresponding to each pedestrian; Use the set containing all the sets of pedestrian images as the training set of the portrait cross-camera tracking algorithm model; Wherein, before obtaining an image containing a pedestrian captured by any one of the first cameras on the shooting path, it further includes: Determine all available paths between a preset first location and a preset second location; Determine the number of straight-line segments in each of the available paths; Judge whether the number of the straight-line segments is greater than a preset number of straight-line segments; If not, retain the available paths; Obtain the positions and orientation angles of each second camera in each of the retained available paths; According to the positions and orientation angles of the second cameras, regard 4 second cameras respectively used to capture the front, back, left side and right side of a pedestrian on the same available path as a set of camera units for the available path; Determine the number of sets of camera units for each of the available paths; Regard the available paths with the number of sets of camera units greater than a preset number of sets of camera units as the shooting path; Wherein, the second camera serves as the first camera.

2. The method for sampling and classifying an image according to claim 1, wherein, If the number of the straight-line segments is greater than the preset number of straight-line segments, delete the available paths.

3. The method for sampling and classifying an image according to claim 2, wherein Before, according to the positions and orientation angles of the second cameras, regarding 4 second cameras respectively used to capture the front, back, left side and right side of a pedestrian on the same available path as a set of camera units for the available path, it further includes: Determine the included angle degree between the center line of the shooting range of the second camera and the center line of the road segment where the second camera is located in a preset coordinate system, the vertical axis of the preset coordinate system is the center line of the road segment, and the horizontal axis is the perpendicular line of the center line of the road segment; Determine the second camera with the included angle degree between a first preset angle and a second preset angle as the second camera for capturing the front of a pedestrian; Determine the second camera with the included angle degree between the second preset angle and a third preset angle as the second camera for capturing the back of a pedestrian; When the position of the second camera is on the left side of the center line of the road segment, determine the second camera with the included angle degree between the third preset angle and a fourth preset angle as the second camera for capturing the left side of a pedestrian; When the position of the second camera is on the right side of the center line of the road segment, determine the second camera with the included angle degree between the fourth preset angle and the first preset angle as the second camera for capturing the right side of a pedestrian.

4. The method for sampling and classifying an image according to claim 1, characterized in that, Before regarding the available paths with the number of sets of camera units greater than a preset number of sets of camera units as the shooting path, it further includes: Respectively determine the sum of the numbers of images captured by all second cameras in each of the available paths within a first preset time; Rank each of the said quantity sums according to the numerical magnitude relationship; Determine whether there is an available path among all the said available paths where the number of the camera groups is greater than the preset number of camera groups and the rank of the quantity sum is before the preset rank; If so, use the available path where the number of the camera groups is greater than the preset number of camera groups and the rank of the quantity sum is before the preset rank as the shooting path; If not, enter the step of using the available path where the number of the camera groups is greater than the preset number of camera groups as the shooting path.

5. The method for sampling and classifying an image according to claim 1, characterized in that, After obtaining the pedestrian image sets corresponding to each of the said pedestrians, it further includes: Respectively determine the quantities of the frontal images, back images, left side images, and right side images of the pedestrians captured in each of the said pedestrian image sets; Judge whether the quantity differences between any two of the frontal image, back image, left side image, and right side image of the pedestrian are all less than the preset difference; If all are less than the preset difference, determine that the missed shot rate of the pedestrian image set is qualified; If there is a quantity difference not less than the preset difference, determine that the missed shot rate of the pedestrian image set is unqualified.

6. The method for sampling and classifying an image according to claim 1, wherein, After obtaining the pedestrian image sets corresponding to each of the said pedestrians, it further includes: Determine the total number of the first cameras; Based on the total number, determine the theoretical image quantity in each of the said pedestrian image sets; Determine the actual image quantity in each of the said pedestrian image sets; Determine the ratio of the actual image quantity to the theoretical image quantity; Determine the average value of all the said ratios; Judge whether the average value is within the preset numerical range; If so, determine that the actual image quantity meets the preset quantity requirement; If not, determine that the actual image quantity does not meet the preset quantity requirement.

7. The method for sampling and classifying an image according to any one of claims 1 to 6, characterized in that, Cluster each of the said images based on the pedestrian features to obtain the pedestrian image sets corresponding to each of the said pedestrians, including: Determine any one of the first cameras in all the said first cameras as the anchor camera; Using the preset similarity relationship formula, respectively determine the feature similarities between the pedestrian features of each of the images captured by the anchor camera and the pedestrian features of the images captured by the other first cameras; Cluster the images with the feature similarities greater than the first preset similarity into the pedestrian image sets; Merge the pedestrian image sets corresponding to the pedestrians with the historical image set to obtain the new historical image set corresponding to the pedestrians; Use the new historical image set as the pedestrian image set.

8. An image sampling and classification system, characterized in that, It includes: An image acquisition unit, configured to acquire an image containing a pedestrian captured by any one of the first cameras on the shooting path; A feature acquisition unit, configured to acquire the pedestrian features in the images captured by all the said first cameras in the current period; A clustering unit, configured to cluster each of the said images based on the pedestrian features to obtain the pedestrian image sets corresponding to each of the said pedestrians; A determination unit, configured to use the set containing all the said pedestrian image sets as the training set of the portrait cross-camera tracking algorithm model; Wherein, before acquiring an image containing a pedestrian captured by any one of the first cameras on the shooting path, it further includes: Determine all available paths between a preset first location and a preset second location; Determine the number of straight-line segments in each of the available paths; Judge whether the number of the straight-line segments is greater than a preset number of straight-line segments; If not, retain the available paths; Obtain the positions and orientation angles of each second camera in each of the retained available paths; According to the positions and orientation angles of the second cameras, use the 4 second cameras respectively for photographing the front, back, left side and right side of a pedestrian on the same available path as a set of camera units for the available path; Determine the number of the set of camera units for each of the available paths; Use the available paths with the number of the set of camera units greater than a preset number of the set of camera units as the shooting paths; Wherein, the second camera serves as the first camera.

9. An image sampling and classification device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the method for sampling and classifying an image according to any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Multi-angle pedestrian image data acquisition system and method

    CN108769574A

  • Portrait clustering method and device and medium

    CN114333039A