Data set construction method and device, intelligent equipment and computer program product
By performing humanoid detection and filtering on surveillance videos, extracting and utilizing humanoid quality features, the problem of poor quality data sets in the existing technology is solved, and the construction of high-quality data sets and the improvement of pedestrian re-identification model performance is achieved.
Patent Information
- Application Number
- CN202510213976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the quality of the data sets required for pedestrian recognition model training is uneven, resulting in the impact of model performance, and manual labeling is dependent and costly.
By detecting and tracking the surveillance videos, humanoid quality characteristics (such as humanoid key points and humanoid quality scores), filtering multi-frame humanoid images, eliminating poor-quality images, and building high-quality humanoid datasets.
It realizes the acquisition of high-quality humanoid data sets at lower labor costs, improves the performance of pedestrian re-identification model, and reduces the dependence and cost of manual annotation.
Smart Images

Figure CN120148112A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of pedestrian re-identification, and particularly relates to a method for constructing a dataset, a device for constructing a dataset, an intelligent device, and a computer program product. Background Art
[0002] Data is one of the important factors affecting the development of deep learning technology. The scale and quality of the dataset will have a crucial impact on the performance of the deep learning model. Person Re-Identification refers to the ability to identify and retrieve the same human figure under multiple cameras. Therefore, the dataset required for training a pedestrian re-identification model is to label the same human figure in the video with the same label.
[0003] Currently, conventional data annotation methods rely too much on manual annotation, and there is no good filtering and screening mechanism during the annotation process, resulting in uneven quality of the obtained human figure datasets, thus affecting the performance of the pedestrian re-identification model to a certain extent. Summary of the Invention
[0004] This application provides a method for constructing a dataset, a device for constructing a dataset, an intelligent device, and a computer program product, which can obtain a high-quality human figure dataset based on a relatively low labor cost.
[0005] In a first aspect, this application provides a method for constructing a dataset, including:
[0006] Performing human figure detection and tracking on a surveillance video to obtain a first human figure trajectory, where the first human figure trajectory includes multiple frames of human figure images;
[0007] Extracting human figure quality features of each frame of human figure image, where the human figure quality features include: human figure key points and / or human figure quality scores;
[0008] Filtering the multiple frames of human figure images based on the human figure quality features to obtain a second human figure trajectory;
[0009] Constructing a human figure dataset based on the second human figure trajectory.
[0010] In a second aspect, this application provides a device for constructing a dataset, including:
[0011] A detection module, configured to perform human figure detection and tracking on a surveillance video to obtain a first human figure trajectory, where the first human figure trajectory includes multiple frames of human figure images;
[0012] An extraction module, configured to extract human figure quality features of each frame of human figure image, where the human figure quality features include: human figure key points and / or human figure quality scores;
[0013] A filtering module, configured to filter multiple frames of humanoid images based on humanoid quality features to obtain a second humanoid trajectory;
[0014] A construction module, configured to construct a humanoid dataset based on the second humanoid trajectory.
[0015] In a third aspect, the present application provides an intelligent device. The intelligent device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method in the first aspect are implemented.
[0016] In a fourth aspect, the present application provides a computer program product. The computer program product includes a computer program. When the computer program is executed by one or more processors, the steps of the method in the first aspect are implemented.
[0017] The beneficial effects of the present application compared with the prior art are as follows: After obtaining the first humanoid trajectory through humanoid detection and tracking of the surveillance video, the present application solution first extracts the humanoid quality features of each frame of humanoid image in the first humanoid trajectory. According to the actual requirements of the application scenario, the humanoid quality features may specifically include humanoid key points and / or humanoid quality scores. Subsequently, based on the extracted humanoid quality features, the humanoid images are filtered, so as to eliminate the humanoid images with poor quality in the first humanoid trajectory, obtain the corresponding second humanoid trajectory, and finally construct a humanoid dataset based on the second humanoid trajectory. The present application solution proposes a complete dataset filtering mechanism, and this dataset filtering mechanism requires little manual intervention and can obtain a high-quality humanoid dataset at a low labor cost.
[0018] It can be understood that the beneficial effects of the second to fifth aspects can refer to the relevant descriptions in the first aspect, and will not be elaborated here. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart of the implementation of the dataset construction method provided by the embodiment of the present application;
[0021] Figure 2 It is a structural block diagram of the dataset construction device provided by the embodiment of the present application;
[0022] Figure 3It is a schematic structural diagram of the intelligent device provided by the embodiment of the present application. Detailed implementation manners
[0023] The embodiments of the technical solution of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solution of the present application more clearly, and therefore are only examples and cannot be used to limit the protection scope of the present application.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.
[0025] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary and secondary relationship of the indicated technical features.
[0026] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0027] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0028] In the description of the embodiments of the present application, the term "plurality" refers to two or more (including two), unless otherwise specifically limited.
[0029] The embodiment of the present application proposes a method for constructing a data set. Among them, this data construction method can be applied to any intelligent device capable of performing data processing operations, and the embodiment of the present application does not limit the type of this intelligent device. Please refer to Figure 1 , Figure 1 The implementation process of this data set construction method is given and is specifically described in detail as follows:
[0030] Step 101, perform human figure detection and tracking on the surveillance video to obtain a first human figure trajectory.
[0031] The surveillance video can come from cameras in multiple different locations, and each camera has a different viewing angle and monitoring area, thereby enriching the diversity of the surveillance scene, so that the multiple surveillance videos collected can contain human targets. Among them, human targets may include but are not limited to crowd activities and the movement of a single person. Just as an example, in the scene of a shopping mall, multiple cameras can be installed. For example, the first camera is located at the entrance of the shopping mall, the second camera is located in front of the shopping mall's elevator, and the third camera is located next to the shopping mall's escalator, etc., which will not be repeated here; each camera has a different viewing angle, so that multiple surveillance videos in the scene of the shopping mall can be collected and imported into smart devices for subsequent processing.
[0032] For the collected surveillance video, the smart device can use the target detection algorithm to detect human figures. In some examples, the target detection algorithm includes YOLO (You Only Look Once), Single Shot Multibox Detector (SSD) or Faster R-CNN, etc. The embodiment of the present application does not limit the target detection algorithm. Through the target detection algorithm, human figures can be accurately detected from the image frames of the surveillance video.
[0033] The smart device can use a humanoid tracking algorithm to track the detected humanoids and assign a unique ID to each humanoid. In some examples, the humanoid tracking algorithm can be an OC-SORT (Online Multiple Object Tracking with SORT) algorithm, which is implemented based on Kalman filtering and the Hungarian algorithm.
[0034] Of course, after using the target detection algorithm to detect human figures, before using the human figure tracking algorithm to track the detected human figures, the smart device can also use the human figure filtering algorithm to perform preliminary human figure filtering, taking into account possible false detection or repeated detection due to light, background changes, or target posture. In some examples, the human figure filtering algorithm can determine whether the detection result meets the human figure conditions from the aspects of size, shape, and motion trajectory, thereby filtering out the detection results that obviously do not meet the human figure conditions in terms of space or motion behavior.
[0035] Through the above steps, the intelligent device can determine the coordinates of the human figures contained in each frame of the surveillance video and the human figure ID to which the human figure belongs. It can be understood that the human figure ID is actually equivalent to the trajectory ID. After the intelligent device crops the human figure image based on the coordinates of the human figure, classifies the human figure image based on the human figure ID, and sorts the classified human figure images in chronological order, multiple human figure trajectories can be obtained, and each human figure trajectory can be stored separately. For ease of description, this human figure trajectory is denoted as the first human figure trajectory. That is, each first human figure trajectory includes the following data: multiple frames of human figure images that belong to the same human figure ID and are sorted in time sequence.
[0036] Step 102, extract the human figure quality features of each frame of the human figure image.
[0037] In the embodiments of the present application, the following two types of human figure quality features are proposed, namely: human figure key points and human figure quality scores. According to actual requirements, the intelligent device can only extract the human figure key points of each frame of the human figure image, or only extract the human figure quality scores of each frame of the human figure image, or extract both the human figure key points and the human figure quality scores of each frame of the human figure image. The embodiments of the present application do not make any limitations in this regard.
[0038] Among them, for any human figure image, its human figure key points can be specifically extracted in the following way: extract the human figure key points through a key point detection model. In some examples, the key point detection model can specifically be the YOLOv8-pose model, which is a human pose estimation model extended based on the YOLOv8 object detection framework and can detect 17 key points of the human body in the image to be detected (such as the human figure image in the embodiments of the present application), namely the head, shoulders, elbows, wrists, hips, knees, and ankles, etc., which will not be elaborated here.
[0039] Among them, for any human figure image, its human figure quality score can be specifically extracted in the following way: determine multiple scoring dimensions, including but not limited to: clarity dimension, brightness dimension, pose dimension, and / or size dimension, etc.; then calculate the score under each scoring dimension to obtain the final human figure quality score.
[0040] Specifically, in the clarity dimension, the intelligent device can evaluate and obtain a clarity score. The process is briefly described as follows: First, the intelligent device can normalize the humanoid image to ensure that differences in the size and brightness of different humanoid images do not affect the calculation results. Then, the intelligent device can use the Sobel operator or other edge detection operators to calculate the gradient map of the normalized humanoid image to represent the image edges of the humanoid image. Next, the intelligent device can calculate the average gradient of the humanoid image based on the gradient map, and this average gradient reflects the degree of change of the humanoid image. Finally, the intelligent device can calculate the clarity score through this average gradient. Generally speaking, the larger the average gradient, the clearer the edges of the humanoid image, and the higher its corresponding clarity score.
[0041] Specifically, in the brightness dimension, the intelligent device can evaluate and obtain a brightness score. The process is briefly described as follows: The intelligent device can calculate the brightness of each pixel of the humanoid image, specifically by calculating based on the pixel values of each color channel of the pixel (which can be denoted as R, G, and B respectively) and a preset brightness calculation formula. Only as an example, the brightness calculation formula is specifically: Y = x1 * R + x2 * G + x3 * B. Among them, x1, x2, and x3 are preset coefficients respectively. Since the human eye is most sensitive to green, less sensitive to red, and least sensitive to blue, the above coefficients can be set based on this principle, and no specific limitation is made here. Then, the average value of the brightness of all pixels in the humanoid image is obtained to get the overall brightness of the humanoid image. The intelligent device can directly use this overall brightness as the brightness score of the humanoid image. That is, the intelligent device can determine the brightness score of the humanoid image based on the pixel values of each color channel of the humanoid image and the preset brightness calculation formula.
[0042] Specifically, in the pose dimension, the intelligent device can evaluate and obtain a pose score. The process is briefly described as follows: A pose detection model is pre-trained. The input of this pose detection model is a humanoid image, and the output has two items, namely: the perspective probability of this humanoid image, and the truncation degree probability of this humanoid image. In some examples, the pose detection model divides the perspective into the following three categories: front, side, and back. Then the perspective probabilities it outputs include: the probability that the humanoid image belongs to the front, the probability that it belongs to the side, and the probability that it belongs to the back, and the sum of these three probabilities is 1. In other examples, the pose detection model divides the truncation degree into the following four categories: complete, truncated by 1 / 3 (that is, 2 / 3 of the humanoid remains), truncated by 1 / 2 (that is, 1 / 2 of the humanoid remains), and truncated by more than 1 / 2 (that is, less than 1 / 2 of the humanoid remains). Then the truncation degree probabilities it outputs include: the probability that the humanoid image is complete, the probability that the humanoid image is truncated by 1 / 3, the probability that the humanoid image is truncated by 1 / 2, and the probability that the humanoid image is truncated by more than 1 / 2, and the sum of the four probabilities is 1. Thus, based on the pose detection model, the intelligent device can obtain the perspective probability and truncation degree probability of the humanoid image; the intelligent device can then determine the pose score of this humanoid image according to a preset pose calculation formula. Among them, the pose score specifically includes: a perspective score, and a truncation score.
[0043] Only as an example, the pose calculation formula can be specifically as follows:
[0044] Perspective score = a * P1 + b * P2 + c * P3
[0045] Truncation score = d * P4 + e * P5 + f * P6 + g * P7
[0046] Among them, P1 is the probability that the humanoid image belongs to the front, P2 is the probability that the humanoid image belongs to the side, P3 is the probability that the humanoid image belongs to the back; P4 is the probability that the humanoid image is complete, P5 is the probability that the humanoid image is truncated by 1 / 3, P6 is the probability that the humanoid image is truncated by 1 / 2, and P7 is the probability that the humanoid image is truncated by more than 1 / 2. It can be understood that P1 to P7 are all the outputs of the pose detection model.
[0047] Among them, a, b, c, d, e, and f are respectively preset different weight values. It can be understood that the quality of a front and complete humanoid is relatively high. Therefore, it can be set that a > b > c and d > e > f > g. Only as an example, a + b + c = 1, d + e + f + g = 2, a can be 0.5, b can be 0.3, c can be 0.2, d can be 1, e can be 0.6, f can be 0.3, and g can be 0.1.
[0048] Specifically, in the dimension of size, the intelligent device can evaluate and obtain a size score. The process is briefly described as follows: determining the size score of the humanoid image based on the size of the humanoid image. In some examples, the intelligent device can calculate the area of the humanoid image based on the size of the humanoid image, and then calculate the proportion of the area of the humanoid image in the original image (i.e., the original entire image, specifically the image frame of the surveillance video). This proportion of the area can be used as the size score. Among them, the size of the original entire image is known and fixed; that is, the area of the original entire image is known and fixed.
[0049] When there are more than two scoring dimensions, the intelligent device can further set corresponding weights for the scores in each scoring dimension, so as to obtain the humanoid quality score of the humanoid image. Only as an example, when the scoring dimensions include the clarity dimension, the brightness dimension, the pose dimension, and the size dimension, the humanoid quality score can be determined by the following formula:
[0050] Humanoid quality score = A * (viewpoint score + truncation score) + B * clarity score + C * size score + D * brightness score
[0051] Among them, A, B, C, and D are the weights corresponding to different scoring dimensions respectively. For example, A can be 0.5, B can be 0.3, C can be 0.15, and D can be 0.05. The embodiments of the present application do not limit the values of these weights.
[0052] Of course, the intelligent device can also set more or fewer scoring dimensions and adaptively adjust the above formula for calculating the humanoid quality score. The embodiments of the present application do not limit this.
[0053] In some embodiments, before extracting the humanoid quality score of the humanoid image, the humanoid image can also be processed by a humanoid segmentation model first, so as to remove the influence of the background, so that the obtained humanoid quality score can more accurately reflect the quality of the humanoid image.
[0054] Step 103, filtering multiple frames of humanoid images based on humanoid quality features to obtain a second humanoid trajectory.
[0055] As previously mentioned, the humanoid quality features extracted by the intelligent device can include humanoid key points or humanoid quality scores. For different humanoid quality features, the intelligent device can adopt different filtering methods.
[0056] Specifically, when the humanoid mass feature includes humanoid key points, the intelligent device can determine whether to eliminate the humanoid image based on the following process: count the number of humanoid key points included in the humanoid image. If the number is less than a preset number threshold, the humanoid image is considered a low-quality image, and thus the humanoid image can be eliminated. Among them, the setting of the number threshold is usually related to the integrity of the humanoid. It can be understood that 17 key points are the ideal number for complete human body recognition. When the target is occluded or truncated, the number of detected humanoid key points will be significantly less than 17. Based on this, it can be known that if the number of detected humanoid key points is less than the preset number threshold, it indicates that the humanoid may be severely blocked and cannot provide sufficient information for accurate subsequent processing. In some examples, the number threshold can be set to 8. Of course, it can also be changed to other values according to actual needs, and the embodiments of the present application do not limit this.
[0057] Specifically, when the humanoid mass feature includes a humanoid mass score, the intelligent device can determine whether to eliminate the humanoid image based on the following process: compare the humanoid mass score of the humanoid image with a preset mass score threshold. If the humanoid mass score is less than the mass score threshold, eliminate the humanoid image. Among them, the preset mass score threshold can be set according to the actual situation and is not limited here. As described above, the humanoid mass score is a judgment of the humanoid image from each scoring dimension affecting the quality. Therefore, when the humanoid mass score is too low, it indicates that the humanoid image is a low-quality image and should be eliminated.
[0058] In some embodiments, due to the different settings of the camera angle and type, there may be a situation where the humanoid presents a horizontal tilt posture in the picture, and this kind of humanoid will have a certain impact on model training. Based on this, to further improve the quality of the humanoid data set, after filtering the humanoid images based on the humanoid key points, the intelligent device can also perform the following processing on each retained frame of the humanoid image:
[0059] Calculate the angle between a specified connection line in the humanoid image and a preset horizontal line. Among them, the value range of the angle is [0, 360°), and the horizontal line is specifically horizontal to the right. In some examples, the specified connection line can be the connection line between the shoulder key points and the hip key points among the humanoid key points of the humanoid image; or, the specified connection line can also be the connection line between the head key point of the humanoid image and the midpoint of the humanoid image. It should be noted that there may be more than one shoulder key point, hip key point, and head key point. Therefore, when determining the specified connection line, the shoulder key point can be optimized to the midpoint of the shoulder key points, the hip key point can be optimized to the midpoint of the hip key points, and the head key point can be optimized to the midpoint of the head key points, which will not be elaborated here.
[0060] In the case that the obtained included angle exceeds the preset included angle range, the intelligent device can correct the humanoid image based on the included angle. It can be understood that for a generally upright humanoid, the included angle obtained by the above method should be about 90°. Therefore, the included angle range can be set to [45°, 135°]. Correspondingly, if the obtained included angle is within the range of [0, 45°) or (315°, 360°), the intelligent device can specifically rotate the humanoid image clockwise by 90° for correction processing; if the obtained included angle is within the range of (135°, 225°], the intelligent device can specifically rotate the humanoid image counterclockwise by 90° for correction processing; if the obtained included angle is within the range of (225°, 315°], the intelligent device can specifically rotate the humanoid image counterclockwise / clockwise by 180° for correction processing. It can be understood that the corrected humanoid image is upright or nearly upright, which is more conducive to the training of the model. In the case where more accurate correction results are desired, the intelligent device can use affine transformation or other technical means for correction, which is not limited here.
[0061] In some embodiments, to save the computing power and resource consumption of the intelligent device, the intelligent device can first extract the humanoid key points of the humanoid image, and perform primary filtering on the humanoid image based on the humanoid key points; then correct the humanoid image retained after the primary filtering based on the key points; finally, extract the humanoid quality score of the humanoid image after the primary filtering and correction, and perform secondary filtering on the humanoid image based on the humanoid quality score, so as to obtain the final second humanoid trajectory.
[0062] It can be understood that compared with the first humanoid trajectory corresponding to the same humanoid ID, the difference of the second humanoid trajectory is that the second humanoid trajectory generally no longer contains humanoid images with too low quality. And overall, the number of the second humanoid trajectories is generally equal to the number of the first humanoid trajectories; of course, in extremely rare cases, the number of the second humanoid trajectories may also be less than the number of the first humanoid trajectories, which may be caused by all the humanoid images in a certain first humanoid trajectory being excluded during filtering.
[0063] Step 104, constructing a humanoid dataset based on the second humanoid trajectory.
[0064] As described above, the obtained second humanoid trajectory generally no longer contains humanoid images with too low quality, so a high-quality humanoid dataset can be constructed based on the second humanoid trajectory. However, considering that there may be situations of humanoid duplication, different humanoids being misidentified as the same humanoid, or the same humanoid being misidentified as different humanoids in the second humanoid trajectory, the intelligent device can also first perform data reorganization on each second humanoid trajectory, and then construct a humanoid dataset based on the reorganized second humanoid trajectory, so as to make the humanoid dataset better. In some embodiments, the data reorganization may include but is not limited to deduplication, splitting, and / or merging.
[0065] Specifically, for any second humanoid trajectory, deduplication can be performed in the following manner:
[0066] Calculate the similarity between the first humanoid image and the second humanoid image in the second humanoid trajectory.
[0067] Specifically, the intelligent device can traverse the non-first frames of the second humanoid trajectory in sequence, determine the currently traversed humanoid image as the first humanoid image, and determine the N adjacent frames of humanoid images before the first humanoid image as the second humanoid images; that is, the first humanoid image refers to any non-first frame humanoid image in the second humanoid trajectory, and the second humanoid images refer to the N frames of humanoid images that are adjacent to the first humanoid image in time sequence and before the first humanoid image in the second humanoid trajectory, where N is a positive integer. Only as an example, N can be 5 or other values, which is not limited here. Of course, if there are less than N frames of humanoid images before the first humanoid image, all the humanoid images before the first humanoid image can be directly determined as the second humanoid images.
[0068] Specifically, the calculation method of the similarity is as follows: Through a preset person re-identification algorithm, extract the re-identification (ReID) features of the first humanoid image and the second humanoid images; calculate the cosine similarity of the re-identification features to obtain the similarity between the first humanoid image and the second humanoid images. It can be understood that when N is equal to 1, the similarity between the first humanoid image and the second humanoid image is the cosine similarity of the re-identification features of the first humanoid image and the second humanoid image; when N is greater than 1, the similarity between the first humanoid image and the second humanoid images specifically refers to the mean value of the cosine similarities of the re-identification features of the first humanoid image and each of the second humanoid images.
[0069] When the obtained similarity is greater than the preset first similarity threshold, it is considered that the first humanoid image is too similar to the previous humanoid images in the second humanoid trajectory, so the first humanoid image is of little significance for model training. The intelligent device can directly remove the first humanoid image from the second humanoid trajectory. Thus, deduplication of the second humanoid trajectory is achieved, which can reduce the redundancy of the subsequent humanoid dataset.
[0070] On this basis, the intelligent device can also perform splitting based on the similarity between the first humanoid image and the second humanoid images. Specifically, when the obtained similarity is less than the preset second similarity threshold, it is considered that the first humanoid image is too dissimilar to the previous humanoid images in the second humanoid trajectory, and there is a possibility that different humanoids are misidentified as the same humanoid; at this time, the intelligent device does not directly remove the first humanoid image, but determines the first humanoid image as an interval point. Among them, the second similarity threshold is less than the first similarity threshold. Only as an example, the first similarity threshold can be 0.9, and the second similarity threshold can be 0.4.
[0071] After traversing the second human-shaped trajectory, that is, each non-first frame in the second human-shaped trajectory is determined to be the first human-shaped image, and the similarity with the corresponding second human-shaped image is calculated, the intelligent device can determine all the interval points in the second human-shaped trajectory. At this time, the intelligent device can split the second human-shaped trajectory based on all the interval points, and the splitting process can be: for two adjacent interval points, if the number of frames of the human-shaped image between the two interval points reaches the preset frame number threshold, then from the previous interval point to the previous frame of the human-shaped image at the next interval point, a new second human-shaped trajectory can be split (the new second human-shaped trajectory will correspond to a new human-shaped ID). On the contrary, if the number of frames of the human-shaped image between the two interval points is less than the frame number threshold, the intelligent device can remove the human-shaped image from the previous interval point to the previous frame of the human-shaped image at the next interval point.
[0072] Specifically, all the second humanoid trajectories obtained can be merged in the following way:
[0073] The smart device calculates the average re-identification features of each second human-shaped trajectory respectively. The process is briefly described as follows: for any second human-shaped trajectory, the re-identification features of each frame of human-shaped image in the second human-shaped trajectory are first calculated, and then the average is calculated to obtain the average re-identification features of the second human-shaped trajectory. Afterwards, the smart device clusters each second human-shaped trajectory based on the average re-identification features of each second human-shaped trajectory. In some examples, the clustering algorithm used can be a density-based spatial clustering method of noise (Density-Based Spatial Clustering of Applications with Noise, DBSCAN), etc., which is not limited here. Finally, according to the clustering results, two or more second human-shaped trajectories in the same class are merged to solve the problem of the same human figure being mistaken for different human figures.
[0074] It can be understood that the above deduplication, splitting and merging process can be repeated until all the second human-shaped trajectories obtained cannot be deduplicated, split and merged any more, which will not be described in detail here. At this point, all the second human-shaped trajectories obtained have neither redundancy nor misconnection, and a higher quality human-shaped dataset can be constructed.
[0075] As can be seen from the above, after obtaining the first humanoid trajectory through humanoid detection and tracking of the surveillance video in the embodiments of the present application, the humanoid quality features of each frame of humanoid image in the first humanoid trajectory will be extracted first. According to the actual requirements of the application scenario, the humanoid quality features may specifically include humanoid key points and / or humanoid quality scores. Subsequently, based on the extracted humanoid quality features, the humanoid images will be filtered, so as to eliminate the humanoid images with poor quality in the first humanoid trajectory, obtain the corresponding second humanoid trajectory, and finally construct a humanoid data set based on the second humanoid trajectory. The solution of the present application proposes a complete data set filtering mechanism, and this data set filtering mechanism does not require too much manual intervention and can obtain a high-quality humanoid data set at a low labor cost.
[0076] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0077] Corresponding to the data set construction method provided above, the embodiments of the present application also provide a data set construction device. Please refer to Figure 2 , the data set construction device 2 in the embodiments of the present application includes:
[0078] A detection module 201, configured to perform humanoid detection and tracking on a surveillance video to obtain a first humanoid trajectory, where the first humanoid trajectory includes multiple frames of humanoid images;
[0079] An extraction module 202, configured to extract humanoid quality features of each frame of humanoid image, where the humanoid quality features include: humanoid key points and / or humanoid quality scores;
[0080] A filtering module 203, configured to filter multiple frames of humanoid images based on the humanoid quality features to obtain a second humanoid trajectory;
[0081] A construction module 204, configured to construct a humanoid data set based on the second humanoid trajectory.
[0082] In some embodiments, when the humanoid quality features include humanoid key points, the filtering module 203 includes:
[0083] A statistics unit, configured to count the number of humanoid key points included in each frame of humanoid image;
[0084] A first elimination unit, configured to eliminate the humanoid image when the number is less than a preset number threshold.
[0085] In some embodiments, the data set construction device 2 further includes:
[0086] A calculation module, configured to calculate, for each retained humanoid image, an angle between a specified connection line in the humanoid image and a preset horizontal line, where the specified connection line is a connection line between a shoulder key point and a hip key point among the humanoid key points of the humanoid image;
[0087] A correction module, configured to correct the humanoid image based on the angle in the case where the angle exceeds a preset angle range.
[0088] In some embodiments, when the humanoid quality feature includes a humanoid quality score, the filtering module 203 includes:
[0089] A comparison unit, configured to compare the humanoid quality score of each humanoid image with a preset quality score threshold;
[0090] A second elimination unit, configured to eliminate the humanoid image in the case where the humanoid quality score is less than the quality score threshold.
[0091] In some embodiments, when the humanoid quality feature includes a humanoid quality score, the extraction module 202 includes:
[0092] A first determination unit, configured to perform gradient calculation on each humanoid image based on a preset gradient operator to determine the clarity score of the humanoid image;
[0093] A second determination unit, configured to determine the brightness score of the humanoid image based on the pixel values of each color channel of the humanoid image and a preset brightness calculation formula;
[0094] A third determination unit, configured to determine the pose score of the humanoid image based on a trained pose detection model and a preset pose calculation formula;
[0095] A fourth determination unit, configured to determine the size score of the humanoid image based on the size of the humanoid image;
[0096] A fifth determination unit, configured to determine the humanoid quality score of the humanoid image according to the clarity score, the brightness score, the pose score, and / or the size score.
[0097] In some embodiments, there are multiple second humanoid trajectories; the construction module 204 includes:
[0098] A recombination unit, configured to perform data recombination on each of the second humanoid trajectories;
[0099] A construction unit, configured to construct a humanoid data set based on the recombined second humanoid trajectories.
[0100] In some embodiments, the recombination unit includes:
[0101] A first computing subunit, configured to calculate, for each second humanoid trajectory, the similarity between a first humanoid image and a second humanoid image in the second humanoid trajectory, where the first humanoid image is a humanoid image of any non-first frame in the second humanoid trajectory, and the second humanoid image is a humanoid image that is adjacent to the first humanoid image in time sequence and before the first humanoid image in the second humanoid trajectory;
[0102] A deduplication subunit, configured to remove the first humanoid image from the second humanoid trajectory when the similarity is greater than a preset first similarity threshold.
[0103] In some embodiments, the recombination unit further includes:
[0104] A determination subunit, configured to determine the first humanoid image as an interval point when the similarity is less than a preset second similarity threshold, where the first similarity threshold is greater than the second similarity threshold;
[0105] A splitting subunit, configured to split the second humanoid trajectory based on all interval points to obtain at least two new second humanoid trajectories when all interval points in the second humanoid trajectory are determined.
[0106] In some embodiments, the recombination unit includes:
[0107] A second computing subunit, configured to calculate the average re-identification feature of each second humanoid trajectory respectively;
[0108] A clustering subunit, configured to cluster each second humanoid trajectory based on the average re-identification feature;
[0109] A merging subunit, configured to merge two or more second humanoid trajectories in the same class according to the clustering result.
[0110] As can be seen from the above, after obtaining the first humanoid trajectory through humanoid detection and tracking of the surveillance video in the embodiments of the present application, the humanoid quality features of each frame of humanoid image in the first humanoid trajectory will be extracted first. According to the actual requirements of the application scenario, the humanoid quality features may specifically include humanoid key points and / or humanoid quality scores. Subsequently, based on the extracted humanoid quality features, the humanoid images will be filtered, so as to remove the humanoid images with poor quality in the first humanoid trajectory, obtain the corresponding second humanoid trajectory, and finally construct a humanoid dataset based on the second humanoid trajectory. The solution of the present application proposes a complete dataset filtering mechanism, and this dataset filtering mechanism requires little manual intervention and can obtain a high-quality humanoid dataset at a low labor cost.
[0111] Corresponding to the dataset construction method provided above, the embodiments of the present application further provide an intelligent device. Please refer to Figure 3, the intelligent device 3 in the embodiments of the present application includes: a memory 301, one or more processors 302 ( Figure 3 only one is shown in the figure), and a computer program stored on the memory 301 and executable on the processor. Among them: the memory 301 is used to store software programs and modules, and the processor 302 executes various functional applications and data processing by running the software programs and units stored in the memory 301 to obtain the resources corresponding to the above preset events. Specifically, when the processor 302 runs the above computer program stored in the memory 301, the following steps are implemented:
[0112] Perform human detection and tracking on the monitoring video to obtain the first human trajectory, and the first human trajectory includes multiple frames of human images;
[0113] Extract the human quality features of each frame of human image, and the human quality features include: human key points and / or human quality scores;
[0114] Filter the multiple frames of human images based on the human quality features to obtain the second human trajectory;
[0115] Construct a human dataset based on the second human trajectory.
[0116] Assume the above is the first possible implementation manner. Then, in the second possible implementation manner provided based on the first possible implementation manner, when the human quality features include human key points, filtering the multiple frames of human images based on the human quality features includes:
[0117] For each frame of human image, count the number of human key points included in the human image;
[0118] If the number is less than the preset number threshold, eliminate the human image.
[0119] In the third possible implementation manner provided based on the above second possible implementation manner, after filtering the multiple frames of human images based on the human quality features, when the processor 302 runs the above computer program stored in the memory 301, the following steps are implemented:
[0120] For each frame of human image that has been retained, calculate the angle between the specified connection line in the human image and the preset horizontal line, where the specified connection line is the connection line between the shoulder key point and the hip key point among the human key points of the human image;
[0121] If the angle exceeds the preset angle range, correct the human image based on the angle.
[0122] In a fourth possible implementation provided based on the above first possible implementation, when the humanoid quality feature includes a humanoid quality score, filtering multiple frames of humanoid images based on the humanoid quality feature includes:
[0123] For each frame of humanoid image, compare the humanoid quality score of the humanoid image with a preset quality score threshold;
[0124] When the humanoid quality score is less than the quality score threshold, discard the humanoid image.
[0125] In a fifth possible implementation provided based on the above first possible implementation, when the humanoid quality feature includes a humanoid quality score, extracting the humanoid quality feature of each frame of humanoid image includes:
[0126] For each frame of humanoid image, perform gradient calculation on the humanoid image based on a preset gradient operator to determine the clarity score of the humanoid image;
[0127] Based on the pixel values of each color channel of the humanoid image and a preset brightness calculation formula, determine the brightness score of the humanoid image;
[0128] Based on a trained pose detection model and a preset pose calculation formula, determine the pose score of the humanoid image;
[0129] Based on the size of the humanoid image, determine the size score of the humanoid image;
[0130] According to the clarity score, brightness score, pose score, and / or size score, determine the humanoid quality score of the humanoid image.
[0131] In a sixth possible implementation provided based on the above first possible implementation, or the above second possible implementation, or the above third possible implementation, or the above fourth possible implementation, or the above fifth possible implementation, there are multiple second humanoid trajectories; constructing a humanoid data set based on the second humanoid trajectories includes:
[0132] Perform data recombination on each of the second humanoid trajectories;
[0133] Construct a humanoid data set based on the recombined second humanoid trajectories.
[0134] In a seventh possible implementation provided based on the above sixth possible implementation, performing data recombination on each of the second humanoid trajectories includes:
[0135] For each second humanoid trajectory, calculate the similarity between the first humanoid image and the second humanoid image in the second humanoid trajectory, where the first humanoid image is any humanoid image in the second humanoid trajectory that is not the first frame, and the second humanoid image is the humanoid image in the second humanoid trajectory that is adjacent to the first humanoid image in time sequence and before the first humanoid image;
[0136] When the similarity is greater than a preset first similarity threshold, remove the first humanoid image from the second humanoid trajectory.
[0137] In an eighth possible implementation manner provided based on the above seventh possible implementation manner, after calculating the similarity between the first humanoid image and the second humanoid image in the second humanoid trajectory, the processor 302 implements the following steps when running the above computer program stored in the memory 301:
[0138] When the similarity is less than a preset second similarity threshold, determine the first humanoid image as an interval point, where the first similarity threshold is greater than the second similarity threshold;
[0139] When all interval points in the second humanoid trajectory are determined, split the second humanoid trajectory based on all interval points to obtain at least two new second humanoid trajectories.
[0140] In a ninth possible implementation manner provided based on the above sixth possible implementation manner, data reorganization of each second humanoid trajectory includes:
[0141] Calculate the average re-identification feature of each second humanoid trajectory respectively;
[0142] Cluster each second humanoid trajectory based on the average re-identification feature;
[0143] According to the clustering result, merge two or more second humanoid trajectories in the same class.
[0144] It should be understood that in the embodiments of the present application, the so-called processor 302 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0145] The memory 301 may include a read-only memory and a random access memory, and provide instructions and data to the processor 302. A part or all of the memory 301 may also include a non-volatile random access memory. For example, the memory 301 may also store information about the device type.
[0146] As can be seen from the above, after obtaining the first humanoid trajectory through humanoid detection and tracking of the surveillance video in the embodiments of the present application, the humanoid quality features of each frame of humanoid image in the first humanoid trajectory will be extracted first. According to the actual requirements of the application scenario, the humanoid quality features may specifically include humanoid key points and / or humanoid quality scores. Subsequently, based on the extracted humanoid quality features, the humanoid images will be filtered, so as to remove the humanoid images with poor quality in the first humanoid trajectory, obtain the corresponding second humanoid trajectory, and finally construct a humanoid dataset based on the second humanoid trajectory. The solution of the present application proposes a complete dataset filtering mechanism, and this dataset filtering mechanism does not require too much manual intervention and can obtain a high-quality humanoid dataset at a low labor cost.
[0147] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments and will not be described in detail here.
[0148] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0149] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of external device software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0150] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the above-mentioned division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0151] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0152] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, it can also be completed by a computer program instructing the relevant hardware. The above-mentioned computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the above-mentioned computer program includes computer program code, and the above-mentioned computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The above-mentioned computer-readable storage medium can include: any entity or device that can carry the above-mentioned computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer-readable memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the above-mentioned computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0153] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for constructing a data set, characterized in that: include: Performing human figure detection and tracking on the surveillance video to obtain a first human figure trajectory, wherein the first human figure trajectory includes multiple frames of human figure images; Extracting human figure quality features of the human figure image of each frame, wherein the human figure quality features include: human figure key points and / or human figure quality scores; Filtering multiple frames of the human-shaped images based on the human-shaped quality feature to obtain a second human-shaped trajectory; A human figure dataset is constructed based on the second human figure trajectory.
2. The method for constructing a data set according to claim 1, wherein: In the case where the human shape quality feature includes human shape key points, filtering multiple frames of human shape images based on the human shape quality feature includes: For each frame of the human image, counting the number of the human key points contained in the human image; When the number is less than a preset number threshold, the human-shaped image is discarded.
3. The method for constructing a data set according to claim 2, wherein: After filtering the multiple frames of human-shaped images based on the human-shaped quality features, the data set construction method further includes: For each frame of the human-shaped image that has been retained, calculating the angle between a designated line in the human-shaped image and a preset horizontal line, wherein the designated line is a line between a shoulder key point and a hip key point among the human-shaped key points of the human-shaped image; When the angle exceeds a preset angle range, the human image is corrected based on the angle.
4. The method for constructing a data set according to claim 1, wherein: In the case where the human shape quality feature includes a human shape quality score, filtering the plurality of frames of human shape images based on the human shape quality feature includes: For each frame of the human-shaped image, comparing the human-shaped quality score of the human-shaped image with a preset quality score threshold; When the human figure quality score is less than the quality score threshold, the human figure image is discarded.
5. The method for constructing a data set according to claim 1, wherein: In the case where the human figure quality feature includes a human figure quality score, extracting the human figure quality feature of each frame of the human figure image includes: For each frame of the human-shaped image, performing gradient calculation on the human-shaped image based on a preset gradient operator to determine a clarity score of the human-shaped image; Determining a brightness score of the human-shaped image based on the pixel values of each color channel of the human-shaped image and a preset brightness calculation formula; Determining a posture score of the human image based on a trained posture detection model and a preset posture calculation formula; determining a size score for the human-shaped image based on a size of the human-shaped image; A human figure quality score of the human figure image is determined according to the clarity score, the brightness score, the posture score and / or the size score.
6. The method for constructing a data set according to any one of claims 1 to 5, characterized in that: There are multiple second human-shaped trajectories; and constructing a human-shaped dataset based on the second human-shaped trajectories includes: Reorganizing the data of each of the second humanoid trajectories; The human figure dataset is constructed based on the reorganized second human figure trajectory.
7. The method for constructing a data set according to claim 6, wherein: The data reorganization of each of the second humanoid trajectories includes: For each second human-shaped trajectory, calculating the similarity between a first human-shaped image and a second human-shaped image in the second human-shaped trajectory, wherein the first human-shaped image is any non-first-frame human-shaped image in the second human-shaped trajectory, and the second human-shaped image is a human-shaped image in the second human-shaped trajectory that is before the first human-shaped image and is adjacent to the first human-shaped image in time sequence; When the similarity is greater than a preset first similarity threshold, the first human-shaped image is removed from the second human-shaped trajectory.
8. The method for constructing a data set according to claim 7, wherein: After calculating the similarity between the first human-shaped image and the second human-shaped image in the second human-shaped trajectory, the data set construction method further includes: In the case where the similarity is less than a preset second similarity threshold, determining the first human-shaped image as an interval point, wherein the second similarity threshold is less than the first similarity threshold; When all interval points in the second human-shaped trajectory are determined, the second human-shaped trajectory is split based on all the interval points.
9. The method for constructing a data set according to claim 6, wherein: The data reorganization of each of the second humanoid trajectories includes: Calculate the average re-identification feature of each of the second human figure trajectories respectively; Clustering each of the second human-shaped trajectories based on the average re-identification feature; According to the clustering result, two or more second humanoid trajectories in the same class are merged.
10. A data set construction device, characterized in that: include: A detection module, configured to detect and track a human figure in a surveillance video to obtain a first human figure trajectory, wherein the first human figure trajectory includes multiple frames of human figure images; An extraction module is used to extract the human figure quality features of each frame of the human figure image, wherein the human figure quality features include: human figure key points and / or human figure quality scores; A filtering module, configured to filter a plurality of frames of the human-shaped images based on the human-shaped quality feature to obtain a second human-shaped trajectory; A construction module is used to construct a human figure dataset based on the second human figure trajectory.
11. An intelligent device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.
12. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Key person tracking method and device based on face recognition and tracking algorithm
CN121392931A