A method, apparatus, device and medium for identifying pedestrians
By dividing the pedestrian video sequence into sub-sequences facing and deviating from the camera, and using the European-style distance and pose similarity weighting method, the problem of low pedestrian re-identification accuracy caused by camera angle differences is solved, and efficient pedestrian recognition is achieved.
Patent Information
- Application Number
- CN202210824180.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-07-14
AI Technical Summary
The existing pedestrian re-identification technology has difficulty in posture alignment due to camera angle, height and hardware differences, which affects recognition accuracy and has a low recognition rate based on color feature.
By dividing the video sequence into incoming sequences and outgoing sequences, and dividing them into sub-sequences facing and deviating from the camera according to the coordinates of the key points of the pedestrian head, the European-style distance matching and pose similarity weighting are used to perform rough and precise matching, reducing the calculation amount and improving the recognition accuracy.
It effectively improves the accuracy and recognition rate of pedestrian recognition, reduces the calculation amount of the matching process, and improves the matching accuracy through posture correction and weighted matching.
Smart Images

Figure CN115131713B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of target detection technology, and in particular, to a method, device, equipment and medium for identifying pedestrians. Background Art
[0002] Person re-identification, also known as pedestrian re-identification, is a technology that uses computer vision technology to determine whether there is a specific pedestrian in an image or video sequence.
[0003] There have been a lot of studies on the problem of pedestrian re-identification, mainly focusing on the establishment of pedestrian correspondence between different cameras. Because different cameras have different angles, heights, and camera hardware, the pedestrian targets being compared are often unable to be aligned. Even if some alignment methods are proposed, they are relatively rough. In addition, due to the different postures and directions of pedestrians, there has been no mature solution to this problem.
[0004] The existing pedestrian re-identification technology cannot align the target's posture due to the differences in camera angles, heights and camera hardware, and the difference in posture will greatly affect the accuracy of re-identification. At this time, the pedestrian re-identification algorithm mainly relies on the color features of the target for recognition, and the recognition rate is low. Summary of the invention
[0005] The embodiments of the present invention provide a method, device, equipment and medium for identifying pedestrians, which effectively improve the accuracy of pedestrian identification.
[0006] In a first aspect, the present invention provides a method for identifying a pedestrian, comprising:
[0007] According to the motion trajectory of the pedestrian in the video sequence, the video sequence is divided into an in-sequence and an out-sequence, wherein the motion trajectory of the pedestrian in the in-sequence is from the bottom of the image to the top of the image; and the motion trajectory of the pedestrian in the out-sequence is from the top of the image to the bottom of the image;
[0008] According to the coordinates of the key points of the pedestrian's head, the in-sequence and the out-sequence are respectively divided into a sub-sequence facing the camera and a sub-sequence away from the camera; wherein, for the in-sequence, the ordinate of the key points of the pedestrian's head in the sub-sequence facing the camera is greater than the ordinate of the center point of the image, and the ordinate of the key points of the pedestrian's head in the sub-sequence away from the camera is less than the ordinate of the center point of the image; for the out-sequence, the ordinate of the key points of the pedestrian's head in the sub-sequence facing the camera is less than the ordinate of the center point of the image, and the ordinate of the key points of the pedestrian's head in the sub-sequence away from the camera is greater than the ordinate of the center point of the image;
[0009] Matching the camera-facing subsequence in the outgoing sequence with the camera-facing subsequence in the incoming sequence, and matching the camera-facing subsequence in the outgoing sequence with the camera-facing subsequence in the incoming sequence;
[0010] During the matching process, for an object to be matched in one of the subsequences to be matched, a first Euclidean distance from a key point of the head of the object to be matched to a center point of the image is determined, and a second Euclidean distance from a key point of the head of each object in a corresponding subsequence to be matched with the subsequence to be matched to a center point of the image is determined;
[0011] According to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, determining, from the corresponding subsequences matched with the subsequence to be matched, a candidate image sequence whose Euclidean distance relationship meets a preset distance requirement;
[0012] A target image sequence is determined from the candidate image sequence, wherein a target object that is the same object as the object to be matched exists in the target image sequence.
[0013] Optionally, according to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, determining, from the corresponding subsequences matched with the subsequence to be matched, a candidate image sequence whose Euclidean distance relationship meets a preset distance requirement, includes:
[0014] Calculating an absolute value of a difference between the first Euclidean distance and each second Euclidean distance, wherein the absolute value is used to represent a Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance;
[0015] Sort the obtained absolute values in ascending order, and select the target absolute values whose values are arranged in the first Q positions from the sorted absolute values; wherein Q is a positive integer;
[0016] The image sequence corresponding to the target absolute value is used as a candidate image sequence in a corresponding subsequence to be matched with the subsequence to be matched.
[0017] Optionally, determining a target image sequence corresponding to a target candidate object that is the same object as the object to be matched includes:
[0018] Determine the cosine distance and Euclidean distance relationship between the image frame corresponding to the object to be matched and each candidate image sequence;
[0019] Determining a matching weight of each candidate image sequence according to the Euclidean distance relationship;
[0020] Calculate the matching distance between the candidate objects in each candidate image sequence and the object to be matched in the sequence to be matched according to the matching weight and the cosine distance, and use the candidate image sequence with the smallest matching distance as the target image sequence.
[0021] Optionally, calculating the matching distance between the candidate objects in each candidate image sequence and the object to be matched in the sequence to be matched includes:
[0022] Calculate the matching distance according to the following formula:
[0023]
[0024] where S i represents the cosine distance between the sequence to be matched and the i-th frame candidate image sequence; represents the matching weight; FD i represents the Euclidean distance relationship between the sequence to be matched and the i-th frame candidate image sequence; j is the j-th candidate object in each candidate image sequence; m represents the sum of the number of video frames of the subsequence facing the camera and the subsequence away from the camera.
[0025] Optionally, the pedestrian head key point coordinates are the corresponding head key point coordinates when the pedestrian image is rotated so that the line connecting the two shoulder key points after rotation is perpendicular to the line connecting the image center and the head key point.
[0026] Optionally, rotate the image according to the following formula:
[0027]
[0028] where tx = x - Cx, ty = y - Cy; θ = θ2 - θ1, representing the angle to be rotated; θ1 = arctan((Sy r -Sy l ) / (Sx r -Sx l )) represents the angle between the line connecting the two shoulder key points and the horizontal line; θ2 = arctan(-(Hy - Cy) / (Hx - Cx)) represents the angle between the perpendicular line connecting the image center and the head key point and the horizontal line; (Hx, Hy) represents the head key point coordinates before rotation; (Sx l , Sy l ) and (Sx r , Sy r(Cx0,Cy0) and (Cx1,Cy1) respectively represent the key point coordinates of the left and right shoulders before image rotation; (Cx,Cy) represents the image center point coordinates; (x,y) represents any point in the image corresponding to the head and shoulder detection rectangle; (x′,y′) represents the point obtained by rotating (x,y) around the image center; tx represents the distance from the abscissa of any point in the image corresponding to the head and shoulder detection rectangle to the abscissa of the image center point; ty represents the distance from the ordinate of any point in the image corresponding to the head and shoulder detection rectangle to the ordinate of the image center point.
[0029] In a second aspect, an embodiment of the present invention further provides a pedestrian recognition device, including:
[0030] A video sequence division module, configured to divide a video sequence into an in-sequence and an out-sequence according to the movement trajectory of a pedestrian in the video sequence, where the movement trajectory of the pedestrian in the in-sequence is from the bottom of the image to the top of the image; the movement trajectory of the pedestrian in the out-sequence is from the top of the image to the bottom of the image;
[0031] A subsequence division module, configured to divide the in-sequence and the out-sequence into a camera-facing subsequence and a camera-away subsequence respectively according to the key point coordinates of the pedestrian's head; where, for the in-sequence, the ordinate of the key point of the head of the pedestrian in the camera-facing subsequence is greater than the ordinate of the image center point where it is located, and the ordinate of the key point of the head of the pedestrian in the camera-away subsequence is less than the ordinate of the image center point where it is located; for the out-sequence, the ordinate of the key point of the head of the pedestrian in the camera-facing subsequence is less than the ordinate of the image center point where it is located, and the ordinate of the key point of the head of the pedestrian in the camera-away subsequence is greater than the ordinate of the image center point where it is located;
[0032] A subsequence matching module, configured to match the camera-facing subsequence in the out-sequence with the camera-facing subsequence in the in-sequence, and match the camera-away subsequence in the out-sequence with the camera-away subsequence in the in-sequence; during the matching process, for an object to be matched in one of the subsequences to be matched, determine the first Euclidean distance from the key point of the head of the object to be matched to the image center point, and determine the second Euclidean distance from the key point of the head of each object in the corresponding subsequence that is matched with the subsequence to be matched to the image center point;
[0033] A candidate image sequence determination module, configured to determine a candidate image sequence whose Euclidean distance relationship meets a preset distance requirement from the corresponding subsequence that is matched with the subsequence to be matched according to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance;
[0034] The target image sequence determination module is configured to determine a target image sequence from the candidate image sequence, wherein the target image sequence contains a target object that is the same object as the object to be matched.
[0035] Optionally, the candidate image sequence determination module is specifically configured to:
[0036] Calculating an absolute value of a difference between the first Euclidean distance and each second Euclidean distance, wherein the absolute value is used to represent a Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance;
[0037] Sort the obtained absolute values in ascending order, and select the target absolute values whose values are arranged in the first Q positions from the sorted absolute values; wherein Q is a positive integer;
[0038] The image sequence corresponding to the target absolute value is used as a candidate image sequence in a corresponding subsequence to be matched with the subsequence to be matched.
[0039] Optionally, the target image sequence determination module is specifically configured as follows:
[0040] a distance determination unit, configured to determine a cosine distance and a Euclidean distance relationship between the image frame corresponding to the to-be-matched object and each candidate image sequence;
[0041] a matching weight determination unit, configured to determine a matching weight of each candidate image sequence according to the Euclidean distance relationship;
[0042] The target image sequence determination unit is configured to calculate the matching distance between the candidate objects in each candidate image sequence and the objects to be matched in the sequence to be matched according to the matching weight and the cosine distance, and take the candidate image sequence with the smallest matching distance as the target image sequence.
[0043] Optionally, the matching distance is calculated according to the following formula:
[0044]
[0045] Among them, S i represents the cosine distance between the sequence to be matched and the i-th frame candidate image sequence; Represents the matching weight; FD i represents the Euclidean distance relationship between the sequence to be matched and the i-th frame candidate image sequence; j is the j-th candidate object in each candidate image sequence; m represents the sum of the number of video frames facing the camera subsequence and the number of video frames away from the camera subsequence.
[0046] Optionally, the pedestrian head key point coordinates are the corresponding head key point coordinates when the pedestrian image is rotated so that the line connecting the two shoulder key points after rotation is perpendicular to the line connecting the image center and the head key point.
[0047] Optionally, rotate the image:
[0048]
[0049] where tx = x - Cx, ty = y - Cy; θ = θ2 - θ1, representing the angle to be rotated; θ1 = arctan((Sy r -Sy l ) / (Sx r -Sx l )) represents the angle between the line connecting the two shoulder key points and the horizontal line; θ2 = arctan(-(Hy - Cy) / (Hx - Cx)) represents the angle between the perpendicular line connecting the image center and the head key point and the horizontal line; (Hx, Hy) represents the head key point coordinates before rotation; (Sx l , Sy l ) and (Sx r , Sy r ) respectively represent the left and right shoulder key point coordinates in the image before rotation; (Cx, Cy) represents the image center point coordinates; (x, y) represents any point in the image corresponding to the head and shoulder detection rectangle; (x′, y′) represents the point after (x, y) is rotated around the image center; tx represents the distance from the abscissa of any point in the image corresponding to the head and shoulder detection rectangle to the abscissa of the image center point; ty represents the distance from the ordinate of any point in the image corresponding to the head and shoulder detection rectangle to the ordinate of the image center point.
[0050] In a third aspect, an embodiment of the present invention further provides a computing device, including:
[0051] A memory storing executable program code;
[0052] A processor coupled to the memory;
[0053] The processor calls the executable program code stored in the memory and executes the pedestrian recognition method provided by any embodiment of the present invention.
[0054] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the pedestrian recognition method provided by any embodiment of the present invention.
[0055] The technical solution provided by the embodiment of the present invention divides the pedestrian recognition process into a rough matching stage and a fine matching stage. In the rough matching stage, according to the movement trajectory of the pedestrian, the video sequence can be divided into an incoming sequence and an outgoing sequence, and the pedestrian sequence can be divided into subsequences facing and away from the camera. By adopting a method of matching subsequences with the same orientation, for example, matching the subsequence facing the camera in the incoming sequence with the subsequence facing the camera in the outgoing sequence, or matching the subsequence away from the camera in the incoming sequence with the subsequence away from the camera in the outgoing sequence, the amount of calculation in the matching process is effectively reduced. In the matching process, by adopting the method of the Euclidean distance from the object to be matched to the center point of the image, a candidate image sequence with a similar target posture to the object to be matched in the same subsequence can be obtained, laying the foundation for the subsequent pedestrian recognition. Compared with the method of using the color features of the pedestrian target for recognition in the prior art, the technical solution provided by this embodiment effectively improves the recognition rate of pedestrians. In the fine matching stage, by adopting a matching distance estimation method based on the weighting of the similarity of postures, the matching accuracy is further improved.
[0056] The innovative features of the embodiments of the present invention include:
[0057] 1. Rotating the pedestrian image so that the line connecting the two shoulder key points after rotation is perpendicular to the line connecting the image center and the head key points is used to correct the pedestrian's posture, thereby improving the accuracy of subsequent pedestrian re-identification. This is one of the innovative features of the embodiment of the present invention.
[0058] 2. By dividing the incoming sequence and the outgoing sequence in the video sequence into a subsequence facing the camera and a subsequence facing away from the camera, and by matching the subsequences with the same direction, the amount of calculation in the matching process is effectively reduced. This is one of the innovative points of the embodiment of the present invention.
[0059] 3. By adopting a matching distance estimation method based on weighted posture similarity, the matching accuracy is further improved, which is one of the innovations of the embodiment of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0061] Figure 1a A flowchart of a pedestrian recognition method provided in Embodiment 1 of the present invention;
[0062] Figure 1bSchematic diagrams of an image before and after rotation provided in Embodiment 1 of the present invention;
[0063] Figure 1c Schematic diagram of the distance from a target in an image to the center of the image provided in Embodiment 1 of the present invention;
[0064] Figure 1d Flowchart of a pedestrian precise matching stage provided in Embodiment 1 of the present invention;
[0065] Figure 2 Structural block diagram of a pedestrian recognition device provided in Embodiment 2 of the present invention;
[0066] Figure 3 Structural block diagram of a computing device provided in Embodiment 3 of the present invention. Detailed implementation manners
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0068] It should be noted that the terms "including" and "having" and any variations thereof in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0069] The embodiments of the present invention disclose a pedestrian recognition method, device, equipment, and medium. The following will be described in detail respectively.
[0070] Embodiment 1
[0071] Figure 1a Flowchart of a pedestrian recognition method provided in Embodiment 1 of the present invention. This method can be applied to processes such as business analysis and passenger flow analysis, especially in the estimation process of pedestrian re-identification in a monocular camera top-down scenario. This method can be completed by a pedestrian recognition device, and the device can be implemented in software and / or hardware. As Figure 1a shown, this method includes:
[0072] S110. Divide the video sequence into an in-sequence and an out-sequence according to the movement trajectory of pedestrians in the video sequence.
[0073] In this embodiment, when detecting pedestrians in a video sequence, an external bounding rectangle corresponding to a pedestrian can be obtained as the detection result. By sending the detection result image into a head and shoulder key point detection model for detection, the position coordinates of the head key point and the coordinates of two shoulder key points can be obtained. Among them, the head key point is the center point of the head, and the shoulder key points are the two endpoints of the shoulders. Among them, the head and shoulder key point detection model for detecting the head key point and the shoulder key points can be pre-trained with a large number of images marked with the positions of the head key point and the two shoulder key points, so that after the training of the detection model is completed, the association relationship between the head key point and the shoulder key points in the image and their position information is established.
[0074] In this embodiment, after determining the positions of the head key point and the shoulder key points of a pedestrian in a video sequence, a detection result sequence of the same pedestrian in the video sequence can be obtained through a tracking algorithm. Among them, the tracking algorithm can be a KCF (kernelized correlation filters, high-speed tracking based on kernelized correlation filters) algorithm, a TLD (Tracking-Learning-Detection) algorithm, a MediaFlow algorithm, etc. This embodiment does not make specific limitations on this.
[0075] In this embodiment, according to the movement trajectory of the same pedestrian, the video sequence can be divided into an entry sequence and an exit sequence. Since the camera is installed in a top view, the lens plane is almost parallel to the ground. Therefore, the camera center is the image center, and the image is basically symmetric up and down and left and right. That is, the sequence can be distinguished by the head key point coordinates and the image center point coordinates. Among them, the movement trajectory of the pedestrian in the entry sequence is from the bottom of the image to the top of the image; the movement trajectory of the pedestrian in the exit sequence is from the top of the image to the bottom of the image. For the installation of cameras on doors with different orientations, the angle of the camera can be rotated, for example, rotated by 90°, to adjust the entry and exit of the pedestrian to the trajectory movement in the vertical direction of the image. Generally, there will be no more tracking after a person enters the room. In this way, a person will get a period of tracking results before entering the room and also a period of tracking results when leaving. The method provided in this embodiment is to determine whether it is the same person based on the two tracking results.
[0076] Preferably, in the overhead camera, the changes in pedestrian posture are mainly caused by the direction of pedestrian movement, and the distance and proximity to the camera. Therefore, in the pedestrian posture alignment, posture correction can be performed first to improve the accuracy of subsequent pedestrian recognition. Specifically, the pedestrian image can be rotated according to the detection results of the pedestrian head and shoulder key points, and the pedestrian is facing or facing away from the camera after rotation. After the rotation, the line connecting the two shoulder key points is perpendicular to the line connecting the center of the image and the head key point. According to the positional relationship between the coordinates of the head key point and the coordinates of the center point of the image after the pedestrian posture correction, the detection result sequence of the same pedestrian is divided into an input sequence and an output sequence.
[0077] Specifically, the image can be rotated according to the following formula:
[0078]
[0079] Among them, tx=x-Cx, ty=y-Cy; θ=θ2-θ1, which indicates the angle to be rotated; θ1=arctan((Sy r -Sy l ) / (Sx r -Sx l )) represents the angle between the line connecting the two key points of the shoulder and the horizontal line; θ2 = arctan (-(Hy-Cy) / (Hx-Cx)) represents the angle between the vertical line connecting the image center and the key points of the head and the horizontal line; (Hx, Hy) represents the coordinates of the key points of the head before rotation; (Sx l ,Sy l ) and (Sx r ,Sy r ) respectively represent the coordinates of the left and right shoulder key points before the image is rotated; (Cx, Cy) represent the coordinates of the center point of the image; (x, y) represents any point in the image corresponding to the head and shoulder detection rectangular box; (x′, y′) represents the point after (x, y) is rotated around the center of the image; tx represents the distance from the horizontal coordinate of any point in the image corresponding to the head and shoulder detection rectangular box to the horizontal coordinate of the center point of the image; ty represents the distance from the vertical coordinate of any point in the image corresponding to the head and shoulder detection rectangular box to the vertical coordinate of the center point of the image.
[0080] Specifically, Figure 1b This is a schematic diagram of an image before rotation and after image selection provided by the first embodiment of the present invention. Figure 1b As shown in the figure, the arrow direction is the direction of the pedestrian's motion trajectory. After rotating the image, the line connecting the two shoulder key points is perpendicular to the line connecting the image center and the head key point.
[0081] S120 . Divide the incoming sequence and the outgoing sequence into a subsequence toward the camera and a subsequence away from the camera according to the coordinates of the key points of the pedestrian's head.
[0082] Among them, for the incoming sequence, the vertical coordinate of the head key point of the pedestrian in the subsequence towards the camera (Fin) is greater than the vertical coordinate of the center point of the image (Hy > Cy), and the vertical coordinate of the head key point of the pedestrian in the subsequence away from the camera (Bin) is less than the vertical coordinate of the center point of the image (Hy < Cy); for the outgoing sequence, the vertical coordinate of the head key point of the pedestrian in the subsequence towards the camera (Fout) is less than the vertical coordinate of the center point of the image (Hy < Cy), and the vertical coordinate of the head key point of the pedestrian in the subsequence away from the camera (Bout) is greater than the vertical coordinate of the center point of the image (Hy > Cy).
[0083] Specifically, the incoming sequence can be expressed as {Fin1, Fin2, …, Fin m1 , Bin1, Bin2, …, Bin m2}, and the outgoing sequence can be expressed as {Fout1, Fout2, …, Fout n1 , Bout1, Bout2, …, Bout n2}.
[0084] S130. Match the subsequence towards the camera in the outgoing sequence with the subsequence towards the camera in the incoming sequence, and match the subsequence away from the camera in the outgoing sequence with the subsequence away from the camera in the incoming sequence; during the matching process, for the object to be matched in one of the subsequences to be matched, determine the first Euclidean distance from the head key point of the object to be matched to the center point of the image, and determine the second Euclidean distance from the head key point of each object in the corresponding subsequence to be matched with the subsequence to be matched to the center point of the image.
[0085] For the re-identification process of pedestrians, it is a process of matching the tracked objects in the incoming sequence {Fin1, Fin2, …, Fin m1 , Bin1, Bin2, …, Bin m2} and the outgoing sequence {Fout1, Fout2, …, Fout n1 , Bout1, Bout2, …, Bout n2}. For example, for the sequences of the same pedestrian A that have been tracked in the subsequence towards the camera in the incoming sequence and the sequences of another same pedestrian B that have been tracked in the subsequence towards the camera in the outgoing sequence, the matching is to determine whether A and B are the same person.
[0086] Exemplarily, the first Euclidean distance from the head key point of the object to be matched in the subsequence facing the camera in the incoming sequence to the center point of the image can be determined based on the incoming sequence, and the second Euclidean distance from the head key points of each object in the subsequence facing the camera in the sequence to the center point of the image can be determined. Then, according to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, a candidate image sequence in the subsequence facing the camera in the sequence whose Euclidean distance relationship meets the preset distance requirement can be determined, where the object to be matched is an object that matches the pose of the object to be matched in the subsequence facing the camera in the incoming sequence.
[0087] Exemplarily, the first Euclidean distance from the head key point of the object to be matched in the subsequence away from the camera in the incoming sequence to the center point of the image can be determined based on the incoming sequence, and the second Euclidean distance from the head key points of each object in the subsequence away from the camera in the sequence to the center point of the image can be determined. Then, according to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, a candidate image sequence in the subsequence away from the camera in the sequence whose Euclidean distance relationship meets the preset distance requirement can be determined, where the candidate object is an object that matches the pose of the object to be matched in the subsequence away from the camera in the incoming sequence.
[0088] Exemplarily, the first Euclidean distance from the head key point of the object to be matched in the subsequence facing the camera in the outgoing sequence to the center point of the image can be determined based on the outgoing sequence, and the second Euclidean distance from the head key points of each object in the subsequence facing the camera in the incoming sequence to the center point of the image can be determined. Then, according to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, a candidate image sequence in the subsequence facing the camera in the incoming sequence whose Euclidean distance relationship meets the preset distance requirement can be determined, where the candidate object is an object that matches the pose of the object to be matched in the subsequence facing the camera in the outgoing sequence.
[0089] Exemplarily, the first Euclidean distance from the head key point of the object to be matched in the subsequence away from the camera in the outgoing sequence to the center point of the image can be determined based on the outgoing sequence, and the second Euclidean distance from the head key points of each object in the subsequence away from the camera in the incoming sequence to the center point of the image can be determined. Then, according to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, a candidate image sequence in the subsequence away from the camera in the incoming sequence whose Euclidean distance relationship meets the preset distance requirement can be determined, where the candidate object is an object that matches the pose of the object to be matched in the subsequence away from the camera in the outgoing sequence.
[0090] In this embodiment, it is assumed that in the incoming sequence, there are m1 frames of targets facing the camera and m2 frames of targets facing away from the camera. In the target outgoing sequence, there are n1 frames of targets facing the camera and n2 frames of targets facing away from the camera. In the existing pedestrian re-identification algorithms, all incoming and outgoing targets need to be matched. Each target corresponds to an image sequence obtained by tracking. If the two image sequences are traversed, (m1 + m2) × (n1 + n2) times of matching are required. If a target image is randomly selected from each sequence for matching, it may lead to inaccurate matching due to large differences in the poses of the target images or image quality problems. In the application scenario of the present invention, since there are differences between the incoming and outgoing sequences in terms of facing and facing away from the camera, the features of the images of the same target facing and facing away from the camera often have significant differences, resulting in false matching and wasting computational resources. In the embodiment of the present invention, through the different natures of the targets facing and facing away from the camera, a rough matching is performed on the target sequences to be matched, deleting unnecessary matches and saving computational resources. Accordingly, in this embodiment, it is stipulated that the targets in the incoming sequence facing the camera can only be matched with the targets in the outgoing sequence facing the camera, and the targets in the incoming sequence facing away from the camera can only be matched with the targets in the outgoing sequence facing away from the camera. In this way, the computational amount is reduced to m1 × n1 + m2 × n2 times of target matching.
[0091] It should be noted that since the camera is almost parallel to the ground, the tracking objects in the image have the property of isotropy, that is, if the pedestrians to be matched all face or face away from the camera, then among these pedestrians to be matched with similar distances to the center of the image, their poses are the closest, regardless of the angle between the target and the horizontal axis. As Figure 1c shown, although the angles between Target 1-4 and the horizontal axis are different, after the image rotation adjustment provided by this embodiment, the differences in the poses or scales of the pedestrians to be matched are mainly related to the distances from the camera. On the image, it is manifested as the Euclidean distance from the head key point to the center of the image.
[0092] S140. According to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, determine a candidate image sequence whose Euclidean distance relationship meets the preset distance requirement from the corresponding subsequence that is matched with the subsequence to be matched.
[0093] In this embodiment, for each object to be matched in the subsequence to be matched, for example, taking the incoming sequence as a reference, Q targets with the closest poses can be calculated, where Q < n1 or Q < n2. Generally, according to experience, Q can be 1 or 2. In this way, the matching calculation can be further reduced to (m1 + m2) × Q.
[0094] Specifically, taking the pedestrian Fin1 in the incoming sequence in the subsequence facing the camera as an example, it is used as the object to be matched. Determine the first Euclidean distance FinE1 from the head key point of the object to be matched to the center point of the image, and determine another corresponding subsequence for matching, that is, each object {Fout1, Fout2, Fout n1} in the outgoing sequence in the subsequence facing the camera to the second Euclidean distance {FoutE1, FoutE2, FoutE n1} from the center of the image. Among them, the first Euclidean distance from the head key point of the object to be matched to the center point of the image is
[0095] Specifically, the absolute value of the difference between the first Euclidean distance and each second Euclidean distance can be calculated. This absolute value is used to represent the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, denoted as FD i = |FinE1 - FoutE i |. The smaller this Euclidean distance relationship is, the closer the postures of the two objects being matched are.
[0096] Sort the obtained absolute values in ascending order, and select the target absolute values with the first Q numerical values arranged in the sorted absolute values. Among them, for two subsequences being matched, for the object to be matched in one of the subsequences to be matched, the object corresponding to the obtained target absolute value in the other corresponding subsequence for matching can be used as the candidate object whose posture matches that of the object to be matched. Among them, Q is a positive integer; generally, according to experience, Q can be taken as 1 or 2.
[0097] According to this target absolute value, determine the image sequence index of the candidate object whose posture matches that of the object to be matched in the other corresponding subsequence. Specifically, the subscripts of the first Q numerical values in the other corresponding subsequence can be taken as the image sequence index of the candidate object in the other corresponding subsequence. For example, when Q = 1, the corresponding target index index is:
[0098]
[0099] Among them, i represents the i-th frame of the video, and n1 represents the total number of video frames.
[0100] S150. Determine the target image sequence from the candidate image sequences, where there is a target object in the target image sequence that is the same object as the object to be matched.
[0101] Among them, when matching two images, a method of feature extraction plus feature comparison can be adopted. Among them, the feature extraction methods include but are not limited to artificially designed features, CNN (Convolutional Neural Network), CNN+LSTM (Long Short-Term Memory), etc. The feature comparison method uses the cosine distance (cosine distance, also known as cosine similarity) to calculate the distance between two feature vectors. The smaller the cosine, the closer the two targets are.
[0102] Specifically, as Figure 1d shown, the following method can be used to determine the target image sequence:
[0103] S210. Determine the cosine distance and Euclidean distance relationship between the image frame corresponding to the object to be matched and each image sequence.
[0104] In this embodiment, the cosine distance between the image frame corresponding to the object to be matched and each image sequence, that is, the i-th frame image feature in the Q-frame image, can be expressed as cos i , 1≤i≤Q.
[0105] S220. Determine the matching weights of each candidate image sequence according to the Euclidean distance relationship.
[0106] Specifically, the matching weights of each candidate image sequence can be calculated by the following formula. The smaller the FD, the closer the position and pose in the target image are, and the more effective the matching result is:
[0107]
[0108] Among them, i represents the i-th frame of video, and FD i =|FinE1 - FoutE i |, and FD i represents the Euclidean distance relationship. For the object to be matched in one of the subsequences to be matched, FinE1 represents the first Euclidean distance from the head key point of the object to be matched to the center point of the image, and FoutE i represents the second Euclidean distance from the head key points of each object in the i-th corresponding subsequence that is matched with the subsequence to be matched to the center point of the image.
[0109] S230. Calculate the matching distance between the candidate object in each candidate image sequence and the object to be matched in the sequence to be matched according to the matching weight and the cosine distance, and use the candidate image sequence with the smallest matching distance as the target image sequence.
[0110] Specifically, the cosine distances between the object to be matched and each candidate object can be calculated to obtain the cosine distance set {SFin1, SFin2, …, SFin m1 , SBin1, SBin2, …, SBin m2} corresponding to the pedestrian target tracking sequence image. For the convenience of representing the above cosine distance set, it is denoted as {S1, S2, …, S m}, and its corresponding {FD1, FD2, …, FD m}, where m = m1 + m2.
[0111] Among them, the matching distance S j between the object to be matched in the sequence to be matched and the j-th candidate object in the candidate image sequence can be expressed as:
[0112]
[0113] Among them, represents the matching weight, S i represents the cosine distance between the sequence to be matched and the i-th frame candidate image sequence; FD i represents the Euclidean distance relationship between the sequence to be matched and the i-th frame candidate image sequence.
[0114] After calculating the matching distance, the candidate image sequence with the minimum matching distance can be used as the target image sequence, and the index corresponding to the target image sequence can be obtained. Specifically, it can be reflected by the following formula:
[0115]
[0116] Among them, S j is the matching distance between the j-th candidate object in the candidate image sequence and the object to be matched in the sequence to be matched, and L is the number of candidate objects in the candidate image sequence. Since the candidate image sequence and the sequence to be matched respectively correspond to the outgoing sequence and the incoming sequence, therefore, by using the above formula, the image sequence with the minimum matching distance in the outgoing sequence and the incoming sequence can be obtained, and this image sequence is the target image sequence in which the same target object (the same pedestrian) exists in the incoming sequence and the outgoing sequence.
[0117] In this embodiment, step S150 is a process of accurately identifying the image target based on S140. By adopting the matching distance estimation method weighted based on the pose similarity degree, the matching accuracy can be improved, and the video sequence corresponding to the same pedestrian can be obtained.
[0118] In this embodiment, the pedestrian recognition process is divided into a rough matching stage and a fine matching stage. In the rough matching stage, according to the movement trajectory of the pedestrian, the video sequence can be divided into an incoming sequence and an outgoing sequence, and the pedestrian sequence can be divided into subsequences facing and away from the camera. By using the same subsequence for matching, for example, matching the subsequence facing the camera in the incoming sequence with the subsequence facing the camera in the outgoing sequence, or matching the subsequence away from the camera in the incoming sequence with the subsequence away from the camera in the outgoing sequence, the amount of calculation in the matching process is effectively reduced. In the matching process, by using the Euclidean distance method from the object to be matched to the center point of the image, a candidate image sequence with a similar posture to the target object to be matched in the same subsequence can be obtained, laying the foundation for the subsequent pedestrian recognition. Compared with the method of using the color features of the pedestrian target for recognition in the prior art, the technical solution provided by this embodiment effectively improves the recognition rate of pedestrians. In the fine matching stage, by using the matching distance estimation method based on the weighting of the posture similarity, the matching accuracy can be effectively improved.
[0119] Embodiment 2
[0120] Figure 2 A structural block diagram of a pedestrian recognition device provided in Embodiment 2 of the present invention, such as Figure 2 As shown, the device includes: a video sequence division module 310, a subsequence division module 320, a subsequence matching module 330, a candidate image sequence determination module 340 and a target image sequence determination module 350; wherein,
[0121] The video sequence division module 310 is configured to divide the video sequence into an in-sequence and an out-sequence according to the motion trajectory of the pedestrian in the video sequence, wherein the motion trajectory of the pedestrian in the in-sequence is from the bottom of the image to the top of the image; and the motion trajectory of the pedestrian in the out-sequence is from the top of the image to the bottom of the image;
[0122] The subsequence division module 320 is configured to divide the incoming sequence and the outgoing sequence into a subsequence toward the camera and a subsequence away from the camera, respectively, according to the coordinates of the key points of the pedestrian's head; wherein, for the incoming sequence, the vertical coordinates of the key points of the pedestrian's head in the subsequence toward the camera are greater than the vertical coordinates of the center point of the image, and the vertical coordinates of the key points of the pedestrian's head in the subsequence away from the camera are less than the vertical coordinates of the center point of the image; for the outgoing sequence, the vertical coordinates of the key points of the pedestrian's head in the subsequence toward the camera are less than the vertical coordinates of the center point of the image, and the vertical coordinates of the key points of the pedestrian's head in the subsequence away from the camera are greater than the vertical coordinates of the center point of the image;
[0123] The subsequence matching module 330 is configured to match the camera-facing subsequence in the outgoing sequence with the camera-facing subsequence in the incoming sequence, and to match the camera-facing subsequence in the outgoing sequence with the camera-facing subsequence in the incoming sequence; in the matching process, for an object to be matched in one of the subsequences to be matched, determine a first Euclidean distance from a key point of the head of the object to be matched to a center point of the image, and determine a second Euclidean distance from a key point of the head of each object in a corresponding subsequence matched with the subsequence to be matched to a center point of the image;
[0124] The candidate image sequence determination module 340 is configured to determine, based on the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, a candidate image sequence whose Euclidean distance relationship meets a preset distance requirement from the corresponding subsequences matched with the subsequence to be matched;
[0125] The target image sequence determination module 350 is configured to determine a target image sequence from the candidate image sequence, wherein the target image sequence contains a target object that is the same object as the object to be matched.
[0126] Optionally, the candidate image sequence determination module 340 is specifically configured to:
[0127] Calculating an absolute value of a difference between the first Euclidean distance and each second Euclidean distance, wherein the absolute value is used to represent a Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance;
[0128] Sort the obtained absolute values in ascending order, and select the target absolute values whose values are arranged in the first Q positions from the sorted absolute values; wherein Q is a positive integer;
[0129] The image sequence corresponding to the target absolute value is used as a candidate image sequence in a corresponding subsequence to be matched with the subsequence to be matched.
[0130] Optionally, the target image sequence determination module 350 is specifically configured to:
[0131] A distance determination unit, configured to determine a cosine distance and a Euclidean distance relationship between the image frame corresponding to the to-be-matched object and each candidate image sequence;
[0132] a matching weight determination unit, configured to determine a matching weight of each candidate image sequence according to the Euclidean distance relationship;
[0133] A target image sequence determination unit, configured to calculate a matching distance between a candidate object in each candidate image sequence and a to-be-matched object in the to-be-matched sequence according to the matching weight and the cosine distance, and use the candidate image sequence with the minimum matching distance as the target image sequence.
[0134] Optionally, the matching distance is calculated according to the following formula:
[0135]
[0136] where S i represents the cosine distance between the to-be-matched sequence and the i-th frame candidate image sequence; represents the matching weight; FD i represents the Euclidean distance relationship between the to-be-matched sequence and the i-th frame candidate image sequence; j is the j-th candidate object in each candidate image sequence; m represents the sum of the number of video frames in the subsequence towards the camera and the subsequence away from the camera.
[0137] Optionally, the pedestrian head key point coordinates are the corresponding head key point coordinates when the pedestrian image is rotated so that the line connecting the two shoulder key points after rotation is perpendicular to the line connecting the image center and the head key point.
[0138] Optionally, rotate the image:
[0139]
[0140] where tx = x - Cx, ty = y - Cy; θ = θ2 - θ1, representing the angle to be rotated; θ1 = arctan((Sy r - Sy l ) / (Sx r - Sx l )) represents the angle between the line connecting the two shoulder key points and the horizontal line; θ2 = arctan(-(Hy - Cy) / (Hx - Cx)) represents the angle between the perpendicular line connecting the image center and the head key point and the horizontal line; (Hx, Hy) represents the head key point coordinates before rotation; (Sx l , Sy l ) and (Sx r , Sy r ) respectively represent the left and right shoulder key point coordinates in the image before rotation; (Cx, Cy) represents the image center point coordinates; (x, y) represents any point in the image corresponding to the head and shoulder detection rectangle; (x′, y′) represents the point after (x, y) is rotated around the image center; tx represents the distance from the abscissa of any point in the image corresponding to the head and shoulder detection rectangle to the abscissa of the image center point; ty represents the distance from the ordinate of any point in the image corresponding to the head and shoulder detection rectangle to the ordinate of the image center point.
[0141] Example 3
[0142] Please refer to Figure 3 , Figure 3 , which is a structural block diagram of a computing device provided by Example 3 of the present invention.
[0143] As Figure 3 shown, the computing device may include:
[0144] A memory 701 storing executable program code;
[0145] A processor 702 coupled to the memory 701;
[0146] Wherein, the processor 702 calls the executable program code stored in the memory 701 and executes the pedestrian recognition method provided by any embodiment of the present invention.
[0147] The embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the pedestrian recognition method provided by any embodiment of the present invention.
[0148] In various embodiments of the present invention, it should be understood that the magnitudes of the serial numbers of the above processes do not necessarily mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0149] In the embodiments provided by the present invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0150] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0151] When the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests for causing a computer device (which can be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute some or all of the steps of the above-mentioned methods in various embodiments of the present invention.
[0152] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.
[0153] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0154] Those of ordinary skill in the art can understand that the modules in the device in the embodiment can be distributed in the device of the embodiment according to the description of the embodiment, or can be correspondingly changed and located in one or more devices different from this embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying pedestrians, characterized in that, include: According to the motion trajectory of the pedestrian in the video sequence, the video sequence is divided into an in-sequence and an out-sequence, wherein the motion trajectory of the pedestrian in the in-sequence is from the bottom of the image to the top of the image; the motion trajectory of the pedestrian in the out-sequence is from the top of the image to the bottom of the image, and the camera shooting the video sequence is installed in a top-down manner, the lens plane is approximately parallel to the ground, the center of the camera is the center of the image, and the image is symmetrical up and down and left and right respectively; According to the coordinates of the key points of the pedestrian's head, the in-sequence and the out-sequence are respectively divided into a sub-sequence facing the camera and a sub-sequence away from the camera; wherein, for the in-sequence, the ordinate of the key points of the pedestrian's head in the sub-sequence facing the camera is greater than the ordinate of the center point of the image, and the ordinate of the key points of the pedestrian's head in the sub-sequence away from the camera is less than the ordinate of the center point of the image; for the out-sequence, the ordinate of the key points of the pedestrian's head in the sub-sequence facing the camera is less than the ordinate of the center point of the image, and the ordinate of the key points of the pedestrian's head in the sub-sequence away from the camera is greater than the ordinate of the center point of the image; Matching the camera-facing subsequence in the outgoing sequence with the camera-facing subsequence in the incoming sequence, and matching the camera-facing subsequence in the outgoing sequence with the camera-facing subsequence in the incoming sequence; During the matching process, for an object to be matched in one of the subsequences to be matched, a first Euclidean distance from a key point of the head of the object to be matched to a center point of the image is determined, and a second Euclidean distance from a key point of the head of each object in a corresponding subsequence to be matched with the subsequence to be matched to a center point of the image is determined; According to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, determining, from the corresponding subsequences matched with the subsequence to be matched, a candidate image sequence whose Euclidean distance relationship meets a preset distance requirement; A target image sequence is determined from the candidate image sequence, wherein a target object that is the same object as the object to be matched exists in the target image sequence.
2. The method according to claim 1, characterized in that, According to the Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance, determining a candidate image sequence whose Euclidean distance relationship meets a preset distance requirement from the corresponding subsequences matched with the subsequence to be matched, comprising: Calculating an absolute value of a difference between the first Euclidean distance and each second Euclidean distance, wherein the absolute value is used to represent a Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance; Sort the obtained absolute values in ascending order, and select the target absolute values whose values are arranged in the first Q positions from the sorted absolute values; wherein Q is a positive integer; The image sequence corresponding to the target absolute value is used as a candidate image sequence in a corresponding subsequence to be matched with the subsequence to be matched.
3. The method according to claim 1 or 2, characterized in that, The step of determining a target image sequence corresponding to a target candidate object that is the same object as the object to be matched includes: Determine the cosine distance and Euclidean distance relationship between the image frame corresponding to the object to be matched and each candidate image sequence; Determining a matching weight of each candidate image sequence according to the Euclidean distance relationship; According to the matching weight and the cosine distance, the matching distance between the candidate object in each candidate image sequence and the object to be matched in the sequence to be matched is calculated, and the candidate image sequence with the smallest matching distance is used as the target image sequence.
4. The method according to claim 3, wherein The calculating the matching distance between the candidate objects in each candidate image sequence and the objects to be matched in the sequence to be matched comprises: The matching distance is calculated according to the following formula: Among them, S i represents the cosine distance between the sequence to be matched and the i-th frame candidate image sequence; Represents the matching weight; FD i represents the Euclidean distance relationship between the sequence to be matched and the i-th frame candidate image sequence; j is the j-th candidate object in the candidate image sequence; m represents the sum of the number of video frames facing the camera subsequence and the number of video frames away from the camera subsequence.
5. The method according to claim 1, wherein The pedestrian head key point coordinates are obtained by rotating the pedestrian image so that the line connecting the two shoulder key points after rotation is perpendicular to the line connecting the image center and the head key point.
6. The method according to claim 5, characterized in that, Rotate the image according to the following formula: where, tx = x - Cx, ty = y - Cy; θ = θ2 - θ1, representing the angle to be rotated; θ1 = arctan((Sy r -Sy l ) / (Sx r -Sx l )) represents the angle between the line connecting the two key points of the shoulder and the horizontal line; θ2 = arctan(-(Hy - Cy) / (Hx - Cx)) represents the angle between the perpendicular line connecting the image center and the key point of the head and the horizontal line; (Hx, Hy) represents the coordinates of the key point of the head before rotation; (Sx l , Sy l ) and (Sx r , Sy r ) respectively represent the coordinates of the key points of the left and right shoulders in the image before rotation; (Cx, Cy) represents the coordinates of the image center point; (x, y) represents any point in the image corresponding to the head and shoulder detection rectangle; (x′, y′) represents the point obtained by rotating (x, y) around the image center; tx represents the distance from the abscissa of any point in the image corresponding to the head and shoulder detection rectangle to the abscissa of the image center point; ty represents the distance from the ordinate of any point in the image corresponding to the head and shoulder detection rectangle to the ordinate of the image center point.
7. An identification device for pedestrians, characterized in that, include: The video sequence division module is configured to divide the video sequence into an in-sequence and an out-sequence according to the motion trajectory of the pedestrian in the video sequence, wherein the motion trajectory of the pedestrian in the in-sequence is from the bottom of the image to the top of the image; the motion trajectory of the pedestrian in the out-sequence is from the top of the image to the bottom of the image, and the camera shooting the video sequence is installed in a top-down manner, the lens plane is approximately parallel to the ground, the center of the camera is the center of the image, and the image is symmetrical up and down and left and right respectively; The subsequence division module is configured to divide the in-sequence and out-sequence into a subsequence toward the camera and a subsequence away from the camera, respectively, according to the coordinates of the key points of the pedestrian's head; wherein, for the in-sequence, the ordinate of the key points of the pedestrian's head in the subsequence toward the camera is greater than the ordinate of the center point of the image, and the ordinate of the key points of the pedestrian's head in the subsequence away from the camera is less than the ordinate of the center point of the image; for the out-sequence, the ordinate of the key points of the pedestrian's head in the subsequence toward the camera is less than the ordinate of the center point of the image, and the ordinate of the key points of the pedestrian's head in the subsequence away from the camera is greater than the ordinate of the center point of the image; The subsequence matching module is configured to match the camera-facing subsequence in the outgoing sequence with the camera-facing subsequence in the incoming sequence, and to match the camera-facing subsequence in the outgoing sequence with the camera-facing subsequence in the incoming sequence; in the matching process, for an object to be matched in one of the subsequences to be matched, determine a first Euclidean distance from a key point of the head of the object to be matched to a center point of the image, and determine a second Euclidean distance from a key point of the head of each object in a corresponding subsequence matched with the subsequence to be matched to a center point of the image; a candidate image sequence determination module configured to determine, based on a Euclidean distance relationship between a first Euclidean distance and a second Euclidean distance, a candidate image sequence whose Euclidean distance relationship satisfies a preset distance requirement from corresponding subsequences matched with the subsequence to be matched; The target image sequence determination module is configured to determine a target image sequence from the candidate image sequence, wherein the target image sequence contains a target object that is the same object as the object to be matched.
8. The device according to claim 7, characterized in that The candidate image sequence determination module is specifically configured as follows: Calculating an absolute value of a difference between the first Euclidean distance and each second Euclidean distance, wherein the absolute value is used to represent a Euclidean distance relationship between the first Euclidean distance and the second Euclidean distance; Sort the obtained absolute values in ascending order, and select the target absolute values whose numerical values are ranked among the top Q from the sorted absolute values; where Q is a positive integer; Use the image sequence corresponding to the target absolute value as the candidate image sequence in the corresponding subsequence for matching with the to-be-matched subsequence.
9. A computing device, characterized in that, The device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the pedestrian recognition method described in any one of claims 1-5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the pedestrian recognition method described in any one of claims 1-5.
Citation Information
Patent Citations
Pedestrian multi-target tracking method and device and computer system
CN111784746A
Top view image-based pedestrian trajectory generation method, storage medium and electronic equipment
CN113033353A