Artificial intelligence-based pedestrian re-identification method, device, equipment, and storage medium

By identifying pedestrian regions and vehicle types in pedestrian re-identification, and combining the shooting time and camera identification, visual similarity and spatiotemporal probability are calculated, thus solving the problem of insufficient accuracy in pedestrian re-identification and achieving higher recognition accuracy.

CN115240226BActive Publication Date: 2026-03-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-10
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately determine whether the same pedestrian appears in images captured by different cameras during pedestrian re-identification, especially lacking effective means to consider the pedestrian's mobility when using vehicles.

Method used

By identifying pedestrian areas and associated vehicle types, and combining the image capture time and camera identification, visual similarity and spatiotemporal probability are calculated, and a weighted calculation method is used for pedestrian re-identification.

Benefits of technology

This improves the accuracy of pedestrian re-identification, ensuring that the spatiotemporal probability better matches the likelihood of a pedestrian appearing in images captured by different cameras, thus enhancing the accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240226B_ABST
    Figure CN115240226B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, device, and storage medium for pedestrian re-identification based on artificial intelligence, relating to the field of artificial intelligence technology, specifically image recognition and video analysis technology, which can be applied in smart cities, urban governance, and public security emergency scenarios. The specific implementation scheme is as follows: determining the pedestrian regions and the types of vehicles associated with the pedestrians in each image to be compared; obtaining the visual similarity between pedestrians appearing in each image to be compared based on the determined pedestrian regions; determining the spatiotemporal probability of the same pedestrian appearing in each image to be compared based on the shooting time, camera identification, and the type of vehicle associated with the pedestrian; and performing pedestrian re-identification based on the visual similarity and the spatiotemporal probability. Applying the scheme provided by the embodiments of this disclosure can improve the accuracy of re-identification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, in particular to image recognition and video analysis technology, which can be applied in smart city, city governance and public security emergency scenarios. BACKGROUND

[0002] In a monitoring scene, a large number of cameras are generally arranged, which form a camera network. In order to better monitor the monitoring area, it is increasingly important to perform correlation analysis on the contents of images captured by each camera in the camera network. In view of this, pedestrian re-identification technology has become one of the current research hotspots.

[0003] Pedestrian re-identification technology can be used to determine whether the pedestrians in images captured by different cameras are the same person. SUMMARY

[0004] The present disclosure provides a pedestrian re-identification method and device based on artificial intelligence, equipment and a storage medium.

[0005] According to an aspect of the present disclosure, a pedestrian re-identification method based on artificial intelligence is provided, comprising:

[0006] determining the pedestrian region in each to-be-compared image and the type of the vehicle associated with the pedestrian;

[0007] obtaining the visual similarity between pedestrians appearing in each to-be-compared image according to the determined pedestrian region;

[0008] determining the space-time probability of the same pedestrian appearing in each to-be-compared image according to the shooting time, camera identifier and type of the vehicle associated with the pedestrian of each to-be-compared image;

[0009] performing pedestrian re-identification based on the visual similarity and the space-time probability.

[0010] According to another aspect of the present disclosure, a model training method is provided, comprising:

[0011] obtaining a sample image pair, wherein the sample image pair includes a positive sample image pair in which the same pedestrian appears in the same space-time in the contained sample image and a negative sample image pair in which the same pedestrian does not appear in the same space-time in the contained sample image;

[0012] obtaining the shooting time, camera identifier and type of the vehicle associated with the pedestrian of each sample image in each sample image pair;

[0013] input the shooting time, the camera identifier and the type of the vehicle associated with the pedestrian of each sample image in each sample image pair into a preset neural network model, to obtain a spatiotemporal probability output by the neural network model, the spatiotemporal probability representing whether the pedestrian in the sample image pair is in the same spatiotemporal space;

[0014] According to the obtained spatiotemporal probability and the positive sample label or the negative sample label corresponding to each sample image pair, the model parameters of the neural network model are adjusted to obtain a spatiotemporal probability estimation model.

[0015] According to another aspect of the present disclosure, a pedestrian re-identification device based on artificial intelligence is provided, comprising:

[0016] A type determination module is configured to determine the pedestrian region in each to-be-compared image and the type of the vehicle associated with the pedestrian.

[0017] A similarity determination module is configured to obtain the visual similarity between the pedestrians appearing in each to-be-compared image according to the determined pedestrian region.

[0018] A spatiotemporal probability determination module is configured to determine the spatiotemporal probability of the same pedestrian appearing in each to-be-compared image according to the shooting time, the camera identifier and the type of the vehicle associated with the pedestrian of each to-be-compared image.

[0019] A re-identification module is configured to perform pedestrian re-identification based on the visual similarity and the spatiotemporal probability.

[0020] According to another aspect of the present disclosure, a model training device is provided, comprising:

[0021] An image pair obtaining module is configured to obtain sample image pairs, wherein the sample image pairs include positive sample image pairs in which the same pedestrian appears in the same spatiotemporal space in the contained sample images and negative sample image pairs in which the same pedestrian does not appear in the same spatiotemporal space in the contained sample images.

[0022] An image information obtaining module is configured to obtain the shooting time, the camera identifier and the type of the vehicle associated with the pedestrian of each sample image in each sample image pair.

[0023] A spatiotemporal probability obtaining module is configured to input the shooting time, the camera identifier and the type of the vehicle associated with the pedestrian of each sample image in each sample image pair into a preset neural network model, to obtain a spatiotemporal probability output by the neural network model, the spatiotemporal probability representing whether the pedestrian in the sample image pair is in the same spatiotemporal space.

[0024] A model obtaining module is configured to adjust the model parameters of the neural network model according to the obtained spatiotemporal probability and the positive sample label or the negative sample label corresponding to each sample image pair, to obtain a spatiotemporal probability estimation model.

[0025] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0026] at least one processor; and

[0027] a memory in communication connection with the at least one processor; wherein

[0028] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned artificial intelligence-based pedestrian re-identification method or model training method.

[0029] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the above-mentioned artificial intelligence-based pedestrian re-identification method or model training method.

[0030] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the above-mentioned artificial intelligence-based pedestrian re-identification method or model training method.

[0031] As can be seen from the above, in the scheme provided by the embodiments of the present disclosure, in addition to obtaining the visual similarity between pedestrians appearing in each to-be-compared image, the spatiotemporal probability is also determined according to the shooting time of each to-be-compared image, the camera identifier, and the type of the vehicle associated with the pedestrian. In this way, the activity ability of the pedestrian after using each type of vehicle is referred to when determining the spatiotemporal probability, so that the obtained spatiotemporal probability is more in line with the possibility of the actual appearance of the pedestrian in the images taken by different cameras, and the obtained spatiotemporal probability is more accurate. Accordingly, the accuracy of pedestrian re-identification based on visual similarity and spatiotemporal probability is improved.

[0032] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0033] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:

[0034] Figure 1 is a flowchart of a first artificial intelligence-based pedestrian re-identification method provided by the embodiments of the present disclosure;

[0035] Figure 2 is a flowchart of a second artificial intelligence-based pedestrian re-identification method provided by the embodiments of the present disclosure;

[0036] Figure 3This is a flowchart illustrating the first model training method provided in this embodiment of the disclosure;

[0037] Figure 4 This is a schematic diagram of a data processing flow provided in an embodiment of this disclosure;

[0038] Figure 5 This is a flowchart illustrating the second model training method provided in this embodiment of the disclosure;

[0039] Figure 6 This is a schematic diagram of the structure of an artificial intelligence-based pedestrian re-identification device provided in an embodiment of this disclosure;

[0040] Figure 7 This is a schematic diagram of the structure of a model training device provided in an embodiment of this disclosure;

[0041] Figure 8 This is a block diagram of an electronic device used to implement the artificial intelligence-based pedestrian re-identification and model training method of the embodiments of this disclosure. Detailed Implementation

[0042] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0043] The following describes the application scenarios of the solutions provided in the embodiments of this disclosure.

[0044] The solutions provided in this disclosure can be applied to public safety and urban security scenarios. For example, when searching for suspicious individuals from a network of surveillance cameras deployed in a city, this solution can be used to re-identify pedestrians appearing in the images captured by each surveillance camera in the network to determine the movement of suspicious individuals, and can trigger an alarm when a suspicious individual is detected.

[0045] Similarly, the solution provided in this disclosure can also be applied to commercial customer flow analysis scenarios. Following the re-identification method described above, the solution provided in this disclosure can also be used to obtain the movement traces of a large number of consumers. In this case, by statistically analyzing the obtained movement traces, consumer spending preferences can be analyzed.

[0046] The solution provided in this disclosure can also be applied to image search scenarios. For example, a user specifies a target image, and the electronic device retrieves images from a stored image library that contain the same pedestrian as the target image, according to the solution provided in this disclosure, as the search result.

[0047] In addition, the solutions provided in this disclosure can also be used in other scenarios such as video image big data processing.

[0048] The pedestrian re-identification method based on artificial intelligence provided in the embodiments of this disclosure will be described below.

[0049] In one embodiment of this disclosure, see [link to embodiment]. Figure 1 The paper provides a flowchart of the first pedestrian re-identification method based on artificial intelligence, which includes the following steps S101-S104.

[0050] Step S101: Determine the pedestrian areas and the types of vehicles associated with the pedestrians in each image to be compared.

[0051] The pedestrian region is determined by the position of the pedestrian in the image to be compared. This type of image region may contain some or all of the pedestrian's image information. The pedestrian region can be a face region or a body region.

[0052] In one embodiment of this disclosure, the image to be compared can be input into a trained detection model to obtain the pedestrian region output by the detection model. Specifically, the pedestrian region can be an image region contained within a face bounding box or a body bounding box. The detection model can be a YOLOv5-based object detection model, a YOLOv3D-based object detection model, etc.

[0053] The vehicles associated with pedestrians can be understood as the vehicles used by pedestrians in the image during their movement.

[0054] In one embodiment of this disclosure, the vehicle associated with a pedestrian can be determined by the distance relationship between the positions of pedestrians and vehicles in an image. For example, after obtaining the position of a vehicle in the image, the distance between the position of the vehicle and the center position of the pedestrian area can be calculated. If the obtained distance is less than a preset threshold, the vehicle is determined to be the vehicle associated with the pedestrian.

[0055] The type of vehicle can be a car, a bicycle, etc.

[0056] The images to be compared can be images captured by cameras in a monitored scene. The number of images to be compared in this step is greater than or equal to two. Pedestrian re-identification is generally performed based on two images to be compared, thus allowing pedestrian re-identification to be completed for each pair of images using the scheme provided in this embodiment.

[0057] Step S102: Based on the determined pedestrian area, obtain the visual similarity between pedestrians appearing in each image to be compared.

[0058] The visual similarity between pedestrians in each image to be compared can be determined based on the similarity between the image features of pedestrian regions in each image to be compared.

[0059] Specifically, features can be extracted from the pedestrian regions in each image to be compared, resulting in image features. The cosine similarity between the feature vectors of each image feature is then calculated, and the resulting cosine similarity is used as the visual similarity. Following this method, pairwise comparisons can be performed on each image to obtain the visual similarity between pedestrians in any image and pedestrians in other images.

[0060] In one scenario, where the pedestrian area is a face area, a pre-trained face recognition model can be used to obtain the image features of the face area. In this case, the face area is input into the face recognition model, and the model outputs image features, specifically, a feature vector of the face.

[0061] In another scenario, where the pedestrian area is actually a human body area, a pre-trained ReID model can be used to obtain the image features of the human body region. In this case, the human body region is input into the ReID model, and the ReID model outputs image features, specifically, a feature vector of the human body.

[0062] Step S103: Determine the spatiotemporal probability of the same pedestrian appearing in each image to be compared based on the shooting time, camera identification, and the type of vehicle associated with the pedestrian.

[0063] The capture time of each image to be compared can be recorded when the camera captures the image; the camera identifier can be pre-set, with each camera corresponding to a unique camera identifier. The capture time and camera identifier mentioned above can be used as additional information for the images to be compared and stored in the same device as the images to be compared. In this way, when retrieving each image to be compared from the device, the capture time and camera identifier of each image to be compared can be obtained at the same time.

[0064] Spatiotemporal probability represents the likelihood of the same person appearing in various images to be compared.

[0065] Because pedestrians have varying mobility when using different modes of transportation, the probability of a pedestrian moving between the locations of the cameras capturing the images to be compared and appearing in the captured images at a specified time varies. For example, pedestrians traveling by car have greater mobility than pedestrians traveling by bicycle; correspondingly, pedestrians traveling by car are more likely to appear in the images to be identified captured by two cameras at a greater distance within a shorter period of time compared to those traveling by bicycle. Therefore, in order to improve the accuracy of the determined spatiotemporal probability, this embodiment of the disclosure introduces the type of transportation associated with the pedestrian as one of the bases for determining the spatiotemporal probability.

[0066] In addition, if the types of transportation used by pedestrians in two images to be compared are different within a short period of time, it can be directly determined that the pedestrians in the images to be compared are likely not the same pedestrian.

[0067] For details on how to determine the spatiotemporal probability, please refer to the following embodiments, which will not be described in detail here.

[0068] Step S104: Perform pedestrian re-identification based on visual similarity and spatiotemporal probability.

[0069] When performing pedestrian re-identification, visual similarity and spatiotemporal probability can be considered together. That is, based on these two types of information, visual similarity and spatiotemporal probability, it can be determined that the pedestrians appearing in each image to be compared are the same pedestrians.

[0070] The following describes the specific implementation method for pedestrian re-identification of any two images to be compared.

[0071] In one embodiment of this disclosure, visual similarity and spatiotemporal probability can be weighted and calculated based on the weight coefficients corresponding to visual similarity and spatiotemporal probability, respectively, to obtain a weighted result; based on the weighted result, it can be determined whether the pedestrians appearing in each image to be compared are the same pedestrian.

[0072] The above weighted calculation process can be represented by the following formula:

[0073] p = α*p1 + (1-α)*p2

[0074] Where p represents the weighted result, p1 represents visual similarity, p2 represents spatiotemporal probability, and α is the weight coefficient corresponding to visual similarity, which can be preset. Specifically, the value range of α can be [0,1].

[0075] Because visual similarity and spatiotemporal probability have different abilities to represent pedestrian identities, their influence on the weighted results is also different. Therefore, in this implementation, by setting a weight system, the influence of visual similarity and spatiotemporal probability on the weighted results can be made more consistent with the actual situation, thereby making the weighted results more accurate.

[0076] When determining whether the pedestrians appearing in each image to be compared are the same pedestrian based on the weighted result, the magnitude of the weighted result can be compared with the preset threshold. If the weighted result is greater than the threshold, then the pedestrians appearing in the images to be compared used to calculate the visual similarity and spatiotemporal probability are determined to be the same pedestrian.

[0077] In another embodiment of this disclosure, corresponding discrimination thresholds can be set for visual similarity and spatiotemporal probability, respectively. When both visual similarity and spatiotemporal probability are greater than the corresponding discrimination thresholds, it is determined that the pedestrians appearing in the comparison images used to calculate the above-mentioned visual similarity and spatiotemporal probability are the same pedestrians.

[0078] As can be seen from the above, the solution provided in this embodiment, in addition to obtaining the visual similarity between pedestrians appearing in each image to be compared, also determines the spatiotemporal probability based on the shooting time of each image to be compared, the camera identifier, and the type of transportation associated with the pedestrian. This means that the determination of the spatiotemporal probability takes into account the pedestrian's mobility after using various types of transportation, making the obtained spatiotemporal probability more consistent with the actual probability of the pedestrian appearing in images taken by different cameras, thus making the obtained spatiotemporal probability more accurate. Accordingly, the accuracy of pedestrian re-identification based on visual similarity and spatiotemporal probability is improved.

[0079] The application process of the solutions provided in the embodiments of this disclosure in different scenarios is described below.

[0080] Scenario 1: Image Search

[0081] In this scenario, the target image is first obtained as the basis for the search. Then, the visual similarity between the target image and pedestrians in each image in the image library is determined. Based on the visual similarity, k preset candidate images are selected from the image library. For each candidate image, the spatiotemporal probability of the same pedestrian appearing in both the target image and the candidate image is determined based on the shooting time, camera identification, and the type of transportation associated with the pedestrian. Then, based on the determined spatiotemporal probabilities and the corresponding visual similarity of each candidate image, images containing the same pedestrian as those in the target image are selected as the search results. This allows for filtering of images in the image library based on visual similarity, significantly improving the speed and accuracy of image searching.

[0082] Scenario 2: Monitoring and Alarm Scenario

[0083] In this scenario, the first step is to obtain the image to be tested. Then, the visual similarity between the image to be tested and pedestrians in each image in the database is determined. For each image in the database, based on the shooting time of both the image to be tested and the images in the database, the camera identification, and the type of vehicle associated with the pedestrian, the spatiotemporal probability of the same pedestrian appearing in both images is determined. Then, based on the determined spatiotemporal probabilities and visual similarity, images from the database that show the same pedestrian as the one in the image to be tested are selected. In this case, whether to trigger an alarm can be determined based on information such as the identity of the pedestrian in the selected image.

[0084] The first implementation method for determining the spatiotemporal probability in step S102 above will be explained below.

[0085] In one implementation, the spatiotemporal probability can be determined according to the following steps D1-D4.

[0086] Step D1: Determine the shooting location of each image based on the camera identification of each image to be compared.

[0087] Cameras can be installed in specific locations, such as on the side of a street or at an intersection. In this case, camera markers can be set based on the camera's location, allowing the image's capture location to be directly determined during pedestrian re-identification. For example, camera markers could be the camera's latitude and longitude, or landmarks where the camera is located, such as a subway entrance.

[0088] In another scenario, camera identification can also be the camera's ID, name, etc. In this case, the correspondence between the camera's identification and its shooting location can be obtained in advance. Thus, when performing pedestrian re-identification, the shooting location of each image to be compared can be determined based on the aforementioned correspondence and the camera's identification.

[0089] Step D2: Obtain the predicted time to travel from one filming location to another using the target type of vehicle.

[0090] The target type is the type of vehicle associated with each pedestrian in the images to be compared.

[0091] In one implementation, the map software's API (Application Program Interface) can be called to predict the estimated travel time from one location to another using a vehicle of the target type.

[0092] Step D3: Determine the shooting interval based on the shooting time of each image to be compared.

[0093] By subtracting the shooting time of each image to be compared from each other, the shooting interval between each pair of images can be obtained.

[0094] Step D4: Determine the spatiotemporal probability of the same pedestrian appearing in each image to be compared based on the prediction duration and shooting interval.

[0095] In one embodiment of this disclosure, the spatiotemporal probability can be calculated in the following manner:

[0096] If the shooting interval is greater than the predicted duration, the spatiotemporal probability of the same pedestrian appearing in each image to be compared is determined to be 1; if the shooting interval is not greater than the predicted duration, the spatiotemporal probability is determined to be the ratio of the shooting interval to the predicted duration.

[0097] The above process can be represented by the following formula:

[0098]

[0099] Where p represents the spatiotemporal probability, t real t represents the time interval. pred Indicates the predicted duration.

[0100] Calculating the spatiotemporal probability in the above manner, on the one hand, if the shooting interval is greater than the predicted duration, it can be assumed that the pedestrian has sufficient time to move between the two shooting locations using the target type of vehicle. In this case, the spatiotemporal probability is 1, indicating that the pedestrian is more likely to appear in the images taken by the two cameras within the calculated time interval. Using 1 as the spatiotemporal probability is more accurate in this case. On the other hand, if the shooting interval is not greater than the predicted duration, the ratio of the shooting interval to the predicted duration is calculated as the spatiotemporal probability. In this way, the closer the shooting interval is to the predicted duration, the greater the value of the spatiotemporal probability. Since the pedestrian's mobility is kept within a certain limit when using the target type of vehicle, the time spent moving between the two shooting locations is more likely to be close to the predicted duration. Therefore, the spatiotemporal probability obtained in the above manner is more accurate.

[0101] As can be seen from the above, the consistency between the shooting interval and the prediction duration is considered when calculating the spatiotemporal probability, so that the spatiotemporal probability is consistent with the possible range of movement of pedestrians in the time interval, and the obtained spatiotemporal probability is more accurate.

[0102] The second implementation method for determining the spatiotemporal probability in step S102 above will be explained below.

[0103] In the second implementation, the shooting time, camera identification, and type of vehicle associated with each image to be compared can be input into a pre-trained spatiotemporal probability prediction model to obtain the spatiotemporal probability output by the spatiotemporal probability prediction model.

[0104] The spatiotemporal probability prediction model is a model used to obtain the spatiotemporal probability of pedestrians being in the same space and time by training a preset neural network model with the shooting time of the sample images, camera identification, and the type of transportation associated with the pedestrians.

[0105] The methods for obtaining the shooting time, camera identification, and type of vehicle associated with the pedestrian in each of the above sample images are similar to those in step S103 above. The only difference is the substitution of names and concepts such as sample image and image to be compared, which will not be described in detail here.

[0106] The specific training method for the above spatiotemporal probability prediction model can be found in the subsequent embodiments, which will not be detailed here.

[0107] As can be seen from the above, in the solution provided by the embodiments of this disclosure, the shooting time of the sample image, the camera identifier, and the type of vehicle associated with the pedestrian are used for training. In this way, the spatiotemporal probability prediction model learns the relationship between spatiotemporal probability and shooting time, camera identifier, and type of vehicle associated with the pedestrian. Thus, when pedestrian re-identification is performed, the spatiotemporal probability obtained by using the spatiotemporal probability prediction model can improve the accuracy of the obtained spatiotemporal probability.

[0108] The following is combined Figure 2 The flowchart shown illustrates the overall process of the AI-based pedestrian re-identification method.

[0109] like Figure 2 As shown, image a and image b are any two images from the images to be compared. Figure 2 The visual feature vector model can take the pedestrian regions of images a and b as input and output the visual similarity between pedestrians appearing in images a and b; the spatiotemporal probability prediction model can take the shooting time of images a and b, camera identification, and the type of transportation associated with the pedestrians as input and output the spatiotemporal probability of the same pedestrian appearing in images a and b.

[0110] The weighted similarity in the figure is the result of a weighted calculation of spatiotemporal probability and visual similarity. For the specific calculation method, please refer to the formula in the aforementioned embodiment, which will not be described in detail here.

[0111] If the weighted similarity is greater than the preset threshold, then the pedestrians in image a and image b are considered to be the same pedestrian; otherwise, the pedestrians in image a and image b are considered to be different pedestrians.

[0112] This disclosure also provides a model training method for training the aforementioned spatiotemporal probability prediction model, which will be described in detail below.

[0113] In one embodiment of this disclosure, see [link to embodiment]. Figure 3 The flowchart of the first model training method is provided, which includes the following steps S301-S304.

[0114] Step S301: Obtain sample image pairs.

[0115] The sample image pairs include: positive sample image pairs in which the same pedestrian appears in the same space and time and negative sample image pairs in which the same pedestrian does not appear in the same space and time.

[0116] In one implementation, the positive sample image pairs and negative sample image pairs can be obtained by manually selecting images from each sample image.

[0117] In another implementation, positive sample image pairs and negative sample image pairs can be obtained based on image clustering. For details, please refer to the following embodiments, which will not be described in detail here.

[0118] Step S302: Obtain the shooting time, camera identification, and type of vehicle associated with the pedestrian in each sample image pair.

[0119] The method of obtaining the shooting time, camera identification and the type of vehicle associated with the pedestrian in step S302 is similar to that in step S103 above. The only difference is the replacement of the names of each sample image and each image to be compared. This will not be described in detail here.

[0120] Step S303: Input the shooting time, camera identification and the type of vehicle associated with the pedestrian in each sample image of each sample image pair into the preset neural network model to obtain the spatiotemporal probability of pedestrians in the same space and time in the sample images of the sample image pair represented by the output of the neural network model.

[0121] The following section combines the network structure of neural networks and... Figure 4 The process of obtaining spatiotemporal probability is explained.

[0122] Specifically, a neural network model can include: an encoding layer, a feature fusion layer, and a fully connected layer.

[0123] The input information for a neural network model includes:

[0124] a: Vehicle type, indicating the type of vehicle associated with the pedestrian in image a to be compared;

[0125] a: Camera ID, which is the camera identifier for the image 'a' to be compared;

[0126] a: Timestamp, indicating the time when image a was captured;

[0127] b: Vehicle type, b: Camera ID, b: Timestamp, representing the type of vehicle associated with the pedestrian in the image to be compared, the camera identifier of the image to be compared, and the shooting time of the image to be compared, respectively.

[0128] These input information are fed into the encoding layer, which encodes the various input information to obtain dense vectors of various input information, and outputs the obtained encoded vectors to the feature fusion layer.

[0129] Then, the feature fusion layer concatenates the received encoded vectors to obtain a concatenated vector, and outputs the concatenated vector to the fully connected layer.

[0130] The fully connected layer uses multiple fully connected connections to perform vector mapping and sigmoid normalization on the concatenated vectors to obtain the spatiotemporal probability.

[0131] Step S304: Based on the obtained spatiotemporal probabilities and the positive or negative sample labels corresponding to each sample image, adjust the model parameters of the neural network model to obtain the spatiotemporal probability prediction model.

[0132] The positive sample label can be 1, and the negative sample label can be 0.

[0133] For a pair of sample images labeled with positive samples, if the same pedestrian appears in each sample image within the spatiotemporal probability representation of the sample image pair based on the sample image pair, it indicates that the spatiotemporal probability predicted by the neural network model is relatively accurate; conversely, if the same pedestrian does not appear in each sample image within the spatiotemporal probability representation of the sample image pair based on the sample image pair, it indicates that the spatiotemporal probability predicted by the neural network model is inaccurate.

[0134] For a pair of sample images labeled with negative samples, if the same pedestrian appears in each sample image within the spatiotemporal probability representation sample image pair obtained based on the sample image pair, it indicates that the spatiotemporal probability predicted by the neural network model is inaccurate; conversely, if the same pedestrian does not appear in each sample image within the spatiotemporal probability representation sample image pair obtained based on the sample image pair, it indicates that the spatiotemporal probability predicted by the neural network model is relatively accurate.

[0135] Given the above, the spatiotemporal probability output by the neural network model can be used to calculate the loss value based on the corresponding positive or negative sample labels of the sample images. The model parameters can then be adjusted according to the loss value to obtain the spatiotemporal probability prediction model.

[0136] As can be seen from the above, in the solution provided by the embodiments of this disclosure, the shooting time of the sample image, the camera identifier, and the type of vehicle associated with the pedestrian are used for model training. In this way, the spatiotemporal probability prediction model learns the relationship between spatiotemporal probability and shooting time, camera identifier, and type of vehicle associated with the pedestrian. Therefore, when performing pedestrian re-identification, the spatiotemporal probability obtained by using the spatiotemporal probability prediction model is more accurate, thus improving the accuracy of the obtained spatiotemporal probability.

[0137] The following describes how the sample image pairs S301 in the aforementioned embodiments are obtained.

[0138] In one embodiment of this disclosure, sample image pairs can be obtained in the following manner.

[0139] Identify pedestrian regions in each sample image; cluster each sample image according to the characteristics of the identified pedestrian regions to obtain multiple image classes; select sample images from the same image class to obtain positive sample image pairs containing the selected images; select sample images from different image classes to obtain negative sample image pairs containing the selected images.

[0140] In the above process, the feature vectors corresponding to the features of the determined pedestrian area can be clustered using clustering algorithms such as dbscan and kmeans to obtain the image class to which each sample image belongs.

[0141] Any two sample images selected from the same image class can form a positive sample image pair; any two sample images selected from different image classes can form a negative sample image pair.

[0142] As can be seen from the above, image classes are obtained through clustering, and positive and negative sample image pairs are obtained accordingly. The annotations of each image pair can be obtained directly from the clustering results, saving the workload of annotating each sample image one by one. That is, it avoids collecting annotation data for each sample image, and the cost of generating annotations is low.

[0143] The following is through Figure 5 The overall process of model training is explained below.

[0144] Obtain sample images, perform full object detection on the sample images, and obtain the corresponding face thumbnails for each sample image and the type of transportation associated with pedestrians in the sample images.

[0145] Then, the image features of each face thumbnail are extracted, that is, face vector extraction, to obtain the face vector corresponding to each face thumbnail.

[0146] Next, image clustering is performed based on each face vector to obtain face clusters.

[0147] During the sample construction process, images are selected from the images included in the same face cluster to form positive sample image pairs, and images are selected from the images included in different face clusters to form negative sample image pairs.

[0148] Using positive and negative sample image pairs as the training set for the spatiotemporal probability prediction model, the neural network model is trained to obtain the spatiotemporal probability prediction model.

[0149] Corresponding to the above-mentioned artificial intelligence-based pedestrian re-identification method, this disclosure also provides an artificial intelligence-based pedestrian re-identification device.

[0150] In one embodiment of this disclosure, see [link to embodiment]. Figure 6 A schematic diagram of a pedestrian re-identification device based on artificial intelligence is provided, including:

[0151] The type determination module 601 is used to determine the pedestrian area and the type of vehicle associated with the pedestrian in each image to be compared;

[0152] The similarity determination module 602 is used to obtain the visual similarity between pedestrians appearing in each image to be compared based on the determined pedestrian region.

[0153] The spatiotemporal probability determination module 603 is used to determine the spatiotemporal probability of the same pedestrian appearing in each image to be compared based on the shooting time of each image to be compared, the camera identification, and the type of vehicle associated with the pedestrian.

[0154] The re-identification module 604 is used to perform pedestrian re-identification based on the visual similarity and the spatiotemporal probability.

[0155] As can be seen from the above, the solution provided in this embodiment, in addition to obtaining the visual similarity between pedestrians appearing in each image to be compared, also determines the spatiotemporal probability based on the shooting time of each image to be compared, the camera identifier, and the type of transportation associated with the pedestrian. This means that the determination of the spatiotemporal probability takes into account the pedestrian's mobility after using various types of transportation, making the obtained spatiotemporal probability more consistent with the actual probability of the pedestrian appearing in images taken by different cameras, thus making the obtained spatiotemporal probability more accurate. Accordingly, the accuracy of pedestrian re-identification based on visual similarity and spatiotemporal probability is improved.

[0156] In one embodiment of this disclosure, the spatiotemporal probability determination module 603 is specifically used to determine the shooting location of each image to be compared based on the camera identifier of each image to be compared; obtain the predicted time for a vehicle of a target type to travel from one shooting location to another, wherein the target type is: the type of vehicle associated with the pedestrian in each image to be compared; determine the shooting interval based on the shooting time of each image to be compared; and determine the spatiotemporal probability of the same pedestrian appearing in each image to be compared based on the predicted time and the shooting interval.

[0157] As can be seen from the above, the consistency between the shooting interval and the prediction duration is considered when calculating the spatiotemporal probability, so that the spatiotemporal probability is consistent with the possible range of movement of pedestrians in the time interval, and the obtained spatiotemporal probability is more accurate.

[0158] In one embodiment of this disclosure, the spatiotemporal probability determination module 603 is specifically configured to: determine the shooting location of each image to be compared based on the camera identifier of each image to be compared; obtain the predicted time for a vehicle of a target type to travel from one shooting location to another, wherein the target type is the type of vehicle associated with the pedestrian in each image to be compared; determine the shooting interval based on the shooting time of each image to be compared; if the shooting interval is greater than the predicted time, determine the spatiotemporal probability of the same pedestrian appearing in each image to be compared as 1; if the shooting interval is not greater than the predicted time, determine the spatiotemporal probability as the ratio of the shooting interval to the predicted time.

[0159] Calculating the spatiotemporal probability in the above manner, on the one hand, if the shooting interval is greater than the predicted duration, it can be assumed that the pedestrian has sufficient time to move between the two shooting locations using the target type of vehicle. In this case, the spatiotemporal probability is 1, indicating that the pedestrian is more likely to appear in the images taken by the two cameras within the calculated time interval. Using 1 as the spatiotemporal probability is more accurate in this case. On the other hand, if the shooting interval is not greater than the predicted duration, the ratio of the shooting interval to the predicted duration is calculated as the spatiotemporal probability. In this way, the closer the shooting interval is to the predicted duration, the greater the value of the spatiotemporal probability. Since the pedestrian's mobility is kept within a certain limit when using the target type of vehicle, the time spent moving between the two shooting locations is more likely to be close to the predicted duration. Therefore, the spatiotemporal probability obtained in the above manner is more accurate.

[0160] In one embodiment of this disclosure, the spatiotemporal probability determination module 603 is specifically used to input the shooting time, camera identification, and type of vehicle associated with each image to be compared into a pre-trained spatiotemporal probability prediction model to obtain the spatiotemporal probability output by the spatiotemporal probability prediction model.

[0161] The spatiotemporal probability prediction model is a model obtained by training a preset neural network model using the shooting time of the sample image, camera identification, and the type of vehicle associated with the pedestrian, to obtain the spatiotemporal probability of pedestrians being in the same space and time.

[0162] As can be seen from the above, in the solution provided by the embodiments of this disclosure, the shooting time of the sample image, the camera identifier, and the type of vehicle associated with the pedestrian are used for training. In this way, the spatiotemporal probability prediction model learns the relationship between spatiotemporal probability and shooting time, camera identifier, and type of vehicle associated with the pedestrian. Thus, when pedestrian re-identification is performed, the spatiotemporal probability obtained by using the spatiotemporal probability prediction model can improve the accuracy of the obtained spatiotemporal probability.

[0163] In one embodiment of this disclosure, the re-identification module 604 is specifically used to perform weighted calculation on the visual similarity and the spatiotemporal probability based on the weight coefficients corresponding to the visual similarity and the spatiotemporal probability, respectively, to obtain a weighted result; and to determine whether the pedestrians appearing in each image to be compared are the same pedestrians based on the weighted result.

[0164] Because visual similarity and spatiotemporal probability have different abilities to represent pedestrian identities, their influence on the weighted results is also different. Therefore, in this implementation, by setting a weight system, the influence of visual similarity and spatiotemporal probability on the weighted results can be made more consistent with the actual situation, thereby making the weighted results more accurate.

[0165] Corresponding to the above model training method, this disclosure also provides a model training apparatus.

[0166] In one embodiment of this disclosure, see [link to embodiment]. Figure 7 A schematic diagram of a model training device is provided, comprising:

[0167] Image pair acquisition module 701 is used to acquire sample image pairs, wherein the sample image pairs include: positive sample image pairs in the included sample images where the same pedestrian appears in the same space and time and negative sample image pairs in the included sample images where the same pedestrian does not appear in the same space and time.

[0168] The image information acquisition module 702 is used to obtain the shooting time, camera identification, and type of vehicle associated with each sample image in each sample image pair;

[0169] The spatiotemporal probability acquisition module 703 is used to input the shooting time, camera identification and the type of vehicle associated with the pedestrian into a preset neural network model to obtain the spatiotemporal probability of pedestrians being in the same space and time in the sample images within the sample image pair, as represented by the output of the neural network model.

[0170] The model acquisition module 704 is used to adjust the model parameters of the neural network model based on the obtained spatiotemporal probabilities and the positive or negative sample labels corresponding to each sample image, so as to obtain a spatiotemporal probability prediction model.

[0171] As can be seen from the above, in the solution provided by the embodiments of this disclosure, the shooting time of the sample image, the camera identifier, and the type of vehicle associated with the pedestrian are used for model training. In this way, the spatiotemporal probability prediction model learns the relationship between spatiotemporal probability and shooting time, camera identifier, and type of vehicle associated with the pedestrian. Therefore, when performing pedestrian re-identification, the spatiotemporal probability obtained by using the spatiotemporal probability prediction model is more accurate, thus improving the accuracy of the obtained spatiotemporal probability.

[0172] In one embodiment of this disclosure, the image pair acquisition module 701 is specifically used to determine pedestrian regions in each sample image; cluster each sample image according to the characteristics of the determined pedestrian regions to obtain multiple image classes; select sample images from the same image class to obtain positive sample image pairs containing the selected images; and select sample images from different image classes to obtain negative sample image pairs containing the selected images.

[0173] As can be seen from the above, image classes are obtained through clustering, and positive and negative sample image pairs are obtained accordingly. The annotations of each image pair can be obtained directly from the clustering results, saving the workload of annotating each sample image one by one. That is, it avoids collecting annotation data for each sample image, and the cost of generating annotations is low.

[0174] The collection, storage, use, processing, transmission, provision, and disclosure of pedestrian personal information in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.

[0175] It should be noted that the images to be compared in this embodiment are from a publicly available dataset.

[0176] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0177] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0178] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0179] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0180] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as AI-based person re-identification or model training methods. For example, in some embodiments, the AI-based person re-identification or model training method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the AI-based person re-identification or model training method described above can be performed. Alternatively, in other embodiments, computing unit 801 may be configured by any other suitable means (e.g., by means of firmware) to perform artificial intelligence-based pedestrian re-identification or model training methods.

[0181] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0182] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0183] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0184] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0185] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0186] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0187] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0188] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A pedestrian re-identification method based on artificial intelligence, comprising: determining a pedestrian region in each image to be compared and a type of a vehicle associated with the pedestrian, the vehicle being used by the pedestrian in the image during movement; obtaining visual similarity between the pedestrians appearing in each image to be compared according to the determined pedestrian region; inputting the shooting time, camera identifier and type of the vehicle associated with the pedestrian in each image to be compared into a pre-trained spatiotemporal probability estimation model to obtain a spatiotemporal probability output by the spatiotemporal probability estimation model, wherein the spatiotemporal probability estimation model is a model for obtaining the spatiotemporal probability that the pedestrians are in the same space-time, which is obtained by training a preset neural network model using the shooting time, camera identifier and type of the vehicle associated with the pedestrian in sample images; performing pedestrian re-identification based on the visual similarity and the spatiotemporal probability; the spatiotemporal probability estimation model is trained in the following manner: obtaining sample image pairs, wherein the sample image pairs include positive sample image pairs in which the same pedestrian appears in the same space-time in the contained sample images and negative sample image pairs in which the same pedestrian does not appear in the same space-time in the contained sample images; obtaining the shooting time, camera identifier and type of the vehicle associated with the pedestrian in each sample image in each sample image pair, the vehicle being used by the pedestrian in the image during movement; inputting the shooting time, camera identifier and type of the vehicle associated with the pedestrian in each sample image in each sample image pair into a preset neural network model to obtain a spatiotemporal probability output by the neural network model, the spatiotemporal probability representing that the pedestrians in the sample images in the sample image pair are in the same space-time; adjusting model parameters of the neural network model according to the obtained spatiotemporal probability and positive sample labels or negative sample labels corresponding to each sample image pair to obtain a spatiotemporal probability estimation model. 2.The method of claim 1, further comprising: determining a shooting location of each image to be compared according to the camera identifier of each image to be compared; obtaining a predicted duration for traveling from one shooting location to another using a vehicle of a target type, wherein the target type is the type of the vehicle associated with the pedestrian in each image to be compared; determining a shooting interval according to the shooting time of each image to be compared; and determining a spatiotemporal probability that the same pedestrian appears in each image to be compared according to the predicted duration and the shooting interval.

3. The method of claim 2, wherein, The determination of the spatiotemporal probability that the same pedestrian appears in each image to be compared according to the predicted duration and the shooting interval comprises: if the shooting interval is greater than the predicted duration, determining the spatiotemporal probability that the same pedestrian appears in each image to be compared as 1; if the shooting interval is not greater than the predicted duration, determining the spatiotemporal probability as a ratio of the shooting interval to the predicted duration.

4. The method of any one of claims 1-3, wherein, The pedestrian re-identification based on the visual similarity and the spatiotemporal probability comprises: performing weighted calculation on the visual similarity and the spatiotemporal probability based on weight coefficients corresponding to the visual similarity and the spatiotemporal probability respectively to obtain a weighted result. According to the weighting result, it is judged whether the pedestrians appearing in each to-be-compared image are the same pedestrian.

5. The method of claim 1, wherein, The sample image pair is obtained by: Determining the pedestrian region in each sample image; According to the characteristics of the determined pedestrian region, clustering each sample image to obtain a plurality of image classes; Selecting sample images from the same image class to obtain a positive sample image pair containing the selected images; Selecting sample images from different image classes to obtain a negative sample image pair containing the selected images.

6. An artificial intelligence-based pedestrian re-identification device, comprising: A type determination module for determining the pedestrian region in each to-be-compared image and the type of the pedestrian-associated vehicle, which is the vehicle used by the pedestrian in the image during movement; A similarity determination module for obtaining the visual similarity between the pedestrians appearing in each to-be-compared image according to the determined pedestrian region; A space-time probability determination module for inputting the shooting time, camera identifier, and type of the pedestrian-associated vehicle of each to-be-compared image into a pre-trained space-time probability estimation model to obtain the space-time probability output by the space-time probability estimation model; wherein the space-time probability estimation model is a model obtained by training a preset neural network model using the shooting time, camera identifier, and type of the pedestrian-associated vehicle of sample images to obtain the space-time probability of the pedestrians being in the same space-time; A re-identification module for performing pedestrian re-identification based on the visual similarity and the space-time probability; The space-time probability estimation model is trained by the following modules in the following manner: An image pair obtaining module for obtaining sample image pairs, wherein the sample image pairs include positive sample image pairs containing sample images in which the same pedestrian appears in the same space-time and negative sample image pairs containing sample images in which the same pedestrian does not appear in the same space-time; An image information obtaining module for obtaining the shooting time, camera identifier, and type of the pedestrian-associated vehicle of each sample image in each sample image pair, the pedestrian-associated vehicle being the vehicle used by the pedestrian in the image during movement; A space-time probability obtaining module for inputting the shooting time, camera identifier, and type of the pedestrian-associated vehicle of each sample image in each sample image pair into a preset neural network model to obtain the space-time probability output by the neural network model, which represents the pedestrians in the sample images being in the same space-time; A model obtaining module for adjusting the model parameters of the neural network model according to the obtained space-time probability and the positive sample label or negative sample label corresponding to each sample image pair to obtain a space-time probability estimation model.

7. The apparatus of claim 6, wherein, The space-time probability determination module is also used for: The photographing locations of the to-be-compared images are determined according to the camera identifiers of the to-be-compared images; a predicted time length for a vehicle of a target type to travel from one photographing location to another photographing location is obtained, wherein the target type is a type of vehicle associated with a pedestrian in the to-be-compared images; photographing intervals are determined according to photographing times of the to-be-compared images; and a spatiotemporal probability of a same pedestrian appearing in the to-be-compared images is determined according to the predicted time length and the photographing intervals.

8. The apparatus of claim 7, wherein, The spatiotemporal probability determination module is specifically configured to determine photographing locations of the to-be-compared images according to camera identifiers of the to-be-compared images; obtain a predicted time length for a vehicle of a target type to travel from one photographing location to another photographing location, wherein the target type is a type of vehicle associated with a pedestrian in the to-be-compared images; determine photographing interval lengths according to photographing times of the to-be-compared images; if the photographing interval is greater than the predicted time length, determine that the spatiotemporal probability of a same pedestrian appearing in the to-be-compared images is 1; and if the photographing interval is not greater than the predicted time length, determine that the spatiotemporal probability is a ratio of the photographing interval to the predicted time length.

9. The apparatus of any one of claims 6-8, wherein, The re-identification module is specifically configured to perform weighted calculation on the visual similarity and the spatiotemporal probability based on weight coefficients corresponding to the visual similarity and the spatiotemporal probability, respectively, to obtain a weighted result; and determine whether the pedestrians appearing in the to-be-compared images are a same pedestrian according to the weighted result.

10. The apparatus of claim 6, wherein, The image pair obtaining module is specifically configured to determine pedestrian regions in the sample images; perform clustering on the sample images according to features of the determined pedestrian regions to obtain a plurality of image classes; select sample images from a same image class to obtain a positive sample image pair containing the selected images; and select sample images from different image classes to obtain a negative sample image pair containing the selected images.

11. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

12. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-5.

13. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Pedestrian re-identification method and system based on space-time joint model of map data

    CN111178284A

  • Model training method, image feature extraction method and target detection method and device

    CN112232384A