Target re-identification method, target registration method and related devices
By extracting single-frame image features and image feature difference information from videos with different imaging methods, and combining them with graph neural networks and location identification information, the problem of insufficient robustness of target re-identification in existing technologies is solved, and higher accuracy target recognition and registration is achieved.
Patent Information
- Application Number
- CN202111484845.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2041-12-07
AI Technical Summary
Existing target re-identification technologies rely on single features, which lack robustness and representational ability, resulting in poor identification results.
Video acquisition using different imaging methods (such as infrared and visible light) extracts single-frame image features and image feature difference information from video frames. These features are then fused using graph neural networks for identification and registration. Finally, a topology graph is constructed using location identification information for feature fusion.
It improves the robustness of target features and the accuracy of re-identification results, and enhances the recognition capability of cross-device target matching.
Smart Images

Figure CN114419532B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer vision, in particular to a target re-identification method, a target registration method and an electronic device. BACKGROUND
[0002] With the rise of intelligent monitoring technology, in recent years it has attracted extensive research and attention of the industry and academia, and target re-identification technology as an important branch of intelligent monitoring is also more and more concerned by researchers.
[0003] The problem to be solved by target re-identification is to search and match the same target between different devices, that is, cross-device target re-identification. However, the current technical means mainly relies on the extraction of static external features of the overall posture of the target, such as clothing, hairstyle, backpack, umbrella, etc. The current various technical solutions for target re-identification task mostly rely on single features, which are not strong in robustness and representation ability, resulting in general final recognition results. SUMMARY
[0004] The present application provides a target re-identification method, a target registration method and an electronic device, which can improve the robustness and representation of target features, thereby improving the accuracy of the re-identification result.
[0005] To solve the above technical problems, one technical solution adopted by the present application is to provide a target re-identification method, which comprises: acquiring a first video and a second video collected in the same time period for a target to be identified, the imaging modes of the first video and the second video being different; extracting second features of the first video and second features of the second video respectively; wherein the second features include single-frame image features of each video frame in the corresponding video, and image feature difference information between different video frames in the corresponding video; and identifying the target to be identified based on the second features of the first video and the second features of the second video.
[0006] To solve the above technical problems, one technical solution adopted by the present application is to provide a target registration method, which comprises: acquiring a first video and a second video collected in the same time period for a candidate target, the imaging modes of the first video and the second video being different; extracting second features of the first video and second features of the second video respectively; wherein the second features include single-frame image features of each video frame in the corresponding video, and image feature difference information between different video frames in the corresponding video; and registering the candidate target based on the second features of the first video and the second features of the second video.
[0007] To solve the above technical problems, a technical solution adopted in this application is: to provide an electronic device, which includes a processor and a memory connected to the processor; the memory stores program instructions; and the processor is used to execute the program instructions stored in the memory to implement the above technical solution.
[0008] In order to solve the above technical problems, a technical solution adopted in this application is: providing a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to execute the above technical solution.
[0009] Compared with the prior art, the beneficial effects of the present application are as follows: through the above-mentioned method, the present application provides a target re-identification method, a target registration method, an electronic device, and a computer-readable storage medium. By acquiring a first video and a second video captured in the same time period for the target to be identified, the first video and the second video have different imaging modes; extracting the second features of the first video and the second video respectively; wherein the second features include the single-frame image features of each video frame in the corresponding video, and the image feature difference information between different video frames in the corresponding video; based on the second features of the first video and the second features of the second video, the target to be identified is identified. This improves the robustness and representation power of the target features and the accuracy of the re-identification results.
[0010] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0012] Figure 1 This is a flow chart of an embodiment of the target re-identification method of the present application;
[0013] Figure 2 This application Figure 1 Specific process diagram of S13;
[0014] Figure 3 This application Figure 1 Specific process diagram of S12;
[0015] Figure 4 This application Figure 3 Specific process diagram of S34;
[0016] Figure 5 is a specific flowchart of S31 and S32 in the application Figure 3
[0017] Figure 6 is a specific flowchart of S33 in the application Figure 3
[0018] Figure 7 is a specific flowchart of S35 in the application Figure 3
[0019] Figure 8 is a network structure diagram for extracting and fusing features in the target re-identification method and the target registration method of the application
[0020] Figure 9 is a flowchart of an embodiment of the target registration method of the application
[0021] Figure 10 is a specific flowchart of S93 in the application Figure 9
[0022] Figure 11 is a structural diagram of an embodiment of the electronic device of the application
[0023] Figure 12 is a structural diagram of an embodiment of the computer readable storage medium of the application DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, but not all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work, fall within the scope of protection of the application.
[0025] Some embodiments of the application will be described in detail below with reference to the drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.
[0026] The terms "first", "second", and the like in the application are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0027] Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all referring to a common embodiment, or an embodiment that is independent of other embodiments. One of ordinary skill in the art will readily recognize from the disclosure herein, that embodiments of the present application can be combined, combined and permuted with other embodiments.
[0028] Figure 1 is a flowchart of an embodiment of the target re-identification method of the present application. It should be noted that the sequence of the flow shown in the embodiment is not limited if there is substantially the same result. As shown in Figure 1 , the embodiment includes: Figure 1
[0029] S11: Obtain a first video and a second video collected in the same time period for a target to be identified, the imaging modes of the first video and the second video being different.
[0030] The imaging mode of the video can be infrared imaging, visible light imaging, or X-ray imaging in a security check scene, and the like, so as to obtain an infrared video stream, a visible light video stream, or an X-ray video stream, and the like, which are not limited herein.
[0031] In the embodiment, the first video is an infrared video stream, and the second video is a visible light video stream. The collection device of the infrared video stream can be an infrared camera, and the collection device of the visible light video stream can be a visible light camera, so as to obtain the infrared video stream and the visible light video stream.
[0032] The single image in the infrared video stream can be a thermal image, a depth image, and the like, which are not limited herein.
[0033] The single image in the visible light video stream can be RGB, HSV, SILTP, and the like, which are not limited herein.
[0034] S12: Extract the second features of the first video and the second features of the second video, respectively; wherein the second features include single-frame image features of each video frame in the corresponding video, and image feature difference information between different video frames in the corresponding video.
[0035] The single-frame image features can be features extracted based on pixel value information in the image data; the image feature difference information can be features extracted based on pixel value difference information between multiple frames of images, and can be between consecutive frames.
[0036] The specific steps can be referred to in detail in Figure 3 .
[0037] S13: identifying the to-be-identified target based on the second feature of the first video and the second feature of the second video.
[0038] The specific steps can be referred to in detail Figure 2 and Figure 3 .
[0039] Figure 2 is a specific flowchart of S13 in the present application Figure 1 . In the present embodiment, in combination with reference to Figure 1 , S13 can be further extended to the following sub-steps:
[0040] Before fusing the second feature of the first video, the second feature of the second video and the location identification information to obtain the fused feature, the target re-identification method further comprises: determining the location identification information between the device for collecting the first video and the device for collecting the second video.
[0041] Among them, the location identification information can represent information indicating whether the devices are located at the same location. Specifically, the location identification information can be obtained based on the device number and / or the location information (for example, longitude and latitude, etc.) of the device.
[0042] S21: fusing the second feature of the first video, the second feature of the second video and the location identification information to obtain the fused feature.
[0043] Fusing the second feature of the infrared video, the second feature of the visible light video and the location identification information of the device for acquiring the infrared video and the device for acquiring the visible light video to obtain the fused feature.
[0044] Specifically, taking the second feature of the first video as an infrared node; taking the second feature of the second video as a visible light node; taking the device information corresponding to the first video and the device information corresponding to the second video as a device node. The infrared node, the visible light node and the device node constitute a second topology graph.
[0045] Among them, the information of the device node includes, but is not limited to, the device number after One-Hot and Embedding, the location (for example, longitude and latitude) of the device, etc.
[0046] After obtaining the second topology graph, inputting the second topology graph into a second graph neural network, i.e., a back-end graph neural network, and inputting the output of the back-end graph neural network into a fully connected neural network. The training loss is a multi-classification cross-entropy loss and is trained on a server until the entire network converges. The training loss can be mean square error, mean absolute error or other training losses, which are not limited here.
[0047] Thus, the fused feature is obtained.
[0048] S22: identifying the target to be identified by using the fusion feature.
[0049] The fusion feature of the target to be identified is compared with the fusion feature of the candidate target in the database to obtain a comparison result. According to the comparison result, a re-identification result of the target to be identified is determined.
[0050] The specific steps can be referred to in detail Figure 7 .
[0051] Figure 3 is a specific flowchart of S12 in the present application Figure 1 . In the present embodiment, in combination with reference to Figure 1 , S12 can be further extended to the following sub-steps:
[0052] The first video and the second video are respectively taken as a video to be processed, and the following operations are performed:
[0053] S31: obtaining multiple frames of first infrared feature maps and multiple frames of first visible light feature maps of the target to be identified.
[0054] The specific steps can be referred to in detail Figure 5 .
[0055] S32: the multiple frames of first infrared feature maps and the multiple frames of first visible light feature maps are respectively cut into multiple first infrared sub-maps and multiple first visible light sub-maps.
[0056] The specific steps can be referred to in detail Figure 5 .
[0057] S33: first infrared topology maps and first visible light topology maps are respectively constructed by taking the first infrared sub-maps and the first visible light sub-maps as nodes.
[0058] The specific steps can be referred to in detail Figure 6 .
[0059] S34: feature extraction is respectively performed on the first infrared topology maps and the first visible light topology maps to respectively obtain second features of the first video and the second video, wherein the second features of the first video and the second video contain static features of the target to be identified and dynamic association information between different frames.
[0060] The specific steps can be referred to in detail Figure 4 .
[0061] S35: identifying the target to be identified based on the second features of the first video and the second features of the second video.
[0062] The specific steps can be referred to in detail Figure 2 and Figure 7 .
[0063] Figure 4 is a specific flowchart of S34 in this application Figure 3 . In this embodiment, in combination with reference to Figure 3 , S34 can be further expanded into the following sub-steps:
[0064] S41: using an average pooling layer to respectively perform average pooling on the first infrared topology graph and the first visible light topology graph to obtain an average-pooled first infrared topology graph and an average-pooled first visible light topology graph respectively.
[0065] The first infrared topology graph is subjected to average pooling through the average pooling layer. Specifically, the first infrared topology graph is divided into a plurality of regions, and the values in each region are subjected to average arithmetic. The average pooling can extract all feature information of the first infrared topology graph, and thus can retain more background information.
[0066] The first visible light topology graph is subjected to average pooling through the average pooling layer. Specifically, the first visible light topology graph is divided into a plurality of regions, and the values in each region are subjected to average arithmetic. The average pooling can extract all feature information of the first visible light topology graph, and thus can retain more background information.
[0067] S42: using a first graph neural network to respectively perform feature extraction on the average-pooled first infrared topology graph and the average-pooled first visible light topology graph to obtain second features of the first video and the second video.
[0068] Specifically, the first graph neural network performs feature extraction on the first infrared topology graph to obtain the second features of the first video, i.e., the graph convolution features of the infrared video modal. The first graph neural network performs feature extraction on the first visible light topology graph to obtain the second features of the first video, i.e., the graph convolution features of the visible light video modal.
[0069] Figure 5 is a specific flowchart of S31 and S32 in this application Figure 3 . In this embodiment, in combination with reference to Figure 3 , S31 and S32 can be further expanded into the following sub-steps:
[0070] S51: obtaining a first video and a second video containing a target to be recognized.
[0071] The first video and the second video containing the target to be recognized are obtained by a camera device. The first video and the second video include an original visible light video stream and an original infrared video stream, wherein the original visible light video stream is obtained by a visible light camera, and the original infrared video stream is obtained by an infrared camera.
[0072] A single image in the visible light video stream can be RGB, HSV, SILTP, etc., which is not limited here.
[0073] The single image in the infrared video stream can be a thermal image, a depth image, etc., which is not limited herein.
[0074] S52: Extracts multiple frames of first original images and multiple frames of second original images from the first video and the second video respectively.
[0075] Extracts single images in the multiple frames of original visible light video streams. Specifically, every N frames in the visible light video stream is extracted as a representative frame, and finally T frames are reserved as key frames of the visible light video stream, i.e., multiple frames of original images of the visible light video stream.
[0076] Extracts single images in the multiple frames of original infrared video streams. Specifically, every N frames in the infrared video stream is extracted as a representative frame, and finally T frames are reserved as key frames of the infrared video stream, i.e., multiple frames of original images of the infrared video stream.
[0077] The multiple frames of original images include the multiple frames of original images of the visible light video stream and the multiple frames of original images of the infrared video stream. N and T are integers greater than 0.
[0078] S53: Performed target recognition and target tracking on the multiple frames of first original images and the multiple frames of second original images respectively to obtain a target detection frame surrounding the to-be-identified target on the first original image and the second original image.
[0079] The multiple frames of original images of the visible light video stream are parsed into a visible light image sequence of consecutive frames through a target detection and target tracking algorithm, and the detected and tracked targets are circled with a target detection frame in the visible light image sequence of consecutive frames.
[0080] The multiple frames of original images of the infrared video stream are parsed into an infrared image sequence of consecutive frames through a target detection and target tracking algorithm, and the detected and tracked targets are circled with a target detection frame in the infrared image sequence of consecutive frames.
[0081] S54: Cropped images within the target detection frame from at least part of the first original image and the second original image, and performed size normalization to obtain multiple frames of first sub-original images and multiple frames of second sub-original images.
[0082] After the detected target in the visible light image sequence is circled with a target detection frame, a sub-image (i.e., a target detection frame part, which can be part or all of the image) with the target in the visible light image sequence is found. Multiple visible light sub-images are cropped from the visible light image sequence to obtain visible light sub-images. The obtained visible light sub-images are size-normalized to obtain standard target visible light images, i.e., sub-original images of the visible light images, thereby obtaining a standard target visible light image sequence, i.e., multiple frames of sub-original images of the visible light image sequence; the size can be set according to the situation, which is not limited herein.
[0083] The target detected in the infrared image sequence is circled with a target detection frame, i.e., a sub-image with a target in the infrared sequence (i.e., a target detection frame part, which can be part or all of an image) is obtained, and the infrared sub-image is cropped from the infrared image sequence to obtain an infrared sub-image. The obtained infrared sub-image is subjected to size normalization to obtain a standard target infrared image, i.e., a sub-original image of an infrared image, so that a standard target infrared image sequence, i.e., a plurality of sub-original images of the infrared image sequence, is obtained. The size can be set according to the situation, and is not limited here.
[0084] The plurality of sub-original images include a plurality of sub-original images of the infrared image sequence and a plurality of sub-original images of the visible light image sequence.
[0085] S55: Feature extraction is performed on the plurality of first sub-original images and the plurality of second sub-original images respectively by using a convolutional neural network to obtain a plurality of first infrared feature maps and a plurality of first visible light feature maps.
[0086] Feature extraction is performed on the plurality of sub-original images of the visible light image sequence obtained. In this embodiment, the feature extraction is performed by using a convolutional neural network, but other embodiments are also possible, and are not limited here. Feature extraction is performed on each visible light image in the standard target visible light image sequence by using a convolutional neural network to obtain a first feature map corresponding to the visible light image, so that a plurality of first feature maps of the visible light image sequence are obtained.
[0087] Feature extraction is performed on the plurality of sub-original images of the infrared image sequence obtained. In this embodiment, the feature extraction is performed by using a convolutional neural network, but other embodiments are also possible, and are not limited here. Feature extraction is performed on the standard target infrared image sequence by using a convolutional neural network to obtain a first feature map corresponding to the infrared image, so that a plurality of first feature maps of the infrared image sequence are obtained.
[0088] The plurality of first feature maps include a plurality of first feature maps of the visible light image sequence (i.e., a plurality of first visible light feature maps) and a plurality of first feature maps of the infrared image sequence (i.e., a plurality of first infrared feature maps).
[0089] Figure 6 is a specific flowchart of S33 in this application Figure 3 . In this embodiment, in combination with Figure 3 , S33 can be further expanded into the following sub-steps:
[0090] S61: The plurality of first infrared feature maps and the plurality of first visible light feature maps are respectively cut into a plurality of first infrared sub-images and a plurality of first visible light sub-images by using the same cutting manner.
[0091] The first feature map of each frame of the visible light image sequence is segmented, and the first feature map is segmented into a plurality of subgraphs from top to bottom. The specific segmentation manner is not limited here. For example, the subgraphs can be horizontally or vertically segmented, but the segmentation manner of each frame is the same.
[0092] The first feature map of each frame of the infrared image sequence is segmented, and the first feature map is segmented into a plurality of subgraphs from top to bottom. The specific segmentation manner is not limited here. For example, the subgraphs can be horizontally or vertically segmented, but the segmentation manner of each frame is the same.
[0093] S62: Establish a connection between the first infrared subgraphs corresponding to the same frame, and establish a connection between the first infrared subgraphs corresponding to the same segmentation position and corresponding to different frames.
[0094] The subgraphs of the first feature map of each visible light image are taken as a node, and the subgraphs in a frame of the first feature map are connected. Then, the nodes at the same position of the first feature maps of multiple frames are connected to obtain a node feature map of multiple frames. The node feature map of multiple frames is input into a front-end graph neural network to obtain fine-grained static features of a target in a single frame of image and dynamic association information between different frames in multiple frames, so as to obtain a second visible light feature map, that is, a first visible light topology map corresponding to the first feature map of the multiple frames of visible light images.
[0095] S63: Establish a connection between the first visible light subgraphs corresponding to the same frame, and establish a connection between the first visible light subgraphs corresponding to the same segmentation position and corresponding to different frames.
[0096] The subgraphs of the first feature map of each infrared image are taken as a node, and the subgraphs in a frame of the first feature map are connected. Then, the nodes at the same position of the first feature maps of multiple frames are connected to obtain a node feature map of multiple frames. The node feature map of multiple frames is input into a front-end graph neural network to obtain fine-grained static features of a target in a single frame of image and dynamic association information between different frames in multiple frames, so as to obtain a second infrared feature map, that is, a first infrared topology map corresponding to the first feature map of the multiple frames of infrared images.
[0097] Figure 7 is a specific flowchart of S35 in the present application Figure 3 . In the present embodiment, in combination with reference to Figure 3 , S35 can be further expanded into the following sub-steps:
[0098] S71: Calculate the similarity of the fusion feature and the fusion feature of one or more candidate targets.
[0099] The similarity between the fusion feature of the target to be identified and the fusion feature of one or more candidate targets is calculated. The similarity can be a cosine similarity, a Euclidean distance, etc., which is not limited here.
[0100] S72: determining the re-identification result of the target to be identified from the one or more candidate targets according to the similarity.
[0101] The obtained similarity is sorted, and the final re-identification result is obtained according to the sorted similarity. The sorting order can be simply sorted according to the size of the similarity, or can be sorted based on the weight of different features as needed.
[0102] For reference Figures 1-7 the steps of the embodiment in Figure 8 is a network structure diagram for extracting the fusion feature in the target re-identification method and the target registration method of the present application.
[0103] As Figure 8 shown, the multi-frame visible light image (i.e., visible light video) is input into a graph convolution network to obtain a visible light feature map; the visible light feature map is input into a front-end graph neural network (GCN, Graph Neural Networks) after average pooling to obtain the modal feature of the multi-frame visible light image, i.e., a visible light modal node.
[0104] The multi-frame infrared image (i.e., infrared video) is input into a graph convolution network to obtain an infrared feature map; the infrared feature map is input into a front-end graph neural network (GCN, Graph Neural Networks) after average pooling to obtain the modal feature of the multi-frame infrared image, i.e., an infrared modal node.
[0105] The device information, such as the device number, etc., is embedded into a single device information modal node after one-hot encoding.
[0106] The three modal nodes are combined to obtain a multi-modal fusion feature input into a back-end graph neural network; the output of the back-end neural network is input into a fully connected layer.
[0107] The training loss of the network structure is a multi-classification cross-entropy loss and is trained on a server until the entire network converges. The training loss can be a mean square error, a mean absolute error, etc., other training losses, which are not limited here.
[0108] Figure 9 is a flowchart of an embodiment of the target registration method of the present application. It should be noted that the flow order shown in Figure 9 is not limited by the embodiment of the present application. As Figure 9 shown, the embodiment includes:
[0109] S91: acquire a first video and a second video collected at the same time period for the candidate target, the first video and the second video being different in imaging mode.
[0110] The imaging mode of the video can be infrared imaging, visible light imaging, X-ray imaging in a security check scene, etc., thereby obtaining an infrared video stream, a visible light video stream, or an X-ray video stream, etc., which are not limited herein.
[0111] In this embodiment, the first video is an infrared video stream, and the second video is a visible light video stream. The acquisition device of the infrared video stream can be an infrared camera, and the acquisition device of the visible light video stream can be a visible light camera, thereby obtaining the infrared video stream and the visible light video stream.
[0112] The single image in the infrared video stream can be a thermal image, a depth image, etc., which are not limited herein.
[0113] The single image in the visible light video stream can be RGB, HSV, SILTP, etc., which are not limited herein.
[0114] S92: extract a second feature of the first video and a second feature of the second video, respectively; wherein the second feature includes single-frame image features of each video frame in the corresponding video, and image feature difference information between different video frames in the corresponding video.
[0115] The single-frame image feature can be a feature extracted based on pixel value information in the image data; the image feature difference information can be a feature extracted based on pixel value difference information between multiple frames of images, which can be between continuous frames.
[0116] S93: register the candidate target based on the second feature of the first video and the second feature of the second video.
[0117] The specific steps can be referred to in detail in Figure 10 .
[0118] Figure 10 is a specific flowchart of S93 in the present application Figure 9 . In this embodiment, in combination with reference to Figure 9 , S93 can be further extended to the following sub-steps:
[0119] Before fusing the second feature of the first video, the second feature of the second video, and the location identification information to obtain the fusion feature, the target re-identification method further includes: determining the location identification information between the device acquiring the first video and the device acquiring the second video.
[0120] The position identification information can represent information indicating whether the devices are located at the same position. Specifically, the position identification information can be obtained based on a device number and / or position information (e.g., longitude and latitude) of the device.
[0121] S101: Fuse the second feature of the first video, the second feature of the second video, and the position identification information to obtain a fused feature.
[0122] Fuse the second feature of the infrared video, the second feature of the visible light video, and the position identification information of the device obtaining the infrared video and the device obtaining the visible light video to obtain a fused feature.
[0123] Specifically, the second feature of the first video is taken as an infrared node; the second feature of the second video is taken as a visible light node; and the device information obtained for the first video and the device information obtained for the second video are taken as a device node. The infrared node, the visible light node, and the device node form a second topology graph.
[0124] The information of the device node includes, but is not limited to, a device number embedded through one-hot encoding, and position information (e.g., longitude and latitude) of the device.
[0125] After obtaining the second topology graph, the second topology graph is input into a second graph neural network, i.e., a backend graph neural network, and an output of the backend graph neural network is input into a fully connected neural network. A training loss is a multi-classification cross-entropy loss and is trained on a server until the entire network converges. The training loss can be a mean square error, a mean absolute error, or other training losses, which are not limited herein.
[0126] Thus, the fused feature is obtained.
[0127] S102: Register the fused feature in a database.
[0128] The fused feature of the candidate target is stored in the database. The fused feature of the candidate target in the database can be classified in a certain manner, which includes, but is not limited to, being marked in chronological order through encoding, or being classified according to obtained device information, so as to be queried and compared subsequently. The device information includes a device number and position information (e.g., longitude and latitude) of the device, which are not limited herein.
[0129] Figure 11 is a structural schematic diagram of an embodiment of an electronic device of the present application. As shown in Figure 10 The electronic device includes a processor 111 and a memory 112 coupled to the processor 111.
[0130] The memory 112 stores program instructions for implementing the method of any of the above embodiments. The processor 111 is configured to execute the program instructions stored in the memory 112 to implement the steps of the above method embodiments. The processor 111 can also be referred to as a CPU (Central Processing Unit). The processor 111 can be an integrated circuit chip having a processing capability. The processor 111 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application-Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general-purpose processor can be a microprocessor or the processor 111 can also be any conventional processor.
[0131] Figure 12 is a structural schematic diagram of an embodiment of the computer readable storage medium of the present application. As shown in Figure 12 the computer readable storage medium stores computer executable instructions for causing a computer to execute the steps of the above method embodiments.
[0132] The computer readable storage medium 120 of the embodiments of the present application stores computer executable instructions 121 which, when executed, implement the method provided by the above embodiments of the present application. The computer executable instructions 121 can form a program file and be stored in the above computer readable storage medium 120 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor executes all or part of the steps of the method of each embodiment of the present application. The aforementioned computer readable storage medium 120 includes a U disk, a mobile hard disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk or an optical disk, etc. various media that can store program codes, or a computer, a server, a mobile phone, a tablet, etc. terminal device.
[0133] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus embodiments described above are merely illustrative, for example, the flow charts and structural charts in the drawings show the possible implementation architecture, function and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flow chart or block diagram can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logic function. It should also be noted that in the alternative implementation, the functions annotated in the blocks can also occur in the order different from that annotated in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can also be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structural chart and / or flow chart, and the combination of blocks in the structural chart and / or flow chart, can be implemented by a dedicated hardware-based system for executing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0134] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above description is merely an implementation of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A target re-identification method, characterized in that, The method comprises: acquiring a first video and a second video collected at the same time period for a target to be identified, the first video and the second video being different in imaging mode; extracting second features of the first video and second features of the second video respectively; wherein the second features comprise single-frame image features of each video frame in the corresponding video, and image feature difference information between different video frames in the corresponding video; identifying the target to be identified based on the second features of the first video and the second features of the second video; the step of extracting the second features of the first video and the second features of the second video respectively comprises: respectively taking the first video and the second video as a video to be processed, and performing the following operations: acquiring a plurality of first infrared feature maps and a plurality of first visible light feature maps of the target to be identified; dividing the plurality of first infrared feature maps and the plurality of first visible light feature maps into a plurality of first infrared sub-maps and a plurality of first visible light sub-maps respectively; constructing a first infrared topological graph and a first visible light topological graph respectively by taking the first infrared sub-maps and the first visible light sub-maps as nodes; performing feature extraction on the first infrared topological graph and the first visible light topological graph respectively to obtain the second features of the first video and the second video respectively, wherein the second features of the first video and the second video contain static features of the target to be identified and dynamic association information between different frames.
2. The method of claim 1, wherein, The method further comprises: determining position identification information between a device for collecting the first video and a device for collecting the second video; the step of identifying the target to be identified based on the second features of the first video and the second features of the second video comprises: fusing the second features of the first video, the second features of the second video and the position identification information to obtain fused features; identifying the target to be identified by using the fused features.
3. The method of claim 1, wherein, the step of performing feature extraction on the first infrared topological graph and the first visible light topological graph respectively to obtain the second features of the first video and the second video respectively comprises: performing average pooling on the first infrared topological graph and the first visible light topological graph respectively by using an average pooling layer to obtain an average-pooled first infrared topological graph and an average-pooled first visible light topological graph respectively; performing feature extraction on the average-pooled first infrared topological graph and the average-pooled first visible light topological graph respectively by using a first graph neural network to obtain the second features of the first video and the second video.
4. The method of claim 1, wherein, the step of acquiring a plurality of first infrared feature maps and a plurality of first visible light feature maps of the target to be identified comprises: acquiring the first video and the second video containing the target to be identified; extracting a plurality of first original images and a plurality of second original images from the first video and the second video respectively; performing target identification and target tracking on the plurality of first original images and the plurality of second original images respectively to obtain a target detection box surrounding the target to be identified on the first original image and the second original image; cropping images within the target detection frame from at least part of the first original image and the second original image respectively, and performing size normalization to obtain a plurality of frames of first sub-original images and a plurality of frames of second sub-original images; performing feature extraction on the plurality of frames of first sub-original images and the plurality of frames of second sub-original images respectively by using a convolutional neural network to obtain a plurality of frames of first infrared feature maps and a plurality of frames of first visible light feature maps.
5. The method of claim 1, wherein, The step of constructing a first infrared topological graph and a first visible light topological graph respectively with the first infrared sub-graph and the first visible light sub-graph as nodes comprises: The plurality of frames of first infrared feature maps and the plurality of frames of first visible light feature maps are respectively divided into a plurality of first infrared sub-graphs and a plurality of first visible light sub-graphs by using the same segmentation method; connections are established between first infrared sub-graphs corresponding to the same frame, and connections are established between first infrared sub-graphs corresponding to the same segmentation position and corresponding to different frames; connections are established between first visible light sub-graphs corresponding to the same frame, and connections are established between first visible light sub-graphs corresponding to the same segmentation position and corresponding to different frames.
6. The method of claim 2, wherein, The step of identifying the target to be identified by using the fusion feature comprises: calculating the similarity of the fusion feature and the fusion feature of one or more candidate targets; determining the re-identification result of the target to be identified from the one or more candidate targets according to the similarity.
7. A target registration method, characterized by, comprises: obtaining a first video and a second video collected for a candidate target in the same time period, the imaging modes of the first video and the second video being different; extracting second features of the first video and second features of the second video respectively; wherein the second features comprise single-frame image features of each video frame in the corresponding video, and image feature difference information between different video frames in the corresponding video; registering the candidate target based on the second features of the first video and the second features of the second video; The step of extracting the second features of the first video and the second features of the second video respectively comprises: respectively taking the first video and the second video as a video to be processed, and performing the following operations: obtaining a plurality of frames of first infrared feature maps and a plurality of frames of first visible light feature maps of the candidate target; dividing the plurality of frames of first infrared feature maps and the plurality of frames of first visible light feature maps into a plurality of first infrared sub-graphs and a plurality of first visible light sub-graphs respectively; constructing a first infrared topological graph and a first visible light topological graph respectively with the first infrared sub-graph and the first visible light sub-graph as nodes; performing feature extraction on the first infrared topological graph and the first visible light topological graph respectively to obtain second features of the first video and the second video respectively, wherein the second features of the first video and the second video contain static features of the candidate target and dynamic association information between different frames.
8. The method of claim 7, wherein, The method further comprises: determining position identification information between a device for collecting the first video and a device for collecting the second video; The step of registering the candidate target based on the second features of the first video and the second features of the second video comprises: fusing the second feature of the first video, the second feature of the second video and the location identification information to obtain a fused feature; registering the fused feature in a database.
9. An electronic device, comprising: comprising a processor and a memory connected to the processor, wherein the memory stores program instructions; the processor is configured to execute the program instructions stored in the memory to implement the method in any one of claims 1-6 or 7-8.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions for causing the computer to execute the method in any one of claims 1-6 or 7-8.
Citation Information
Patent Citations
Behavior recognition method and device and terminal equipment
CN110633630A
Target object recognition method and device, electronic equipment and storage medium
CN110781711A