Target matching method and target matching device

By correcting and supplementing feature dimensions in multiple devices with different computing power, the accuracy and efficiency of cross-lens target matching are solved, and efficient target tracking and tracking are achieved.

CN120279062APending Publication Date: 2025-07-08ENTROPY CLOUD BRAIN MACHINE (HANGZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510397797.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In multiple devices with different computing power, how to efficiently and accurately match and track targets, especially when the feature extraction dimensions are inconsistent, achieve cross-lens target correlation.

Method used

By obtaining the video feature extraction results of multiple terminal devices, the correction process is performed to make the feature dimensions consistent, and then the target matching is determined based on the modified feature results, and feature supplementation and deletion are used to use preset correspondence and feature importance sorting to achieve cross-device target tracking.

Benefits of technology

It improves the accuracy and efficiency of target matching, improves the effect of cross-lens tracking, adapts to equipment with different computing power, and improves the equipment operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279062A_ABST
    Figure CN120279062A_ABST
Patent Text Reader

Abstract

The invention discloses a target matching method and a target matching device. The method comprises the following steps: acquiring first feature extraction results of detection targets in videos collected by a plurality of terminal devices for the same picture scene; if the feature dimensions in the first feature extraction results are different, correcting at least one first feature extraction result to obtain a target feature extraction result; and based on the target feature extraction result, determining a matching result of the detection target in the video collected by each terminal device. Thus, under the condition that the feature dimensions in the first feature extraction results are different, the features extracted by the terminal devices with different computing capacities can be associated by performing correction processing on one or more first feature extraction results, so that target feature extraction results with the same feature dimensions are obtained, efficient and accurate target matching is realized, and the target matching efficiency is improved. And the target tracking effect and efficiency are improved based on the target matching result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent transportation technologies, and in particular, to a target matching method and a target matching device. Background Art

[0002] The real-time monitoring and tracking technologies of pedestrians, vehicles, etc. are one of the important technologies in intelligent transportation systems. Without human intervention or with only a small amount of intervention, the tracking algorithm analyzes the video sequence recorded by the camera to achieve the detection and tracking of pedestrians, vehicles, etc. The combination of multi-lens videos can integrate the monitoring lenses, realize the tracking of the same target under different lenses, greatly improve the information transmission efficiency, provide data for cross-video intelligent analysis, and achieve cross-lens (lens) tracking more efficiently. It can perform global real-time monitoring of large scenes and quickly retrieve historical events.

[0003] However, the most important thing in cross-lens tracking is that the same target can be correctly matched. Although extracting more features can more accurately represent the target information, in order to meet the efficiency requirements, each device is limited by computing power and cannot extract feature information of the same dimension. Therefore, how to correlate the features extracted from multiple devices with different computing powers for target matching and judge target tracking based on the target matching result is an urgent problem to be solved. Summary of the Invention

[0004] The embodiments of this application provide a target matching method and a target matching device. The method can achieve efficient and accurate target matching based on the feature extraction results of multiple terminal devices with different computing powers, so as to improve the effect and efficiency of target tracking based on the target matching result.

[0005] The technical solution of this application is implemented as follows:

[0006] This application provides a target matching method, including:

[0007] Obtaining first feature extraction results of detection targets in videos respectively collected by multiple terminal devices for the same picture scene; the first feature extraction results include multi-dimensional features of the detection targets;

[0008] If the feature dimensions in each of the first feature extraction results are different, performing correction processing on at least one of the first feature extraction results to obtain target feature extraction results; the feature dimensions in each of the target feature extraction results are the same;

[0009] Based on the target feature extraction results, determining matching results of detection targets in the videos respectively collected by each terminal device.

[0010] This application provides a target matching device, including:

[0011] An acquisition module, configured to acquire first feature extraction results of detection targets in videos respectively collected by multiple terminal devices for the same picture scene; the first feature extraction results include multi-dimensional features of the detection targets;

[0012] A correction module, configured to, if the feature dimensions in each of the first feature extraction results are different, perform correction processing on at least one of the first feature extraction results to obtain target feature extraction results; the feature dimensions in each of the target feature extraction results are the same;

[0013] A determination module, configured to determine matching results of detection targets in videos respectively collected by each terminal device based on the target feature extraction results. Description of the Drawings

[0014] Figure 1 It is a schematic flowchart of a target matching method provided by an embodiment of the present application;

[0015] Figure 2 It is a schematic structural diagram of a cross-device tracking system provided by an embodiment of the present application;

[0016] Figure 3 It is a schematic flowchart of a method for feature extraction and fusion of a computing power-unbalanced edge device provided by an embodiment of the present application;

[0017] Figure 4 It is a schematic flowchart of a method for target matching based on a fused feature code provided by an embodiment of the present application;

[0018] Figure 5 It is a schematic composition structure diagram of a target matching device provided by an embodiment of the present application. Detailed Embodiments

[0019] In order to be able to understand the features and technical content of the embodiments of the present application in more detail, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are only for reference and explanation, and are not used to limit the embodiments of the present application.

[0020] Based on the problems in the related art, the present application provides a target matching method, which can be applied to edge devices, such as edge servers, network video recorders (NVR), intelligent analysis boxes, etc. As Figure 1 shown, the method includes:

[0021] S101. Acquire first feature extraction results of detection targets in videos respectively collected by multiple terminal devices for the same picture scene.

[0022] It should be noted that the detection target can be pedestrians, vehicles, etc. that need to be concerned about during the target tracking process. The first feature extraction result includes multi-dimensional features of the detection target, such as the appearance features of pedestrians (eyes, eyebrows, facial contours, etc.), clothing features (clothing color, style, type, etc.), etc., and the color, type, brand, license plate number, etc. of vehicles.

[0023] In some embodiments, the first feature extraction result can be a one-dimensional array including features of multiple dimensions, or a feature vector including multiple features. Here, the dimension can represent the number of features, and the higher the dimension of the feature, the more features it represents.

[0024] In some embodiments, the terminal device can be an electronic device with image acquisition and image processing functions (such as feature extraction), such as a camera, a webcam, etc. Each terminal device can capture the same picture scene and perform feature extraction on the detection target in the video obtained by each of them, obtaining multiple first feature extraction results.

[0025] Here, due to the different computing capabilities of different terminal devices, in the same time period, the feature dimensions in the first feature extraction results corresponding to each terminal device may be different. The dimension of the features extracted by the terminal device with higher computing power is higher than that of the features extracted by the terminal device with lower computing power. For example, if the computing power of device A is higher than that of device B, the first feature extraction result of device A may include features of 96 dimensions, and the first feature extraction result of device B may include features of 60 dimensions. The dimensions of the features extracted by devices with different computing capabilities here are only exemplary descriptions, and the present application does not limit this.

[0026] In some embodiments, it is possible to control multiple terminal devices to simultaneously start shooting the same picture scene and perform feature extraction on the video obtained by the shooting. After a period of time, obtain the first feature extraction results after each terminal device performs feature extraction.

[0027] S102. If the feature dimensions in each of the first feature extraction results are different, perform correction processing on at least one of the first feature extraction results to obtain a target feature extraction result.

[0028] In some embodiments, the fact that the feature dimensions in each of the first feature extraction results are different may mean that the feature dimension of at least one of the feature extraction results among the multiple first feature extraction results is different from the feature dimensions of other feature extraction results.

[0029] Here, the reason for the different feature dimensions in each of the first feature extraction results may be the difference in the computing capabilities of each terminal device, or it may be because the captured video image frames are not clear, resulting in the feature extraction model being unable to extract effective or correct features.

[0030] In some embodiments, when it is determined that there are differences in the feature dimensions among multiple first feature extraction results and those of other feature extraction results, correction processing can be performed on one or more first feature extraction results. This correction processing can involve supplementing, deleting, replacing features, etc., so that the dimensions of the obtained target feature extraction results after the correction processing are all the same, facilitating target matching.

[0031] S103. Based on the target feature extraction results, determine the matching results of the detection targets in the videos collected by each terminal device.

[0032] In some embodiments, when the target feature extraction result is a feature vector, the similarity of the feature vectors can be calculated to determine whether the detection targets in the videos collected by different terminals match or are the same. When it is determined that the targets detected by each terminal device match, the trajectory of the detection target can be determined to facilitate tracking of the detection target.

[0033] In the embodiments of the present application, obtain the first feature extraction results of the detection targets in the videos collected by multiple terminal devices for the same picture scene; if the feature dimensions in each of the first feature extraction results are different, perform correction processing on at least one of the first feature extraction results to obtain the target feature extraction results; based on the target feature extraction results, determine the matching results of the detection targets in the videos collected by each terminal device. In this way, when the feature dimensions in the first feature extraction results are different, by performing correction processing on one or more first feature extraction results, the features extracted by terminal devices with different computing capabilities can be associated to obtain target feature extraction results with the same feature dimensions, achieving efficient and accurate target matching, and improving the effect and efficiency of target tracking based on the target matching results.

[0034] In some embodiments of the present application, obtain the first feature extraction results of the detection targets in the videos collected by multiple terminal devices for the same picture scene, that is, step S101 can be implemented through the following steps S1011 to S1013, and each step is described below.

[0035] S1011. Obtain the real-time computing capabilities of multiple terminal devices.

[0036] In some embodiments, when starting a target tracking event, such as when starting to shoot the picture scene corresponding to the detection target, the real-time computing capabilities of each terminal device can be obtained. The real-time computing capabilities of the terminal device can be determined based on information such as the operating parameters (such as the number of cores, frequency, etc.) of the processor of the terminal device, the memory capacity, and the network bandwidth.

[0037] S1012. Based on the preset correspondence between the computing power of the terminal device and the feature extraction dimension, determine the target feature extraction dimension corresponding to the specific terminal device with the real-time computing power being the first computing power.

[0038] Among them, the specific terminal device can be any one of multiple terminal devices. The preset correspondence between the computing power of the terminal device and the feature extraction dimension can be pre-established. One computing power corresponds to one feature extraction dimension, and this preset correspondence can be reflected in the form of a preset correspondence table. The higher the computing power, the larger the corresponding feature extraction dimension.

[0039] In some embodiments, the real-time computing power of the specific terminal device can be compared with each computing power in the preset comparison relationship, so as to determine the target feature extraction dimension that matches the real-time computing power of the specific terminal device.

[0040] S1013. Based on the target feature extraction dimension, control the specific terminal device to extract features from the detection target in the video image frames collected by itself.

[0041] In some embodiments, the edge device can send the target feature extraction dimension corresponding to the specific terminal device to the specific terminal device, so that the specific terminal device can extract features from the detection target in the video image frames collected by itself based on this target feature extraction dimension.

[0042] It can be understood that based on the preset correspondence between the computing power of the terminal device and the feature extraction dimension, determining the target feature extraction dimension corresponding to the specific terminal device with the real-time computing power being the first computing power, and controlling the specific terminal device to extract features from the detection target in the video image frames collected by itself based on the target feature extraction dimension can make the feature extraction of the specific terminal device adapt to its own computing power, thereby improving the efficiency of feature extraction.

[0043] In some embodiments of the present application, the "performing correction processing on at least one first feature extraction result" in step S102 can be performing correction processing on the second feature extraction results in each of the first feature extraction results.

[0044] Among them, the second feature extraction result is the feature extraction result with a feature dimension smaller than the largest feature dimension in each of the first feature extraction results. The second feature extraction result can be one or more. For example, the second feature extraction result can be the feature extraction result with the smallest feature dimension in each of the first feature extraction results and the other feature extraction results with feature dimensions smaller than the largest feature dimension in each of the first feature extraction results.

[0045] Exemplarily, if there are three first feature extraction results, and the feature dimensions of each first feature extraction result are 48, 96, and 128 respectively, then the first feature extraction results with feature dimensions of 48 and 96 can be determined as the second feature extraction results.

[0046] In some embodiments, the correction process for the second feature extraction results can perform operations such as supplementing and replacing the features in the second feature extraction results, so that the feature dimensions in the corrected second feature extraction results are the same as the maximum feature dimension in each first feature extraction result.

[0047] In some embodiments, the "correcting at least one first feature extraction result" in step S102 can also be to correct the third feature extraction results in each first feature extraction result.

[0048] Among them, the third feature extraction result is the feature extraction result with a feature dimension greater than the minimum feature dimension in each first feature extraction result. The third feature extraction result can be one or more. For example, the third feature extraction result can be the first feature extraction result with the largest feature dimension in each first feature extraction result and the first feature extraction results with other feature dimensions greater than the minimum feature dimension in each first feature extraction result.

[0049] Exemplarily, if there are five first feature extraction results, and the feature dimensions of each first feature extraction result are 32, 48, 96, 128, and 132 respectively, then the four first feature extraction results with feature dimensions of 48, 96, 128, and 132 can be determined as the third feature extraction results.

[0050] In some embodiments, the correction process for the third feature extraction results can delete the features in the third feature extraction results, so that the feature dimensions in the corrected third feature extraction results are the same as the minimum feature dimension in each first feature extraction result.

[0051] In other embodiments, the "correcting at least one first feature extraction result" in step S102 can also be to correct the second feature extraction results in each first feature extraction result, and to correct the third feature extraction results in each first feature extraction result.

[0052] Here, the second feature extraction result can be the feature extraction result with the smallest feature dimension among all the first feature extraction results, or it can be the feature extraction result with the smallest feature dimension among all the first feature extraction results and the feature extraction results with other feature dimensions less than the maximum feature dimension among all the first feature extraction results. For example, if there are 3 first feature extraction results, and the feature dimensions of each first feature extraction result are 48, 96, and 128 respectively, then the first feature extraction result with a feature dimension of 48 can be determined as the second feature extraction result, or both the first feature extraction results with a feature dimensions of 48 and 96 can be determined as the second feature extraction result.

[0053] The third feature extraction result can be the feature extraction result with the largest feature dimension among all the first feature extraction results, or it can be the feature extraction result with the largest feature dimension among all the first feature extraction results and the feature extraction results with other feature dimensions greater than the minimum feature dimension among all the first feature extraction results. For example, if there are 5 first feature extraction results, and the feature dimensions of each first feature extraction result are 32, 48, 96, 128, and 132 respectively, then the first feature extraction result with a feature dimension of 132 can be determined as the third feature extraction result, or the four first feature extraction results with feature dimensions of 48, 96, 128, and 132 can be determined as the third feature extraction result.

[0054] In some embodiments, in a scenario where the second feature extraction result and the third feature extraction result are simultaneously corrected, the correction processes for the second feature extraction result and the third feature extraction result can each be at least one of supplementing, replacing, and deleting features in the second feature extraction result.

[0055] It can be understood that by correcting the second feature extraction result and / or processing the third feature extraction result, the feature dimension of the second feature extraction result can be aligned with the minimum feature dimension, and the feature dimension of the third extraction result can be aligned with the maximum feature dimension, so that the feature extraction results obtained after the correction process have the same dimension, facilitating target matching.

[0056] In some embodiments of the present application, each first feature extraction result is the feature extraction result of each terminal device for the image frame corresponding to the same timestamp in the video collected by each terminal device, that is, the feature extraction result of each terminal device for the same detection target in the video image frame corresponding to the same timestamp. Based on this, the correction process for the second feature extraction result in each first feature extraction result may be to determine the first feature type missing in the second feature extraction result relative to the feature extraction result with the largest feature dimension in each first feature extraction result; based on the feature extraction results of multiple terminal devices for the image frames corresponding to multiple timestamps, supplement the first feature type in the second feature extraction result.

[0057] In some embodiments, each feature in the second feature extraction result can be compared with each feature in the first feature extraction result one by one to determine the first feature type missing in the second feature extraction result. The first feature type can represent the features of the detection target. For example, when the detection target is a pedestrian, the first feature type may be the texture, style, etc. of the clothes.

[0058] In some embodiments, the first feature type in the second feature extraction result can be supplemented based on the feature extraction result of the video image frame corresponding to the current timestamp or other timestamps different from the current timestamp in the video collected by the current terminal device (corresponding to the second feature extraction result), and / or based on the feature extraction results of the video image frames corresponding to multiple timestamps in the video collected by other terminal devices. Here, the current timestamp can be the timestamp of the image frame corresponding to the second feature extraction result, and other terminal devices can be other devices different from the terminal device corresponding to the second feature extraction result.

[0059] It can be understood that by determining the first feature type missing in the second feature extraction result relative to the feature extraction result with the largest feature dimension in each first feature extraction result and supplementing the first feature type in the second feature extraction result based on the feature extraction results of the terminal device for the image frames corresponding to multiple timestamps, the second feature extraction result can be further improved, thereby improving the accuracy of target detection.

[0060] In some embodiments of the present application, the features of the first feature type in the second feature extraction result can also be predicted based on other features in the second feature extraction result different from the first feature type and a pre-established feature prediction model to obtain predicted features; the predicted features are determined as the features of the first feature type in the second feature extraction result.

[0061] Among them, the pre-established feature prediction model can be a neural network model. Through the training of the neural network model, the correlation relationships between different features can be determined, and the unknown or missing features can be inferred based on different types of features.

[0062] In some embodiments, other features whose second feature extraction results are different from the first feature type can be input into the pre-established feature prediction model, so as to predict the predicted features of the first feature type, and use the predicted features as the features of the first feature type in the second feature extraction results.

[0063] It can be understood that based on the pre-established feature prediction model and other features whose second feature extraction results are different from the first feature type, the features of the first feature type can be predicted, and the missing features of the first feature type in the second feature extraction results can be quickly obtained.

[0064] In some embodiments of the present application, in the process of supplementing the features of the first feature type in the second feature extraction results based on the feature extraction results of multiple terminal devices for the image frames corresponding to multiple timestamps, the features of the first feature type in the second feature extraction results can be supplemented based on the feature extraction results of the first terminal device corresponding to the first image frame for the second feature extraction results.

[0065] Among them, the first image frame is other image frames collected by the first terminal device with timestamps different from the image frames corresponding to the second feature extraction results. The first image frame can include one or more. The first image frame can be adjacent frames of the image frames corresponding to the second feature extraction results, or can be the previous few frames or the next few frames of the image frames corresponding to the second feature extraction results. The features of the first feature type can be obtained from the feature extraction results of one or more first image frames, and the features of the first feature type in the second feature extraction results can be supplemented based on the features.

[0066] In some embodiments, in the process of supplementing the features of the first feature type in the second feature extraction results based on the feature extraction results of multiple terminal devices for the image frames corresponding to multiple timestamps, the features of the first feature type in the second feature extraction results can also be supplemented based on the feature extraction results of the second terminal device for the image frames corresponding to multiple timestamps in the video collected by itself.

[0067] Among them, the second terminal device is the other terminal devices except the first terminal device among the multiple terminal devices. For the time stamps of multiple image frames in the video collected by the second terminal device itself, they may be different from the time stamps of the corresponding image frames of the second feature extraction result, or may be the same as the time stamps of the corresponding image frames of the second feature extraction result. The features of the first feature type in the second feature extraction result can be supplemented based on the feature extraction results of the image frames with multiple time stamps by the second terminal device.

[0068] It can be understood that by supplementing the features of the first feature type in the second feature extraction result based on the feature extraction result of the first terminal device corresponding to the second feature extraction result for the first image frame, or based on the feature extraction results of the image frames corresponding to multiple time stamps in the video collected by the second terminal device itself, the feature dimension of the second feature extraction result can be made consistent with the maximum feature dimension in each first feature extraction result, which is convenient for subsequent target matching.

[0069] In some embodiments of the present application, when supplementing the features of the feature type in the second feature extraction result based on the feature extraction result of the first terminal device corresponding to the second feature extraction result for the first image frame, the first detection target in the first image frame can be determined; if the first detection target is the same as the detection target corresponding to the second feature extraction result, the features of the first feature type in the second feature extraction result are supplemented based on the feature extraction result of the first detection target.

[0070] Here, the first detection target in each first image frame and the detection target corresponding to the second feature extraction result can be compared to determine whether they are the same detection target. For example, according to the position information of the detection frame of the detection target in the first image frame and the position information of the detection frame corresponding to the detection target corresponding to the second feature extraction result, it is determined whether the overlap degree of the two detection frames is greater than a preset threshold. If it is greater than the preset threshold, it is considered that the two detection frames are consistent, that is, the first detection target in the first image frame and the detection target corresponding to the second feature extraction result are the same detection target.

[0071] In the case where it is determined that the first detection target in the first image frame and the detection target corresponding to the second feature extraction result are the same detection target, the features of the first feature type in the second feature extraction result can be supplemented based on the feature extraction result of the first detection target in any one first image frame, or the features of the first feature type in the second feature extraction result can be supplemented based on the feature extraction results of the first detection targets in multiple first image frames.

[0072] It can be understood that when it is determined that the first detection target in the first image frame is the same as the detection target corresponding to the second feature extraction result, by using the feature extraction result of the first detection target to supplement the features of the first feature type missing in the second feature extraction result, the accuracy of the supplemented features can be ensured, thereby improving the correct rate of subsequent target matching.

[0073] In some embodiments of the present application, the implementation manner of supplementing the features of the feature type in the second feature extraction result based on the feature extraction result of the first detection target may be to determine the features of the first feature type in the feature extraction result of the first detection target as the features of the first feature type in the second feature extraction result.

[0074] In some embodiments, the implementation manner of supplementing the features of the feature type in the second feature extraction result based on the feature extraction result of the first detection target may also be to determine the average value of the features of the first feature type extracted from multiple first image frames as the features of the first feature type in the second feature extraction result. Among them, the multiple first image frames may be 2, 3,..., N first image frames, and N is less than the total number of image frames of the video corresponding to the second feature extraction result.

[0075] In other embodiments, the weighted average value of the features of the first feature type extracted from multiple first image frames may also be determined as the features of the first feature type in the second feature extraction result, where the weights of the features of the first feature type in different image frames may be determined based on the distance between the first image frame and the image frame corresponding to the second feature extraction result.

[0076] It can be understood that by determining the features of the first feature type in the feature extraction result of the first detection target as the features of the first feature type in the second feature extraction result, or determining the average value of the features of the first feature type extracted from multiple first image frames as the features of the first feature type in the second feature extraction result, the implementation of supplementing based on the same type of features can improve the restoration degree of the supplemented features in the second feature extraction result.

[0077] In some embodiments of the present application, in the process of supplementing the features of the first feature type in the second feature extraction result based on the feature extraction results of multiple timestamp corresponding image frames collected by the second terminal device for itself, the relative position information of each second terminal device relative to the detection target corresponding to the second feature extraction result may be determined; based on the relative position information, the candidate features of the first feature type are determined from the feature extraction results of the second image frames in the videos collected by each second terminal device for itself; and the candidate features are determined as the features of the first feature type in the second feature extraction result.

[0078] Among them, the position information of each second terminal device may be unchanged, and the position information of the detection target corresponding to the second feature extraction result may change during the shooting process. The position information of the second terminal device relative to the detection target can be determined according to the azimuth information of the detection target (the detection target corresponding to the second feature extraction result) in the video shot by the second terminal device.

[0079] Exemplarily, if the detection target is a pedestrian and the pedestrian is in a standing sideway state in the video shot by the second terminal device, it can be determined that the pedestrian is facing the second terminal device sideways; if the pedestrian is in a walking forward state in the video shot by the second terminal device, it can be determined that the pedestrian is facing the terminal device directly.

[0080] In some embodiments, the second image frame is an image frame with the same timestamp as the image frame corresponding to the second feature extraction result. The candidate feature may be a feature of the first feature type collected and detected by the second terminal device when the detection target is in a state of facing the second terminal device directly. In implementation, the candidate terminal device facing the detection target can be determined according to the position information of each second terminal device relative to the detection target, and the feature of the first feature type detected by this terminal device is determined as the candidate feature.

[0081] In other embodiments, if the target terminal device facing the target to be detected can be determined in advance from each terminal device, the extraction result of the frontal feature of the target terminal device for the target to be detected can be directly used as the extraction result of the frontal feature of other terminal devices for the target to be detected, so that the computational amount of feature extraction by other terminal devices can be reduced, and thus the efficiency and accuracy of feature extraction can be improved.

[0082] It can be understood that since the candidate feature may be a feature extracted when the detection target is in a state of facing the second terminal device directly, the accuracy of the candidate feature is relatively high. Therefore, determining the candidate feature as the feature of the first feature type in the second feature extraction result can improve the accuracy of the supplemented features in the second feature extraction result.

[0083] In some embodiments of the present application, an implementation manner of correcting the third feature extraction result in each first feature extraction result may be to determine the reference dimension of the feature in the reference feature extraction result with the smallest feature dimension in each first feature extraction result; based on the pre-established feature importance ranking information, sort the features in the third feature extraction result to obtain the fourth feature extraction result; based on the reference dimension, delete some features with relatively low rankings in the fourth feature result.

[0084] The pre-established feature importance ranking information may be a feature importance ranking determined for a detection target, where the higher the feature importance, the higher the ranking, and the more important the feature, the greater the role played by the feature in target detection. For example, if the detection target is a pedestrian, the corresponding pedestrian feature importance ranking from front to back may be: facial features, hairstyle features, clothing features.

[0085] In some embodiments, the features with the same dimension as the reference dimension that are ranked higher in the fourth feature extraction result can be retained, while the remaining features ranked lower can be deleted, so that the feature dimension of the corrected third feature extraction result is the same as the feature dimension of the other first feature extraction results.

[0086] It can be understood that by sorting the features in the third feature extraction result based on the pre-established feature importance sorting information, the features in the fourth feature extraction result can be arranged in order of importance from high to low. After deleting some features with lower ranking in the fourth feature result based on the reference dimension, some features with higher importance can still be retained, thereby improving the accuracy of target matching.

[0087] In some embodiments of the present application, one implementation method of "based on the target feature extraction result, matching the detection target in the video collected by each terminal device to obtain a matching result" in step S103 may be to determine the weight of each feature in each target feature extraction result based on pre-established feature importance ranking information; based on each weight, determine the similarity of each target feature extraction result; based on the similarity, determine the matching result of the detection target in the video collected by each terminal device.

[0088] Among them, the weight of each feature can be determined according to the feature importance ranking information. The more important the feature is, the larger the corresponding weight is. Based on the various weights, the weighted inner product of each target feature extraction result can be determined, and the similarity of each target feature extraction result is determined based on the calculation result of the weighted inner product. If the similarity is greater than a preset threshold, the match of the detection target in the video collected by each terminal device is determined; otherwise, the match of the detection target in the video collected by each terminal device is determined.

[0089] Exemplarily, when the target feature extraction result is represented by a feature vector, the similarity cos(x, y) of each target feature extraction result can be calculated based on the following formula (1):

[0090]

[0091] Among them, X w ,Y wrespectively represent two target feature extraction results obtained by matching based on the weights determined by feature importance, w i represents the weight corresponding to the i-th feature, x i , y i respectively represent the eigenvalues in the feature vectors corresponding to the two target feature extraction results, and n represents the total number of features of the detection target.

[0092] In some embodiments, it is also possible to directly perform a direct inner product multiplication based on the two target feature extraction results, that is, without considering the sorting result of feature importance, and use the cosine distance to evaluate the similarity between the two targets. The larger the cosine distance value, the more similar the detection targets corresponding to the two target feature extraction results are. On the contrary, the smaller the distance value, the lower the similarity between the two targets. The formula for the cosine distance can be expressed by formula (2):

[0093]

[0094] where X and Y respectively represent the two target feature extraction results for matching, x i , y i respectively represent the eigenvalues in the feature vectors corresponding to the two target feature extraction results, and n represents the total number of features of the detection target.

[0095] It can be understood that by determining the weights of each feature in each target feature extraction result based on the pre-established feature importance sorting information, the similarity of each target feature extraction result can be determined, and the role of features with high importance in target matching can be reflected, thereby improving the accuracy of target matching.

[0096] In the embodiments of the present application, the first feature extraction results of the detection targets in the videos respectively collected by multiple terminal devices for the same video scene are obtained; if the feature dimensions in each first feature extraction result are different, at least one first feature extraction result is corrected to obtain the target feature extraction result; based on the target feature extraction result, the matching results of the detection targets in the videos respectively collected by each terminal device are determined. In this way, when the feature dimensions in the first feature extraction results are different, by correcting one or more first feature extraction results, the features extracted by terminal devices with different computing capabilities can be associated to obtain target feature extraction results with the same feature dimension, realizing efficient and accurate target matching, and improving the effect and efficiency of target tracking based on the target matching results.

[0097] Next, the implementation process of the application embodiments in actual application scenarios will be introduced.

[0098] The present application provides a method for feature extraction and fusion of end-side devices with unbalanced computing power, which can be applied to a cross-device tracking system, such as Figure 2As shown in the figure, the cross-device tracking system 200 includes a plurality of end devices 201, a plurality of edge devices 202, and a cloud device 203.

[0099] Among them, the end device 201 can be different cameras, that is, devices representing different computing powers, which can extract the feature information of the target and convert it into corresponding feature codes. Its characteristic is to extract feature with different dimensional sizes according to the size of its own computing power; the edge device 202 refers to different edge-side devices, such as NVR intelligent analysis boxes, etc., which are mainly used to fuse the features extracted from the end side. Because its computing power is stronger than that of the end side, it can synthesize various features to analyze whether the targets recognized by the end device 201 are the same target; the cloud device 203 refers to the cloud device, which has a greater computing power and can fuse more powerful features to identify whether they are the same target. Feature data can be transmitted and used mutually between different devices through the network.

[0100] As Figure 3 shown, it is a schematic flow chart of a method for feature extraction and fusion of end-side devices with unbalanced computing power provided by this application. The method includes:

[0101] S301. The edge device (equivalent to the "terminal device" in other embodiments) adaptively extracts the features of the target based on its own computing power (equivalent to the "computing ability" in other embodiments).

[0102] For the case where the target is a pedestrian, the extracted features include: appearance features, clothing information, ornaments, personal belongings, etc.; for the case where the target is a vehicle, the extracted features include: vehicle inherent attributes, driver information, personalized information.

[0103] Among them, more detailed features for feature extraction include face features and face area detail features, such as information on whether wearing glasses or a mask, etc. The device where the target is located can be correspondingly encoded into a device code for determining the location of the target.

[0104] Pedestrian structured information can include one or more of the following information: appearance features, clothing information, ornaments, personal belongings, etc.; appearance features include one or more of the following features: gender (male, female), age group (child, youth, middle-aged, elderly), hairstyle, hair color, beard, etc.; clothing information includes one or more of the following features: colors, textures, styles, types of upper and lower body clothing, etc., ornaments include one or more of the following features: whether wearing glasses, a mask, a hat, etc.; personal belongings information includes one or more of the following features: whether holding an umbrella, carrying a child, pulling a suitcase, carrying a backpack, carrying a handbag, etc. When extracting pedestrian structured information, the appearance features, clothing information, ornaments, personal belongings, etc. of the pedestrian are digitized to generate a re-authentication feature code for the pedestrian, which is used for cross-lens tracking, trajectory search, etc.

[0105] The inherent attributes of a vehicle include: one or more of the following features: vehicle color, brand, model, year, license plate number, vehicle type, etc. When identifying the color of a vehicle, the vehicle color is divided into 13 types, including: black, white, silver, gray, cyan, blue, green, yellow, gold, red, purple, pink, and brown. When identifying the brand and model of a vehicle, vehicles with the same appearance will be merged, and the front and rear information will be distinguished to achieve 360° full-view recognition of the brand, model, and year of the vehicle, of which the front information includes more than 5,500 types and the rear information includes more than 3,500 types.

[0106] In addition to the character information on the license plate, the color and type of the license plate can also be identified when identifying the license plate information. The license plate colors include: blue plate (blue background with white characters), yellow plate (yellow background with black characters), white plate (white background with black characters), black plate (black background with white characters), and green plate (green background with black characters). The license plate types corresponding to each license plate color include: ordinary small cars, large cars, police cars, Hong Kong and Macao entry and exit vehicles, and new energy vehicles. Vehicle types are divided into 21 categories according to the national standard, including sedans, small trucks, large trucks, light buses, small passenger cars, large passenger cars, vans, pickup trucks, off-road vehicles, commercial vehicles, trailers, concrete mixer trucks, tank trucks, truck cranes, fire trucks, muck trucks, escort trucks, engineering repair vehicles, rescue vehicles, flatbed trucks, and tricycles.

[0107] Driver information includes: the main driver not wearing a seat belt, talking on the phone while driving, the front passenger not wearing a seat belt, and the faces of the main and front passenger. In order to overcome the impact of fake and cloned license plates on vehicle identification, each vehicle is given a unique identifier, and the vehicle's personalized information is also identified, including one or more of the following features: annual vehicle inspection mark, sun visor, pendant, ornaments, tissue box, sunroof, luggage rack, spare tire, and collision marks. When extracting pedestrian structured information, the vehicle's inherent attributes, driver information, personalized information and other information are digitized to generate a vehicle re-authentication feature code, which is used for cross-lens tracking, image search, and driving trajectory reproduction.

[0108] Other feature information that can be extracted includes information such as the positional relationship between cameras.

[0109] The main features are extracted for the end devices with different computing power. Due to the limitation of computing power of some end devices, only partial features can be extracted. For example, the computing power of end device 1 is relatively small, so in order to meet the efficiency requirements, only 96-dimensional features are extracted. The computing power of end device 2 is relatively large compared with that of device 1, so 128-dimensional features are extracted. The dimension of the feature will define an optimal efficiency according to the computing power of the device. At the same time, each end device also needs to extract as many features as possible.

[0110] S302: Encode the extracted target features.

[0111] According to the diverse features extracted from different devices, all the extracted features are correspondingly converted into feature codes corresponding to the features. Finally, the feature code information of the target is output, and the length of the corresponding feature code corresponds to the dimension of the feature.

[0112] The feature code can be a numerical representation of the feature. For example, it can be numerically represented by formula (3):

[0113]

[0114] Where F represents the feature code of this feature, x represents the feature, if the feature is present, it is 1, and if the feature is not recognized or there is no such attribute feature or other situations, the feature code is 0.

[0115] S303. Perform feature fusion on the encoded features to obtain the fused feature code.

[0116] Since the feature dimensions extracted from different devices are different, in order to compare features of different dimensions to determine whether they are the same target, and based on the judgment result for cross-border tracking of the target. According to the previously extracted features, feature fusion can be performed on the encoded features. Missing value processing can be performed on encoded features with different dimensions.

[0117] Here, there are various reasons for missing values (missing features), mainly two reasons. One is that due to factors such as insufficient computing power and time limit of the end device, only the main features of the target can be inferred, and other features cannot be inferred, resulting in the missing of some features. The missing features are attributed to the reason of end device computing power limitation. The other is that due to unclear target images, model recognition errors and other reasons, the recognition of this feature is incorrect, resulting in the missing of the feature. The missing features are attributed to the reason of model recognition ability limitation. The loss of its features will affect subsequent feature processing and judgment to a certain extent.

[0118] The processing methods for missing values include the following three ways:

[0119] Method 1: Directly use the two target ID numbers in the tracking information for processing. If they are the same, directly use the non-missing values, that is, the feature information calculated by other end devices, and directly assign it to the target with missing values. If the ID numbers in the tracking cannot determine whether the target is the same target, the missing value is directly judged as other, that is, directly set to 0.

[0120] Method 2: Missing value filling strategy. The missing values can be interpolated and filled with the most likely values. The interpolation methods include:

[0121] (1) Mean interpolation. If there are tracking targets with the same ID number in multiple frames and multiple devices, the communication between devices can be used to determine the feature information in a certain dimension. Mean interpolation is performed by extracting information from the feature codes extracted from multiple frames and multiple devices, so as to comprehensively utilize the multi-frame information, as shown in formula (4) below:

[0122] X d = mean(∑ j x i,j,d ) (4);

[0123] Among them, d represents the d-th dimension feature, and the features of other dimensions are processed similarly. i represents the i-th target, j represents the corresponding features extracted from the same ID on different frames or different end devices. The available features are averaged, and finally the features of the missing frame target are determined.

[0124] (2) Modeling and predicting features. Predict the features of this dimension based on the features of other dimensions of the target, that is, take the missing features as the prediction target and predict by establishing a model. The model prediction is not limited to the following methods. For example, in pedestrian features, if the pedestrian is male, the information about the hair length is more likely to be short hair, and this dimension feature can be directly set to 1.

[0125] (3) Multiple imputation features. Due to the different dimensions of features extracted by different end devices due to computing power limitations, and for the positional relationship of the installed devices, it is possible to know in advance the effectiveness of extracting target feature information from the perspectives of each device. For some end devices, the extracted obvious features have higher accuracy. For example, when extracting whether a pedestrian wears a mask from the front, the accuracy is higher than that from the back. Then the feature of whether the pedestrian wears a mask extracted by this camera is used as the feature information in this dimension. On the contrary, the feature extracted from the back image has a higher degree of untrustworthiness, so this feature does not need to be extracted, which reduces the feature extraction time to a certain extent. And this feature uses the feature extracted by another device as the feature in this dimension. Similar features in other dimensions can be processed similarly.

[0126] Overall, it is equivalent to some devices extracting some features and other devices extracting other features, which reduces the computational complexity of a single end device extracting all features to a certain extent, improves the effect of extracting target features while reducing the inference delay rate, and fills in the missing values of the features in the corresponding dimensions of the target, facilitating subsequent feature comparison.

[0127] After supplementing the missing eigenvalues during reading, the encoded features can be fused. For example, the final target feature code can be output by sorting and integrating the target features according to their importance order. Each feature placeholder occupies a feature code. If the feature exists, the corresponding position of the feature code is 1; if the feature does not exist, the feature code at that position is 0. The specific fusion methods can include:

[0128] (1) Use one-hot encoding, and use N-bit states to correspond to feature information. For example, the male and female genders of a pedestrian target each occupy one bit. If both positions are 0, it means the feature information is missing and feature missing value processing is required. If the missing value cannot be processed, the feature code at that position is 0.

[0129] (2) The position where the feature code is located also represents the importance of the feature relative to other features. The higher the feature importance, the more it is placed in the front. This is mainly considering the computing power differences of various terminal devices. The length of the extracted feature code is not necessarily all-dimensional features. At the same time, it is impossible to compensate for all position feature code information. In the case of different lengths of feature codes, only the shortest feature code length can be used as the maximum length to intercept the corresponding longer feature code length.

[0130] (3) Common feature code lengths. For example, the shortest feature code length can be set to 96 dimensions, and the longest feature code length can be set to 1024 dimensions. The size of the dimension is determined according to the computing power and effect of different devices, and it can also be set to other sizes.

[0131] (4) After the feature codes are fused, they are bound to the target tracking information and then updated and sent to other devices for fusion or use according to the situation.

[0132] S304. Perform target matching based on the fused feature codes.

[0133] As Figure 4 shown, step S304 can be implemented through the following steps S3041 to S3045:

[0134] S3041. Obtain multiple fused feature codes.

[0135] S3042. Determine whether the dimensions of each fused feature code are consistent.

[0136] If they are consistent, execute the following step S3043; otherwise, execute the following steps S3044 to S3045.

[0137] S3043. Perform target matching based on each fused feature code.

[0138] It is possible to determine whether two targets are the same target based on the fused feature codes. The determination methods can include the direct inner product method, the weighted inner product method, and the truncated weighted inner product method, which can be selectively used according to the size of the feature code dimension and the specific scenario effect. Among them, the direct inner product method is shown in formula (2), and the weighted inner product method is shown in formula (1).

[0139] S3044. Truncate the fused feature codes to obtain the truncated feature codes.

[0140] S3045. Perform target matching based on the truncated feature codes.

[0141] For the case where the dimensions of the fused feature codes are inconsistent, the truncated weighted inner product method can be used. Since there are actually feature vectors with inconsistent lengths in practice, if the similarity is calculated for two vectors with different dimensions, the corresponding dimensions of the features in the front can be truncated, and then the direct inner product method or the weighted inner product method similar to the above can be directly used for processing. The truncation method is to perform truncation processing according to the corresponding positions. For example, if the length of one feature code is 512 dimensions and the length of another device feature code is 128 dimensions, the truncation method is to select the first 128 dimensions, and finally the dimensions of both are 128 dimensions. Finally, the inner product processing is performed on the truncated feature codes to calculate the similarity between the two targets.

[0142] An end-side device feature extraction and fusion method with unbalanced computing power provided by an embodiment of the present application can adapt to devices with different computing powers, improve the device operation efficiency, and at the same time can obtain more effective features by combining features in various ways, fuse the features extracted by different devices, realize cross-border target tracking, and the obtained combined features can also be used for other purposes, such as pedestrian query, and match according to the combined features to find the movement trajectory of the target.

[0143] The present application provides a target matching device. Figure 5 As shown in the composition structure diagram of a target matching device provided by an embodiment of the present application, Figure 5 as shown, the target matching device 500 includes:

[0144] An acquisition module 501, configured to acquire the first feature extraction results of the detection targets in the videos respectively collected by multiple terminal devices for the same picture scene; the first feature extraction results include the multi-dimensional features of the detection targets;

[0145] A correction module 502, configured to, if the feature dimensions in each of the first feature extraction results are different, perform correction processing on at least one of the first feature extraction results to obtain target feature extraction results; the feature dimensions in each of the target feature extraction results are the same;

[0146] A determination module 503, configured to determine a matching result of a detection target in videos collected by each terminal device respectively based on the target feature extraction result.

[0147] In some embodiments, the correction module 502 includes:

[0148] A first correction sub-module, configured to perform correction processing on the second feature extraction results in the respective first feature extraction results, and / or perform correction processing on the third feature extraction results in the respective first feature extraction results;

[0149] Wherein, the second feature extraction result is a feature extraction result whose feature dimension is smaller than the maximum feature dimension in the respective first feature extraction results; the third feature extraction result is a feature extraction result whose feature dimension is larger than the minimum feature dimension in the respective first feature extraction results.

[0150] In some embodiments, the first correction sub-module includes:

[0151] A first determination unit, configured to determine a first feature type missing in the second feature extraction result with respect to the feature extraction result having the maximum feature dimension in the respective first feature extraction results;

[0152] A first supplementation unit, configured to supplement the first feature type in the second feature extraction result based on the feature extraction results of the multiple terminal devices for the image frames corresponding to multiple timestamps.

[0153] In some embodiments, the first supplementation unit includes:

[0154] A first supplementation sub-unit, configured to supplement the features of the first feature type in the second feature extraction result based on the feature extraction result of the first terminal device corresponding to the first image frame for the second feature extraction result; the first image frame is an other image frame whose timestamp collected by the first terminal device is different from the image frame corresponding to the second feature extraction result; or,

[0155] Supplement the features of the first feature type in the second feature extraction result based on the feature extraction results of the second terminal device for the image frames corresponding to multiple timestamps in the video collected by the second terminal device itself; the second terminal device is other terminal devices except the first terminal device among the multiple terminal devices.

[0156] In some embodiments, the first supplementation sub-unit is further configured to determine a first detection target in the first image frame; if the first detection target is the same as the detection target corresponding to the second feature extraction result, supplement the features of the first feature type in the second feature extraction result based on the feature extraction result of the first detection target.

[0157] In some embodiments, the first supplementary subunit is further configured to determine, as the feature of the first feature type in the second feature extraction result, the feature of the first feature type in the feature extraction result of the first detection target; or, determine, as the feature of the first feature type in the second feature extraction result, the average value of the features of the first feature type extracted from a plurality of the first image frames.

[0158] In some embodiments, the first supplementary unit is further configured to determine the relative position information of each second terminal device with respect to the detection target corresponding to the second feature extraction result; based on the relative position information, determine candidate features of the first feature type from the feature extraction results of the second image frames in the videos collected by each second terminal device for itself; the second image frame is an image frame with a timestamp the same as that of the image frame corresponding to the second feature extraction result. Determine the candidate features as the features of the first feature type in the second feature extraction result.

[0159] In some embodiments, the first determination unit is further configured to predict the feature of the first feature type in the second feature extraction result based on other features of the second feature extraction result that are different from the first feature type and a pre-established feature prediction model, to obtain a predicted feature; determine the predicted feature as the feature of the first feature type in the second feature extraction result.

[0160] In some embodiments, the first correction sub-module further includes:

[0161] A second determination unit, configured to determine the reference dimension of the features in the reference feature extraction result with the smallest feature dimension among the respective first feature extraction results;

[0162] A sorting unit, configured to sort the features in the third feature extraction result based on pre-established feature importance sorting information, to obtain a fourth feature extraction result;

[0163] A deletion unit, configured to delete some features with a lower ranking in the fourth feature result based on the reference dimension.

[0164] In some embodiments, the matching module 503 includes:

[0165] A first determination sub-module, configured to determine the weight of each feature in each target feature extraction result based on pre-established feature importance sorting information;

[0166] A second determination sub-module, configured to determine the similarity of each target feature extraction result based on each weight;

[0167] A third determination sub-module, configured to determine a matching result of a detection target in the videos collected by the respective terminal devices based on the similarity.

[0168] In some embodiments, the obtaining module 501 includes:

[0169] A first obtaining sub-module, configured to obtain the real-time computing capabilities of the multiple terminal devices;

[0170] A fourth determination sub-module, configured to determine a target feature extraction dimension corresponding to a specific terminal device with a real-time computing capability of a first computing capability based on a preset correspondence between the computing capabilities of the terminal devices and the feature extraction dimensions; the specific terminal device is any one of the multiple terminal devices;

[0171] A first control sub-module, configured to control the specific terminal device to extract features of a detection target in a video image frame collected by itself based on the target feature extraction dimension.

[0172] It should be noted that the description of the target matching device in the embodiments of the present application is similar to the description of the above-mentioned target matching method embodiments, and has similar beneficial effects as the method embodiments, so details will not be repeated. For the technical details not disclosed in the embodiments of this device, please refer to the description of the method embodiments of the present application for understanding.

[0173] Correspondingly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the target matching method provided in the above embodiments is implemented.

[0174] It should be noted that in the embodiments of the present application, if the above-mentioned target matching method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related solutions, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0175] It should be noted that in this text, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including at least one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes such element.

[0176] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical functional division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0177] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0178] In addition, each functional unit in the embodiments of this application can be fully integrated into one processing unit, or each unit can be separately a unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0179] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, ROMs, magnetic disks, or optical discs that can store program codes.

[0180] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a product to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as removable storage devices, ROMs, magnetic disks, or optical discs that can store program codes.

[0181] As described above, the above are only the implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application.

Claims

1. A target matching method, characterized in that, Including: Obtaining first feature extraction results of detection targets in videos respectively collected by multiple terminal devices for the same screen scene; the first feature extraction results include multi-dimensional features of the detection targets; If the feature dimensions in each of the first feature extraction results are different, performing correction processing on at least one of the first feature extraction results to obtain target feature extraction results; the feature dimensions in each of the target feature extraction results are the same; Based on the target feature extraction results, determining matching results of detection targets in videos respectively collected by each terminal device.

2. The method according to claim 1, wherein The performing correction processing on at least one of the first feature extraction results includes: Performing correction processing on second feature extraction results in each of the first feature extraction results, and / or performing correction processing on third feature extraction results in each of the first feature extraction results; Wherein, the second feature extraction results are feature extraction results with feature dimensions smaller than the maximum feature dimension in each of the first feature extraction results; the third feature extraction results are feature extraction results with feature dimensions larger than the minimum feature dimension in each of the first feature extraction results.

3. The method according to claim 2, characterized in that Each of the first feature extraction results is a feature extraction result of an image frame corresponding to the same timestamp in the videos respectively collected by each terminal device; The performing correction processing on the second feature extraction results in each of the first feature extraction results includes: Determining first feature types missing in the second feature extraction results relative to the feature extraction results with the maximum feature dimension in each of the first feature extraction results; Based on the feature extraction results of the multiple terminal devices for image frames corresponding to multiple timestamps, supplementing the first feature types in the second feature extraction results.

4. The method according to claim 3, wherein The supplementing the features of the first feature types in the second feature extraction results based on the feature extraction results of the multiple terminal devices for image frames corresponding to multiple timestamps includes: Based on the feature extraction result of a first terminal device corresponding to a first image frame for the second feature extraction result, supplementing the features of the first feature types in the second feature extraction result; the first image frame is an other image frame collected by the first terminal device with a timestamp different from that of the image frame corresponding to the second feature extraction result; or, Based on the feature extraction results of a second terminal device for image frames corresponding to multiple timestamps in the video collected by itself, supplementing the features of the first feature types in the second feature extraction result; the second terminal device is other terminal devices except the first terminal device among the multiple terminal devices.

5. The method according to claim 4, wherein The supplementing the features of the first feature types in the second feature extraction result based on the feature extraction result of the first terminal device corresponding to the first image frame for the second feature extraction result includes: Determining a first detection target in the first image frame; If the first detection target is the same as the detection target corresponding to the second feature extraction result, supplement the features of the first feature type in the second feature extraction result based on the feature extraction result of the first detection target.

6. The method according to claim 5, characterized in that, The supplementing the features of the feature type in the second feature extraction result based on the feature extraction result of the first detection target includes: Determining the features of the first feature type in the feature extraction result of the first detection target as the features of the first feature type in the second feature extraction result; or, Determining the average value of the features of the first feature type extracted from multiple first image frames as the features of the first feature type in the second feature extraction result.

7. The method according to claim 4, characterized in that The supplementing the features of the first feature type in the second feature extraction result based on the feature extraction result of the multiple image frames corresponding to multiple timestamps collected by the second terminal device for itself includes: Determining the relative position information of each second terminal device with respect to the detection target corresponding to the second feature extraction result; Based on the relative position information, determining candidate features of the first feature type from the feature extraction results of the second image frames in the videos collected by each second terminal device for itself; the second image frame is an image frame with a timestamp the same as the timestamp of the image frame corresponding to the second feature extraction result; Determining the candidate features as the features of the first feature type in the second feature extraction result.

8. The method according to claim 3, wherein The method further includes: Predicting the features of the first feature type in the second feature extraction result based on the other features different from the first feature type in the second feature extraction result and a pre-established feature prediction model to obtain predicted features; Determining the predicted features as the features of the first feature type in the second feature extraction result.

9. The method according to claim 2, wherein The correcting the third feature extraction result in each first feature extraction result includes: Determining the reference dimension of the features in the reference feature extraction result with the smallest feature dimension in each first feature extraction result; Sorting the features in the third feature extraction result based on the pre-established feature importance ranking information to obtain a fourth feature extraction result; Deleting some features with a lower ranking in the fourth feature result based on the reference dimension.

10. The method according to claim 1, wherein The obtaining the matching result by matching the detection targets in the videos collected by each terminal device for itself based on the target feature extraction result includes: Determining the weights of the features in each target feature extraction result based on the pre-established feature importance ranking information; Determining the similarity of each target feature extraction result based on each weight; Determining the matching result of the detection targets in the videos collected by each terminal device for itself based on the similarity.

11. The method according to claim 1, wherein The obtaining the first feature extraction result of the detection target in the videos collected by multiple terminal devices for the same picture scene includes: Obtaining the real-time computing capabilities of the multiple terminal devices; Based on a preset correspondence relationship between the computing power of the terminal device and the feature extraction dimension, determine the target feature extraction dimension corresponding to a specific terminal device with a real-time computing power of the first computing power; the specific terminal device is any one of the multiple terminal devices; Based on the target feature extraction dimension, control the specific terminal device to extract features of a detection target in a video image frame collected by itself.

12. A target matching device, characterized in that, It includes: An acquisition module, configured to acquire first feature extraction results of a detection target in videos respectively collected by multiple terminal devices for the same picture scene; the first feature extraction results include multi-dimensional features of the detection target; A correction module, configured to, if the feature dimensions in each of the first feature extraction results are different, perform correction processing on at least one of the first feature extraction results to obtain a target feature extraction result; the feature dimensions in each of the target feature extraction results are the same; A determination module, configured to determine a matching result of a detection target in videos respectively collected by each terminal device based on the target feature extraction result.