Target tracking method and device, computer equipment and storage medium

By acquiring both intra-domain and cross-domain facial images and combining matching and scaling techniques, the problem of lost interactive objects in the external interaction of intelligent connected vehicles has been solved, achieving stable and efficient tracking of external interactions.

CN120997254AActive Publication Date: 2025-11-21CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511119200.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

In existing technologies, in the external interaction scenarios of intelligent connected vehicles, the interaction objects are easily lost and difficult to recapture, resulting in the interruption of external interaction functions.

Method used

By acquiring facial images from both the same and different domains, facial image matching technology is used to re-identify the tracked object. A combination of same-domain and cross-domain matching methods, along with scaling and image mean filling techniques, is employed to improve matching accuracy.

Benefits of technology

It effectively recovers the tracked object, ensures the continuity and stability of external interactions, enhances the user interaction experience, and improves the robustness and matching accuracy of the tracking algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997254A_ABST
    Figure CN120997254A_ABST
Patent Text Reader

Abstract

The invention relates to a target tracking method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a first detection image, and tracking a tracking object according to the first detection image; in response to the loss of the tracking object, obtaining a second detection image, the second detection image being a same-domain image of the first detection image; and matching an object in the second detection image according to the face image so as to re-identify the successfully matched object as a tracking object, the face image comprising a first face image in the first detection image and / or a preset second face image, and the second face image being a cross-domain image of the first detection image. By adopting the method, the problem of tracking object loss in the prior art can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and in particular to a target tracking method and device, a computer device, and a storage medium. BACKGROUND

[0002] With the rapid development of intelligent and connected vehicle technology, user off-board interaction scenarios are becoming an important development direction for improving user experience. In the process of realizing off-board interaction, accurately and stably tracking the interaction object is one of the key technologies.

[0003] In actual application scenarios, the interaction object is easily lost, and it is difficult to recapture the interaction object in related technologies, which may cause the off-board interaction function to be interrupted. SUMMARY

[0004] Therefore, a target tracking method and device, a computer device, and a storage medium are provided to improve the problem of tracking object loss in the prior art.

[0005] In one aspect, a target tracking method is provided, including:

[0006] obtaining a first detection image, and tracking a tracking object according to the first detection image;

[0007] in response to loss of the tracking object, obtaining a second detection image, the second detection image being a same-domain image of the first detection image;

[0008] matching objects in the second detection image according to a face image to re-identify the objects that are successfully matched as the tracking object, the face image including a first face image in the first detection image and / or a preset second face image, wherein the second face image is a cross-domain image of the first detection image, and wherein the same-domain image is an image obtained according to a first shooting device under a first shooting condition, and the cross-domain image is an image obtained according to a second shooting device under a second shooting condition.

[0009] In one embodiment, the matching objects in the second detection image according to the face image includes:

[0010] in a case where the first detection image includes the first face image of the tracking object, performing same-domain matching of the objects in the second detection image according to the first face image;

[0011] otherwise, performing cross-domain matching of the objects in the second detection image according to the second face image.

[0012] In one embodiment, after the same-domain matching of the objects in the second detection image according to the first face image, the method further includes:

[0013] In response to the intra-domain matching failure, performing cross-domain matching on the object in the second detection image according to the second face image.

[0014] In one embodiment, the performing cross-domain matching on the object in the second detection image according to the second face image comprises:

[0015] scaling the second face image based on a scaling ratio, and performing cross-domain matching based on the scaled second face image;

[0016] The scaling ratio comprises a first scaling ratio or a second scaling ratio, the first scaling ratio being a scaling ratio based on a face scale of the first face image, and the second scaling ratio being a scaling ratio based on an overall scale of the tracked object.

[0017] In one embodiment, the scaling the second face image based on a scaling ratio comprises:

[0018] In a case where the first face image of the tracked object exists in the first detection image, scaling the second face image according to the first scaling ratio, and performing cross-domain matching based on the scaled second face image;

[0019] The first scaling ratio is obtained from first candidate ratios according to a first optimization algorithm, an optimization target of the first optimization algorithm being a matching accuracy of face samples after scaling processing according to the first candidate ratios in a first data set, the first data set being an image data set same as the first detection image, the face samples being image data same as the second face image, and the first candidate ratios being scaling ratios of the face samples based on face scales of sample objects in the first data set.

[0020] In one embodiment, the scaling the second face image based on a scaling ratio comprises:

[0021] In a case where the first face image of the tracked object is not detected in the first detection image, scaling the second face image according to the second scaling ratio, and performing cross-domain matching based on the scaled second face image;

[0022] The second scaling ratio is obtained from second candidate ratios according to a second optimization algorithm, an optimization target of the second optimization algorithm being a matching accuracy of face samples after scaling processing according to the second candidate ratios in a second data set, the second data set being an image data set same as the first detection image, the face samples being image data same as the second face image, and the second candidate ratios being scaling ratios of the face samples based on overall scales of sample objects in the second data set.

[0023] In one embodiment, the cross-domain matching of the object in the second detection image according to the second face image further comprises:

[0024] According to the second face image, an image mean of the second face image is obtained;

[0025] According to the image mean of the second face image, the second face image is filled to make the size of the second face image consistent with the second detection image.

[0026] In another aspect, a target tracking device is provided, the device comprising:

[0027] An acquisition module is configured to acquire a first detection image, and in response to loss of a tracking object, acquire a second detection image, the second detection image being a co-domain image of the first detection image;

[0028] A tracking module is configured to track a tracking object according to the first detection image, and determine whether the tracking object is lost;

[0029] A re-identification module is configured to, in the case that the tracking object is lost, match an object in the second detection image according to a face image, to re-identify the matched object as the tracking object, the face image comprising a first face image in the first detection image and / or a preset second face image, wherein the second face image is a cross-domain image of the first detection image, the co-domain image being an image according to a first shooting device under a first shooting condition, and the cross-domain image being an image according to a second shooting device under a second shooting condition.

[0030] In yet another aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor implementing the method when executing the computer program.

[0031] A computer readable storage medium is also provided, having a computer program stored thereon, the computer program being executable by a processor to implement the method.

[0032] The above target tracking method, device, computer device and storage medium, acquire a first detection image, track according to the first detection image, after loss of a tracking object, acquire a second detection image, the second detection image being a co-domain image of the first detection image, and match an object in the second detection image according to a face image, the face image comprising a first face image in the first detection image and / or a preset second face image, the second face image being a cross-domain image of the first detection image, the above process re-finding the tracking object by co-domain or cross-domain recognition of a face in the second detection image. Attached Figure Description

[0033] Figure 1 This is a flowchart illustrating a target tracking method in one embodiment;

[0034] Figure 2 This is a flowchart illustrating the target tracking method in another embodiment;

[0035] Figure 3 This is a structural block diagram of a target tracking device in one embodiment;

[0036] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0038] Intelligent connected vehicles can provide users with easier and more convenient services through external interaction technologies, such as activating various vehicle functions from outside the car and achieving contactless operation. By allowing users to interact with the vehicle using gestures and movements outside the car, entertainment functions such as human-vehicle interaction and external games can be realized, significantly enhancing the convenience and fun of interaction between users and vehicles.

[0039] Vehicle-to-everything (V2X) interaction technology relies on precise and stable tracking of objects. Only by accurately locking onto the user can the vehicle capture the user's interaction commands in a timely manner, thus achieving efficient human-vehicle interaction.

[0040] In related technologies, tracking algorithms are used to achieve real-time tracking of objects outside the vehicle, such as Bytetrack. These tracking algorithms can predict and update the position of the tracked object and can achieve effective tracking of the target with low computational resource consumption in normal scenarios.

[0041] However, real-world application scenarios are often complex and varied. When the tracked object is temporarily occluded by pedestrians, obstacles, or other obstacles, existing tracking algorithms face the problem of difficulty in re-capturing the target after tracking is lost.

[0042] This invention provides a target tracking method that uses facial images to recapture the tracked object.

[0043] In one embodiment, the target tracking method is as follows: Figure 1 As shown, it includes the following steps:

[0044] Step 110: Obtain the first detection image and track the target object based on the first detection image.

[0045] The embodiment is applied to the vehicle external interaction technology, and can obtain continuous external images of the vehicle as first detection images through an external camera of the vehicle, and uses a tracking algorithm such as ByteTrack to identify, lock, and predict the position of each object (such as a pedestrian) in the continuous first detection images, so as to realize tracking of the object.

[0046] In a feasible implementation manner, a user opens the external interaction function, and then a vehicle end detects all pedestrians under the current external camera, traverses the detected pedestrians, uses a motion recognition algorithm (such as a yolopose algorithm) to locate key points, locks the user by a specified motion (such as raising a hand), and then takes the locked user as a tracking object, and uses a tracking algorithm (such as ByteTrack) to perform real-time tracking, which is to frame the whole tracking object and mark a portrait bounding box (bbox), the portrait bbox is a rectangular frame in computer vision, and the boundary of the rectangular frame is used to describe the position and size of an object, and the function is to mark a portrait in an image, so as to realize tracking of the target object.

[0047] The portrait bbox is assigned a tracking ID (Identity Document, identity) by the tracking algorithm, and the tracking ID assigned to the same tracking object remains unchanged before the tracking object is lost, and the tracking ID is used as a basis for subsequent tracking.

[0048] It can be understood that the tracking algorithm performs real-time tracking based on prediction of the position and state of the tracking object.

[0049] In step 120, a second detection image is obtained in response to loss of the tracking object.

[0050] If a tracking object in a frame is blocked by a pedestrian or an object, the tracking ID of the tracking object disappears, at this time, it is judged that the tracking object has lost tracking, and if the tracking object is continuously blocked for several frames, and a new pedestrian appears in the picture in the several frames, the pedestrian may be assigned the ID of the original tracking object by the tracking algorithm, and a mis-tracking situation occurs.

[0051] In the embodiment, when the tracking ID of the tracking object disappears, the matching judgment of the tracking ID of the tracking algorithm on the portrait bbox detected by the tracking algorithm is stopped, and the re-identification process in the following text is used instead.

[0052] After the tracking object is lost, the vehicle end continues to obtain images outside the vehicle as second detection images.

[0053] It can be understood that for the images continuously acquired at the vehicle end, the images before the tracking object is lost are the first detection images, and the images continuously acquired after the tracking object is lost are the second detection images.

[0054] The second detection image and the first detection image are same-domain images. Same-domain images refer to images from the same data distribution or an image set with highly similar features. Same-domain images have high consistency or similarity in visual style, content type, imaging condition, resolution, illumination, background, object posture, sensor characteristics, etc. In this embodiment, the first detection image and the second detection image are continuous image frames captured at the same time period and by the same camera. Therefore, the second detection image and the first detection image are same-domain images.

[0055] In step 130, the object in the second detection image is matched according to the face image, so as to re-identify the object matched successfully as the tracking object.

[0056] After the matching judgment of the stop tracking ID, the face image is used to re-identify the tracked user.

[0057] In an implementation, the face image comes from the previously collected first detection image. For example, before the user is determined as the tracking object, the user indicates himself as the tracking object through a specified action. When the specified action of the user is detected, the first frame image in which the user is recognized to make the specified action is recorded at the vehicle end.

[0058] When the portrait bbox is marked, the face bbox can also be marked. For the portrait bbox recorded in the first frame image, the face bbox of the current tracking object is detected based on a face detection model (such as Yunet), and after the tracking object is lost, the image region framed by the face bbox is taken as the first face image, which is used as the matching reference.

[0059] Since the first detection image and the second detection image are same-domain images, they have similar image feature distribution. Therefore, the object in the second detection image can be matched according to the first face image, that is, after the first frame image and the subsequent second detection image are subjected to the same preprocessing process (such as background elimination, face extraction, smoothing processing, etc.), they are input into a matching model for matching. The matching model is further described below.

[0060] Same-domain matching is essentially face matching between the first detection image with the face image of the tracking object and the second detection image to be matched.

[0061] Same-domain images are highly similar in image features such as illumination, resolution, viewing angle, noise pattern, etc. The feature ambiguity caused by the difference in data domains is reduced, and therefore the same-domain matching is more accurate.

[0062] Possibly, for the first frame image, the portrait bbox of the tracked object can be detected, but the face bbox cannot be detected (the tracked object can face away or sideways to the external camera), in this case, a preset second face image is used for matching.

[0063] For example, when the user opens the external interaction function on the vehicle side, a tracking recovery function is provided, if the user checks the function, the user is requested to take a face selfie image. For another example, the user pre-records a face selfie image in his own account, and thereafter the face selfie image can be used as the second face image.

[0064] Possibly, for the user's face selfie image, a face detection model (such as Yunet) can be used to detect the face bbox, and the area outside the face bbox can be removed to exclude interference. The area framed by the face bbox is used as the second face image.

[0065] It can be known that the second face image has different shooting conditions, sensor parameters, and the like from the first detection image or the second detection image, and therefore the second face image is a cross-domain image of the first detection image and the second detection image.

[0066] Cross-domain matching of the second face image to the object in the second detection image increases the data source, which is conducive to re-identifying the tracked object when the first detection image does not detect a face.

[0067] Cross-domain matching can also be based on a matching model.

[0068] As described above, the face image can be a same-domain first face image or a cross-domain second face image. In one possible embodiment, in the case where one type of face image fails to match, the other type of face image is used for further matching to increase the probability of successful matching.

[0069] For example, in one embodiment, after obtaining the first detection image, it is identified whether the first detection image contains the first face image (for example, the first frame image is identified to determine whether the face bbox of the tracked object can be identified), and in the case where the first detection image contains the first face image of the tracked object, the object in the second detection image is matched based on the first face image; otherwise, the object in the second detection image is matched based on the second face image.

[0070] Preferentially, it is identified whether the conditions for same-domain matching are met, so that the same-domain image with higher matching accuracy is used for re-identification, otherwise cross-domain matching is used to ultimately achieve the purpose of re-identification.

[0071] In some cases, the intra-domain matching can fail (for example, the first face image is of poor quality), and in the case of intra-domain matching failure, switching to cross-domain matching of the second detection image according to the second face image.

[0072] The target tracking method provided by the application re-finds the tracking object for tracking by matching the object in the second detection image with the face image after the tracking object is lost, so that the off-vehicle interaction can continue.

[0073] The cross-domain image and the intra-domain image are further described below.

[0074] For the off-vehicle interaction scenario of the application, the off-vehicle environment is complex and it is difficult to achieve stable shooting of the environment, but by limiting the shooting device, the domain offset between images can be reduced, so for the intra-domain image, it is defined as an image shot by the same external camera as the first detection image, or an image with the same shooting device parameters (image size, resolution, etc.) as the shooting device when the first detection image is shot.

[0075] Possibly, the off-vehicle interaction scenario has certain requirements for the in-view portrait, such as requiring the in-view portrait to be complete as a whole, or requiring the in-view portrait to meet certain conditions relative to the proportion of the portrait as a whole, so as to ensure the completeness of the in-view portrait. The image that meets the shooting device condition and / or the in-view portrait condition can be regarded as an image in the same data domain.

[0076] The cross-domain image has significant differences in data distribution, statistical characteristics, or resolution, etc. relative to the intra-domain image. In the application, the image that does not meet the shooting device condition and the in-view portrait condition can be regarded as a cross-domain image. In particular, in the above embodiment, the cross-domain second face image is a face selfie image shot by a user's selfie device or an in-vehicle shooting device, and in terms of image distribution, it mainly focuses on the user's own face area.

[0077] In summary, for the off-vehicle interaction technology of the application, the cross-domain image and the intra-domain image can be exemplarily defined as follows: the image under the first shooting condition according to the first shooting device (for example, the external camera of the vehicle) is the intra-domain image, and the image under the second shooting condition according to the second shooting device (for example, the user's selfie device) is the cross-domain image. The first shooting condition is the in-view portrait condition in which the portrait size meets the preset requirement, and the second shooting condition is the in-view face condition in which the face size meets the preset requirement.

[0078] The cross-domain matching is further described below.

[0079] The difficulty of cross-domain matching lies in that the distribution difference between the source domain and the target domain destroys the visual consistency relied on by traditional matching algorithms. For example, in the above embodiment, the main area of the second face image serving as the benchmark is a face, while in the second detection image to be matched, the face only occupies a small area, that is, the face scale (size) in the second face image is different from the face scale in the first detection image or the second detection image, and the difference in the face scale leads to a decrease in the matching accuracy of cross-domain matching.

[0080] The present application proposes an adaptive scaling mechanism based on image scale perception and optimization algorithm driving to solve the problem of inconsistent face scale. By performing scale adjustment processing on the second face image of the user's selfie, the second face image is made consistent with the second detection image outside the vehicle in terms of spatial structure, face size, etc., thereby effectively improving the matching accuracy of cross-domain images.

[0081] In one embodiment, cross-domain matching of an object in the second detection image based on the second face image includes:

[0082] scaling the second face image based on the scaling ratio, and performing cross-domain matching based on the scaled second face image;

[0083] The scaling ratio includes a first scaling ratio or a second scaling ratio, the first scaling ratio being a scaling ratio based on the face scale of the first face image, and the second scaling ratio being a scaling ratio based on the overall scale of the tracked object.

[0084] In this process, the first scaling ratio or the second scaling ratio is pre-configured, and in the application process, scaling is performed based on the size of the first face image or based on the size of the overall tracked object, and a suitable scaling ratio is dynamically selected, so as to adjust the second face image to a suitable scale and increase the matching accuracy.

[0085] Further explanation is as follows:

[0086] As described above, in the case where the face bbox is detected in the first frame of the first detection image, the first face image framed by the face bbox is used for same-domain matching, but there is a case where the same-domain matching fails. At this time, the first scaling ratio can be used to scale the second face image (for example, the height of the second face image is adjusted to 1.5 times the height of the face bbox of the first frame, and the width is still adjusted according to the original aspect ratio).

[0087] On the other hand, in the case that the first face image of the tracking object is not detected in the first detection image, the second face image is scaled according to the second scaling ratio, for example, when the first frame does not detect a face bbox, since there is no face bbox as a reference, the second face image is scaled by the second scaling ratio (for example, the height of the second face image is adjusted to 25% of the portrait bbox in the first frame, and the width is still adjusted according to the original aspect ratio).

[0088] It can be understood that in the embodiment, the size of the marked face bbox is taken as the face scale of the first face image, and the size of the portrait bbox is taken as the overall scale of the tracking object.

[0089] By flexibly selecting the scaling manner, the second face image is scaled to a scale size that is more matched with the actual face of the current tracking object based on the face bbox in the case of the face bbox, or based on the portrait bbox otherwise, which is beneficial to subsequent cross-domain matching.

[0090] It should be noted that the scaling manner provided by the present application is different from the face alignment manner in the related art. The target of the conventional face alignment is to realize the standardization of the facial feature structure through key point positioning (such as eyes, nose, mouth, etc.), which depends on key point detection and affine transformation. The scaling manner provided by the present application is an image processing strategy based on cross-domain information perception, which uses the scale information in the external image (face bbox or portrait bbox) to guide the scaling of the current image (second face image), and has obvious cross-domain guiding characteristics.

[0091] The scaling ratio is further described below.

[0092] The scaling ratio (such as 1.5 times or 25%) is not set based on fixed empirical values, but is obtained based on an optimization parameter combination obtained in an optimization process of an optimization algorithm.

[0093] Exemplarily, according to the actual application scene of the present application, two types of optimization tasks are constructed. The first type of optimization task faces the case that the face bbox can be detected in the first frame image, and the first scaling ratio is obtained by optimization. The second type of optimization task faces the case that the face bbox cannot be detected in the first frame image, and the second scaling ratio is obtained by optimization.

[0094] For the first type of optimization task, a face sample and a first data set are respectively constructed for matching. The face sample is image data in the same domain as the second face image, that is, the face sample is a face selfie image. The first data set is a set of image data in the same domain as the first detection image and the second detection image, that is, the first data set is a set of image data that meets the mirror-in portrait condition and is captured by using the same external camera (or a camera with the same parameters as the external camera). The images in the first data set contain sample objects (that is, there are portraits for matching in the images), and the sample objects can be detected for faces (for example, the sample objects are facing the camera when being captured). At least one sample object corresponds to the object indicated by the face sample, so that the face sample can be successfully matched with at least one image in the first data set.

[0095] The first scaling ratio is obtained from the first candidate ratio according to the first optimization algorithm. The optimization target of the first optimization algorithm is the matching accuracy of the face sample after scaling processing according to the first candidate ratio in the first data set. The first candidate ratio is a ratio of scaling the face sample based on the face scale (for example, the size of the face bbox) of the sample object in the first data set.

[0096] The first optimization algorithm is, for example, a Whale Optimization Algorithm (WOA). The WOA simulates the searching and chasing behavior of a whale group, so as to find a global optimal solution. In the first type of optimization task, the search range of the optimization parameter (that is, the first scaling ratio) is set to the interval [0.5, 2] (which can be calibrated), and the matching accuracy of the face sample in the first data set is used as the objective function of the WOA algorithm.

[0097] The following exemplary description of the optimization process of the first type of optimization task is as follows:

[0098] An image containing a sample object A in the first data set is selected, and the sample object A can be detected for a face bbox. The image is used as the first detection image in a simulated actual scenario in which a face bbox is detected. The selfie image of the sample object A is used as the face sample.

[0099] A target function suitable for the first type of optimization task is constructed. In the present application, the target function of the first type of optimization task is the overall matching success rate when the face sample is matched with the first data set.

[0100] An initial solution (the first candidate ratio) is generated based on the WOA algorithm. The scaling ratio of the face sample is adjusted based on the initial solution, and the objects in the first data set are matched. By matching different face samples, the overall matching accuracy of the cross-domain matching is calculated.

[0101] According to the WOA algorithm, the position of the whale (solution) in each generation is adjusted, the objective function value of the new solution is calculated, if the objective function value of the new solution is greater than the objective function value of the previous solution, the new solution is updated as the new optimal solution, until the termination condition is met, the matching accuracy is maximized, and the final optimal solution is taken as the first scaling ratio. The calculation process of the WOA algorithm for calculating a new solution is described in the prior art, and will not be described here.

[0102] In the above process, the face sample is used to simulate the second face image in the actual situation, the first data set is used to simulate the second detection image, and the first optimization algorithm is used for optimization to determine the best scaling ratio for scaling the second face image, so that the second face image can be adjusted to the most suitable scale for cross-domain matching in the actual situation, thereby improving the accuracy of cross-domain matching.

[0103] For the second type of optimization task, a face sample for matching and a second data set are constructed respectively. The face sample is image data in the same domain as the second face image, that is, the face sample is a face selfie image. The second data set is an image data set in the same domain as the first detection image and the second detection image, that is, the second data set is an image data set that meets the mirror-in portrait condition and is taken by the same external camera (or a camera with the same parameters as the external camera).

[0104] Unlike the first type of optimization task, the second type of optimization task is aimed at the situation where the external camera cannot provide an effective first face image for same-domain matching. To simulate this situation, the second data set includes two types of image data for the same sample object: the first type of image data cannot detect a face bbox (the sample object is facing or facing away from the camera), corresponding to the first detection image without detecting the first face image; the second type of image data can detect a face bbox (the sample object is facing the camera), corresponding to the second detection image to be matched.

[0105] The second scaling ratio is obtained from the second candidate ratio according to the second optimization algorithm. The optimization target of the second optimization algorithm is the matching accuracy of the face sample after scaling processing according to the second candidate ratio in the second data set. The second candidate ratio is the scaling ratio of the face sample based on the overall scale (portrait bbox) of the sample object in the second data set.

[0106] The second optimization algorithm can refer to the first optimization algorithm and use the WOA algorithm. The search range of the optimization parameter (i.e., the second scaling ratio) is set to the interval [0.1, 1] (which can be calibrated) as an example. The optimization process of the second type of optimization task is described as an example:

[0107] An image containing sample object A in the second data set is selected, and sample object A cannot detect a face bbox, which is used as a first detection image that cannot detect a face bbox in an actual scene. A selfie image of sample object A is used as a face sample.

[0108] A target function suitable for the second type of optimization task is constructed. In the present application, the target function of the second type of optimization task is the overall matching accuracy based on matching the second data set with the face sample.

[0109] An initial solution (second candidate ratio) is generated based on the WOA algorithm. The scaling ratio of the face sample is adjusted based on the initial solution, and the sample object in the second data set is matched. By matching different face samples, the overall matching accuracy of cross-domain matching is calculated.

[0110] According to the WOA algorithm, the position of the whale (solution) in each generation is adjusted, and the target function value of the new solution is calculated. If the target function value of the new solution is greater than that of the previous solution, the new solution is updated as the new optimal solution. Until the termination condition is met, the matching accuracy is maximized.

[0111] In another possible implementation, the first optimization algorithm and the second optimization algorithm can be different. The optimization algorithm used can be determined based on the actual situation, and the grey wolf optimization algorithm, the ant colony algorithm, etc. can be selected.

[0112] The above optimization process finds a second scaling ratio suitable for adjusting the overall scale of the face image based on the tracking object. In the absence of a face bbox, the portrait bbox is used to scale the face image, thereby improving the accuracy of cross-domain matching.

[0113] In another embodiment, after the scale adjustment, the second face image is filled. According to the second face image, the image mean of the second face image is obtained. According to the image mean of the second face image, the second face image is filled so that the size of the second face image is consistent with that of the second detection image.

[0114] For example, the edge region of the second face image is filled with the image mean, so that the final image size is consistent with that of the second detection image taken outside the vehicle.

[0115] In some possible implementations, the position of the second face image is adjusted with reference to the image taken outside the vehicle. For example:

[0116] In the case where the first detection image can detect a face bbox, the center of the face bbox is taken as the center of the second face image, and then the filling is performed.

[0117] In the case that the first detection image cannot detect the face bbox, the center of the face bbox in the second detection image is taken as the center of the second face image, and then padding is performed.

[0118] Alternatively, whether the first detection image detects the face bbox or not, the center of the face bbox in the second detection image is taken as the center of the second face image, and then padding is performed.

[0119] The image mean is used for padding to ensure a uniform spatial layout before inputting into the matching model. On the basis of the second face image after scale adjustment, the overall scale of the image is changed. In this way, it can be ensured that the feature extraction has more consistent distribution characteristics with the out-of-vehicle image, and the detection result will not be affected due to excessive scaling.

[0120] The purpose of image mean padding is to:

[0121] Good visual continuity: The image mean can better simulate the overall tone of the image, making the padding area naturally integrated with the original image and reducing the interference caused by edge mutation.

[0122] Conform to image statistical characteristics: Compared with extreme values such as black (0) or white (255), the image mean is closer to the distribution of the image itself, which helps to improve the stability of the model.

[0123] The matching model is described as follows:

[0124] The matching model uses a face recognition network (such as SFace) to extract a feature vector and calculate a feature similarity (for example, using cosine similarity or Euclidean distance) to determine whether it is the same object. For intra-domain matching, the first detection image with the first face image and the second detection image to be matched are input into the matching model, the face bbox in both images is extracted, and the feature similarity between the face bbox is calculated. The object that meets the similarity condition (for example, the feature similarity is greater than a similarity threshold) is re-identified as a tracking object. For cross-domain matching, the scale-adjusted second face image and the second detection image to be matched are input into the matching model, the face bbox in both images is extracted, and the feature similarity between the face bbox is calculated. The object that meets the similarity condition is re-identified as a tracking object.

[0125] If the interval between the current frame captured by the external camera and the frame in which the tracking object is lost is within the set frame threshold (for example, 50 frames), the next iteration is entered. If the frame threshold is reached, the current out-of-vehicle interaction session is closed, and the user needs to reopen the function.

[0126] For example, Figure 2As shown, a flowchart of the target tracking method in an embodiment is provided, in which tracking object determination is performed according to action recognition, and the first frame when the specified action is made is recorded, when the tracking object is lost, and within the frame number threshold, if the first frame detects a face bbox, same-domain matching is performed, otherwise cross-domain matching is performed, and when the same-domain matching fails, the cross-domain matching is switched.

[0127] The target tracking method provided by the application is suitable for external complex environments, supports tracking recovery means within a frame number threshold range when the user's actual use scene frequently appears to be blocked, ensures the continuity of the vehicle interaction task outside the vehicle, greatly enhances the user's interaction experience with the vehicle in the vehicle scene outside the vehicle, enhances the robustness of the tracking algorithm, and the vehicle end can actually deploy a lightweight tracking algorithm, which only needs to ensure the tracking consistency when the user appears in the picture, and greatly reduces the requirements for the tracking algorithm.

[0128] In another aspect, a tracking recovery strategy is proposed, which is based on the same-domain face image and the cross-domain face image respectively. The same-domain face image has a higher matching accuracy because it is in the same data domain as the tracking object, but because the same-domain image collected may appear situations such as facing away from the camera and being blocked, the cross-domain face image is used to enhance the robustness of the tracking recovery.

[0129] In view of the problem of inconsistent face scales in cross-domain images, an adaptive scaling mechanism based on reference image scale perception and optimization algorithm driving is proposed. Through scale adjustment and padding processing of the user's selfie, it is made consistent with the vehicle image in terms of spatial structure, face size, etc., thereby effectively improving the matching accuracy of cross-domain matching.

[0130] Possibly, the object indicated by the preset second face image is different from the object making the specified action outside the vehicle, for example, user a pre-records a face selfie image, but user b makes the specified action outside the vehicle and is locked as the tracking object, and after re-identification, user a may be confirmed as a new tracking object. In order to avoid this situation, when the user starts the vehicle interaction function, a pop-up window reminds the user to pay attention to whether other people make the specified action.

[0131] It should be understood that, although Figure 1 , Figure 2 The steps in the flowchart are displayed in sequence according to the direction of the arrow, but these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, Figure 1 Figure 2 ​At least one of the steps in the method can comprise a plurality of sub-steps or a plurality of stages, which sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the order of the sub-steps or stages is not necessarily sequential, but can be performed alternately or in rotation with other steps or sub-steps or stages of other steps.

[0132] In one embodiment, as shown in FIG. 1, a target tracking device is provided, comprising an acquisition module 201, a tracking module 202 and a re-identification module 203, wherein: Figure 3

[0133] The acquisition module 201 is configured to acquire a first detection image, and in response to a loss of a tracking object, acquire a second detection image, the second detection image being a co-domain image of the first detection image.

[0134] The tracking module 202 is configured to track the tracking object according to the first detection image, and determine whether the tracking object is lost.

[0135] The re-identification module 203 is configured to, in the case that the tracking object is lost, match objects in the second detection image according to a face image, to re-identify a matched object as the tracking object, the face image comprising a first face image in the first detection image and / or a preset second face image, wherein the second face image is a cross-domain image of the first detection image, and wherein the co-domain image is an image according to a first shooting device under a first shooting condition, and the cross-domain image is an image according to a second shooting device under a second shooting condition.

[0136] In one embodiment, the re-identification module 203, in the case that the first face image of the tracking object exists in the first detection image, performs co-domain matching of the objects in the second detection image according to the first face image; otherwise, performs cross-domain matching of the objects in the second detection image according to the second face image.

[0137] In one embodiment, the re-identification module 203, in response to a failure of the co-domain matching, performs cross-domain matching of the objects in the second detection image according to the second face image.

[0138] The re-identification module 203 is further configured to scale the second face image based on a scaling ratio, and perform the cross-domain matching based on the scaled second face image; the scaling ratio comprising a first scaling ratio or a second scaling ratio, the first scaling ratio being a scaling ratio based on the first face image, and the second scaling ratio being a scaling ratio based on the tracking object as a whole.

[0139] ​The re-identification module 203, in a case where the first detection image contains the first face image of the tracked object, scales the second face image according to a first scaling ratio, and performs cross-domain matching based on the scaled second face image.

[0140] The first scaling ratio is obtained from a first candidate ratio according to a first optimization algorithm, an optimization target of the first optimization algorithm is a matching accuracy of face samples after scaling processing according to the first candidate ratio, the first data set is an image data set in the same domain as the first detection image, the face samples are image data in the same domain as the second face image, and the first candidate ratio is a scaling ratio of the face samples based on a face scale of a sample object in the first data set.

[0141] The re-identification module 203, in a case where the first detection image does not contain the first face image of the tracked object, scales the second face image according to a second scaling ratio, and performs cross-domain matching based on the scaled second face image.

[0142] The second scaling ratio is obtained from a second candidate ratio according to a second optimization algorithm, an optimization target of the second optimization algorithm is a matching accuracy of face samples after scaling processing according to the second candidate ratio, the second data set is an image data set in the same domain as the first detection image, the face samples are image data in the same domain as the second face image, and the second candidate ratio is a scaling ratio of the face samples based on an overall scale of a sample object in the second data set.

[0143] When performing cross-domain matching, the re-identification module 203 obtains an image mean of the second face image according to the second face image, and fills the second face image according to the image mean of the second face image, so that the size of the second face image is consistent with that of the second detection image.

[0144] The specific limitations of the target tracking device can be referred to the limitations of the target tracking method in the foregoing, and will not be described herein. Each module in the target tracking device can be realized by software, hardware, and a combination thereof, in whole or in part. Each module can be embedded in or independent of a processor in a computer device in a hardware form, or can be stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform operations corresponding to each module.

[0145] In an embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram of the computer device can be as shown in Figure 4The computer device includes a processor, a memory, a network interface, a display screen and an input device connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement a target tracking method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0146] Those skilled in the art can understand that, Figure 4 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. A specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0147] In one embodiment, a computer device is provided, including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the following steps when executing the computer program:

[0148] A first detection image is acquired, and a tracking object is tracked according to the first detection image;

[0149] In response to the loss of the tracking object, a second detection image is acquired, the second detection image being a same-domain image of the first detection image;

[0150] An object in the second detection image is matched according to a face image to re-identify the object that is successfully matched as the tracking object, the face image including a first face image in the first detection image and / or a preset second face image, wherein the second face image is a cross-domain image of the first detection image, and wherein the same-domain image is an image according to a first shooting device under a first shooting condition, and the cross-domain image is an image according to a second shooting device under a second shooting condition.

[0151] In one embodiment, the processor further implements the following steps when executing the computer program:

[0152] In a case where the first detection image includes a first face image of the tracking object, the object in the second detection image is matched according to the first face image.

[0153] Otherwise, cross-domain matching is performed on the object in the second detection image according to the second face image.

[0154] In one embodiment, the processor, when executing the computer program, further implements the following steps:

[0155] In response to the intra-domain matching failure, cross-domain matching is performed on the object in the second detection image according to the second face image.

[0156] In one embodiment, the processor, when executing the computer program, further implements the following steps:

[0157] The second face image is scaled based on the scaling ratio, and the cross-domain matching is performed based on the scaled second face image.

[0158] The scaling ratio includes a first scaling ratio or a second scaling ratio, the first scaling ratio is a scaling ratio based on the first face image, and the second scaling ratio is a scaling ratio based on the whole of the tracked object.

[0159] In one embodiment, the processor, when executing the computer program, further implements the following steps:

[0160] In a case where the first face image of the tracked object exists in the first detection image, the second face image is scaled according to the first scaling ratio, and the cross-domain matching is performed based on the scaled second face image.

[0161] The first scaling ratio is obtained from the first candidate ratio according to a first optimization algorithm, an optimization target of the first optimization algorithm is a matching accuracy of face samples after scaling processing according to the first candidate ratio, the first data set is an image data set in the same domain as the first detection image, the face samples are image data in the same domain as the second face image, and the first candidate ratio is a ratio of scaling the face samples based on the face scale of a sample object in the first data set.

[0162] In one embodiment, the processor, when executing the computer program, further implements the following steps:

[0163] In a case where the first face image of the tracked object is not detected in the first detection image, the second face image is scaled according to the second scaling ratio, and the cross-domain matching is performed based on the scaled second face image.

[0164] The second scaling ratio is obtained from the second candidate ratio according to a second optimization algorithm, an optimization target of the second optimization algorithm is a matching accuracy of face samples after scaling processing according to the second candidate ratio, the second data set is an image data set in the same domain as the first detection image, the face samples are image data in the same domain as the second face image, and the second candidate ratio is a ratio of scaling the face samples based on the whole scale of a sample object in the second data set.

[0165] In one embodiment, the processor, when executing the computer program, also implements the following steps:

[0166] According to the second face image, an image mean of the second face image is obtained;

[0167] According to the image mean of the second face image, the second face image is filled to make the size of the second face image consistent with the second detection image.

[0168] In one embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the following steps:

[0169] Obtaining a first detection image, and tracking a tracked object according to the first detection image;

[0170] In response to a loss of the tracked object, a second detection image is obtained, and the second detection image is a same-domain image of the first detection image;

[0171] Matching objects in the second detection image according to face images to re-identify the matched objects as the tracked object, the face images including a first face image in the first detection image and / or a preset second face image, wherein the second face image is a cross-domain image of the first detection image, and wherein the same-domain image is an image obtained according to a first shooting device under a first shooting condition, and the cross-domain image is an image obtained according to a second shooting device under a second shooting condition.

[0172] In one embodiment, the computer program, when executed by the processor, also implements the following steps:

[0173] In a case where the first detection image includes a first face image of the tracked object, performing same-domain matching of objects in the second detection image according to the first face image;

[0174] Otherwise, performing cross-domain matching of the objects in the second detection image according to the second face image.

[0175] In one embodiment, the computer program, when executed by the processor, also implements the following steps:

[0176] In response to a failure of the same-domain matching, performing cross-domain matching of the objects in the second detection image according to the second face image.

[0177] In one embodiment, the computer program, when executed by the processor, also implements the following steps:

[0178] Scaling the second face image based on a scaling ratio, and performing the cross-domain matching based on the scaled second face image;

[0179] The scaling ratio includes a first scaling ratio or a second scaling ratio, the first scaling ratio being a scaling ratio based on the first face image, and the second scaling ratio being a scaling ratio based on the tracking object as a whole.

[0180] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0181] In a case where the first face image of the tracking object exists in the first detection image, the second face image is scaled according to the first scaling ratio, and cross-domain matching is performed based on the scaled second face image;

[0182] The first scaling ratio is obtained from the first candidate ratio according to a first optimization algorithm, an optimization target of the first optimization algorithm being a matching accuracy of face samples after scaling processing according to the first candidate ratio in a first data set, the first data set being an image data set in the same domain as the first detection image, the face samples being image data in the same domain as the second face image, and the first candidate ratio being a ratio of scaling the face samples based on a face scale of a sample object in the first data set.

[0183] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0184] In a case where the first face image of the tracking object is not detected in the first detection image, the second face image is scaled according to the second scaling ratio, and cross-domain matching is performed based on the scaled second face image;

[0185] The second scaling ratio is obtained from the second candidate ratio according to a second optimization algorithm, an optimization target of the second optimization algorithm being a matching accuracy of face samples after scaling processing according to the second candidate ratio in a second data set, the second data set being an image data set in the same domain as the first detection image, the face samples being image data in the same domain as the second face image, and the second candidate ratio being a ratio of scaling the face samples based on an overall scale of a sample object in the second data set.

[0186] In one embodiment, the computer program, when executed by the processor, further implements the following steps:

[0187] According to the second face image, an image mean of the second face image is obtained;

[0188] According to the image mean of the second face image, the second face image is padded so that a size of the second face image is consistent with the second detection image.

[0189] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiment methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0190] Any combination of the technical features of the above embodiments can be made, and in order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0191] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A target tracking method characterized by, The method comprises: acquiring a first detection image, and tracking a tracked object according to the first detection image; in response to a loss of the tracked object, acquiring a second detection image, the second detection image being a same-domain image of the first detection image; matching an object in the second detection image according to a face image to re-identify the object that is successfully matched as the tracked object, the face image comprising a first face image in the first detection image and / or a preset second face image, wherein the second face image is a cross-domain image of the first detection image; wherein the same-domain image is an image acquired according to a first shooting device under a first shooting condition, and the cross-domain image is an image acquired according to a second shooting device under a second shooting condition.

2. The object tracking method of claim 1, wherein, The matching of the object in the second detection image according to the face image comprises: in a case where the first face image of the tracked object exists in the first detection image, performing same-domain matching of the object in the second detection image according to the first face image; otherwise, performing cross-domain matching of the object in the second detection image according to the second face image.

3. The target tracking method according to claim 2, characterized in that, After the same-domain matching of the object in the second detection image according to the first face image, the method further comprises: in response to a failure of the same-domain matching, performing cross-domain matching of the object in the second detection image according to the second face image.

4. The object tracking method according to any one of claims 2-3, characterized in that, The cross-domain matching of the object in the second detection image according to the second face image comprises: scaling the second face image based on a scaling ratio, and performing cross-domain matching based on the scaled second face image; the scaling ratio comprises a first scaling ratio or a second scaling ratio, the first scaling ratio being a scaling ratio based on a face scale of the first face image, and the second scaling ratio being a scaling ratio based on an overall scale of the tracked object.

5. The target tracking method according to claim 4, characterized in that, The scaling of the second face image based on the scaling ratio comprises: in a case where the first face image of the tracked object exists in the first detection image, scaling the second face image according to the first scaling ratio, and performing cross-domain matching based on the scaled second face image; wherein the first scaling ratio is obtained from a first candidate ratio according to a first optimization algorithm, an optimization target of the first optimization algorithm being a matching accuracy rate of a face sample after scaling processing according to the first candidate ratio on a first data set, the first data set being a same-domain image data set of the first detection image, the face sample being image data that is same-domain with the second face image, and the first candidate ratio being a scaling ratio of the face sample based on a face scale of a sample object in the first data set.

6. The target tracking method according to claim 4, characterized by, The scaling of the second face image based on the scaling ratio comprises: in a case where the first face image of the tracked object is not detected in the first detection image, scaling the second face image according to the second scaling ratio, and performing cross-domain matching based on the scaled second face image; The second scaling ratio is obtained from a second candidate scaling ratio according to a second optimization algorithm, an optimization target of the second optimization algorithm being a matching accuracy of the face sample after scaling processing according to the second candidate scaling ratio on a second data set, the second data set being an image data set in the same domain as the first detection image, the face sample being image data in the same domain as the second face image, and the second candidate scaling ratio being a scaling ratio of the face sample based on an overall scale of a sample object in the second data set.

7. The object tracking method according to any one of claims 2-3, characterized in that, The cross-domain matching of the object in the second detection image according to the second face image further includes: obtaining an image mean of the second face image according to the second face image; filling the second face image according to the image mean of the second face image, so that the size of the second face image is consistent with the second detection image.

8. A target tracking device, characterized by, The apparatus includes: an acquisition module configured to acquire a first detection image, and in response to loss of a tracked object, acquire a second detection image, the second detection image being a same-domain image of the first detection image; a tracking module configured to track the tracked object according to the first detection image, and determine whether the tracked object is lost; a re-identification module configured to, in a case where the tracked object is lost, match an object in the second detection image according to a face image, to re-identify the matched object as the tracked object, the face image including a first face image in the first detection image and / or a preset second face image, wherein the second face image is a cross-domain image of the first detection image; wherein the same-domain image is an image obtained according to a first shooting device under a first shooting condition, and the cross-domain image is an image obtained according to a second shooting device under a second shooting condition.

9. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method of any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • A pedestrian re-identification method and device and equipment

    CN109784130A

  • Object tracking method, device and system and storage medium

    CN110706250A

  • Target tracking method and system based on air-ground cooperation

    CN117218157A

  • Target tracking method and device, intelligent selling cabinet and storage medium

    CN119672363A

  • Face-Tracking Method with High Accuracy

    US20130094768A1