Human image tracking method and device based on lost tracking YOLOV8 model

By adding the L1 Loss loss function and replacing the dynamic allocation strategy of simOTA in the last 10 epochs of the YOLOV8 model, combining face clustering, recognition and verification technology, quickly filtering and tracking the lost people's action routes, and ensuring the rapid transmission of information through classified push messages, the problem of excessive time-consuming search process is solved, and the speed and accuracy of searching is significantly improved.

CN117671726BActive Publication Date: 2025-06-06GUANGXI UNIV FOR NATITIES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311463118.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-06-06
Estimated Expiration
2043-11-06

AI Technical Summary

Technical Problem

The prior art takes too long to find a person, especially in areas with low camera coverage, which makes it difficult to accurately track the lost person's movement route and send lost information quickly and accurately.

Method used

The character image tracking method based on the YOLOV8 model is adopted, and the accuracy and speed of the model are improved by adding the L1 Loss loss function to the last 10 epochs of the model and replacing the Taskaligned Assigner with simOTA. At the same time, face clustering technology, recognition technology and verification technology are used to quickly filter out the camera picture collection where the lost person is located, and draw the lost person's movement trajectory through the map. Finally, by pushing messages in a classified manner, ensure that the information can be delivered to relevant personnel quickly and accurately.

Benefits of technology

It significantly improves the speed and accuracy of the search process, can effectively track the lost people's movement routes in areas with low camera coverage, and ensures the rapid and accurate transmission of information, reducing the possibility of danger.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117671726B_ABST
    Figure CN117671726B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for tracking human images based on a lost tracking YOLOV8 model, and relates to the field of image recognition technology. A lost tracking YOLOV8 model for finding lost people is constructed; a camera captured picture set is collected, and the lost tracking YOLOV8 model is trained to obtain a complete picture set of people; a portrait of the lost person is compared with a complete picture set of people to obtain a complete picture set of the lost person; a moving trajectory is obtained; and a message is pushed to a first target and a second target. In order to solve the technical problems of model instability, slow training speed, and inability to track lost people, the present invention improves the stability, accuracy, and speed of the model by improving the dynamic distribution strategy and adding a loss function to the existing YOLOV8 model; and the way of sending missing person messages is optimized, so as to make greater use of the power of the masses, shorten the time of tracking lost people, and avoid the possibility of more dangers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a method and device for tracking human images based on a lost tracking YOLOV8 model. Background Art

[0002] For example, a Chinese patent application number PCT / CN2020 / 112694 discloses an image classification method and device that uses a neural network to classify images, remove noise, and locate lost persons. Combining convolutional neural networks with image recognition technology can effectively identify characteristic information such as faces, clothing, and vehicles, and compare them with lost persons to determine the movement trajectory of lost persons and track the target. While combining multiple technologies, the YOLOV8 model of the convolutional neural network is optimized to further avoid the problem of difficulty in locating the target person due to crowds. In such scenarios, the speed of finding people is a very important factor. Finding the lost person as soon as possible can avoid the possibility of many dangerous situations, so the YOLOV8 model needs to be optimized to increase the speed again.

[0003] The video information obtained by the YOLOV8 algorithm is obtained through cameras. Although the camera coverage rate in first-tier cities has almost reached 100%, most cities cannot achieve high coverage. In sections without cameras, the route of the missing person will be lost, and the public needs to be called in to provide help. After the algorithm calculates the missing route, how to quickly and accurately send the missing information is the key to quickly finding the missing person. Since the unobvious notification is easy to be ignored by people, using unblocked message push and separately reminding passers-by who have had contact with the missing person at the time closest to the current time can improve the speed of finding people and provide information for finding people more efficiently.

[0004] In order to solve the problem that the process of searching for people takes too long, a person image tracking method and device based on the lost tracking YOLOV8 model are proposed. Summary of the invention

[0005] The purpose of the present invention is to provide a person image tracking method based on the lost tracking YOLOV8 model to solve the above technical problems, wherein the person and object information analyzed by face recognition and image recognition technology has been authorized by the user, and the use of this technology complies with the relevant national laws and regulations.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A person image tracking method based on the lost tracking YOLOV8 model, characterized by comprising:

[0008] S1. Build a YOLOV8 model for missing person tracking;

[0009] The S1 specifically includes:

[0010] S11. Add the loss function L1 Loss to the last 10 epochs of the training set of the lost tracking YOLOV8 model;

[0011] The training set includes 500 epochs;

[0012] The first 490 epochs of the training set were first enhanced using the Mosaic module, and then using HSV data enhancement;

[0013] The Mosaic module enhancement is turned off for the last 10 epochs of the training set, and the L1 Loss function is added after the HSV data enhancement is completed;

[0014] S12. Replace the Taskaligned Assigner in the YOLOV8 model with simOTA;

[0015] The specific implementation steps are as follows:

[0016] S121. Pre-screen the anchor point;

[0017] S122. Calculate the classification loss and position loss of each anchor point, and obtain the cost matrix and iou matrix;

[0018] S123. Select a candidate box according to the iou matrix;

[0019] S124. Dynamically assign the anchor point with the lowest cost value to the gt box;

[0020] S125. Perform secondary screening on the gt box to which multiple anchor points are repeatedly assigned;

[0021] S126. Classify all remaining anchor points as negative samples and calculate the loss of the filtered prediction box.

[0022] Furthermore, the YOLOV8 model improves the Backbone, Neck and Head of the YOLOV5 model; in the Backbone part, the core idea of ​​CSP is still used, the C2f module is used to replace the C3 module, and the k of the convolution layer is changed from 6 to 3; in the Neck part, the convolution in the top down upsampling stage is directly deleted; in the Head part, the coupled-head is changed to a decoupled-head, and the objectness branch is deleted.

[0023] Furthermore, the top down is a target detection method, which first detects the whole object, and then refines it into each step to detect a single object. The CSP architecture is mainly used to solve the information flow problem in deep CNN, which is hindered by the disappearance of gradients and a large number of parameters. The CSP architecture divides the problem into problems involving low spatial resolution and high spatial resolution, and finally connects the feature map of low spatial resolution and the feature map of high spatial resolution.

[0024] Furthermore, the L1 Loss loss function is added to the part of the YOLOV8 model where Mosaic is cancelled; the L1 Loss loss function is the mean absolute error, which is used to measure the mean absolute difference MSE between the model prediction value and the true value, following the formula:

[0025]

[0026] f(x i ) is the predicted value of the i-th sample, y i is the true value of the i-th sample, and n is the total number of samples. The accuracy of the training results of the last 10 epochs is evaluated by L1 Loss.

[0027] Furthermore, a loss function is added in the model training stage, and a predicted value is output after each training data passes through the model. The loss function can calculate the error between the true value and the predicted value, and update each parameter in reverse through the error, thereby reducing the loss between the true value and the predicted value and improving the performance of the model.

[0028] Furthermore, the Mosaic module is an enhancement strategy, which includes a Cutmix submodule; the Cutmix submodule randomly extracts a part of the area in the sample, and replaces and fills the area with the pixel values ​​of the unextracted area.

[0029] Furthermore, the Taskaligned Assigner is a positive and negative sample dynamic assignment strategy, which obtains the weighted scores of the Cls score and Reg Score predicted for all pixels, selects positive samples by sorting the weighted scores, and the specific implementation steps include anchor points falling inside the gt boxes as preliminary screened positive samples; further screening to remove possible negative samples; ensuring that an anchor point is only assigned to one gt box; and obtaining the screened labels.

[0030] Furthermore, the Cls score is a classification score, which is obtained during the target detection process. After the network predicts the bounding boxes of all objects in the image, a classification score is assigned to each bounding box, indicating the probability that the object in the box belongs to each predefined category, and the class with the highest probability is assigned to the bounding box as the predicted class. The classification score is between 0 and 1, and the higher the score, the higher the accuracy of the predicted category. The Reg score is a regression score, which refers to the regression score assigned to each predicted bounding box, representing the accuracy of the predicted bounding box. The regression score is between 0 and 1, and the higher the score, the higher the accuracy of the predicted category. The model calculates the regression score based on the IoU indicator to measure the overlap between the predicted bounding box and the true bounding box, and then uses the regression score to filter and refine the final bounding box prediction.

[0031] Furthermore, the loss is calculated according to the following formula:

[0032]

[0033] L cls is the classification loss, L reg is the positioning loss, L obj is the obj loss, γ is the balance coefficient of positioning loss, and N is the total number of anchor points.

[0034] Furthermore, the loss function is used in the YOLOV8 model to calculate the difference between the predicted bounding box and the true bounding box and the class probability. The loss function is divided into positioning loss, classification loss and objectivity loss. The positioning loss is calculated using the mean square error, the classification loss is calculated using the cross entropy loss, and the objectivity loss is calculated using the binary cross entropy loss. The loss function of the YOLOV8 model is ultimately calculated by the weighted average of the three, which is the above loss function formula.

[0035] Furthermore, the objectness branch is cancelled in the YOLOV8 model, L obj is 0.

[0036] Furthermore, the objectness represents the probability that the bounding box contains the object to be studied, not just the background or noise. The objectness is calculated based on the probability that the center of the bounding box falls within the object of interest and the probability that the bounding box has the correct aspect ratio and size. The objectness is used to filter out low-confidence detections and refine the final bounding box prediction. The objectness can be used to filter out the bounding boxes with low accuracy and practicality and refine the prediction of the bounding box.

[0037] Furthermore, the anchor points are a set of anchor points to which each pixel belongs; the gt boxes, also known as ground truth, are real annotation boxes; the cls is used to determine the type of each feature point; and the reg obtains the predicted box by determining and adjusting the regression parameters of each feature point.

[0038] Furthermore, the IOU stands for Intersection over Union, which conforms to the following formula:

[0039]

[0040] Box 1 and Box 2 both contain anchor points.

[0041] Furthermore, the dynamic allocation is to dynamically allocate anchor points according to each GT box, based on the principle of lower cost and higher priority. Unlike static allocation, which reserves space in advance, it is allocated immediately according to needs.

[0042] S2. Collect the image set captured by the camera and train the missing tracking YOLOV8 model to obtain the image set of all people.

[0043] Furthermore, the camera-captured picture set includes camera pictures near the lost location; and the all-character picture set includes all character portrait pictures extracted from the camera-captured picture set.

[0044] Furthermore, based on the video captured by the camera near the last place where the missing person appeared, the clips containing the portraits are extracted, and the partial pictures containing clear faces are captured as the camera captured picture set. The lost tracking YOLOV8 model is trained with the camera captured picture set to obtain a complete person picture set containing all the portraits.

[0045] S3. Compare the portrait of the missing person with the picture collection of all persons to obtain the picture collection of the missing person.

[0046] Furthermore, the portrait of the lost person includes a life photo of the lost person and / or an ID photo of the lost person; the photo collection of the lost person includes photos of the lost person in the photo collection of all people.

[0047] Furthermore, the comparison includes: S31. using face clustering technology such as Chinese Whispers to group the faces in the entire character picture set into identity groups; S32. using face recognition technology and a method based on geometric features to extract a high-similarity group from the identity group that contains one or more high-similarity features with the portrait of the lost person; if the high-similarity group is not included, repeating steps S2-S3 in a different geographical area; S33. using face verification technology to determine whether the high-similarity group belongs to the lost person through personal features composed of iris and face; the face verification technology includes comparing a preset threshold to determine whether the high-similarity group matches the portrait of the lost person.

[0048] Furthermore, the Chinese Whispers algorithm is a clustering algorithm for machine learning and natural language processing, also known as a hierarchical agglomerative clustering algorithm or a nearest neighbor chain algorithm. The Chinese Whispers algorithm assigns labels to data points based on the nearest label through iteration. In each iteration, the algorithm randomly selects a data point, analyzes the k nearest neighbors around the data point, and then assigns the most frequently occurring label contained in the k nearest neighbors to the randomly selected data point, and iterates continuously until all data points are assigned label features.

[0049] Furthermore, since the identity information of people other than the missing person is not important, the face clustering technology does not require manual identity labeling, and only needs to ensure that all portraits in the group belong to the same person.

[0050] Furthermore, due to the large number of people captured by the camera and the existence of a large number of cameras, the content of the camera picture collection near the lost location is extremely large. Direct face recognition takes too long and is prone to overload. First, use face clustering technology to group them, and then extract a photo from each group for face recognition, which can greatly reduce the workload and improve the comparison speed.

[0051] S4. Obtain movement trajectory.

[0052] Furthermore, the step of obtaining the moving trajectory includes obtaining trajectory information, sorting the missing person picture set in chronological order; drawing the trajectory, locating the camera position corresponding to the missing person picture set, and connecting the camera positions according to the sorting; predicting the missing path, and predicting the untracked path based on the trajectory.

[0053] Furthermore, the missing path includes a route drawn after the first camera, with the position of the first camera as the starting point and based on a third feature; the first camera includes the camera to which the latest camera picture belongs; the third feature includes the road structure and / or terrain direction around the first camera.

[0054] S5. Pushing messages to the first target and the second target; the first target includes people passing through the moving track; the push message specifically includes: S51. Training the lost tracking YOLOV8 model through the latest camera image to obtain the second target; the latest camera image is the last image in the chronological order of the lost person image set; S52. Obtaining the personal information of the second target;

[0055] S53. Push messages to the first target and the second target in a classified manner; the classified push includes a first feature and a second feature; the first feature includes mute and long-time display; the second feature includes a prompt sound and a long-time display.

[0056] Furthermore, a missing person message is silently sent to the first target, and the message is synchronously displayed on at least half of the main screen and the lock screen, and the lock screen mode does not partially shield the message content display.

[0057] Furthermore, the personal information of the second target is compared with the information database to confirm the information of the second target, and a separate message prompt is sent. The message is accompanied by a highly penetrating prompt sound, and the prompt sound is automatically turned off after the message is read. The message is synchronously displayed on at least half of the area of ​​the main screen and the lock screen, and the lock screen mode does not partially block the display of the message content.

[0058] A person image tracking device based on the lost tracking YOLOV8 model, characterized by comprising:

[0059] Preparation module for building the missing person tracking YOLOV8 model for missing person finding;

[0060] An acquisition module is used to collect a set of pictures captured by a camera and train the missing tracking YOLOV8 model to obtain a set of pictures of all people;

[0061] A positioning module is used to compare the portrait of the missing person with the entire person picture set to obtain the missing person picture set;

[0062] Drawing module, used to obtain movement trajectory;

[0063] The message push module is used to push messages to the first target and the second target.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] 1. When the YOLOV8 model cancels Mosaic in the last 10 epochs, add the L1 Loss function to evaluate the accuracy of stopping Mosaic training. YOLOV8 stops Mosaic enhancement in the last 10 epochs based on the previous YOLOV5 model, which improves the accuracy. Adding a loss function on this basis can further improve the accuracy and avoid omissions when extracting portraits from camera photos.

[0066] 2. Replace the Taskaligned Assigner of the YOLOV8 model with simOTA, use the lowest value in the cost matrix to assign gt boxes, and then filter the duplicated gt boxes, which is faster than the sequential assignment of Taskaligned Assigner. SimOTA simplifies OTA and eliminates the manual parameter adjustment step, which can adaptively perform dynamic assignment, improve the speed of face extraction, and avoid time waste.

[0067] 3. After obtaining the movement trajectory of the missing person, the YOLOV8 model for missing person tracking is trained again and the face recognition technology is used to obtain the information of passers-by who have contact with the missing person. The missing person information is released to people near the trajectory and passers-by who have contact with the missing person respectively, and a mandatory display mechanism is adopted. At the same time, the content of the missing person message is displayed on the lock screen to prevent people from ignoring the message. The movement trajectory of the missing person can also be obtained from passers-by on roads without cameras, combining manual and algorithm to maximize the value. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 A flow chart of a method for tracking human images based on a lost tracking YOLOV8 model provided by an embodiment of the present invention;

[0069] Figure 2 A schematic diagram of YOLOV8 model training provided in an embodiment of the present invention;

[0070] Figure 3 A schematic diagram of image recognition results of the YOLOV8 model for lost tracking provided in an embodiment of the present invention;

[0071] Figure 4 Flow chart of the simOTA implementation method provided by the embodiment of the present invention;

[0072] Figure 5 A flow chart of a person image tracking device based on the lost tracking YOLOV8 model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0073] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0074] See also Figures 1 to 5 The present invention provides a method and device for tracking human images based on the YOLOV8 model for lost tracking. The technical solution is as follows:

[0075] Specifically, refer to Figure 1 As shown, the present invention provides a person image tracking method based on the lost tracking YOLOV8 model, which can be performed by a device, and the device can be implemented by software and / or hardware. In a specific implementation, it includes steps S1 to S5.

[0076] Specifically, S1. construct a lost person tracking YOLOV8 model for finding lost persons; S2. collect a set of pictures captured by a camera, train the lost person tracking YOLOV8 model to obtain a set of pictures of all persons; S3. compare the portrait of the lost person with the set of pictures of all persons to obtain a set of pictures of the lost person; S4. obtain a moving trajectory; S5. push a message to the first target and the second target.

[0077] S1. Constructing a lost person tracking YOLOV8 model for finding lost people; S1 specifically includes:

[0078] S11. Add the loss function L1 Loss to the last 10 epochs of the training set of the lost tracking YOLOV8 model; the training set includes 500 epochs; the first 490 epochs of the training set are first enhanced by the Mosaic module, and then the HSV data is enhanced; the last 10 epochs of the training set are closed for the Mosaic module enhancement, and the L1 Loss loss function is added after the HSV data enhancement is completed;

[0079] S12. Replace the Taskaligned Assigner in the YOLOV8 model with simOTA; the specific execution steps are as follows, refer to Figure 4As shown: S121. Pre-screen the anchor points; S122. Calculate the classification loss and position loss of each anchor point, and obtain the cost matrix and iou matrix; S123. Select the candidate box according to the iou matrix; S124. Dynamically assign the anchor point with the lowest cost value to the gt box; S125. Perform secondary screening on the gt box that repeatedly assigns multiple anchor points; S126. Classify all the remaining anchor points as negative samples and calculate the loss of the screened prediction box.

[0080] Specifically, the YOLOV8 model improves the Backbone, Neck and Head of the YOLOV5 model; in the Backbone part, the core idea of ​​CSP is still used, the C2f module is used to replace the C3 module, and the k of the convolutional layer is changed from 6 to 3; in the Neck part, the convolution in the top down upsampling stage is directly deleted; in the Head part, the coupled-head is changed to a decoupled-head, and the objectness branch is deleted.

[0081] Specifically, the top down is a target detection method that first detects the whole object, and then refines it into each step to detect a single object. The CSP architecture is mainly used to solve the information flow problem in deep CNN, which is hindered by the disappearance of gradients and a large number of parameters. The CSP architecture divides the problem into problems involving low spatial resolution and high spatial resolution, and finally connects the feature map of low spatial resolution and the feature map of high spatial resolution.

[0082] Specifically, the L1 Loss loss function is added to the part of the YOLOV8 model where Mosaic is cancelled; the L1 Loss loss function is the mean absolute error, which is used to measure the mean absolute difference MSE between the model prediction value and the true value, following the formula:

[0083]

[0084] f(x i ) is the predicted value of the i-th sample, y i is the true value of the i-th sample, and n is the total number of samples. The accuracy of the training results of the last 10 epochs is evaluated by L1 Loss.

[0085] Specifically, a loss function is added during the model training phase, and a predicted value is output after each training data passes through the model. The loss function can calculate the error between the true value and the predicted value, and update each parameter in reverse based on the error, thereby reducing the loss between the true value and the predicted value and improving the performance of the model.

[0086] Specifically, if Figure 2 As shown, since YOLOV8 is improved on the basis of YOLOV5, the original use of Mosaic enhancement was changed to the final cancellation. This change has an impact on the running speed and produces instability. According to the comparison results of the embodiment, the five sub-models n, s, l, m, and x, except l and x, the other three have increased from 1.9M, 7.2M, and 21.2M to 3.2M, 11.2M, and 25.9M respectively; the accuracy of the five models has increased from 4.5B, 16.5B, 49.0B, 109.1B, and 205.7B to 8.7B, 28.6B, 78.9B, 165.2B, and 257.8B respectively. Adding the L1 Loss loss function can enhance stability by real-time monitoring of the mean absolute error value. Since the loss speed is exchanged for the improvement of accuracy, the speed cannot be improved under this condition.

[0087] Specifically, the Mosaic module is an enhancement strategy that includes a Cutmix submodule; the Cutmix submodule randomly extracts a portion of the sample and replaces the region with the pixel values ​​of the unextracted region. Mosaic enhances local image recognition and positioning capabilities to avoid unnatural mixing of images.

[0088] Specifically, the replacement includes replacing the dynamic allocation strategy, namely Taskaligned Assigner, with simOTA.

[0089] Specifically, the Taskaligned Assigner is a dynamic assignment strategy for positive and negative samples, which obtains the weighted scores of Cls score and Reg Score predicted for all pixels, selects positive samples by sorting the weighted scores, and the specific implementation steps include anchor points falling inside gt boxes as preliminary screened positive samples; further screening to remove possible negative samples; ensuring that an anchor point is only assigned to one gt box; and obtaining the screened labels.

[0090] Specifically, the Cls score is a classification score, which is obtained during the target detection process. After the network predicts the bounding boxes of all objects in the image, a classification score is assigned to each bounding box, indicating the probability that the object in the box belongs to each predefined category. The class with the highest probability is assigned to the bounding box as the predicted class. The classification score is between 0 and 1, and the higher the score, the higher the accuracy of the predicted category. The Reg score is a regression score, which refers to the regression score assigned to each predicted bounding box, representing the accuracy of the predicted bounding box. The regression score is between 0 and 1, and the higher the score, the higher the accuracy of the predicted category. The model calculates the regression score based on the IoU indicator to measure the overlap between the predicted bounding box and the true bounding box, and then uses the regression score to filter and refine the final bounding box prediction.

[0091] Specifically, the loss is calculated according to the following formula:

[0092]

[0093] L cls is the classification loss, L reg is the positioning loss, L obj is the obj loss, γ is the balance coefficient of positioning loss, and N is the total number of anchor points.

[0094] Specifically, the loss function is used in the YOLOV8 model to calculate the difference between the predicted bounding box and the true bounding box and the class probability. The loss function is divided into positioning loss, classification loss and objectivity loss. The positioning loss is calculated using the mean square error, the classification loss is calculated using the cross entropy loss, and the objectivity loss is calculated using the binary cross entropy loss. The loss function of the YOLOV8 model is ultimately calculated by the weighted average of the three, which is the above loss function formula.

[0095] Specifically, the objectness branch is cancelled in the YOLOV8 model, L obj is 0.

[0096] Specifically, the objectness represents the likelihood that the bounding box contains the object of interest, not just the background or noise. The objectness is calculated based on the probability that the center of the bounding box falls within the object of interest and the probability that the bounding box has the correct aspect ratio and size. The objectness is used to filter out low-confidence detections and refine the final bounding box prediction. The objectness can be used to filter out bounding boxes with low accuracy and practicality and refine the prediction of the bounding box.

[0097] Specifically, the anchor points are a set of anchor points to which each pixel belongs; the gt boxes, also known as ground truth, are real annotation boxes; the cls is used to determine the type of each feature point; and the reg obtains the predicted box by determining and adjusting the regression parameters of each feature point.

[0098] Specifically, the IOU stands for Intersection over Union, which conforms to the following formula:

[0099]

[0100] Box 1 and Box 2 both contain anchor points.

[0101] Specifically, the dynamic allocation is to dynamically allocate anchor points according to each GT box, according to the principle of lower cost and higher priority. Unlike static allocation, which reserves space in advance, it is allocated immediately according to needs.

[0102] Specifically, since the changes based on YOLOV5 have sacrificed speed to a certain extent, changing the dynamic allocation strategy from Taskaligned Assigner to simOTA can make up for this loss. In the case of training 300 epochs, the training time will increase by 25%. For this reason, it is replaced by SimOTA and manually set rules are used to shorten the training time.

[0103] Adding the L1 Loss function and changing the dynamic allocation strategy from Taskaligned Assigner to simOTA improves the computing speed and model stability while preserving accuracy.

[0104] S2. Collect the image set captured by the camera and train the missing tracking YOLOV8 model to obtain the image set of all people.

[0105] Specifically, the camera-captured picture set includes camera pictures near the lost location; the all-character picture set includes all character portrait pictures extracted from the camera-captured picture set.

[0106] Specifically, based on the video captured by the camera near the last place where the missing person appeared, extract the clips containing the portraits, cut out the parts containing clear faces to obtain the pictures as the camera captured pictures set, and use the camera captured pictures set to train the missing tracking YOLOV8 model to obtain the complete person picture set containing all the portraits. Figure 3 As shown, Figure 3 This is the image recognition result obtained by training the YOLOV8 model.

[0107] S3. Compare the portrait of the missing person with the picture collection of all persons to obtain the picture collection of the missing person.

[0108] Specifically, the portrait of the lost person includes a life photo of the lost person and / or an ID photo of the lost person; the picture collection of the lost person includes photos in the picture collection of all people that belong to the same person as the portrait of the lost person.

[0109] Specifically, the comparison includes: S31. Using face clustering technology such as Chinese Whispers to group the faces in the entire character picture set according to identity to obtain identity groups; S32. Using face recognition technology and a method based on geometric features, extracting a high-similarity group from the identity group that contains one or more high-similarity features with the portrait of the lost person; if the high-similarity group is not included, repeating steps S2-S3 in a different geographical area; S33. Using face verification technology, judging whether the high-similarity group belongs to the lost person through personal features composed of iris and face; the face verification technology includes comparing a preset threshold to determine whether the high-similarity group matches the portrait of the lost person.

[0110] Specifically, the Chinese Whispers algorithm is a clustering algorithm for machine learning and natural language processing, also known as a hierarchical agglomerative clustering algorithm or a nearest neighbor chain algorithm. The Chinese Whispers algorithm assigns labels to data points based on the nearest label through iteration. In each iteration, the algorithm randomly selects a data point, analyzes the k nearest neighbors around the data point, and then assigns the most frequently occurring label contained in the k nearest neighbors to the randomly selected data point, and iterates continuously until all data points are assigned label features.

[0111] Specifically, since the identity information of people other than the missing person is not important, the face clustering technology does not require manual identity labeling, and only needs to ensure that all portraits in the group belong to the same person.

[0112] Specifically, due to the large number of people captured by the cameras and the large number of cameras, the content of the camera picture collection near the lost location is extremely large. Direct face recognition takes too long and is prone to overload. First, use face clustering technology to group them, and then extract a photo from each group for face recognition. This can greatly reduce the workload and increase the comparison speed.

[0113] S4. Obtain movement trajectory.

[0114] Specifically, the method includes the following steps: obtaining trajectory information, sorting the missing person picture set in chronological order; drawing a trajectory, locating the camera position corresponding to the missing person picture set, and connecting the camera positions according to the sorting; predicting the missing path, and predicting the untracked path based on the trajectory.

[0115] Specifically, the missing path includes a route drawn after the first camera, with the position of the first camera as the starting point, and based on a third feature; the first camera includes the camera to which the latest camera picture belongs; the third feature includes the road structure and / or terrain direction around the first camera.

[0116] S5. Push messages to the first target and the second target; the first target includes people passing through the moving trajectory; the push message specifically includes: S51. Train the lost tracking YOLOV8 model through the latest picture of the camera to obtain the second target; the latest picture of the camera is the last picture in the chronological order of the lost person picture set; S52. Obtain the personal information of the second target; S53. Classify and push messages to the first target and the second target; the classified push includes a first feature and a second feature; the first feature includes mute and long-time display; the second feature includes a prompt sound and a long-time display.

[0117] Specifically, a missing person message is silently sent to the first target, and the message is synchronously displayed on at least half of the main screen and the lock screen, and the lock screen mode does not partially block the display of the message content.

[0118] Specifically, the personal information of the second target is compared with the information database, the information of the second target is confirmed, and a message prompt is sent separately. The message is accompanied by a highly penetrating prompt sound. The prompt sound is automatically turned off after the message is read. The message is synchronously displayed on at least half of the area of ​​the main screen and the lock screen. The lock screen mode does not partially block the display of the message content.

[0119] Specifically, this message notification method refers to Amber Alert, which is used to find abducted children. It notifies a wide range of people through a long alarm to help provide information, and the abducted children can be found within only three hours. However, there are many disadvantages. There are many complaints that the amber alert disturbs the public, and it is not applicable to China's national conditions. Due to the concentrated residence of people, sending missing person information accompanied by a penetrating reminder sound will cause excessive noise pollution, and people who are not on the movement track cannot provide assistance to the search, so the information can only be released to people who may have been in contact. In particular, for people who appear at the same time as the lost person was last photographed by the camera, more effective information can be provided, and information should be collected from these people.

[0120] Tracking lost people in indoor scenes of shopping malls includes:

[0121] Close all entrances and exits as soon as the person is found missing, obtain the photo and / or clothing features and / or appearance features of the missing person, obtain the video clips containing portraits taken by all cameras in the room, and cut out the parts containing clear faces to obtain a picture set. Substitute the picture set into the lost tracking YOLOV8 model to obtain all portrait pictures in the picture set, and then use face clustering technology such as Chinese Whispers, face recognition technology similar to the method based on geometric features, and face verification technology based on personal features such as iris and facial composition to compare all portrait pictures with the features of the missing person, and extract the camera picture set containing the missing person. Sort the camera picture set containing the missing person in chronological order. If the missing person is not found near the last position in the sorting, connect the camera positions according to the time sorting to obtain the movement trajectory, and predict the future movement trajectory according to the structure of the shopping mall.

[0122] While completing the above operations, the photo and / or clothing features and / or appearance features of the missing person are sent to all shops with a strong penetrating prompt sound, and are sent silently to all people in the mall.

[0123] Missing persons who have been tracked for less than 24 hours include:

[0124] Obtain the photos and / or clothing features and / or appearance features of the missing person, obtain the videos shot by all cameras within a radius of 10 kilometers with the last place where the missing person appeared as the center, extract the clips containing the portrait, and cut out the part containing the clear face to obtain the picture set. Substitute the picture set into the lost tracking YOLOV8 model to obtain all the portrait pictures in the picture set, and then use the face clustering technology such as Chinese Whispers, face recognition technology similar to the method based on geometric features, and face verification technology based on personal features such as iris and facial composition to compare all the portrait pictures with the features of the missing person, and extract the camera picture set containing the missing person. Sort the camera picture set containing the missing person in chronological order, obtain the camera position of each photo, use the map as the background, connect the camera positions in chronological order, and obtain the historical movement trajectory of the missing person; according to the urban planning and / or terrain direction near the last position sorted at the time point, connect all possibilities starting from the last position photographed by the camera to draw the future movement trajectory.

[0125] Substitute the last camera image sorted by time point into the missing person tracking YOLOV8 model again to obtain all the portraits contained in the photo, and perform face recognition to obtain the shape information of the person who had the last contact with the missing person. Send the missing person message by category, with a strong penetrating prompt sound to the person who had the last contact with the missing person, and send it silently to people near the movement track. The message will be displayed on the main screen and lock screen in an area exceeding half of the screen.

[0126] Missing persons tracked for more than 24 hours include:

[0127] Obtain the photo and / or clothing features and / or appearance features of the missing person, obtain the video clips containing portraits taken by all cameras within a radius of 10 kilometers with the last appearance of the missing person as the center, and intercept the part containing clear faces to obtain a picture set. Substitute the picture set into the missing tracking YOLOV8 model to obtain all portrait pictures in the picture set, and then use face clustering technology such as Chinese Whispers, face recognition technology similar to the method based on geometric features, and face verification technology based on personal features such as iris and facial composition to compare all portrait pictures with the features of the missing person, and extract the camera picture set containing the missing person. If the picture set does not contain the missing person, expand the search radius and repeat the training of the missing tracking YOLOV8 model and face recognition until the picture containing the missing person is found. Repeat the use of the camera containing the picture of the missing person as the last appearance of the missing person, and repeat the above operation within the circular range with this place as the center until the time difference between the picture containing the missing person and the current time does not exceed 3 hours, sort the camera picture set containing the missing person in chronological order, and obtain the camera position to which each photo belongs. With the map as the background, the camera positions are connected in chronological order to obtain the historical movement trajectory of the missing person; based on the urban planning and / or terrain direction near the last position sorted at the time point, all possibilities are connected starting from the last position captured by the camera to draw the future movement trajectory.

[0128] Substitute the last camera image sorted by time point into the missing person tracking YOLOV8 model again to obtain all the portraits contained in the photo, and perform face recognition to obtain the shape information of the person who had the last contact with the missing person. Send the missing person message by category, with a strong penetrating prompt sound to the person who had the last contact with the missing person, and send it silently to people near the movement track. The message will be displayed on the main screen and lock screen in an area exceeding half of the screen.

[0129] In summary, the present invention first improves the accuracy by adding the L1 Loss loss function in the last 10 epochs of the training process. Since Mosaic enhancement was used in the previous training process and was only turned off in the last 10 epochs, although it improved the accuracy while turning it off, it was not conducive to the stability of the model. The addition of the L1 Loss loss function can further improve the accuracy while improving the stability, laying a good foundation for the next step of taking speed improvement as the highest level goal and appropriately giving up accuracy. Then, the dynamic allocation strategy Taskaligned Assigner was replaced with simOTA. SimOTA can automatically analyze how many positive samples each gt box should have, reducing the time-consuming manual addition; it can also automatically determine the feature map to be detected corresponding to each gtbox. Compared with OTA and Taskaligned Assigner, simOTA has a faster computing speed and can avoid additional hyperparameters. Next, by collecting the video captured by the camera near the place where the lost person last appeared, extract the clip containing the portrait, intercept the part containing the clear face to obtain the picture, and collect the camera pictures near the place where the lost person last appeared to obtain the camera picture set near the lost place. Through the lost tracking YOLOV8 model, all the faces contained in the picture set are extracted. This operation can collect all the portrait pictures of the lost person near the lost place, which is convenient for analysis. Then, by using the face clustering technology, the faces in the camera picture set are grouped, and the face pictures belonging to the same person are grouped together; by using the face recognition technology, the picture group that may be the same person as the portrait of the lost person is located; by using the face verification technology, it is determined whether the photo group belongs to the lost person. Through the three technologies, grouping is first performed to reduce the workload and working time, and avoid unnecessary time waste. Then, through feature extraction and comparison, find out whether the lost person exists in the photo set. If so, the camera position where the lost person was captured can be obtained. If not, change the area search until the lost person is found. According to the camera locations of all the photos that have been identified as containing the missing person, the map is marked, and the existing points are connected into lines in chronological order to obtain the movement trajectory, and the future movement trajectory is obtained according to the urban road planning or terrain direction. The movement trajectory can be predicted in advance to help find the missing person as soon as possible. Finally, the camera photo containing the missing person closest to the current time is substituted into the lost tracking YOLOV8 model, the portraits of passers-by around the lost person are extracted, compared with the personal information database, and the information of the passers-by who have contact is confirmed. The missing person message is sent in two different forms to people near the track and passers-by who have contact. Because passers-by who have contact may have a greater impression of the lost person and can provide more help, the missing person message is sent with a penetrating prompt sound, and the prompt sound is sent silently to people near the track to avoid panic caused by the prompt sound.With the help of the public, we can obtain the trajectory of people who are lost in sections of the road that are not fully covered by cameras, and provide information for finding people more efficiently.

[0130] Figure 5 Flow chart of the person image tracking device based on the lost tracking YOLOV8 model provided by the embodiment of the present invention, refer to Figure 5 As shown, specifically, the technical solution is as follows:

[0131] S100. Preparation module, used to build a lost tracking YOLOV8 model for finding lost persons; S200. Acquisition module, used to collect a set of pictures captured by a camera, train the lost tracking YOLOV8 model to obtain a set of pictures of all persons; S300. Positioning module, used to compare the portrait of the lost person with the set of pictures of all persons, and obtain a set of pictures of the lost person; S400. Drawing module, used to obtain a moving trajectory; S500. Message push module, used to push messages to the first target and the second target.

[0132] S100. A preparation module is used to construct a lost person tracking YOLOV8 model for finding lost persons. The lost person tracking YOLOV8 model specifically includes: S11. Adding the loss function L1 Loss to the last 10 epochs of the YOLOV8 model training set; S12. Changing the dynamic distribution strategy of the YOLOV8 model from Taskaligned Assigner to simOTA.

[0133] Specifically, the L1 Loss function is added to the part of the YOLOV8 model that cancels Mosaic; the L1 Loss loss function is the mean absolute error, which is used to measure the average value of the absolute difference between the model prediction value and the true value, and the L1 Loss loss function is used to evaluate the accuracy of stopping Mosaic training. simOTA can automatically analyze how many positive samples each gt box should have, reducing the time-consuming manual addition; it can also automatically determine the feature map to be detected corresponding to each gt box.

[0134] Specifically, a loss function is added during the model training phase, and a predicted value is output after each training data passes through the model. The loss function can calculate the error between the true value and the predicted value, and update each parameter in reverse based on the error, thereby reducing the loss between the true value and the predicted value and improving the performance of the model.

[0135] S200. An acquisition module is used to collect a set of pictures captured by a camera and train the missing tracking YOLOV8 model to obtain a set of pictures of all people.

[0136] Specifically, the camera extracts fragments containing portraits in a circular area with a radius of 10 kilometers and the lost location as the center, and the parts containing clear faces are captured as the camera captured picture set. The lost tracking YOLOV8 model is trained with the camera captured picture set to obtain a complete person picture set containing all portraits.

[0137] S300. A positioning module is used to compare the portrait of the missing person with a collection of all person pictures to obtain a collection of pictures of the missing person.

[0138] Specifically, the faces in the entire person picture set are grouped by using face clustering technology such as Chinese Whispers; the grouping that matches the portrait of the missing person is extracted by face recognition technology, similar to the method based on geometric features; and the grouping is determined to belong to the missing person by face verification technology based on personal features such as iris and facial composition. Since the identity information of people other than the missing person is not important, the face clustering technology does not require manual identity labeling, and only needs to ensure that all portraits in the group belong to the same person.

[0139] Specifically, due to the large number of people captured by the cameras and the large number of cameras, the content of the camera picture collection near the lost location is extremely large. Direct face recognition takes too long and is prone to overload. First, use face clustering technology to group them, and then extract a photo from each group for face recognition. This can greatly reduce the workload and increase the comparison speed.

[0140] S400. A drawing module is used to obtain a moving trajectory.

[0141] Specifically, by obtaining the camera positions containing the missing person's photo set, all cameras are connected in chronological order with the map as the background to obtain a historical trajectory map. The camera position with the latest time point is found, and the future trajectory map of the missing person is determined based on the terrain direction and urban road planning near the position.

[0142] S500. A message push module, used to push messages to a first target and a second target.

[0143] Specifically, a classified push method is adopted to send silent message reminders to people who have passed the location included in the historical trajectory map of the lost person, and a message reminder accompanied by a prompt sound is used for people who appear in the last picture taken by the camera at the same time as the lost person.

[0144] Specifically, a missing person message is silently sent to people near the track, and the message is synchronously displayed on at least half of the main screen and the lock screen, and the lock screen mode does not partially shield the message content display.

[0145] Specifically, the personal information of the contacted passerby is compared with the information database, the information of the passerby is confirmed, and a message prompt is sent separately. The message is accompanied by a highly penetrating prompt sound. The prompt sound is automatically turned off after the message is read. The message is synchronously displayed on at least half of the main screen and the lock screen, and the lock screen mode does not partially shield the message content display. Specifically, the working principle of this device is based on the above-mentioned character image tracking method based on the lost tracking YOLOV8 model. Although the embodiments of the present invention have been shown and described, it can be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the attached claims and their equivalents.

Claims

1. A person image tracking method based on the lost tracking YOLOV8 model, It is characterized in that include: S1. Build a YOLOV8 model for missing person tracking; The S1 specifically includes: S11. Add the loss function L1 Loss to the last 10 epochs of the training set of the lost tracking YOLOV8 model; the training set includes 500 epochs; the first 490 epochs of the training set are first enhanced by the Mosaic module, and then the HSV data is enhanced; the last 10 epochs of the training set are closed for the Mosaic module enhancement, and the L1 Loss loss function is added after the HSV data enhancement is completed; S12. Replace the Taskaligned Assigner in the YOLOV8 model with simOTA; The specific implementation steps are as follows: S121. Pre-screen the anchor point; S122. Calculate the classification loss and position loss of each anchor point, and obtain the cost matrix and iou matrix; S123. Select a candidate box according to the iou matrix; S124. Dynamically assign the anchor point with the lowest cost value to the gt box; S125. Perform secondary screening on the gt box to which multiple anchor points are repeatedly assigned; S126. All the anchor points that are not assigned to the gt box are classified as negative samples and the loss of the filtered prediction box is calculated; S2. Collect the camera captured picture set, train the missing tracking YOLOV8 model to obtain the whole person picture set; S3. Compare the portrait of the missing person with the picture collection of all people to obtain the picture collection of the missing person; S4. Obtaining the movement trajectory; the movement trajectory of S4 includes: Obtaining trajectory information, and sorting the missing person photo collection in chronological order; Draw a trajectory, locate the camera position corresponding to the missing person picture set, and connect the camera positions according to the order; Predict missing paths, predicting untracked paths based on the trajectory; S5 push message to the first target and the second target; the S5 includes: the first target includes a person passing through the movement trajectory; The push message specifically includes: S51. Train the lost tracking YOLOV8 model through the latest camera image to obtain the second target; The latest camera image is the last image in the lost person image collection sorted in chronological order; S52. Obtaining personal information of the second target; S53. Classify and push messages to the first and second targets; The classification push includes a first feature and a second feature; The first feature includes mute and long-time display; The second feature includes a prompt sound and a long-time display.

2. The method according to claim 1, It is characterized in that The S11 includes: The L1 Loss loss function is the mean absolute error, which is used to measure the mean absolute difference MSE between the model prediction value and the true value, following the formula: ; is the predicted value of the i-th sample, is the true value of the i-th sample, and n is the total number of samples; The accuracy of the training results of the last 10 epochs is evaluated using L1 Loss.

3. The method according to claim 1, It is characterized in that The S11 Mosaic module includes: The Mosaic module is an enhancement strategy, which includes a Cutmix submodule; the Cutmix submodule randomly extracts a part of the area in the sample, and replaces and fills the area with the pixel values ​​of the unextracted area.

4. The method according to claim 1, It is characterized in that The S2 includes: The camera captured picture set includes camera pictures near the lost location; The entire character picture set includes all character portrait pictures extracted from the camera captured picture set.

5. The method according to claim 1, It is characterized in that The S3 includes: The portrait of the missing person includes a photo of the missing person in his / her daily life and / or a photo of the missing person’s ID card; The missing person picture collection includes photos of the missing person in the entire person picture collection.

6. The method according to claim 1, It is characterized in that The S3 comparison includes: S31. Using face clustering technology such as Chinese Whispers, the faces in the entire character picture set are grouped according to identity to obtain identity groups; S32. Using face recognition technology and a method based on geometric features, extracting a highly similar group from the identity group that contains one or more highly similar features to the portrait of the missing person; If the highly similar group is not included, then change the geographical area and repeat steps S2-S3; S33. Through face verification technology, the personal features of the iris and face are used to determine whether the highly similar group belongs to the missing person; The face verification technology includes comparing a preset threshold to determine whether the high similarity group and the portrait of the missing person are similar. combine.

7. The method according to claim 1, It is characterized in that The missing paths include: After the first camera, a route is drawn based on the third feature, taking the position of the first camera as a starting point; The first camera includes a camera that captures the latest camera image; The third feature includes the road structure and / or terrain direction around the first camera.

8. A person image tracking device based on the lost tracking YOLOV8 model, which adopts the method as claimed in claim 1, It is characterized in that include: Preparation module for building the missing person tracking YOLOV8 model for missing person finding; The S1 specifically includes: S11. Add the loss function L1 Loss to the last 10 epochs of the training set of the lost tracking YOLOV8 model; the training set includes 500 epochs; the first 490 epochs of the training set are first enhanced by the Mosaic module, and then the HSV data is enhanced; the last 10 epochs of the training set are closed for the Mosaic module enhancement, and the L1 Loss loss function is added after the HSV data enhancement is completed; S12. Replace the Taskaligned Assigner in the YOLOV8 model with simOTA; The specific implementation steps are as follows: S121. Pre-screen the anchor point; S122. Calculate the classification loss and position loss of each anchor point, and obtain the cost matrix and iou matrix; S123. Select a candidate box according to the iou matrix; S124. Dynamically assign the anchor point with the lowest cost value to the gt box; S125. Perform secondary screening on the gt box to which multiple anchor points are repeatedly assigned; S126. All the anchor points that are not assigned to the gt box are classified as negative samples and the loss of the filtered prediction box is calculated; An acquisition module is used to collect a set of pictures captured by a camera and train the missing tracking YOLOV8 model to obtain a set of pictures of all people; A positioning module is used to compare the portrait of the missing person with the entire person picture set to obtain the missing person picture set; A drawing module is used to obtain a moving trajectory; the moving trajectory of S4 includes: Obtaining trajectory information, and sorting the missing person photo collection in chronological order; Draw a trajectory, locate the camera position corresponding to the missing person picture set, and connect the camera positions according to the order; Predict missing paths, predicting untracked paths based on the trajectory; A message push module is used to push messages to a first target and a second target; the S5 includes: the first target includes a person passing through the moving track; The push message specifically includes: S51. Train the lost tracking YOLOV8 model through the latest camera image to obtain the second target; The latest camera image is the last image in the lost person image collection sorted in chronological order; S52. Obtaining personal information of the second target; S53. Classify and push messages to the first and second targets; The classification push includes a first feature and a second feature; The first feature includes mute and long-time display; The second feature includes a prompt sound and a long-time display.

Citation Information

Patent Citations

  • Pomegranate fruit detection method before fruit thinning based on improved YOLOv8s

    CN116958962A

  • Method for person re-identification based on deep model with multi-loss fusion training strategy

    US20200285896A1