Pedestrian motion recognition method, device and system, and storage medium
By using the YOLOv8 model to perform rectangular annotation and data enhancement on the monitoring video, training and application for real-time recognition, the problems of low accuracy and low efficiency of pedestrian motion recognition in the prior art are solved, and high-precision and high-efficiency pedestrian motion recognition are achieved.
Patent Information
- Application Number
- CN202510355071.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has problems such as slow speed, low accuracy and insufficient intelligence in pedestrian motion recognition, which leads to inaccurate recognition of recognition results and multiple errors reduce accuracy and efficiency.
The YOLOv8 model is used to identify pedestrian motions. By obtaining surveillance videos, rectangle annotation and data enhancement, the YOLOv8 model is trained, and the video is input into the model for identification in real-time scenarios.
It improves the accuracy and efficiency of pedestrian motion recognition, enhances the stability of the model, reduces the classification error rate, and has wide practicality.
Smart Images

Figure CN120198968A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a pedestrian motion recognition method, device, system, and storage medium. Background Art
[0002] With the rapid development of fields such as intelligent transportation, intelligent security, and human-computer interaction, higher requirements are put forward for the accurate recognition and positioning of pedestrian motion states. Pedestrian motion recognition not only requires being able to accurately detect pedestrians in videos or images, but also further determining their positions, motion trajectories, and behavior patterns. This technology has been widely applied in multiple scenarios such as public place monitoring, pedestrian navigation, autonomous driving assistance, and virtual reality interaction. Therefore, by means of advanced computer vision technology, automatically recognizing traffic modes through surveillance videos has become a current research and application hotspot.
[0003] In order to recognize different pedestrian motions, methods such as video analysis and machine learning can be adopted. Traditional video analysis methods mainly rely on image processing technology, and detect and recognize pedestrian motions by means of frame difference method, background subtraction method, and edge detection on surveillance videos. Machine learning-based methods recognize pedestrian motions by training classifiers (such as support vector machines, random forests, etc.). These methods rely on feature extraction and classifier training. However, using video analysis methods for recognition has problems such as slow speed, low accuracy, and low intelligence; the results of using traditional machine learning methods to recognize pedestrian motions are not accurate enough, and multiple misidentifications caused thereby will directly reduce accuracy and efficiency. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a pedestrian motion recognition method, device, system, and storage medium.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A pedestrian motion recognition method includes:
[0007] Step S1, obtaining surveillance scene videos collected in different scenarios;
[0008] Step S2, performing rectangular annotation on the video frame images of the surveillance scene videos to generate a VOC format pedestrian motion data set;
[0009] Step S3, performing data augmentation on the pedestrian motion data set;
[0010] Step S4, training a YOLOv8 model according to the pedestrian motion data set after data augmentation;
[0011] Step S5: Input the surveillance scene video in the real-time scene into the trained YOLOv8 model to realize the recognition of pedestrian movement in the real-time scene.
[0012] Preferably, the loss function L of the trained YOLOv8 model CIoU is:
[0013]
[0014] where ρ represents the distance in the Euclidean space, and (b, b gt ) represent the center points of B and B gt respectively, c is the diagonal length of the smallest closed box covering the two bounding boxes, B = (x, y, w, h) is the predicted regression box, and B gt = (x gt , y gt , w gt , h gt ) is the true regression box;
[0015] Preferably, in step S5, use the trained YOLOv8 model to judge whether the video frame image in the surveillance scene video in the real-time scene contains and which kinds of pedestrian movements it contains. If so, mark it with a bounding box; if not, do nothing.
[0016] The present invention also provides a pedestrian movement recognition device, including:
[0017] An acquisition module for acquiring surveillance scene videos collected in different scenes;
[0018] A labeling module for performing rectangular labeling on the video frame images of the surveillance scene video to generate a VOC-format pedestrian movement data set;
[0019] An enhancement module for performing data enhancement on the pedestrian movement data set;
[0020] A training module for training the YOLOv8 model according to the pedestrian movement data set after data enhancement;
[0021] A recognition module for inputting the surveillance scene video in the real-time scene into the trained YOLOv8 model to realize the recognition of pedestrian movement in the real-time scene.
[0022] Preferably, the loss function L of the trained YOLOv8 model CIoU is:
[0023]
[0024] where ρ represents the distance in the Euclidean space, and (b, bgt ) represent the center points of B and B gt respectively, c is the diagonal length of the smallest enclosing box covering the two bounding boxes, B=(x, y, w, h) is the predicted regression box, and B gt =(x gt , y gt , w gt , h gt ) is the true regression box;
[0025] Preferably, the recognition module uses the trained YOLOv8 model to determine whether the video frame image of the monitored scene video in the real-time scene contains and which kinds of pedestrian movements are included. If so, it marks them with bounding boxes, and if not, it does not perform any processing.
[0026] A computer program is stored on the memory and run by the processor. When the computer program is run by the processor, it executes the pedestrian movement recognition method.
[0027] An embodiment of the present invention also provides a storage medium, on which a computer program is stored. When the computer program runs, it executes the pedestrian movement recognition method.
[0028] By introducing the YOLOv8 model, the present invention can effectively detect and identify various pedestrian movements of users, with a simple mechanism and high detection accuracy, and can be used in various intelligent terminals. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0030] Figure 1 It is a flowchart of the pedestrian movement recognition method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0032] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] Example 1:
[0034] As Figure 1 shown, an embodiment of a pedestrian motion recognition method of the present invention includes:
[0035] Step S1, obtaining surveillance scene videos collected under different scenarios;
[0036] Step S2, performing rectangular annotation on the video frame images of the surveillance scene videos to generate a pedestrian motion dataset in VOC format;
[0037] Step S3, performing data augmentation on the pedestrian motion dataset;
[0038] Step S4, training a YOLOv8 model according to the data-augmented pedestrian motion dataset;
[0039] Step S5, inputting the surveillance scene video in the real-time scene into the trained YOLOv8 model to realize the recognition of pedestrian motion in the real-time scene.
[0040] As an implementation manner of the embodiment of the present invention, in step S2, the LabelImg image annotation tool is used to perform rectangular annotation on the video frame images of the obtained surveillance scene videos to accurately define the area related to pedestrian motion, and a pedestrian motion dataset in VOC format is generated in a specified folder.
[0041] As an implementation manner of the embodiment of the present invention, in step S3, brightness adjustment, scaling, flipping, and Mosaic are adopted to expand the dataset by transforming the images to improve the adaptability to different scenarios and changes, and an enhanced pedestrian motion dataset is obtained.
[0042] As an implementation manner of the embodiment of the present invention, the loss function L of the trained YOLOv8 model CIoU is:
[0043]
[0044] where ρ represents the distance in Euclidean space, and (b, b gt ) represent the center points of B and B gt respectively. B = (x, y, w, h) is the predicted regression box, and B gt = (x gt , y gt , w gt , h gt) is the real regression box; x represents the ratio of the central abscissa to the image width; y represents the ratio of the central ordinate to the image height; w represents the ratio of the bbox width to the image width; h represents the ratio of the bbox height to the image height; bbox represents the predicted box; B gt The parameters in it are similar;
[0045] Among them, for B and B gt Find the center points of the two boxes, and the coordinates of the corresponding center points are the ordered pair (b, b gt );For B and B gt Find the minimum enclosing bounding box for the two boxes. Denote c as the diagonal length of the minimum enclosing box covering the two box frames; α is a parameter related to the angle, v is a parameter related to the angle;
[0046] As an implementation manner of the embodiment of the present invention, in step S5, use the trained YOLOv8 model to determine whether the video frame image of the monitored scene video in the real-time scene contains and which types of pedestrian movements are included. If so, mark it with a bounding box; if not, do nothing.
[0047] Further, the specific method of using the bounding box for marking is as follows:
[0048] Use the five values output by the YOLOv8 model: (class, x, y, w, h), where class represents the class name. For example: 0 represents escalator, 1 represents lift, 2 represents cycling; x represents the ratio of the central abscissa to the image width; y represents the ratio of the central ordinate to the image height; w represents the ratio of the bbox width to the image width; h represents the ratio of the bbox height to the image height; bbox represents the predicted box.
[0049] Assume that the state classes class = {1, 2, 3, 4}, corresponding to {walking, cycling, taking a vehicle, running} respectively. When using the trained YOLOv8 for recognition, if the probabilities of recognizing class as 1, 2, 3, 4 are 0.02, 0.01, 0.92, 0.03 respectively, and the remaining 0.02 is the probability of not being all of the above, then the recognition result will be taking a vehicle, and the recognition result and the corresponding probability will be marked on the display screen. For video types, each frame will be repeated and alternated according to the above process until the video ends.
[0050] The technical effects of the present invention are as follows:
[0051] 1. Improve data utilization rate: Based on the pedestrian movement recognition of the YOLOv8 model, it can effectively utilize the collected image frame data, thus greatly improving the data rate.
[0052] 2. Enhance model stability: Compared with conventional methods, the present invention uses image enhancement to train the initial dataset multiple times, thereby enhancing the stability of the model.
[0053] 3. Reduce the classification error rate: Compared with conventional methods, this method uses the loss function L CIoU designed a fault tolerance mechanism for the classification model, and maximized the correction of misclassifications by setting the confidence level.
[0054] 4. Strong practicality: It can be widely applied to personnel positioning, pattern recognition, traffic control, etc., so it has the characteristic of strong practicality.
[0055] Example 2:
[0056] The embodiment of the present invention also provides a pedestrian motion recognition device, including:
[0057] An acquisition module, configured to acquire surveillance scene videos collected under different scenarios;
[0058] A labeling module, configured to perform rectangular labeling on the video frame images of the surveillance scene video to generate a pedestrian motion dataset in VOC format;
[0059] An enhancement module, configured to perform data enhancement on the pedestrian motion dataset;
[0060] A training module, configured to train the YOLOv8 model according to the pedestrian motion dataset after data enhancement;
[0061] A recognition module, configured to input the surveillance scene video in the real-time scene into the trained YOLOv8 model to realize the recognition of pedestrian motion in the real-time scene.
[0062] As an implementation manner of the embodiment of the present invention, the loss function L of the trained YOLOv8 model CIoU is:
[0063]
[0064] where ρ represents the distance in the Euclidean space, (b, b gt ) respectively represent the center points of B and B gt , c is the diagonal length of the smallest enclosing box covering the two bounding boxes, B = (x, y, w, h) is the predicted regression box, and B gt = (x gt , y gt , w gt , h gt ) is the true regression box;
[0065] As an implementation manner of an embodiment of the present invention, the recognition module uses the trained YOLOv8 model to determine whether the video frame image of the monitored scene video in the real-time scene contains and which types of pedestrian movements are included. If so, it is marked with a bounding box; if not, no processing is performed.
[0066] Embodiment 3:
[0067] The embodiment of the present invention further provides a pedestrian movement recognition system, including: a memory and a processor. A computer program is stored on the memory and run by the processor. When the computer program is run by the processor, it executes the pedestrian movement recognition method.
[0068] Embodiment 4:
[0069] The embodiment of the present invention further provides a storage medium, on which a computer program is stored. When the computer program runs, it executes the pedestrian movement recognition method.
[0070] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solution of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A pedestrian motion recognition method, characterized in that: include: Step S1, obtaining surveillance scene videos collected in different scenes; Step S2: perform rectangular annotation on the video frame images of the monitoring scene video to generate a pedestrian motion dataset in VOC format; Step S3, performing data enhancement on the pedestrian motion dataset; Step S4, training a YOLOv8 model based on the pedestrian motion dataset after data enhancement; Step S5: input the monitoring scene video in the real-time scene into the trained YOLOv8 model to realize the recognition of pedestrian movement in the real-time scene.
2. The pedestrian motion recognition as claimed in claim 1, characterized in that: The loss function L of the trained YOLOv8 model CIoU for: Where ρ represents the distance in Euclidean space, (b,b gt ) represent B and B respectively gt The center point of the two boxes, c is the diagonal length of the smallest enclosing box covering the two boxes, B = (x, y, w, h) is the predicted regression box, B gt =(x gt ,y gt ,w gt ,h gt ) is the true regression box; 3. The pedestrian motion recognition as claimed in claim 1, characterized in that: In step S5, the trained YOLOv8 model is used to determine whether and which types of pedestrian motion are contained in the video frame image of the monitoring scene video in the real-time scene. If so, it is marked with a bounding box, otherwise no processing is performed.
4. A pedestrian motion recognition device, characterized in that: include: An acquisition module is used to acquire surveillance scene videos collected in different scenarios; The annotation module is used to perform rectangular annotation on the video frame images of the surveillance scene video and generate a pedestrian motion dataset in VOC format; Enhancement module, used to perform data enhancement on pedestrian motion dataset; The training module is used to train the YOLOv8 model based on the data-enhanced pedestrian motion dataset; The recognition module is used to input the surveillance scene video in the real-time scene into the trained YOLOv8 model to realize the recognition of pedestrian movement in the real-time scene.
5. The pedestrian motion recognition device according to claim 4, characterized in that: The loss function L of the trained YOLOv8 model CIoU for: Where ρ represents the distance in Euclidean space, (b,b gt ) represent B and B respectively gt The center point of the two boxes, c is the diagonal length of the smallest enclosing box covering the two boxes, B = (x, y, w, h) is the predicted regression box, B gt =(x gt ,y gt ,w gt ,h gt ) is the true regression box; 6. The pedestrian motion recognition device according to claim 5, characterized in that: The recognition module uses the trained YOLOv8 model to determine whether and which types of pedestrian motion are contained in the video frame image of the monitoring scene video in the real-time scene. If so, it is marked with a bounding box, otherwise no processing is performed.
7. A pedestrian motion recognition system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the pedestrian motion recognition method according to any one of claims 1 to 3 is executed.
8. A storage medium, characterized in that: The storage medium stores a computer program, which executes the pedestrian motion recognition method according to any one of claims 1 to 3 when running.