Scooter riding detection method and system for multiple persons and electronic equipment

By using pre-trained YOLOv8 model and dynamic interchange ratio calculation, the scooter and human body candidate boxes are detected and multi-person riding phenomenon is judged, and the problems of inefficient detection efficiency and insufficient accuracy in the prior art are solved, and accurate identification and rapid response in complex environments are achieved.

CN120047864APending Publication Date: 2025-05-27TERMINUS GENERAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411863800.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is inefficient when detecting multiple people riding scooters, it is difficult to cover a wide range of areas, and is easily disturbed by human factors. The accuracy and real-time nature of traditional video surveillance systems are difficult to meet actual needs.

Method used

The pre-trained YOLOv8 model is used to combine dynamic interchange comparison calculations. By continuously obtaining video stream data, detecting scooter and human body candidate boxes, calculating the intersection and union area of ​​candidate boxes, combining the speed of scooter and human body, calculating dynamic interchange comparisons, determining whether there is a multi-person riding phenomenon, and sending the relevant data to the traffic management department and shared scooter service provider.

Benefits of technology

Accurately identify multi-person riding behavior in complex environments, improve traffic safety and the operation efficiency of shared scooters, reduce misidentification and missed inspections, and enhance the accuracy and real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047864A_ABST
    Figure CN120047864A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a detection method and system for multi-person riding of a scooter and electronic equipment. The method is applied to the technical field of traffic safety, and comprises the following steps: determining a scooter candidate frame and a scooter ID of a scooter in each piece of video frame data through a pre-trained model, and obtaining a human body candidate frame and a human body ID in each piece of video frame data; determining the scooter speed of each scooter target and the human body speed of each human body target; calculating a dynamic intersection-union ratio of a candidate frame pair formed by each scooter candidate frame and each human body candidate frame in each video frame data; and determining a target video frame data set according to a preset time window and the video frame data, and if the dynamic intersection-to-union ratio of at least two human body candidate frames to the same scooter candidate frame in each video frame data of the target video frame data set is greater than a preset threshold, determining that a multi-person riding phenomenon exists. According to the scheme, multi-person riding behaviors can be accurately identified in a complex environment, quick response is achieved, and the traffic safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of traffic safety, and particularly to a method, a system and an electronic device for detecting multiple people riding a scooter. Background Art

[0002] With the popularization of shared scooter services, the phenomenon of multiple people sharing a single scooter has been increasing. This not only violates traffic rules, increases the risk of traffic accidents, but also affects the urban traffic order. Existing detection means mainly rely on manual inspections or simple video monitoring. Manual detection includes on-site inspections and video monitoring assistance. Inspectors need to regularly or irregularly inspect key areas, observe and record violations; at the same time, video monitoring devices can monitor the usage of scooters in real time and provide evidence for violations.

[0003] However, manual inspections are inefficient, difficult to cover a wide area, and are easily interfered by human factors, such as omissions by inspectors or biases in subjective judgments. Secondly, although traditional video monitoring systems can provide continuous image data, in the detection of multiple people riding a scooter, due to environmental complexity, lighting changes, occlusion problems, and the difficulty of target tracking in dynamic backgrounds, it is difficult to meet the actual requirements in terms of accuracy and real-time performance. Summary of the Invention

[0004] To solve the deficiencies of the prior art, the present disclosure provides a method, a system and an electronic device for detecting multiple people riding a scooter. The present disclosure solves the technical problems such as the low efficiency of current manual inspections, the difficulty in covering a wide area, and being easily interfered by human factors, as well as the difficulty of traditional video monitoring systems to meet the actual requirements in terms of accuracy and real-time performance.

[0005] According to a first aspect of the present disclosure, there is provided a method for detecting multiple people riding a scooter, including: continuously acquiring video stream data, continuously determining video frame data according to the video stream data, inputting the video frame data into a pre-trained YOLOv8 model to obtain scooter candidate boxes and scooter IDs of scooters in each video frame data, and obtaining human candidate boxes and human IDs of humans in each video frame data; wherein, the number of both the scooter candidate boxes and the human candidate boxes is at least one;

[0006] Determining the scooter position data of the scooter target corresponding to each scooter ID in each video frame data, and determining the human position data of the human target corresponding to each human ID in each video frame data, determining the scooter speed of each scooter target according to the scooter position data, and determining the human speed of each human target according to the human position data;

[0007] Determine the intersection area and union area between each scooter candidate box and each human candidate box in each video frame data. According to the intersection area, the union area, the scooter speed, the human speed, and a preset dynamic intersection-over-union calculation formula, calculate the dynamic intersection-over-union of each candidate box pair composed of each scooter candidate box and each human candidate box in each video frame data;

[0008] Determine a target video frame data set according to a preset time window and the video frame data. If, in each video frame data of the target video frame data set, there are at least two human candidate boxes whose dynamic intersection-over-union with the same scooter candidate box is greater than a preset threshold, it is determined that there is a phenomenon of multiple people riding, and the target video frame data set is sent to the traffic management department and the shared scooter service provider.

[0009] According to a second aspect of the present disclosure, there is provided a multiple-person scooter riding detection system for executing the method as described in the first aspect, including: an identification module for continuously acquiring video stream data, continuously determining video frame data according to the video stream data, and inputting the video frame data into a pre-trained YOLOv8 model to obtain the scooter candidate boxes and scooter IDs of the scooters in each video frame data, and, obtaining the human candidate boxes and human IDs in each video frame data; wherein, the number of scooter candidate boxes and human candidate boxes is at least one;

[0010] A speed confirmation module for determining the scooter position data of the scooter target corresponding to each scooter ID in each video frame data, and, determining the human position data of the human target corresponding to each human ID in each video frame data, determining the scooter speed of each scooter target according to the scooter position data, and, determining the human speed of each human target according to the human position data;

[0011] A dynamic intersection-over-union calculation module for determining the intersection area and union area between each scooter candidate box and each human candidate box in each video frame data. According to the intersection area, the union area, the scooter speed, the human speed, and a preset dynamic intersection-over-union calculation formula, calculate the dynamic intersection-over-union of each candidate box pair composed of each scooter candidate box and each human candidate box in each video frame data;

[0012] A judgment module for determining a target video frame data set according to a preset time window and the video frame data. If, in each video frame data of the target video frame data set, there are at least two human candidate boxes whose dynamic intersection-over-union with the same scooter candidate box is greater than a preset threshold, it is determined that there is a phenomenon of multiple people riding, and the target video frame data set is sent to the traffic management department and the shared scooter service provider.

[0013] According to a third aspect of the present disclosure, there is provided an electronic device, which includes: a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the methods described above are implemented.

[0014] In a method, system, and device for detecting multi-person scooter riding provided as above, embodiments of the present disclosure can accurately identify multi-person riding behaviors in complex environments and respond quickly by combining advanced object detection and tracking technologies, dynamic intersection over union calculation, and an intelligent early warning mechanism, helping to improve traffic safety and the operation efficiency of shared scooters. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 Shows a schematic flowchart of a method for detecting multi-person scooter riding according to an embodiment of the present disclosure;

[0017] Figure 2 Shows a schematic flowchart of a method for detecting multi-person scooter riding according to an embodiment of the present disclosure;

[0018] Figure 3 Shows a schematic block diagram of a system for detecting multi-person scooter riding according to an embodiment of the present disclosure;

[0019] Figure 4 Shows a block diagram of an exemplary electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] Now, various exemplary embodiments of the present disclosure will be described in detail with reference to the drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0021] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them. It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more. It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, without clear limitation or contrary indication in the context, it can generally be understood as one or more. In addition, the term "and / or" in the present disclosure is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after. It should also be understood that the present disclosure emphasizes the differences between the various embodiments, and their similarities can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0022] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present disclosure and its application or use. Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the specification. It should be noted that similar reference numerals and letters denote similar items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0023] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0024] Figure 1 It is a schematic flow chart of a method for detecting a multi-person scooter provided for the embodiments of the present disclosure.

[0025] As Figure 1 shown, the method includes:

[0026] S101. Continuously obtain video stream data, continuously determine video frame data based on the video stream data, input the video frame data into a pre-trained YOLOv8 model to obtain scooter candidate boxes and scooter IDs for each video frame data, and obtain human candidate boxes and human IDs for each video frame data; where the number of scooter candidate boxes and human candidate boxes is at least one.

[0027] An edge computing device can refer to a computing device that processes, stores, and analyzes data at the edge of the network (i.e., closer to the data source or terminal device). Different from traditional cloud computing centers, edge computing devices usually directly process data where the data is generated to reduce the latency and bandwidth consumption of data transmission to the cloud and improve the response speed of the system.

[0028] Video stream data can be continuous video information collected in real time through devices such as cameras, consisting of a series of image frames arranged in chronological order.

[0029] Video frame data can be a single static image extracted from a video stream. Each frame of data can be regarded as a snapshot of the video stream at a certain point in time.

[0030] YOLOv8 can be an advanced real-time object detection model that can quickly locate and classify objects in images. The pre-trained YOLOv8 model means that the model has been trained on a dataset (such as the COCO dataset) and has learned a large number of object detection features.

[0031] A scooter candidate box can be a rectangular box output by the YOLOv8 model, used to represent the area of the scooter target. It contains information such as position (x, y coordinates, width, and height) and classification confidence.

[0032] A scooter ID can be a unique identifier assigned to each scooter target, facilitating cross-frame tracking of scooters.

[0033] A human candidate box can be a rectangular box output by the YOLOv8 model, used to represent the area of the human target.

[0034] A human ID can be a unique identifier assigned to each human target, used to identify and track humans.

[0035] High-definition cameras can be deployed to capture video streams in real time. Connect to the camera through streaming media protocols (such as RTSP or HTTP) to obtain video data. Then extract continuous video frame data from the video stream, usually at a fixed frame rate, and convert the video frames into a format that the model can process. Then adjust the size of the video frames (such as adjusting to the input size required by the YOLOv8 model). Perform normalization to scale the pixel values to the range [0, 1]. Input the preprocessed frame data into the YOLOv8 model, and the model outputs scooter candidate boxes and human candidate boxes. YOLOv8 supports object tracking through the track command, which essentially combines the YOLOv8 detection results with external object tracking algorithms, such as: SORT (Simple Online and Realtime Tracking), DeepSORT (Deep Learning-based SORT), and ByteTrack (a high-performance multi-object tracking algorithm). These algorithms run independently of the YOLOv8 model, but through the integration of toolkits, they can be seamlessly applied to the detection results to obtain scooter IDs and human IDs. Taking DeepSORT as an example, the steps are as follows:

[0036] (1) Object matching: The object tracking algorithm matches the detection results in the new frame with the tracking results in the previous frame: Then calculate the association degree: Use the Hungarian algorithm or the greedy matching algorithm, based on the geometric features and motion characteristics between objects. Common matching methods include IOU matching: Compare the overlap degree between the current detection box and the tracking box in the previous frame. Feature matching: Extract the target appearance feature vectors through a pre-trained ReID (Re-identification) network, and calculate the cosine similarity or Euclidean distance.

[0037] (2) Assign a unique ID

[0038] If the detection box matches an existing tracking box: Assign the same object ID, indicating it is the same object.

[0039] If the detection box does not match any tracking boxes: Generate a new ID, indicating that this is a newly emerged object.

[0040] If an existing tracking box is not matched in the current frame: It may be considered that the object has left the field of view, and its ID is temporarily suspended.

[0041] (3) Trajectory prediction and update

[0042] Use methods such as the Kalman Filter to predict the position of the object in the current frame.

[0043] Update the motion trajectory of the object to maintain the continuity of the object ID.

[0044] The specific logic of ID generation is as follows: Initial frame: Detect all targets and generate a unique ID for each target (such as ID_1, ID_2).

[0045] Subsequent frames: Compare the detection results in the current frame with the tracking results of the previous frame.

[0046] If the match is successful, retain the original ID.

[0047] If there is no match, assign a new ID.

[0048] ID disappearance: If a target does not appear for a long time (exceeding the preset time window), its ID is recycled.

[0049] If a target is partially occluded or lost for a short time, the ReID feature can help restore the ID. When one target becomes multiple targets, the tracking algorithm generates new IDs for each new target. When multiple targets merge into one, one of the IDs may be retained.

[0050] S102, determine the scooter position data of the scooter targets corresponding to each scooter ID in each video frame data, and determine the human position data of the human targets corresponding to each human ID in each video frame data. Determine the scooter speed of each scooter target according to the scooter position data, and determine the human speed of each human target according to the human position data.

[0051] The scooter target can refer to the scooter entity detected in the video. It is a specific scooter instance represented by the candidate box recognized by the YOLOv8 detection model and assigned a unique ID through a multi-object tracking algorithm (such as DeepSORT).

[0052] The scooter position data can be the spatial position of the scooter target in each video frame, usually the coordinate data of the candidate box.

[0053] The human target can be the human entity detected in the video. Similar to the scooter target, it is recognized and assigned a unique ID through a detection model and a tracking algorithm.

[0054] The human position data can be the spatial position of the human target in each video frame, with the same format as the scooter position data, which is the coordinate data of the candidate box.

[0055] The scooter speed can be the moving speed of the scooter target in the video stream, calculated based on the change of the scooter position data over time, and the unit is usually pixels / second.

[0056] The human speed can represent the moving speed of the human target in the video stream.

[0057] The candidate box coordinates corresponding to the scooter target in each frame can be extracted, such as [x_min, y_min, x_max, y_max], representing the upper left and lower right coordinates of the scooter target. The center point of the scooter target is calculated by the following formula, and this center point is used to calculate the speed:

[0058]

[0059] Then, extract the candidate box coordinates corresponding to the human target in each frame (with the same position data format as the scooter), and the calculation method of the center point is the same.

[0060] The scooter speed is calculated based on the change of the scooter position data in consecutive frames:

[0061]

[0062] where (x c1 , y c1 ) is the center point coordinate of the scooter target in the current video frame; (x c2 , y c2 ) is the center point coordinate of the scooter target in the next video frame; Δt is the time interval between the current video frame and the next video frame. The calculation method of the human speed is the same as that of the scooter speed, based on the change of the center point of the human target.

[0063] S103. Determine the intersection area and union area of each scooter candidate box and each human candidate box in the video frame data. According to the intersection area, the union area, the scooter speed, the human speed, and a preset dynamic intersection-over-union calculation formula, calculate the dynamic intersection-over-union of the candidate box pairs composed of each scooter candidate box and each human candidate box in the video frame data.

[0064] The intersection area can be the overlapping area of two candidate boxes.

[0065] The union area can be the total coverage area of two candidate boxes, including the overlapping part and the non-overlapping part.

[0066] A candidate box pair can be a combination composed of one scooter candidate box and one human candidate box.

[0067] The dynamic intersection-over-union can be a measure of the dynamic spatial association between the scooter and the human by combining the intersection area, the union area, the scooter speed, and the human speed.

[0068] The preset dynamic intersection-over-union calculation formula can be an improved measure for describing the spatio-temporal correlation of two candidate boxes (scooter and human candidate boxes) in the video stream.

[0069] The left boundary value on the right can be found in the scooter frame and the human body frame. The upper boundary value on the lower side can be found in the scooter frame and the human body frame. The right boundary value on the left can be found in the scooter frame and the human body frame. The lower boundary value on the higher side can be found in the scooter frame and the human body frame. Using the intersection rectangle boundary determined in the first step, calculate the width and height of the intersection rectangle. If the width or height is negative, the intersection area is zero. Otherwise, multiply the width and height to obtain the area of the intersection. Then, using the intersection rectangle boundary determined in the first step, calculate the width and height of the intersection rectangle. If the width or height is negative, the intersection area is zero. Otherwise, multiply the width and height to obtain the area of the intersection. Combine the intersection area, the union area, the scooter speed, the human body speed with a preset dynamic intersection-over-union calculation formula to obtain the dynamic intersection-over-union of the candidate box pairs formed by each scooter candidate box and each human candidate box in each video frame data.

[0070] Based on the above technical solution, optionally, according to the intersection area, the union area, the scooter speed, the human body speed, and a preset dynamic intersection-over-union calculation formula, calculating the dynamic intersection-over-union of the candidate box pairs formed by each scooter candidate box and each human candidate box in each video frame data includes:

[0071] Calculate the speed difference between the scooter speed and the human body speed, and calculate the speed influence factor coefficient according to the speed difference and a preset speed influence factor coefficient calculation formula;

[0072] According to the intersection area, the union area, the speed influence factor coefficient, and a preset dynamic intersection-over-union calculation formula, calculate the dynamic intersection-over-union of the candidate box pairs formed by each scooter candidate box and each human candidate box in each video frame data; where the preset dynamic intersection-over-union calculation formula is:

[0073]

[0074] where, A Union is the union area; α is a preset adjustment factor for adjusting the influence degree of speed on the intersection-over-union, V Impact is the speed influence factor coefficient for measuring the influence of the speed difference between targets on the dynamic intersection-over-union; A Interestion is the intersection area; Dynamic Iou is the dynamic intersection-over-union of the candidate box pairs formed by each scooter candidate box and each human candidate box.

[0075] In this solution, the speed difference can be the speed difference between the scooter target and the human body target. The speed difference is obtained by calculating the difference between their respective speed values.

[0076] The preset calculation formula for the speed influence factor coefficient can be to adjust the dynamic intersection-over-union ratio based on the speed difference. The speed difference affects the interaction behavior between targets. Especially in the detection of multiple people riding scooters, the speed difference indicates whether they are on the same trajectory or whether one rider is following another.

[0077] The speed influence factor coefficient can be an adjustment factor used to adjust the calculation of the dynamic intersection-over-union ratio to reflect the impact of the speed difference between the scooter and human targets on their temporal continuity. The larger its value, the greater the impact of the speed difference on the intersection-over-union ratio; the smaller the value, the smaller the impact of the speed difference.

[0078] The speeds of the scooter and the human body can be calculated based on the continuous frame data in the video stream. The speed is determined by the displacement amount of the target between two frames and the time interval, and the speed difference is the absolute difference between these two speeds. Then, using the preset calculation formula for the speed influence factor coefficient, the speed influence factor coefficient is obtained based on the speed difference calculated in the previous step. Finally, the intersection area, union area, scooter speed, and human body speed are substituted into the preset dynamic intersection-over-union ratio calculation formula to calculate the dynamic intersection-over-union ratio of the candidate box pairs formed by each scooter candidate box and each human candidate box in each video frame data.

[0079] In this solution, by introducing the speed difference to calculate the speed influence factor, it is possible to accurately distinguish those targets that are actually related to each other and those that are not related. This can reduce misidentifications and missed detections and improve accuracy.

[0080] Based on the above technical solution, optionally, the preset calculation formula for the speed influence factor coefficient is:

[0081] V Impact = 1 + v max + γ·Δv;

[0082] where v max is the maximum speed of the two targets included in each candidate box pair; γ is the preset speed difference weight coefficient; and Δv is the speed difference.

[0083] In this solution, v max is the maximum speed of the two targets included in each candidate box pair, that is, the maximum speed of the targets included in the scooter candidate box and the human candidate box in each candidate box pair.

[0084] S104. Determine the target video frame data set according to the preset time window and the video frame data. If in the video frame data of the target video frame data set, for each video frame data, there is at least one dynamic intersection-over-union ratio of two human candidate boxes and the same scooter candidate box greater than the preset threshold, it is determined that there is a phenomenon of multiple people riding, and each video frame data is sent to the traffic management department and the shared scooter service provider.

[0085] The preset time window can refer to a time period during which the system continuously tracks and evaluates video data to check for the presence of multiple riders. The length of this time window can be set according to actual needs, such as ranging from a few seconds to several minutes. The video frames within the time window will be collected and analyzed to determine if abnormal behavior occurs.

[0086] The target video frame dataset can refer to all the video frame data extracted from the video stream within the set time window. These video frame data contain the detection results of scooters and humans (such as candidate boxes and IDs), and will be analyzed in subsequent dynamic intersection over union calculations.

[0087] The preset threshold can refer to the critical value of the dynamic intersection over union ratio used to determine whether there is a multiple-rider phenomenon between a scooter and a human. Only when the dynamic intersection over union ratio of two or more human candidate boxes and the same scooter candidate box in each video frame data of the target video frame dataset exceeds this threshold will the system consider that there is a multiple-rider phenomenon.

[0088] The traffic management department can be the official agency responsible for urban traffic management, and they are responsible for supervising traffic behaviors.

[0089] The shared scooter service provider can be a company or platform that provides shared scooter services, and is usually responsible for managing the use, maintenance, and scheduling of scooters.

[0090] Video frame data can be filtered according to the set preset time window (such as 5 seconds) and stored in the target video frame dataset. For each frame of data, check if there are at least two candidate boxes with a dynamic intersection over union ratio greater than the preset threshold with the same scooter candidate box (that is, the intersection over union ratio between at least two human candidate boxes and the same scooter candidate box exceeds the threshold). If the condition is met, it is determined that there is a multiple-rider phenomenon. Once a multiple-rider phenomenon is detected, the target video frame dataset is sent to the traffic management department and the shared scooter service provider. These data can be sent through network interfaces (such as APIs), emails, text messages, etc. for relevant departments to make subsequent processing.

[0091] In the embodiments of the present application, video stream data is continuously acquired, video frame data is continuously determined according to the video stream data, and the video frame data is input into a pre-trained YOLOv8 model to obtain scooter candidate boxes and scooter IDs of scooters in each video frame data, and human candidate boxes and human IDs in each video frame data; wherein, the number of scooter candidate boxes and human candidate boxes is at least one; determining scooter position data of scooter targets corresponding to each scooter ID in each video frame data, and determining human position data of human targets corresponding to each human ID in each video frame data, determining the scooter speed of each scooter target according to the scooter position data, and determining the human speed of each human target according to the human position data; determining the intersection area and union area between each scooter candidate box and each human candidate box in each video frame data, and calculating the dynamic intersection over union of candidate box pairs formed by each scooter candidate box and each human candidate box in each video frame data according to the intersection area, the union area, the scooter speed, the human speed, and a preset dynamic intersection over union calculation formula; determining a target video frame data set according to a preset time window and the video frame data, if in each video frame data of the target video frame data set, there are at least two human candidate boxes whose dynamic intersection over union with the same scooter candidate box is greater than a preset threshold, it is determined that there is a phenomenon of multiple people riding, and the target video frame data set is sent to the traffic management department and the shared scooter service provider. Through the above method for detecting multiple people riding a scooter, by combining advanced object detection and tracking technologies, dynamic intersection over union calculation, and an intelligent early warning mechanism, it is possible to accurately identify the behavior of multiple people riding in a complex environment and respond quickly, helping to improve traffic safety and the operation efficiency of shared scooters.

[0092] Based on the above technical solution, optionally, the training steps of the pre-trained YOLOv8 model include:

[0093] Obtain historical video stream data, and determine historical video frame data, scooter target box labels of each annotated historical video frame data, scooter ID labels of each scooter target box label, human target box labels of each annotated historical video frame data, and human ID labels of each human target box label according to the historical video stream data;

[0094] Train the YOLOv8 model according to the historical video frame data, scooter target box labels of each annotated historical video frame data, scooter ID labels of each scooter target box label, human target box labels of each annotated historical video frame data, and human ID labels of each human target box label until the YOLOv8 model reaches a preset YOLOv8 model training standard.

[0095] In this solution, the historical video stream data can refer to video files recorded in the scooter usage scenario, including a sequence of multiple frames of images in a dynamic environment. The source can be a surveillance camera, a dash cam, etc.

[0096] The historical video frame data can be static image frames extracted from the historical video stream for subsequent annotation and model training.

[0097] The scooter target box label can be used to mark the position of the scooter in the video frame with a bounding box, and a unique scooter ID label is assigned to each scooter target box to distinguish different scooters.

[0098] The human target box label can be used to mark the position of the human in the video frame in the same format as the scooter label. A unique human ID label is assigned to each human target box to distinguish different humans.

[0099] A suitable variant of YOLOv8 can be selected according to the application scenario: YOLOv8-nano / tiny: High real-time requirements and limited hardware resources. YOLOv8-large / x: High accuracy requirements. Then use the above-mentioned annotated historical video frame data and target box labels. The data is split into a training set (e.g., 80%) and a validation set (e.g., 20%). Then use the annotated training data to train the YOLOv8 model through supervised learning. During the training process, the model will learn how to automatically identify scooter and human targets from the input images and assign corresponding IDs to each target. During the training process, the performance of the model is evaluated using the validation set, calculating the mAP and loss values after each round of training until the preset training criteria are met. Use a lightweight classification head and spatial-channel decoupled downsampling technology to optimize the YOLOv8 model. The lightweight classification head can reduce the computational amount and is suitable for scenarios with high real-time requirements; spatial-channel decoupled downsampling can improve the computational efficiency of the model. Techniques such as model pruning and quantization can reduce the number of model parameters and computational complexity, improve the inference speed, and are suitable for deployment on edge computing devices. Then select a suitable YOLOv8 model for deployment according to the application scenario. For devices with limited computing resources (such as edge computing devices or embedded devices), a lightweight model of YOLOv8 (such as YOLOv8-nano or YOLOv8-tiny) can be selected.

[0100] For high-performance computing platforms (such as GPU servers), a larger YOLOv8 model (such as YOLOv8-large) can be selected to ensure detection accuracy. Use a suitable inference framework (such as TensorRT, OpenVINO, etc.) to accelerate the model inference process. Before deployment, evaluate the accuracy and real-time performance of the model by using a new test dataset (i.e., data not used in training). Evaluation metrics can include: accuracy (AP), recall, F1 score, inference time, etc.

[0101] Based on the above technical solution, optionally, after sending the target video frame dataset to the traffic management department and the shared scooter service provider, the method further includes:

[0102] If recognition error information is received from the traffic management department and the shared scooter service provider, determine an adjusted dataset according to the recognition error information, and retrain the preset YOLOv8 model according to the adjusted dataset until the preset YOLOv8 model meets the preset YOLOv8 model training standard; wherein, the adjusted dataset includes adjusted video frames, scooter target box labels for each adjusted video frame data that has been annotated, scooter ID labels for each scooter target box label, and human target box labels for each adjusted video frame data that has been annotated.

[0103] In this solution, the adjusted dataset can be a correction and supplement to the original dataset based on feedback information in order to improve the accuracy and performance of the model after receiving the recognition error information.

[0104] The adjusted video frames can be those selected from the original video stream that have recognition errors or need improvement, and data correction or supplementation is performed. These frames may be misjudged frames or frames after enhancement processing.

[0105] The scooter target box labels can be the scooter targets in the adjusted video frames, ensuring that the bounding box positions (upper left and lower right coordinates) of each target are accurate and that the class labels are correct.

[0106] The scooter ID labels can be correct ID labels assigned to the scooter targets. By ensuring that the same scooter target has a consistent ID in different frames, the model can be prevented from repeatedly recognizing or mis-tracking the same target.

[0107] The human target box labels can be similar to the scooter target box labels, referring to the bounding box annotations of human targets in the adjusted video frames. Ensure that the bounding box positions and sizes of human targets are accurate and that the class labels are correct.

[0108] Identification error information can be received from the traffic management department and shared scooter service providers. The error information usually includes the error type, the video frame or area of the specific error, the error in the target box annotation, etc. Analyze the recognition problems of the model based on the error information. For example, check the incorrect labels of the scooter and human target boxes, identify incorrect ID assignments, etc. Identify the specific video frame data where misrecognition or correction may be required. Then select the video frames that need to be corrected from the original video data, or supplement new video frames for problematic scenarios to ensure data diversity and coverage. And modify the target box labels: Scooter target box label: Correct the bounding box and class label of the scooter to ensure its accurate annotation in all video frames. Human target box label: Adjust the annotation of the human target box to ensure the accuracy of the human box and avoid misjudgment. Scooter ID label: Ensure that each scooter ID label remains consistent in different video frames by accurately calibrating the scooters in the video frames. Then retrain the YOLOv8 model using the adjusted dataset. This step requires adding new data (including the adjusted video frames and labels) to the training set and training the model through supervised learning. Finally, after the model is retrained, perform tests on the validation set to verify whether the corrected model can reduce the previous errors and improve the accuracy of object detection. If the performance on the validation set still does not meet the standards, the process of adjusting the dataset and model training can be repeated until the expected effect is achieved.

[0109] In this solution, by receiving identification error information and adjusting the dataset, the errors that occur in the actual application of the model can be effectively corrected, thereby improving the accuracy of object detection and reducing the situations of misrecognition or missed detection.

[0110] Figure 2 It is a schematic flowchart of a method for detecting multiple people riding scooters provided by an embodiment of the present disclosure. As Figure 2 shown, the method includes:

[0111] S201, Continuously obtain video stream data, continuously determine video frame data according to the video stream data, input the video frame data into a pre-trained YOLOv8 model to obtain scooter candidate boxes and scooter IDs of scooters in each video frame data, and obtain human candidate boxes and human IDs of humans in each video frame data; wherein, the number of scooter candidate boxes and human candidate boxes is at least one.

[0112] S202, Determine the scooter position data of the scooter target corresponding to each scooter ID in each video frame data, and determine the human position data of the human target corresponding to each human ID in each video frame data. Determine the scooter speed of each scooter target according to the scooter position data, and determine the human speed of each human target according to the human position data.

[0113] S203. Determine the intersection area and union area between each scooter candidate box and each human candidate box in each video frame data. According to the intersection area, the union area, the scooter speed, the human speed, and a preset dynamic intersection-over-union calculation formula, calculate the dynamic intersection-over-union of each candidate box pair composed of each scooter candidate box and each human candidate box in each video frame data.

[0114] S204. Determine a target video frame data set according to a preset time window and the video frame data. Input the video frame data corresponding to each human candidate box in the target video frame data set into a preset action recognition model to determine the human actions of the human targets corresponding to each human ID.

[0115] The preset action recognition model can be a trained deep learning model dedicated to recognizing human actions in videos. It is usually trained based on convolutional neural networks (CNNs), recurrent neural networks (RNNs), or their variants (such as 3D CNNs, long short-term memory networks (LSTMs), etc.). It can extract time series features related to actions from the input continuous video frame data, thereby recognizing the specific actions of humans.

[0116] Human actions can refer to specific behaviors or postures performed by humans recognized in videos. For example, common human actions include: walking, running, jumping, waving, sitting, standing, squatting, etc. Each action corresponds to a specific class label, and the model is trained on these classes to determine the actions of the people in the video frames.

[0117] The target video frame dataset is a set of processed frame data collected from a real-time video stream. Each frame of data has undergone object detection and contains human candidate bounding boxes and corresponding IDs (i.e., the unique identifiers for each human object). In this dataset, there is at least one human candidate bounding box and human ID. Then, the coordinate information (such as the upper-left and lower-right coordinates) of the detected human candidate bounding boxes is extracted from each frame, and it is ensured that they correspond one-to-one with the human IDs. According to the coordinates of each candidate box, the human regions are cropped from the video frame images. These regions will serve as the original data for subsequent processing and input into the action recognition model. Ensure that the data format input into the model is consistent with the training input of the model. Usually, the input of the action recognition model is an image or a video frame. Therefore, the human regions need to be obtained by cropping (for example, obtaining the regions within each candidate box). If it is a temporal-based action recognition model (such as LSTM or 3D CNN), then multiple frames of data need to be serialized to ensure that the model can process the temporal dimension information. Then, each cropped human region is used as input and fed into a preset action recognition model. The model will analyze the spatio-temporal features of these regions and output the corresponding action categories (such as "walking", "running", "sitting down"). Based on the output of the action recognition model, the actions of each human object (based on the ID) are determined. For example, the model may output results such as "ID 1: cycling", "ID 2: walking fast", etc. The model will assign corresponding action labels to each human ID to ensure that each human object in the video can be labeled with an appropriate action.

[0118] S205, determine whether the human action belongs to the scooter action category according to the preset action rules.

[0119] The preset action rules can refer to a set of pre-set standards or conditions used to determine whether a specific action conforms to the definition of scooter-related actions. This rule is usually designed based on the motion characteristics, action patterns, or specific action labels related to scooters. For example, the riding action rule: If the detected human action is "riding" or "walking fast" and is accompanied by the presence of a scooter at the same time, it is considered that this action belongs to the scooter action category. The standing action rule: If the detected human action is "standing" and the position is close to the scooter, it may conform to the action pattern of "standing on the scooter" and is regarded as a scooter action. The smooth movement rule: The action recognition model may output certain actions (such as "walking" or "running"). If this action occurs beside the scooter and conforms to the riding motion trajectory, it can be considered as part of the scooter action category.

[0120] The scooter action categories can refer to specific action patterns related to scooters. Specifically, they can include: Riding: The action pattern when riding a scooter, usually manifested as rapid paces and alternating movements of both feet. Standing: The posture of a person standing on a scooter, commonly seen when the scooter has stopped. Gliding: The sliding or smooth driving of a scooter without obvious external propulsion, usually accompanied by slight posture changes or low-speed movements.

[0121] The human body action tags can be compared with preset action rules to determine whether the action belongs to the scooter action category. The specific rules can include: Action type judgment: If the recognized action is "riding" or "standing", and the distance between the human body target and the scooter target conforms to the rules (for example, the human body candidate box is close to the scooter candidate box), then this human body action can be regarded as a scooter action. Movement trajectory judgment: If the human body action pattern is similar to riding (such as rapid movement, accompanied by the movement of the scooter), then this action can be classified as a scooter-related action. Then, based on temporal continuity, it is determined whether an action conforms to the continuous characteristics of scooter actions within a period of time. For example, if a human body is shown as "riding" in multiple video frames and is accompanied by the scooter target at the same time, then it can be further confirmed that this human body action belongs to the scooter action category. According to the preset rules, it is finally determined whether the action of each human body target belongs to the scooter action category.

[0122] Based on the above technical solution, optionally, after determining whether the human body action belongs to the scooter action category according to the preset action rules, the method further includes:

[0123] If in each video frame data of the target video frame dataset, there is only one dynamic intersection-over-union ratio between the human body candidate box and the same scooter candidate box that is greater than the preset threshold, and the human body actions all belong to the scooter action category, then it is determined as a single-person riding phenomenon, and the target video frame dataset is sent to the shared scooter service provider.

[0124] In this solution, if each frame in the target video frame dataset meets the following conditions, it can be determined as a "single-person riding phenomenon": There is only one human body candidate box and one scooter candidate box in each frame. The dynamic intersection-over-union ratio between the human body candidate box and the scooter candidate box is greater than the preset threshold. The action of this human body belongs to the scooter-related action category (such as riding, standing, etc.). Once it is confirmed that all video frames in the target video frame dataset meet the above conditions, the video frame dataset can be marked as a "single-person riding" phenomenon and sent to the shared scooter service provider for further processing or monitoring. Specifically, the data can be sent to the service provider through methods such as APIs and message queues, or stored through file upload, cloud storage, etc. for subsequent processing.

[0125] In this solution, by combining the dynamic intersection over union (IoU) and human action recognition, it is possible to more accurately determine whether it is a single-person riding phenomenon, rather than relying solely on the appearance or static information of the image. And it can process the video stream in real time, quickly judge and respond to the riding behavior, so as to provide timely data feedback for shared scooter service providers.

[0126] S206. If, in each video frame data of the target video frame dataset, there are at least two human candidate boxes whose dynamic IoU with the same scooter candidate box is greater than a preset threshold, and the human actions all belong to the scooter action category, it is determined that there is a multi-person riding phenomenon, and the target video frame dataset is sent to the traffic management department and the shared scooter service provider.

[0127] In this embodiment, by using a pre-trained action recognition model, the system can more accurately recognize the actions of humans in the video stream. It is more precise than traditional image processing methods, which can reduce misjudgment and missed judgment. Based on the action recognition model, it can clearly distinguish multiple human action categories and classify them into scooter-related action categories, improving the accuracy of multi-person riding behavior recognition.

[0128] Based on the above technical solution, optionally, the training steps of the preset action recognition model include:

[0129] Obtain historical human image data and historical human actions, and label the human action labels of each historical human image data according to the historical human actions; where the historical human image data contains the corresponding historical human ID;

[0130] Construct an action recognition model, and train the action recognition model according to the historical human image data and the human action labels until the action recognition model reaches the preset action recognition model training standard.

[0131] In this solution, the historical human image data can be a human image dataset collected in the past, and each image or video frame contains at least one human target. They are the basic data for training the action recognition model, usually from surveillance cameras, video records, or sensor devices. Each image data should contain the feature information of the human body, such as position, shape, etc., and may also contain candidate box data (such as the position area of the human body in the image).

[0132] The historical human actions can be the actions or behaviors shown by humans in these historical human image data. These actions can be different postures or activities such as walking, running, riding, standing, etc. The labels of the actions are usually obtained through manual annotation or other detection methods. Each action annotation corresponds to one or more image frames and is assigned a corresponding action category.

[0133] The human body action label can refer to the action category assigned to each piece of historical human body image data. By annotating the actions in the historical image data, the human body action label can tell the action recognition model what action the image corresponds to. The action label is the target output in supervised learning, such as: "walking", "riding", "running", etc. Each label is associated with specific image data and is used for the training and testing of the model. For each image, there may be multiple labels (if it involves the case of multiple targets).

[0134] The historical human body ID can be the unique identifier of the human body target in each human body image. This ID is used to distinguish the human body objects detected in different video frames or image data. The role of the historical human body ID is to maintain the tracking of the same human body in multiple video frames or images, ensuring that the action label can be correctly matched with a specific human body target. Through the historical human body ID, the action recognition model can distinguish which actions belong to the same person and which are different human body targets.

[0135] The preset training criteria for the action recognition model can be the criteria for the model to reach a certain performance. Specifically, it can include accuracy / precision: the accuracy of the model on the test set reaches a certain level, such as above 90%. Loss function: the value of the loss function of the model reaches a predetermined standard during the training process, usually below a certain threshold. Generalization ability: the model performs well on different data sets and can effectively recognize unseen action types. Number of training times: the model undergoes a certain number of training epochs to ensure its full training and avoid underfitting.

[0136] Historical human body image data can be collected from cameras, sensors, or existing video databases to ensure that the data covers a variety of human motion scenarios (e.g., walking, running, cycling, sitting, etc.). Action labels are assigned to each image or video frame through manual annotation or automated methods (e.g., using existing action recognition models). Each human object in an image requires a unique ID for subsequent tracking and classification. This ID can be generated by object detection algorithms (e.g., YOLO, Faster R-CNN, etc.). Then, a deep learning model suitable for human action recognition is selected, such as CNN (Convolutional Neural Network) or RNN (Recurrent Neural Network), or a complex model that combines the two (e.g., 3D CNN, two-stream network, etc.). According to the historical image data and action labels, the data is divided into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to adjust hyperparameters, and the test set is used to evaluate the final performance of the model. The historical human body image data and the corresponding human action labels are input into the model. The data is preprocessed (e.g., normalized, data augmented, etc.) according to the characteristics of the data (such as image size, number of channels, etc.). The historical human body image data is used for training to optimize the model parameters and minimize the loss function (e.g., cross-entropy loss, mean squared error, etc.). The training process is monitored through the loss function and accuracy to ensure that the performance of the model on the training set continues to improve and overfitting is avoided. After training is completed, the test set is used to evaluate the model to ensure that the model can also meet the preset training standards on unseen data.

[0137] The above is the introduction of the method embodiment. The following further illustrates the solution of the present disclosure through a system embodiment.

[0138] Figure 3 The following is a schematic block diagram of a multi-person scooter detection system provided by an embodiment of the present disclosure. As Figure 3 shown, the system includes:

[0139] An identification module 301, configured to continuously obtain video stream data, continuously determine video frame data according to the video stream data, input the video frame data into a pre-trained YOLOv8 model, obtain scooter candidate boxes and scooter IDs of scooters in each video frame data, and obtain human candidate boxes and human IDs of humans in each video frame data; wherein, the number of scooter candidate boxes and human candidate boxes is at least one;

[0140] A speed confirmation module 302, configured to determine scooter position data of scooter targets corresponding to each scooter ID in each video frame data, and determine human position data of human targets corresponding to each human m in each video frame data, determine the scooter speed of each scooter target according to the scooter position data, and determine the human speed of each human target according to the human position data;

[0141] The dynamic intersection-over-union calculation module 303 is configured to determine the intersection area and the union area between each scooter candidate box and each human candidate box in each video frame data, and calculate the dynamic intersection-over-union of each candidate box pair formed by each scooter candidate box and each human candidate box in each video frame data according to the intersection area, the union area, the scooter speed, the human speed, and a preset dynamic intersection-over-union calculation formula;

[0142] The determination module 304 is configured to determine a target video frame data set according to a preset time window and the video frame data. If, in each video frame data of the target video frame data set, the dynamic intersection-over-union of at least two human candidate boxes and the same scooter candidate box is greater than a preset threshold, it is determined that there is a phenomenon of multiple people riding, and the target video frame data set is sent to the traffic management department and the shared scooter service provider.

[0143] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0144] Figure 4 FIG. shows a schematic block diagram of an electronic device 400 that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0145] The electronic device 400 includes a computing unit 401, which can execute various appropriate actions and processes according to a computer program stored in the ROM 402 or a computer program loaded from the storage unit 408 into the RAM 404. In the RAM 404, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 404 are connected to each other through a bus 404. The I / O interface 405 is also connected to the bus 404.

[0146] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0147] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the described method for detecting multiple people riding a scooter. For example, in some embodiments, the described method for detecting multiple people riding a scooter can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 404 and executed by the computing unit 401, one or more steps of the described method for detecting multiple people riding a scooter can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the described method for detecting multiple people riding a scooter in any other suitable way (e.g., by means of firmware).

[0148] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chips (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0149] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0150] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0152] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0153] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server that incorporates a blockchain.

[0154] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0155] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for detecting multiple people riding a scooter, characterized in that: The method is performed by an edge computing device, and the method includes: Continuously acquiring video stream data, continuously determining video frame data according to the video stream data, inputting the video frame data into a pre-trained YOLOv8 model, obtaining a scooter candidate frame and a scooter ID of a scooter in each video frame data, and obtaining a human candidate frame and a human ID in each video frame data; wherein the number of the scooter candidate frame and the number of the human candidate frame are both at least one; Determine scooter position data of a scooter target corresponding to each scooter ID in each video frame data, and determine human position data of a human target corresponding to each human ID in each video frame data, determine a scooter speed of each scooter target according to the scooter position data, and determine a human speed of each human target according to the human position data; Determine the intersection area and union area of ​​each scooter candidate frame and each human candidate frame in each video frame data, and calculate the dynamic intersection-and-union ratio of the candidate frame pairs composed of each scooter candidate frame and each human candidate frame in each video frame data according to the intersection area, the union area, the scooter speed, the human speed and a preset dynamic intersection-and-union ratio calculation formula; A target video frame data set is determined according to a preset time window and the video frame data. If, in each video frame data of the target video frame data set, a dynamic intersection-and-union ratio of at least two human candidate frames and the same scooter candidate frame is greater than a preset threshold, it is determined that multiple people are riding, and the target video frame data set is sent to the traffic management department and the shared scooter service provider.

2. The method according to claim 1, characterized in that: in, According to the intersection area, the union area, the scooter speed, the human body speed and a preset dynamic intersection-and-union ratio calculation formula, a dynamic intersection-and-union ratio of a candidate frame pair consisting of each scooter candidate frame and each human body candidate frame in each video frame data is calculated, including: Calculate the speed difference between the scooter speed and the human body speed, and calculate the speed influence factor coefficient according to the speed difference and a preset speed influence factor coefficient calculation formula; According to the intersection area, the union area, the speed influence factor coefficient and the preset dynamic intersection-and-union ratio calculation formula, the dynamic intersection-and-union ratio of the candidate frame pairs consisting of each scooter candidate frame and each human candidate frame in each video frame data is calculated; wherein the preset dynamic intersection-and-union ratio calculation formula is: Among them, A Union is the union area; α is the preset adjustment factor, which is used to adjust the influence of speed on the intersection-union ratio, V Impact is the speed influence factor coefficient, which is used to measure the impact of the speed difference between targets on the dynamic intersection ratio; A Interestion is the intersection area; DynamicIou is the dynamic intersection-union ratio of the candidate frame pairs composed of each scooter candidate frame and each human candidate frame.

3. The method according to claim 2, characterized in that in, The preset speed influence factor coefficient calculation formula is: V Impact =1+v max +γ·Δv; Among them, v max is the maximum speed of the two targets contained in each candidate frame pair; γ is the preset speed difference weight coefficient; Δv is the speed difference.

4. The method according to claim 1, characterized in that: in, After determining the target video frame data set according to the preset time window and the video frame data, the method further includes: Input each video frame data corresponding to each human candidate frame in the target video frame data set into a preset action recognition model to determine the human action of the human target corresponding to each human ID; Determine whether the human body motion belongs to the scooter motion category according to the preset motion rules; Correspondingly, if in each video frame data of the target video frame data set, the dynamic intersection-and-union ratio of at least two human candidate frames and the same scooter candidate frame is greater than the preset threshold, it is determined that there is a multi-person riding phenomenon, and the target video frame data set is sent to the traffic management department and the shared scooter service provider, including: If in each video frame data of the target video frame data set, the dynamic intersection-union ratio of at least two human candidate frames and the same scooter candidate frame is greater than a preset threshold, and the human movements all belong to the scooter movement category, it is determined that there is a multi-person riding phenomenon, and the target video frame data set is sent to the traffic management department and the shared scooter service provider.

5. The method according to claim 4, characterized in that in, The training steps of the preset action recognition model include: Acquire historical human image data and historical human movements, and annotate human movement labels of each historical human image data according to the historical human movements; wherein the historical human image data includes a corresponding historical human ID; Constructing an action recognition model, and training the action recognition model according to the historical human image data and the human action labels until the action recognition model reaches a preset action recognition model training standard.

6. The method according to claim 4, characterized in that in, After determining whether the human body motion belongs to the scooter motion category according to the preset motion rule, the method further includes: If, in each video frame data of the target video frame data set, there is only one human candidate frame and the dynamic intersection-union ratio of the same scooter candidate frame is greater than a preset threshold, it is judged as a single-person riding phenomenon, and the target video frame data set is sent to the shared scooter service provider.

7. The method according to claim 1, characterized in that in, The training steps of the pre-trained YOLOv8 model include: Acquire historical video stream data, and determine historical video frame data, scooter target frame labels of each annotated historical video frame data, scooter ID labels of each scooter target frame label, human target frame labels of each annotated historical video frame data, and human ID labels of each human target frame label according to the historical video stream data; A YOLOv8 model is trained according to the historical video frame data, the scooter target frame labels of each annotated historical video frame data, the scooter ID labels of each scooter target frame label, the human target frame labels of each annotated historical video frame data, and the human ID labels of each human target frame label, until the YOLOv8 model reaches a preset YOLOv8 model training standard.

8. The method according to claim 7, characterized in that in, After sending the target video frame dataset to the traffic management department and the shared scooter service provider, the method further includes: If recognition error information is received from the traffic management department and the shared scooter service provider, an adjustment data set is determined according to the recognition error information, and a preset YOLOv8 model is retrained according to the adjustment data set until the preset YOLOv8 model reaches the preset YOLOv8 model training standard; wherein the adjustment data set includes adjustment video frames, scooter target frame labels of each adjusted video frame data that has been annotated, scooter ID labels of each scooter target frame label, and human target frame labels of each adjusted video frame data that has been annotated.

9. A system for detecting multiple people riding a scooter, used to execute the method according to any one of claims 1 to 8, characterized in that: The system is configured on an edge computing device, and the system includes: an identification module, for continuously acquiring video stream data, continuously determining video frame data according to the video stream data, inputting the video frame data into a pre-trained YOLOv8 model, obtaining a scooter candidate frame and a scooter ID of a scooter in each video frame data, and obtaining a human candidate frame and a human ID in each video frame data; wherein the number of the scooter candidate frame and the number of the human candidate frame are both at least one; A speed confirmation module, used to determine the scooter position data of the scooter target corresponding to each scooter ID in each video frame data, and determine the human position data of the human target corresponding to each human ID in each video frame data, determine the scooter speed of each scooter target according to the scooter position data, and determine the human speed of each human target according to the human position data; A dynamic intersection-and-union ratio calculation module is used to determine the intersection area and union area of ​​each scooter candidate frame and each human candidate frame in each video frame data, and calculate the dynamic intersection-and-union ratio of the candidate frame pairs composed of each scooter candidate frame and each human candidate frame in each video frame data according to the intersection area, the union area, the scooter speed, the human speed and a preset dynamic intersection-and-union ratio calculation formula; The judgment module is used to determine the target video frame data set according to the preset time window and the video frame data. If, in each video frame data of the target video frame data set, the dynamic intersection and union ratio of at least two human candidate frames and the same scooter candidate frame is greater than a preset threshold, it is determined that there is a phenomenon of multiple people riding, and the target video frame data set is sent to the traffic management department and the shared scooter service provider.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8.