A Monitoring Method for Mobile Vendors Based on Rotated Object Detection and Pedestrian Statistics
By adopting mobile vendor monitoring methods based on rotary target detection and pedestrian statistics in urban management, the problem of inaccurate mobile vendor detection is solved, and a higher accuracy of early warning of violations and illegal behaviors and the accuracy of smart city management is achieved.
Patent Information
- Application Number
- CN202310058276.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-01-15
AI Technical Summary
In the prior art, mobile vendors are inaccurately tested, resulting in an increase in urban management workload and it is difficult to predict violations and laws.
The mobile vendor monitoring method based on rotation target detection and pedestrian statistics is adopted, video information is collected through cameras, and characteristic information of mobile vendors and pedestrians is extracted using the target detection model. Combined with the rotating target frame and pedestrian monitoring area, the illegal business situation of mobile vendors is judged.
It improves the accuracy of early warning of illegal and illegal behavior of mobile vendors, reduces the workload of urban management, and improves the accuracy of smart city management.
Smart Images

Figure CN116311030B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vendor management, and particularly to a mobile vendor monitoring method based on rotated object detection and pedestrian statistics. Background Art
[0002] With the continuous development of deep learning technology, it has become possible to quickly and effectively extract useful semantic information from videos or images, and it has been widely applied in fields such as video surveillance and image recognition. The deep learning technology continuously extracts high-dimensional feature information in images through convolutional neural networks, and infers the corresponding object categories and related information through the probability of the feature information. In the management of smart cities, by adopting deep learning technology, violations and illegal acts in the urban management process can be effectively detected. By predicting violations and illegal acts, early warning information can be sent to relevant management personnel and units in advance, providing effective decision-making information for relevant management personnel and units to guide and persuade against violations in advance, and improving the efficiency of urban management and reducing the ineffective cruising time of urban management personnel.
[0003] Currently, the method of manual patrol is often adopted in urban management. For the management of street vending and mobile vendors, the method of combining manual video surveillance and on-site patrol is often used. Although this method can achieve certain results, it will consume a large amount of manpower and material resources for monitoring and cruising, with low management efficiency, and it cannot well predict violations and illegal acts, so it is easy to be discovered only after vendors gather, thus increasing the management difficulty and workload. In response to such problems, currently, in the management process of some smart cities, the method of object detection is adopted to detect mobile vendors in images. Although this method can detect some violations and illegal acts, it also alarms for passing pedestrians and temporary parking that appear on the road, which invisibly increases the workload and difficulty of urban management.
[0004] Therefore, how to provide a mobile vendor monitoring method that can accurately determine whether the mobile vendor has committed illegal acts such as occupying the road for business, so as to improve the accuracy of early warning of mobile vendor violations and illegal acts and the accuracy of smart city management. Summary of the Invention
[0005] For this purpose, the present invention provides a mobile vendor monitoring method based on rotated object detection and pedestrian statistics to solve the problem of increasing the workload of urban management due to inaccurate detection of mobile vendors in the prior art.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A mobile vendor monitoring method based on rotated object detection and pedestrian statistics, comprising the following steps:
[0008] Step 1: Collect video information through a camera and input it into the target detection model;
[0009] Step 2: Extract feature information about street vendors and pedestrians in the image through the feature extraction network of the target detection model, and input it into the feature fusion network for feature fusion and the prediction network for class determination;
[0010] Step 3: Draw corresponding target boxes in the image based on the feature information of the street vendors and pedestrians, and set up a pedestrian monitoring area extending outward from the target box of the street vendor. Record the movement trajectory of the street vendor in the image according to the center point coordinates of the target box of the street vendor;
[0011] Step 4: Record the movement trajectory information corresponding to each pedestrian according to the predicted feature information of the pedestrians;
[0012] Step 5: Judge whether the street vendor has illegal business operations according to the movement trajectories of the street vendor and pedestrians, and send the statistical information to the back-end monitoring system and urban management personnel.
[0013] Further, the generation of the target box in step 3 specifically includes the following steps:
[0014] Step 301: The feature information of the street vendor and pedestrians is further extracted through a convolutional network and input into the prediction network;
[0015] Step 302: The target center prediction network Center net predicts the center point coordinates of the target box, the target box side length prediction network Side net predicts the side length of the target box, the rotation prediction network Rotation net predicts the rotation angle of the target box, and the classification prediction network Classes net detects the category to which the target object belongs;
[0016] Step 303: Draw the target box of the target object according to the predicted center point coordinates, side length and rotation angle of the target box.
[0017] Further, the judgment of whether the street vendor has illegal business operations in step 5 specifically includes the following steps:
[0018] Step 501: Judge whether the street vendor stays in the parking area according to the center point coordinates and movement trajectory of the street vendor;
[0019] Step 502: If the mobile vendor stays in the parking area for more than a first preset stay time, the pedestrians in the pedestrian monitoring area are counted; if the mobile vendor stays in the non-parking area for more than a second preset stay time, the pedestrians in the pedestrian monitoring area are counted;
[0020] Step 503: If a pedestrian stays in the pedestrian monitoring area for a time exceeding a first stay threshold, it is determined that the mobile vendor is operating illegally; if two or more pedestrians stay in the pedestrian monitoring area for a time exceeding a second stay threshold, it is determined that the mobile vendor is operating illegally.
[0021] Furthermore, the method of inputting the feature information of the mobile vendor into the feature fusion network for feature fusion is as follows:
[0022] The feature extraction network extracts high-dimensional feature information of three scales (F3, F4, F5);
[0023] The high-dimensional feature information is input into the feature fusion network for bidirectional feature fusion. In the bidirectional feature fusion stage, five bidirectional feature fusion units are used for series fusion. Information of each scale is fused within the bidirectional feature fusion unit by using an addition operation.
[0024] The expression of the addition operation is
[0025]
[0026] in and are the input and output feature maps of the kth layer, ω n is the corresponding learnable weight;
[0027] The Softmax function is used to limit the weight of each layer, and its expression is
[0028]
[0029] Among them, ε is a very small parameter greater than 0, with a value of 0.00001, and i and j are the number of feature maps in each layer.
[0030] Furthermore, when performing bidirectional feature fusion, a multi-scale fusion network is used for multi-scale fusion, and a concatenate operation is used during the fusion process to connect information of different scales.
[0031] Furthermore, the feature information of the mobile vendor after feature fusion is input into the prediction network for prediction, and the prediction network performs predictions including target category prediction, center point prediction, side length prediction, rotation angle prediction and confidence prediction.
[0032] Further, when predicting the side length, the border is predicted by regressing the deviation between the true target box and the preset box:
[0033] The output value P=(x p , y p , w p , h p ) of the side length prediction network is
[0034]
[0035] Then the output border is
[0036]
[0037] Among them, B=(x, y, w, h) is the true border, G=(x g , y g , w g , h g ) is the preset border, x and y are pixel coordinates in pixels, w and h are the border width and length in pixels, and s is the downsampling ratio in the feature extraction network.
[0038] Further, the rotation angle formula applied in the rotation angle prediction is
[0039]
[0040] Among them, (x0, y0) is the center point coordinate of the target box, (x i , y i ) are the four vertex coordinates of the target box, satisfying x2 - x1 = w, y3 - y1 = h, y1 = y2, y3 = y4, x1 = x3, x2 = x4, (x, y) are the corresponding coordinates after vertex rotation, and θ is the rotation angle.
[0041] Further, the target detection model is obtained by training with a labeled dataset. According to the requirements of the monitoring object, the roLabellmg annotation tool is used to annotate the dataset with the target to be detected.
[0042] Further, the target detection model is a convolutional neural network model built using pytorch.
[0043] The present invention has the following advantages:
[0044] 1. The mobile vendor detection method currently widely used in urban management often uses a simple rectangular frame to select the target, resulting in invalid areas being added to the calculation. The present invention uses a target detection method with a rotation angle to detect mobile vendors in the image. By rotating the target frame, the mobile vendors in the image can be more accurately framed, so that the framed area can be further processed efficiently.
[0045] 2. The commonly used mobile vendor identification method based on target detection is a method that simply detects mobile vendors through multiple images to obtain the residence time. The present invention predicts the movement trajectory and continuous residence time of mobile vendors by using video stream methods to predict illegal behaviors, and monitors fast-moving vendors and temporary parking behaviors in the field of view without warning, thereby improving the accuracy of smart city management.
[0046] 3. In order to solve the problem of false alarms for temporary parking and roadside parking by judging illegal behaviors based on the stay time of mobile vendors, the present invention defines an effective pedestrian counting area around the target detection area, and determines whether the mobile vendor has committed illegal behaviors such as occupying the road by counting the changes and stay time of pedestrians in the area, thereby improving the accuracy of early warning of mobile vendors' illegal behaviors and the accuracy of smart city management. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.
[0048] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.
[0049] Figure 1 A flow chart of the mobile vendor monitoring method provided by the present invention;
[0050] Figure 2 A specific flow chart of step 3 in the method provided by the present invention;
[0051] Figure 3Specific flowchart of step 5 in the method provided by the present invention;
[0052] Figure 4 Specific implementation flowchart of the monitoring method provided by the present invention;
[0053] Figure 5 Schematic diagram of image annotation provided by the present invention;
[0054] Figure 6 Schematic diagram of the detection network in the object detection model provided by the present invention;
[0055] Figure 7 Schematic diagram of pedestrian detection area and pedestrian tracking simulation provided by the present invention;
[0056] Figure 8 Actual detection result diagram of pedestrian detection area and pedestrian tracking provided by the present invention. Detailed implementation mode
[0057] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] To solve the problem of increasing the workload of urban management caused by inaccurate detection of mobile vendors in the prior art, the present invention provides a mobile vendor monitoring method based on rotating object detection and pedestrian statistics, as Figure 1 shown, including the following steps:
[0059] Step 1: Collect video information through a camera and input it into the object detection model;
[0060] Step 2: Extract feature information about mobile vendors and pedestrians in the image through the feature extraction network of the object detection model, and input it into the feature fusion network for feature fusion and the prediction network for category determination;
[0061] Step 3: Draw corresponding target boxes in the image according to the feature information of mobile vendors and pedestrians, and set up a pedestrian monitoring area outside the target box of the mobile vendor. Record the movement trajectory of the mobile vendor in the image according to the center point coordinates of the target box of the mobile vendor;
[0062] Step 4: Record the movement trajectory information corresponding to each pedestrian according to the predicted feature information of the pedestrian;
[0063] Step 5: Determine whether the mobile vendor is operating illegally based on the movement trajectories of the mobile vendor and pedestrians, and send the statistical information to the backend monitoring system and urban management personnel.
[0064] In step 3, the generation of the target box, as Figure 2 shown, specifically includes the following steps:
[0065] Step 301: The feature information of the mobile vendor and pedestrians is further extracted through a convolutional network and input into the prediction network.
[0066] Step 302: The target center prediction network Centernet predicts the center point coordinates of the target box, the target box side length prediction network Side net predicts the side length of the target box, the rotation prediction network Rotation net predicts the rotation angle, and the classification prediction network Classes net detects the category to which the target object belongs.
[0067] Step 303: Draw the target box of the target object according to the predicted center point coordinates, side length, and rotation angle of the target box.
[0068] In step 5, determine whether the mobile vendor is operating illegally, as Figure 3 shown, specifically includes the following steps:
[0069] Step 501: Determine whether the mobile vendor stays in the parking area according to the center point coordinates and movement trajectory of the mobile vendor.
[0070] Step 502: If the mobile vendor stays in the parking area for more than the first preset stay time, count the pedestrians in the pedestrian monitoring area; if the mobile vendor stays in a non-parking area for more than the second preset stay time, count the pedestrians in the pedestrian monitoring area.
[0071] Step 503: If a pedestrian stays in the pedestrian monitoring area for more than the first stay threshold, it is determined that the mobile vendor is operating illegally; if two or more pedestrians stay in the pedestrian monitoring area for more than the second stay threshold, it is determined that the mobile vendor is operating illegally.
[0072] To facilitate further description of the monitoring method for mobile vendors in the present invention, as Figure 4 shown, specific embodiments are proposed.
[0073] First, collect a data set through a monitoring camera and process it into a picture sequence. As Figure 5 shown, use roLabellmg to label the mobile vendors and pedestrians in the image to form a data set. As Figure 6As shown, a detection network for mobile vendors is built using PyTorch and trained on the dataset for 500 epochs.
[0074] After training, the trained model is deployed to the monitoring platform for online monitoring of mobile vendors. The monitoring video is input into the mobile vendor detection network, and the network outputs the ID of each mobile vendor, its corresponding center point coordinates, the size of the bounding box area, and the rotation angle of the target box. At the same time, it outputs the ID of the pedestrians in the video and the pixel trajectory of their center points.
[0075] For the detection of mobile vendors, based on the detected center point coordinates of the mobile vendors, record the movement trajectories of the mobile vendors and determine whether they stay. When the center point of the mobile vendor stays continuously within 8 pixels for more than 1 minute, determine whether it stays in the planned parking area.
[0076] Set the first preset stay time for mobile vendors to 5 minutes and the second preset stay time to 3 minutes.
[0077] If the mobile vendor is in the planned parking area, continue to monitor whether it stays continuously for more than 5 minutes. If it exceeds the set stay time, plan the corresponding pedestrian detection area as shown in Figure 7 and conduct pedestrian statistics to determine whether there is any illegal business operation. For mobile vendors in non-parking areas, directly determine whether there is illegal parking based on whether they stay continuously for more than 3 minutes. If it exceeds the set time threshold, directly plan their pedestrian detection area and conduct pedestrian statistics.
[0078] For the monitoring of pedestrians, by continuously recording the coordinate values corresponding to the center points of each pedestrian, draw their continuous movement trajectories within 3 minutes. After completing the planning of the pedestrian detection area for mobile vendors and pedestrian trajectory tracking, determine whether there is any business operation by statistically calculating the continuous stay time of pedestrians in the pedestrian detection area of each mobile vendor.
[0079] Set the first stay threshold for pedestrians in the pedestrian detection area to 2 minutes and the second stay threshold to 1 minute.
[0080] When the stay time of the first pedestrian exceeds 2 minutes, regardless of whether there are subsequent pedestrians entering, it is determined as illegal business operation. When the stay time of two or more pedestrians exceeds 1 minute simultaneously, it is determined as illegal business operation. After determining that the mobile vendor is engaged in illegal business operation, the system immediately generates illegal warning message information and image information for the corresponding detection results and sends them to the backend smart city management system for warning, so that the management personnel can handle them effectively.
[0081] As Figure 5The figure shows the annotation of street vendors and pedestrians in the image. The annotation of all images is carried out using roLabellmg for the annotation of rotated targets. The annotation format for a single target is id, x, y, long_side, width_side, Rotation, where id is the category of street vendors and pedestrians, (x,y) is the pixel coordinate of the corresponding center point, long_side and width_side are the widths of the long side and the short side respectively, and Rotation is the rotation angle of the target box, and the angle range is (0, 180°). During the annotation process, since the target box of pedestrians is much smaller than that of street vendors, the rotation angle of all pedestrian target boxes is 90°.
[0082] As Figure 6 The figure shows the schematic diagram of the street vendor detection network proposed by the present invention. The network mainly consists of three parts: feature extraction, feature fusion, and prediction. Among them, the feature extraction network is constructed by a convolutional neural network based on the SE attention mechanism. Through the attention mechanism, the extraction of key feature information by the entire network is improved, and the efficiency of feature extraction is higher than that of the original network.
[0083] After extracting high-dimensional feature information (F3, F4, F5) at three scales through the feature extraction network, the feature information is input into the feature fusion network for bidirectional feature fusion. In the bidirectional feature fusion stage, 5 bidirectional feature fusion units blocks are used for cascaded fusion, and the information at each scale is fused by an addition operation inside the bidirectional feature fusion unit.
[0084] The expression of its addition operation is
[0085]
[0086] Where and are the input and output feature maps of the k-th layer respectively, and ω n is the corresponding learnable weight;
[0087] To balance the relationship between various items and prevent divergence during the training process, the Softmax function is used to limit the weights of each layer, and its expression is
[0088]
[0089] Where ε is a very small parameter greater than 0, and its value is 0.00001, and i and j are the number of feature maps in each layer.
[0090] After performing bidirectional feature fusion, to retain multi-dimensional scale information so as to have the same detection effect on targets of different sizes, a multi-scale fusion network is adopted for multi-scale fusion, and the concatenate operation is used to connect different scale information during the fusion process.
[0091] Finally, the fused feature belongs to the prediction network for prediction. The prediction network mainly includes the prediction of the target category, the prediction of the center point, the prediction of the side length ratio, the prediction of the rotation angle, and the prediction of the confidence level.
[0092] In the side length prediction network, the border is predicted by regressing the deviation between the real target box and the preset box. Suppose the real border is B=(x,y,w,h), and the preset border is G=(x g ,y g ,w g ,h g ), then the output value P=(x p ,y p ,w p ,h p ) of the side length prediction network is expressed as follows:
[0093]
[0094] Then the output border can be expressed as:
[0095]
[0096] Among them, x and y are pixel coordinates in pixels, w and h are the border width and length in pixels, and s is the downsampling ratio in the feature extraction network. The downsampling ratios of the three scales are 8, 16, and 32 respectively.
[0097] In the prediction network, the target detection results of the three scales are predicted separately, so the average is obtained for the center point, side length, and rotation angle.
[0098] The category selects the category with the highest confidence level among the prediction results of the three scales as the output category.
[0099] In the post-processing of the detection results of street vendors, since the prediction network outputs the center point coordinates, side length, and rotation angle, the target box needs to be rotated according to the rotation angle, and its rotation formula is as follows:
[0100]
[0101] Among them, (x0,y0) is the center point coordinate of the target box, (x i ,y i) are the coordinates of the four vertices of the target box, satisfying x2 - x1 = w, y3 - y1 = h, y1 = y2, y3 = y4, x1 = x3, x2 = x4, (x, y) are the corresponding coordinates after the vertex rotation, and θ is the rotation angle.
[0102] For the processing of the target boxes of pedestrians, since the target boxes of pedestrians are all small and all rotation angles are 90°, therefore, the target boxes of pedestrians do not need to be processed and can be directly calculated through the length and width.
[0103] Figure 7 It is a schematic diagram of the pedestrian detection area and pedestrian tracking simulation provided by the present invention. Figure 8 It is the actual detection result diagram of the pedestrian detection area and pedestrian tracking provided by the present invention.
[0104] As Figure 7 shown, 2 mobile vendors C1 and C2 are found through the mobile vendor detection network, and their lengths and widths are (l1, w1) and (l2, w2) respectively, and the rotation angles are θ1 and θ2 respectively. The lengths and widths of the pedestrian monitoring areas of the two mobile vendors can be obtained by scaling the corresponding lengths and widths by d times, which are (dl1, dw1) and (dl2, dw2) respectively, as shown by the dotted boxes in the figure.
[0105] The network detects 7 pedestrians (P1, …, P7). Among them, pedestrian P1 is far away from mobile vendors C1 and C2 and does not participate in the discrimination, and pedestrian P2 is stationary and does not participate in the judgment either. Pedestrians P3 and P5 enter the pedestrian detection area of mobile vendor C1 through movement (the movement trajectories are represented by dotted lines) and stay for more than the set 1-minute threshold. At the same time, pedestrian P4 stays in this area for a long time. Therefore, there is an illegal business situation for mobile vendor C1. For mobile vendor C2, although pedestrian P7 passes through the detection range of C2, the stay time does not exceed the set threshold and is not counted in the statistics. However, pedestrian P6 moves near C2, and its movement trajectories are all within the detection range of C2, and its stay time exceeds the set 2-minute threshold, and there is also an illegal business situation.
[0106] Regarding the mobile vendor detection method that is currently widely used in urban management, simple object detection methods are often adopted. Such object detection methods usually use rectangular boxes parallel to the image border in the image to frame the target area. However, the positions of mobile vendors in the image have various angles. This method of using rectangular boxes to frame the target leads to an enlarged monitoring area and reduces the computational amount of the detection area screening. The present invention adopts an object detection method with a rotation angle to detect mobile vendors in the image. By rotating the target detection box, the mobile vendors in the image can be more accurately framed, so as to perform further efficient processing on the framed area.
[0107] The currently commonly used method for judging mobile vendors based on object detection determines by detecting the staying time of mobile vendors in multiple images, which leads to the same monitoring and early warning for behaviors such as passing-by vendors and temporary parking, resulting in false alarms and ineffective monitoring in actual application processes. The present invention predicts the illegal behaviors by predicting the movement trajectory and continuous staying time of mobile vendors by using the method of video stream, monitors behaviors such as vendors moving quickly and temporary parking in the field of view without giving early warnings, and improves the accuracy of smart city management.
[0108] The currently commonly used method for judging mobile vendors based on object detection determines by the staying time of mobile vendors, and such situations are often interfered by situations such as temporary parking and parking in roadside parking spaces, thus triggering false alarms and other situations. The present invention demarcates an effective pedestrian quantity statistical area around the object detection area, and determines whether the mobile vendor has committed illegal behaviors such as occupying the road for business by counting the changes in the number of pedestrians and the staying time in this area, thereby improving the accuracy of early warning for illegal behaviors of mobile vendors and the accuracy of smart city management.
[0109] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A mobile vendor monitoring method based on rotating object detection and pedestrian statistics, characterized in that, The following steps are involved: Step 1: Collect video information through the camera and input it into the target detection model; Step 2: extracting feature information about street vendors and pedestrians in the image through the feature extraction network of the target detection model, and inputting it into the feature fusion network for feature fusion and the prediction network for category determination; Step 3: According to the characteristic information of the mobile vendor and the pedestrian, a corresponding target frame is drawn in the image, and a pedestrian monitoring area is extended outward from the target frame of the mobile vendor, and the movement trajectory of the mobile vendor in the image is recorded according to the coordinates of the center point of the target frame of the mobile vendor; Step 4: Based on the predicted pedestrian feature information, record the motion trajectory information corresponding to each pedestrian; Step 5: judging whether the mobile vendors are operating illegally based on the movement trajectories of the mobile vendors and pedestrians, and sending statistical information to the back-end monitoring system and city management personnel; The method of inputting the feature information of the mobile vendor into the feature fusion network for feature fusion is as follows: The feature extraction network extracts high-dimensional feature information of three scales (F3, F4, F5); The high-dimensional feature information is input into the feature fusion network for bidirectional feature fusion. In the bidirectional feature fusion stage, five bidirectional feature fusion units are used for series fusion. Information of each scale is fused within the bidirectional feature fusion unit by using an addition operation. The expression of the addition operation is wherein and are the input and output feature maps of the k-th layer respectively, and ω n is the corresponding learnable weight; The Softmax function is used to limit the weight of each layer, and its expression is Where ε is a very small parameter greater than 0, with a value of 0.00001, i and j are the number of feature maps in each layer; Inputting the feature information of the mobile vendor after feature fusion into the prediction network for prediction, wherein the prediction network performs predictions including target category prediction, center point prediction, side length prediction, rotation angle prediction and confidence prediction; When predicting the side length, the border is predicted by regressing the deviation between the real target frame and the preset frame: The output value P of the side length prediction network is P=(x p , y p , w p , h p ), where The output border is Among them, B=(x, y, w, h) is the true bounding box, and G=(x g , y g , w g , h g ) is the preset bounding box, where x and y are pixel coordinates in pixels, w and h are the width and length of the bounding box in pixels, and s is the downsampling ratio in the feature extraction network; The rotation angle formula used in the rotation angle prediction is: Among them, (x0, y0) is the center point coordinate of the target box, and (x i , y i ) are the four vertex coordinates of the target box, satisfying x2 - x1 = w, y3 - y1 = h, y1 = y2, y3 = y4, x1 = x3, x2 = x4, (x, y) are the corresponding coordinates after vertex rotation, and θ is the rotation angle.
2. The mobile vendor monitoring method based on rotating object detection and pedestrian statistics according to claim 1, characterized in that The generation of the target frame in step 3 specifically includes the following steps: Step 301: The feature information of the mobile vendors and pedestrians is further extracted through a convolutional network and input into a prediction network; Step 302: The target center prediction network Center net predicts the coordinates of the center point of the target box, the target box side length prediction network Side net predicts the side length of the target box, the rotation prediction network Rotation net predicts the rotation angle of the target box, and the classification prediction network Classes net detects the category to which the target object belongs; Step 303: Draw a target box of the target object according to the predicted target box center point coordinates, target box side length and rotation angle.
3. The mobile vendor monitoring method based on rotating object detection and pedestrian statistics according to claim 1, characterized in that, In step 5, determining whether the mobile vendor is operating in violation of regulations specifically includes the following steps: Step 501: judging whether the mobile vendor stays in the parking area according to the center point coordinates and movement trajectory of the mobile vendor; Step 502: If the mobile vendor stays in the parking area for more than the first preset stay time, count the pedestrians in the pedestrian monitoring area; If the mobile vendor stays in a non-parking area for more than the second preset stay time, count the pedestrians in the pedestrian monitoring area; Step 503: If a pedestrian stays in the pedestrian monitoring area for more than the first stay threshold, it is determined that the mobile vendor is operating illegally; if two or more pedestrians stay in the pedestrian monitoring area for more than the second stay threshold, it is determined that the mobile vendor is operating illegally.
4. The mobile vendor monitoring method based on rotating object detection and pedestrian statistics according to claim 1, wherein When performing two-way feature fusion, a multi-scale fusion network is used for multi-scale fusion, and the concatenate operation is used to connect different-scale information during the fusion process.
5. The mobile vendor monitoring method based on rotating object detection and pedestrian statistics according to claim 1, wherein The object detection model is trained using a labeled dataset. According to the requirements of the monitoring object, the roLabellmg annotation tool is used to annotate the dataset with the objects to be detected.
6. The mobile vendor monitoring method based on rotating object detection and pedestrian statistics according to claim 1, wherein The object detection model is a convolutional neural network model built using pytorch.
Citation Information
Patent Citations
Vehicle monitoring and warning system
DE202022100944U1
Method, and System for Vehicle Recognition Tracking Based on Deep Learning
KR102373753B1