Video monitoring image processing method and system

By using surveillance cameras in scenic spots to collect videos and perform frame extraction processing and image recognition, and identify and remind tourists of uncivilized behavior, the problem of low detection efficiency in the existing technology is solved, and efficient detection of violations and timely reminders are achieved.

CN120147966AActive Publication Date: 2025-06-13JIANGSU FENGPAN TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510262866.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-13
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

In the prior art, scenic spots have low efficiency in detecting uncivilized behaviors of tourists and cannot promptly remind them.

Method used

By obtaining videos collected by the scenic spot surveillance camera, using frame extraction and image recognition technology to identify violations, extract the characteristics of the perpetrator and send them to the reminder robot for target tracking and reminding.

Benefits of technology

It improves the detection efficiency of violations, realizes timely reminders to tourists, reduces detection costs, and extends the service life of reminder robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147966A_ABST
    Figure CN120147966A_ABST
Patent Text Reader

Abstract

The invention discloses a video monitoring image processing method and system, and belongs to the technical field of image processing and computer vision. The method comprises the following steps: acquiring a monitoring video acquired by a monitoring camera installed at a preset position of a scenic spot; frame extraction processing is carried out on the monitoring video at a preset time interval to obtain a monitoring image, and the preset time interval is related to the visitor flow corresponding to the preset position; identifying illegal behaviors in the monitoring image by using an image identification technology; if the violation behavior belongs to the first category, performing forward identification on the monitoring video according to a time node corresponding to a monitoring image where the violation behavior is located so as to determine an actor feature corresponding to the violation behavior; if the illegal behavior belongs to the second category, extracting actor characteristics in the monitoring image; and sending the actor characteristics to a reminding robot, so that the reminding robot performs target tracking and violation reminding on the actor based on the actor characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of image processing and computer vision, and particularly to a method and system for processing video surveillance images. Background Art

[0002] With the continuous improvement of people's living standards, there are more and more travel plans, and the number of tourists received by each scenic area is also increasing continuously.

[0003] The increasing number of tourists received by scenic areas year by year brings new challenges to the environmental management of scenic areas. For example, some scenic areas choose to completely ban smoking for tourists because there are many factors such as trees and flowers that are prone to fires; and in order to maintain the environment of the scenic area, tourists are usually prohibited or reminded not to litter. However, a small number of tourists still have uncivilized behaviors such as smoking and littering, which have an adverse impact on the environment of the scenic area and the sightseeing environment of other tourists. Therefore, it is particularly important to promptly discover and remind tourists of their uncivilized behaviors.

[0004] Currently, some scenic areas collect surveillance videos of tourists through installed cameras and arrange staff to observe whether there are uncivilized behaviors in these surveillance videos. However, this method has a low detection efficiency and cannot promptly remind tourists; some scenic areas also use patrol robots to detect behaviors such as tourists' smoking and littering and promptly remind them. However, this method has a high input cost, and the patrol routes of the patrol robots and the sightseeing routes of the uncivilized behavior perpetrators may not be the same, resulting in a low detection efficiency. Summary of the Invention

[0005] This application provides a method and system for processing video surveillance images, which are used to solve the technical problems that the current detection efficiency of uncivilized behaviors of tourists in each scenic area is low and cannot be promptly reminded.

[0006] This application adopts the following technical solutions:

[0007] On the one hand, a method for processing video surveillance images, the method includes: obtaining a surveillance video collected by a surveillance camera installed at a preset position in a scenic area; performing frame extraction on the surveillance video at a preset time interval to obtain surveillance images, where the preset time interval is related to the tourist flow rate corresponding to the preset position; using image recognition technology to identify illegal behaviors existing in the surveillance images; if the illegal behavior belongs to the first category, then based on the time node corresponding to the surveillance image where the illegal behavior is located, performing forward recognition on the surveillance video to determine the perpetrator characteristics corresponding to the illegal behavior; if the illegal behavior belongs to the second category, then extracting the perpetrator characteristics from the surveillance image; and sending the perpetrator characteristics to a reminder robot, so that the reminder robot performs target tracking and illegal reminder on the perpetrator based on the perpetrator characteristics.

[0008] In a possible implementation manner of the present application, before identifying the illegal behaviors existing in the monitoring image, the method further includes: inputting the monitoring image into a pre-trained face target annotation model to use the pre-trained face target annotation model to annotate the tourist faces existing in the monitoring image; extracting tourist face slices in the monitoring image according to the annotation result; identifying the illegal behaviors belonging to the second category through the tourist face slices, and the illegal behaviors belonging to the second category at least include smoking behaviors; performing image segmentation processing on the remaining part of the monitoring image after extracting the tourist face slices to obtain a ground slice corresponding to the monitoring image; identifying the illegal behaviors belonging to the first category through the ground slice, and the illegal behaviors belonging to the first category at least include littering behaviors.

[0009] In a possible implementation manner of the present application, identifying the illegal behaviors belonging to the second category through the tourist face slices includes: correspondingly converting the tourist face slices into the YCrCb color space; calculating the mean value of the color layout of the tourist face slices in the Y channel in the YCrCb color space to construct a mean value sequence, and the length of the mean value sequence is equal to the number of the tourist face slices; performing forward difference on the mean value sequence to obtain a difference sequence corresponding to the tourist face slices; if the values of all elements in the difference sequence are within a preset fluctuation range, it is determined that there are no illegal behaviors belonging to the second category in the monitoring image; if there are two consecutive elements in the difference sequence whose values exceed the preset fluctuation range, extracting the common elements corresponding to the two consecutive elements whose values exceed the preset fluctuation range in the mean value sequence, and determining that the tourist face slices corresponding to the common elements have illegal behaviors belonging to the second category.

[0010] In a possible implementation manner of the present application, extracting the actor features in the monitoring image includes: extracting a tourist full-body slice corresponding to the tourist face slice with illegal behaviors belonging to the second category in the monitoring image; inputting the tourist full-body slice into a pre-trained tourist feature extraction model to extract the appearance features and / or body features corresponding to the tourist full-body slice, and the appearance features at least include clothes color, hat color, shoe color, and backpack color, and the body features at least include gender, age range, height, and hair length; constructing the actor features through the appearance features and / or body features.

[0011] In a possible implementation manner of the present application, performing image segmentation processing on the remaining part of the monitoring image after extracting the tourist face slices includes: obtaining the coordinates (x i , y i), i ranges from 1 to n, and n is the number of tourist face slices extracted according to the annotation results; determine the coordinates of the central pixel point (x i ,y i ) in y i The maximum value of y i The maximum value of is taken as the reference point, and a segmentation line L passing through the reference point and parallel to the upper and lower boundaries of the monitoring image is generated; the monitoring image is segmented using the segmentation line L, and the portion of the segmentation result containing the lower boundary of the monitoring image is determined as the ground slice.

[0012] In a possible implementation of the present application, identifying violations belonging to the first category through the ground slice includes: extracting a full body slice of the tourist corresponding to the face slice of the tourist from the ground slice, and removing the full body slice of the tourist from the ground slice to obtain a ground slice to be identified; processing the ground slice to be identified using a pre-trained garbage detection model to determine the garbage features to be identified corresponding to the ground slice to be identified; matching the garbage features to be identified with a garbage feature library, and calculating the cosine similarity between any element in the garbage feature library and the garbage features to be identified; and determining that there is a violation belonging to the first category in the surveillance image when the cosine similarity is greater than a preset similarity threshold.

[0013] In a possible implementation of the present application, the surveillance video is forward identified based on the time node corresponding to the surveillance image where the illegal behavior is located to determine the characteristics of the person who behaves in the illegal behavior, including: taking the time node corresponding to the surveillance image as the starting point and the preset time interval as the total frame extraction time, the surveillance video is forward framed, and the frame extraction frequency is one frame per second; the reverse frame sequence obtained by frame extraction is input into a pre-trained garbage detection model to generate a garbage reverse trajectory in the frame sequence; based on the garbage reverse trajectory, the whole body slice of the tourist corresponding to the illegal behavior belonging to the first category is determined in the reverse frame sequence; and the feature extraction of the whole body slice of the tourist is performed using the pre-trained tourist feature extraction model to obtain the characteristics of the person.

[0014] In a possible implementation manner of the present application, the reminder robot performs target tracking on the actor based on the actor characteristics, including: the reminder robot travels to a preset position where the monitoring camera is installed based on the received actor characteristics; calculates the time difference between the current time and the time node corresponding to the monitoring image with a violation, and calculates the travel route of the actor relative to the preset position corresponding to the monitoring camera according to the average walking speed of the human body and the time difference, and the travel route at least includes the route length; takes the scenic tour direction corresponding to the preset position as the travel direction, and tracks the actor on the travel route through the actor characteristics.

[0015] In a possible implementation manner of the present application, after the reminder robot performs target tracking and violation reminder on the actor based on the actor characteristics, the method further includes: processing the real-time monitoring video frame sequence corresponding to the actor through a built-in edge computing module to determine whether the violation behavior corresponding to the actor disappears, and if so, returning to the nearest robot docking point; wherein, the edge computing module is built with a tourist behavior judgment model, and the tourist behavior judgment module adopts a lightweight network and can at least identify the turning behavior and smoking behavior of the actor.

[0016] On the other hand, the present application also provides a video surveillance image processing system, the system includes: an acquisition module, which acquires the surveillance video collected by the surveillance camera installed at a preset position in the scenic area; the acquisition module performs frame extraction processing on the surveillance video at a preset time interval to obtain surveillance images, and the preset time interval is related to the tourist flow corresponding to the preset position; an identification module, which uses image recognition technology to identify the violation behaviors existing in the surveillance images; the identification module, if the violation behavior belongs to the first category, then performs forward identification on the surveillance video according to the time node corresponding to the surveillance image where the violation behavior is located to determine the actor characteristics corresponding to the violation behavior; the identification module, if the violation behavior belongs to the second category, then extracts the actor characteristics from the surveillance image; a sending module, which sends the actor characteristics to the reminder robot so that the reminder robot performs target tracking and violation reminder on the actor based on the actor characteristics.

[0017] The video surveillance image processing method and system provided by the present application have the following beneficial effects:

[0018] This application uses the surveillance videos collected by the original cameras in the scenic area. After extracting frames from the surveillance videos, image recognition technology is used to identify illegal behaviors in the scenic area. On the one hand, no additional costs are incurred, reducing the detection cost of illegal behaviors. On the other hand, using image recognition technology, such as face detection models, etc., compared with the traditional method of staff viewing surveillance videos, it can significantly improve the detection efficiency of illegal behaviors, and the situations of missed detection and misdetection can be effectively avoided, and the labor cost is also reduced, achieving the efficient detection of illegal behaviors of scenic area tourists without additional costs.

[0019] After detecting an illegal behavior, the perpetrator of the illegal behavior is locked by extracting the perpetrator's features, and a reminder robot is used to perform target tracking and reminder on the perpetrator. This not only realizes the timely reminder of the perpetrator of the illegal behavior but also avoids the problems of low detection efficiency or untimely reminder caused by the inconsistent routes between the reminder robot and the perpetrator of the illegal behavior. At the same time, compared with traditional patrol robots, the reminder robot in this application does not need to perform multi-route patrols. It only needs to perform target tracking and reminder according to the extracted perpetrator's features. By reducing the usage frequency of the reminder robot and extending the service life of the robot, the detection cost of illegal behaviors is further reduced.

[0020] Furthermore, an edge computing module is installed on the reminder robot of this application, and a lightweight network is deployed on it. Compared with traditional patrol robots, the network cost used is also significantly reduced. Moreover, there is no need to upload video frames, which also avoids the frame rate drop caused during the video frame encoding and decoding process, which may affect the detection efficiency of illegal behaviors. Brief Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0022] Figure 1 It is a flowchart of a method for processing video surveillance images provided by this application;

[0023] Figure 2 It is an architecture diagram of a system for processing video surveillance images provided by this application. Detailed Embodiments

[0024] To enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in this application in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0025] The following will detail the method in this application through the accompanying drawings.

[0026] Figure 1 It is a flowchart of a method for processing video surveillance images provided by this application. As Figure 1 shown, the image processing method in this application at least includes the following execution steps:

[0027] Step 101: Obtain the surveillance video collected by the surveillance cameras installed at the preset positions in the scenic area.

[0028] The surveillance video image processing method in this application is applied to the scenario of identifying illegal behaviors in the scenic area. It aims to process the video stream data collected by the original cameras in the scenic area, process the video images, identify the existing illegal behaviors, so as to achieve the efficient detection of illegal behaviors in the scenic area without increasing additional equipment costs, improve the detection efficiency, and use the reminder robot for timely reminder.

[0029] First, obtain the surveillance video collected by the surveillance cameras in the scenic area. These surveillance cameras are installed at the preset positions in the scenic area, and the preset positions are preferably the positions that need to be key monitored in the scenic area, such as narrow roads, etc.

[0030] Step 102: Perform frame extraction processing on the surveillance video at preset time intervals to obtain surveillance images.

[0031] After receiving the surveillance video uploaded by the surveillance cameras, perform frame extraction processing on the surveillance video at preset time intervals. For example, at an interval of 5 seconds, extract one frame every 5 seconds. After the extraction is completed, obtain the surveillance images. This application uses the method of frame extraction to obtain the surveillance images, rather than analyzing and processing each frame. On the one hand, it can reduce the number of image frames to be analyzed and processed, improve the analysis and processing efficiency of the images, that is, improve the analysis efficiency of the surveillance video, so as to more quickly identify illegal behaviors and improve the detection efficiency of illegal behaviors. On the other hand, since illegal behaviors such as smoking and littering in the scenic area have a certain time persistence. For example, when the illegal behavior of smoking occurs, it generally lasts for several minutes. At the same time, when the illegal behavior of littering occurs, the garbage on the ground will also remain stationary for a period of time. Therefore, even if the method of frame extraction is used to analyze the surveillance video, it can ensure the effective detection and identification of illegal behaviors.

[0032] In one example, the time interval for frame extraction from the monitored video may be related to the tourist flow at the preset location where the monitoring camera is located. For example, when the tourist flow at the preset location is large, it means that the probability of occurrence of illegal behaviors is high. At this time, the frame extraction frequency can be increased, that is, the preset time interval is shortened. Moreover, when the tourist flow is large, increasing the frame extraction frequency also helps to reduce the occurrence of missed detections. When the tourist flow at the preset location is small, it means that the probability of occurrence of illegal behaviors is low. At this time, the frame extraction frequency can be decreased, that is, the preset time interval is relatively increased.

[0033] Step 103: Use image recognition technology to identify illegal behaviors existing in the monitored images.

[0034] In a possible implementation manner of the present application, after obtaining the monitored images, first input the monitored images into a pre-trained face target annotation model. Using this model, the faces of tourists existing in the monitored images can be annotated. In one example, the annotation result here is preferably in the form of a rectangular annotation box. That is, after the monitored images processed by this model are output, they are images with several annotation boxes, and each annotation box includes at least one face. Then, use these annotation boxes to intercept the monitored images to obtain tourist face slices. Through the tourist face slices, identify the illegal behaviors belonging to the second category. Here, the illegal behaviors of the second category refer to smoking behaviors.

[0035] The present application uses tourist face slices to identify smoking behaviors because when tourists are smoking, due to the influence of the diffusion of smoke, the number of white pixel points around the face will increase significantly. Using this feature to identify smoking behaviors can, compared with the solutions for detecting cigarette sticks or the burning red dots of cigarette sticks, better ensure the accuracy of identification or detection, thereby improving the detection efficiency. Moreover, the present application detects smoking behaviors by adopting the following solution, avoiding the detection method of extracting features using a complex model, saving computing power resources, thereby also helping to reduce the recognition cost of illegal behaviors and enabling rapid detection.

[0036] Specifically, the tourist face slices intercepted from the surveillance image are correspondingly converted to the YCrCb color space. In the YCrCb color space, the white pixel points around the face due to the influence of smoke diffusion are more obvious in the Y channel. Therefore, calculate the mean value of the color layout of the tourist face slices converted to the YCrCb color space in the Y channel to obtain a mean value sequence. In the mean value sequence, each element corresponds to a tourist face slice. Therefore, the length of the mean value sequence here is equal to the number of tourist face slices. After that, perform forward differencing on the obtained mean value sequence. When differencing to the last element, difference the last element and the first element to obtain a difference sequence. In the difference sequence, if the values of all elements are within the preset fluctuation range, it means that there is no tourist face slice with a more obvious Y channel, that is, it can be considered that there is no smoking violation. However, if there are two consecutive elements in the obtained difference sequence whose values exceed the preset fluctuation range, then according to the common differencing object corresponding to these two elements, determine the mean value corresponding to both of them in the mean value sequence, and determine that there is a smoking violation in the tourist face slice corresponding to this mean value. This is because if a tourist has a smoking violation in the scenic area, then there is at least one tourist face slice among the extracted tourist face slices whose color layout in the Y channel is abnormal, that is, the corresponding mean value is abnormal. And an abnormal mean value in the mean value sequence will cause two consecutive difference results to be abnormal. Therefore, the tourist face slice with a smoking violation can be determined through the difference sequence.

[0037] Further, after determining that there is a violation belonging to the second category in the surveillance image through the above process, the following process can be used to identify whether there is a violation belonging to the first category in the surveillance image:

[0038] First, perform image segmentation on the surveillance image from which the tourist face slices have been extracted, and extract the ground slice corresponding to the surveillance image from the segmentation result. In a possible implementation manner of the present application, in the surveillance image with several annotation frames output by the face target annotation model, extract the central pixel point coordinates (x i , y i ) of each annotation frame, where i ranges from 1 to n, and n is the number of annotation frames existing in the surveillance image, that is, the number of tourist face slices extracted subsequently; then, taking the maximum value of y i as the reference point, generate a segmentation line L passing through this reference point and parallel to the upper and lower boundaries of the surveillance image; the segmentation line L divides the surveillance image into upper and lower parts, and the lower half part containing the lower boundary is determined as the ground slice. In this process, taking y iThe maximum value of is the reference point, which can include all the tourists collected in the monitoring image in the ground slice. Therefore, it can be considered that the ground slice can observe all the garbage throwing behaviors of tourists, which are the violations of the first category, and it is also convenient to find the corresponding violators based on the detected garbage.

[0039] Then, in the ground slice, the face annotation frame is extended to the full body annotation frame of the tourist. In one example, the extension here refers to the extension of the annotation frame horizontally and vertically. For example, the annotation frame is extended horizontally by 1 / 2 of the width of the annotation frame on the left and right sides, and extended vertically by 6 to 8 times the height of the annotation frame, so that the extended annotation frame can include the full body of the tourist. The full body slice of the tourist is removed from the ground slice using the extended annotation frame to obtain the ground slice to be identified.

[0040] Finally, the ground slice to be identified is input into the pre-trained garbage detection model, and the garbage that may exist in the ground slice to be identified is detected, so as to determine the garbage features to be identified corresponding to the ground slice to be identified; the garbage features to be identified are matched with the garbage feature library, and the cosine similarity between any element in the garbage feature library and the garbage features to be identified is calculated; when the cosine similarity is greater than the preset similarity threshold, it is determined that there is a violation belonging to the first category in the monitoring image, that is, there is a behavior of littering. In this process, the model is used for feature extraction and feature comparison to realize garbage detection, thereby realizing the detection of violations belonging to the first category, which can not only ensure the accuracy of the detection, but also use the existing model to realize it, saving the training cost of the model, thereby realizing efficient detection of violations at a low cost.

[0041] Step 104: If the illegal behavior belongs to the first category, forward recognition is performed on the surveillance video according to the time node corresponding to the surveillance image where the illegal behavior is located to determine the characteristics of the person corresponding to the illegal behavior.

[0042] Further, after determining that there is a violation of littering in the surveillance image, the surveillance video is forward framed with the time node corresponding to the surveillance image as the starting point. When extracting frames, it is preferred to extract one frame every 1 second, and the total frame extraction time is the preset time interval. For example, the surveillance image is extracted by extracting 1 frame every 5 seconds. Then, when there is a violation belonging to the second category in this surveillance image, the surveillance image is forward framed with the time node corresponding to the surveillance image as the starting point, and the frame extraction frequency is 1 frame per second to obtain a reverse frame sequence. This frame extraction method will neither cause too many frames to be extracted, which will bring pressure on image recognition to the process of identifying the actor, nor will it cause too few frames to be extracted, making it impossible to identify the actor. It should be noted that when garbage is detected in the surveillance image, it means that the actor discarded the garbage between two frames of surveillance images. Therefore, extracting a reverse frame sequence between two frames of surveillance images can ensure the identification of the actor.

[0043] Afterwards, the extracted reverse frame sequence is input into the aforementioned pre-trained garbage detection model to identify garbage in each frame of the image, so as to obtain the reverse trajectory of the garbage. This reverse trajectory refers to the trajectory of the garbage from the ground to the hands of the perpetrator. Through this trajectory, the full-body slice of the tourist corresponding to the illegal act of littering can be determined. Moreover, during the process of the garbage detection model detecting garbage in the reverse frame sequence, the model can be adjusted to take the extracted garbage features as input and search for the garbage with such features in the reverse frame sequence, thereby reducing the difficulty of garbage detection.

[0044] Finally, the full-body slice of the tourist is input into the pre-trained tourist feature extraction model for feature extraction to obtain the corresponding perpetrator features. This process can refer to the process of extracting the perpetrator features of the smoking behavior, and the two are the same or similar, so this application will not elaborate here.

[0045] Step 105: If the illegal act belongs to the second category, extract the perpetrator features from the surveillance image.

[0046] Furthermore, after obtaining the face slice of the tourist with a smoking violation, the full-body slice of the tourist corresponding to this face slice is extracted from the original surveillance image. This process can be achieved by simple annotation box coding. That is, when annotating the tourists' faces in the surveillance image, each annotation box corresponding to a face is numbered or coded so that the tourists' faces and the numbers or codes of the annotation boxes are uniquely corresponding, and this number or code is correspondingly migrated to the intercepted face slice of the tourist. Then, the face slice of the tourist carrying the code is processed and the illegal act is identified, so that the mean sequence and / or difference sequence used in the identification process also correspondingly carry the aforementioned code or number. Thus, when there is an illegal act in a certain face slice of the tourist, the corresponding full-body slice of the tourist can be found in the surveillance image through its corresponding number or code. The obtained full-body slice of the tourist is input into the pre-trained tourist feature extraction model, and the appearance features and / or human body features corresponding to the full-body slice of the tourist are extracted by this feature extraction model to construct the perpetrator features. In one example, the appearance features extracted here at least include the clothes color, hat color, shoe color, and backpack color, and the human body features extracted at least include gender, age range, height, and hair length. The types of features here are only for illustrative purposes and are not limited in quantity or type. In actual use, as long as the subsequent reminder robot can lock the corresponding perpetrator based on these features.

[0047] It should be noted that the face target annotation model, garbage detection model, and tourist feature extraction model used in the foregoing process can all be obtained by training existing neural network models or deep learning models, and their training processes can all be implemented through existing algorithms. The focus of this application is on how to use these models to finally detect violations in the scenic area, and the training process of the models is not the key concern. Therefore, even if the process of obtaining or training the models is ignored in this application, those skilled in the art can obtain and use these models through corresponding algorithms or technical means.

[0048] Step 106: Send the actor's features to the reminder robot so that the reminder robot can perform target tracking and violation reminder on the actor based on the actor's features.

[0049] Whether there is a violation belonging to the first category or a violation belonging to the second category in the monitored image, the corresponding actor's features will be sent to the reminder robot.

[0050] In a possible implementation manner of this application, after receiving the actor's features, the reminder robot travels to the preset position corresponding to the monitoring camera based on the identifier carried by the monitoring camera that sent the actor's features to it, calculates the time difference between the current time and the time node corresponding to the monitored image with a violation, and then uses the average walking speed of the human body and this time difference to calculate the travel route of the violator relative to the preset position corresponding to the monitoring camera. The travel route at least includes the route direction and the route length. The route direction is the guided tour direction at the preset position corresponding to the monitoring camera, generally opposite to the video acquisition direction of the monitoring camera. Finally, the reminder robot moves forward along this travel route to track the actor.

[0051] During the process of the reminder robot tracking the actor, the collected tourist images are subjected to feature extraction and feature matching with the actor's features, and the corresponding actor is found based on the matching result. Of course, the reminder robot can either issue a voice reminder during the process of finding the actor or perform a voice reminder during the target tracking process after finding the actor to remind the actor to end or remedy the violation.

[0052] In a possible implementation manner of this application, during the process of the reminder robot performing target tracking on the actor, it collects the video corresponding to the actor in real time and processes the real-time monitored video frame sequence corresponding to the actor through the built-in edge computing module to determine whether the violation behavior corresponding to the actor has disappeared. If so, it returns to the nearest robot docking point. The edge computing module is built with a tourist behavior judgment model, and this tourist behavior judgment module uses a lightweight network and can at least identify the actor's turning behavior (that is, it is considered that the actor has a behavior of picking up garbage) and smoking behavior.

[0053] In this application, a pre-trained model is used to process the monitored images obtained by frame extraction, identify possible violations in the monitored images, and after successfully identifying a violation, use a reminder robot embedded with an edge computing module to perform target tracking and reminder on the perpetrator, and determine whether the violation has disappeared. On the premise of ensuring the efficient implementation of scenic area violations, timely reminder and supervision to eliminate violations after reminder are also achieved. At the same time, compared with traditional inspection robots, the reminder robot in this application uses a lightweight network for its built-in computing module, which can directly perform image processing and calculation. On the one hand, it reduces the computing resources occupied by image processing, and on the other hand, it can achieve edge computing without uploading the video for detection, which helps to achieve instant detection and improve efficiency.

[0054] Based on the same inventive concept, this application also provides a processing system for video monitoring images, and its architecture is as Figure 2 shown.

[0055] Figure 2 It is an architecture diagram of a processing system for video monitoring images provided by this application. As Figure 2 shown, the processing system 200 for video monitoring images in this application specifically includes: an acquisition module 201, which acquires the monitored video collected by a monitoring camera installed at a preset position in the scenic area; the acquisition module 201 performs frame extraction on the monitored video at a preset time interval to obtain monitored images, and the preset time interval is related to the tourist flow corresponding to the preset position; an identification module 202, which uses image recognition technology to identify violations existing in the monitored images; the identification module 202, if the violation belongs to the first category, then performs forward identification on the monitored video according to the time node corresponding to the monitored image where the violation is located to determine the perpetrator characteristics corresponding to the violation; the identification module 202, if the violation belongs to the second category, then extracts the perpetrator characteristics from the monitored image; a sending module 203, which sends the perpetrator characteristics to the reminder robot so that the reminder robot performs target tracking and violation reminder on the perpetrator based on the perpetrator characteristics.

[0056] Each embodiment in this application is described in a progressive manner. The same or similar parts between each embodiment can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0057] The system and method provided by this application correspond one by one. Therefore, the system also has beneficial technical effects similar to those of its corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system will not be elaborated here.

[0058] Those skilled in the art should understand that the embodiments of this application can be provided as methods, devices, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0059] It should also be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the element.

[0060] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A method for processing video surveillance images, characterized in that: The method comprises: Obtain surveillance videos collected by surveillance cameras installed at preset locations in the scenic area; Performing frame extraction processing on the monitoring video at a preset time interval to obtain a monitoring image, wherein the preset time interval is related to the flow of tourists corresponding to the preset location; Using image recognition technology to identify illegal behaviors in the monitoring image; If the violation belongs to the first category, forward recognition is performed on the surveillance video according to the time node corresponding to the surveillance image where the violation occurs, so as to determine the characteristics of the person who committed the violation; If the illegal behavior belongs to the second category, extracting the characteristics of the person who committed the behavior from the monitoring image; The actor's characteristics are sent to a reminder robot, so that the reminder robot can track the actor and remind him of violations based on the actor's characteristics.

2. The method for processing video surveillance images according to claim 1, characterized in that: Before identifying the illegal behavior in the monitoring image, the method further includes: Inputting the surveillance image into a pre-trained face target annotation model, so as to annotate the faces of tourists in the surveillance image using the pre-trained face target annotation model; Extracting the facial slices of the tourists from the surveillance images according to the labeling results; identifying a second category of violations through the facial slice of the tourist, wherein the second category of violations at least includes smoking; Performing image segmentation processing on the remaining part of the monitoring image after the facial slice of the tourist is extracted, so as to obtain a ground slice corresponding to the monitoring image; Violations belonging to a first category are identified through the ground slice, where the violations of the first category at least include littering.

3. The method for processing video surveillance images according to claim 2, characterized in that: Violations belonging to the second category are identified through the facial slice of the visitor, including: Convert the tourist face slices to YCrCb color space; In the YCrCb color space, the mean of the color layout of the tourist face slices in the Y channel is calculated to construct a mean sequence, wherein the length of the mean sequence is equal to the number of the tourist face slices; Performing forward differentiation on the mean sequence to obtain a differential sequence corresponding to the tourist face slice; If the values ​​of the elements in the differential sequence are all within the preset fluctuation range, it is determined that there is no violation belonging to the second category in the monitoring image; If there are two consecutive elements whose values ​​exceed the preset fluctuation range in the differential sequence, the common element corresponding to the two consecutive elements whose values ​​exceed the preset fluctuation range in the mean sequence is extracted, and it is determined that the tourist face slice corresponding to the common element has a violation belonging to the second category.

4. The method for processing video surveillance images according to claim 3, characterized in that: Extracting the behavioral person's features from the monitoring image includes: Extracting, from the monitoring image, a whole body slice of a tourist corresponding to a face slice of a tourist who has committed a violation belonging to the second category; Inputting the whole body slice of the tourist into a pre-trained tourist feature extraction model to extract appearance features and / or body features corresponding to the whole body slice of the tourist, wherein the appearance features at least include clothing color, hat color, shoe color and backpack color, and the body features at least include gender, age range, height and hair length; The actor characteristics are constructed by using the appearance characteristics and / or body characteristics.

5. The method for processing video surveillance images according to claim 2, characterized in that: Performing image segmentation processing on the remaining part of the monitoring image after extracting the tourist face slice, including: Obtain the center pixel coordinates (x i ,y i ), i ranges from 1 to n, and n is the number of tourist face slices extracted according to the annotation results; Determine the center pixel coordinates (x i ,y i ) in y i The maximum value of According to the y i The maximum value of is the reference point, generating a segmentation line L passing through the reference point and parallel to the upper and lower boundaries of the monitoring image; The monitoring image is segmented using the segmentation line L, and a portion of the segmentation result that includes the lower boundary of the monitoring image is determined as the ground slice.

6. A method for processing video surveillance images according to claim 5, characterized in that: Violations belonging to the first category are identified through the ground slice, including: Extracting the whole body slice of the tourist corresponding to the face slice of the tourist from the ground slice, and removing the whole body slice of the tourist from the ground slice to obtain the ground slice to be identified; Processing the to-be-identified ground slice using a pre-trained garbage detection model to determine the to-be-identified garbage features corresponding to the to-be-identified ground slice; Matching the to-be-identified garbage feature with a garbage feature library, and calculating the cosine similarity between any element in the garbage feature library and the to-be-identified garbage feature; When the cosine similarity is greater than a preset similarity threshold, it is determined that there is a violation belonging to the first category in the monitoring image.

7. A method for processing video surveillance images according to claim 6, characterized in that: According to the time node corresponding to the surveillance image where the illegal behavior occurs, the surveillance video is forward identified to determine the characteristics of the person who acts in accordance with the illegal behavior, including: Taking the time node corresponding to the monitoring image as the starting point and the preset time interval as the total frame extraction time, the monitoring video is frame extracted forward, and the frame extraction frequency is one frame per second; Inputting the reverse frame sequence obtained by extracting frames into a pre-trained garbage detection model to generate a garbage reverse trajectory in the frame sequence; According to the garbage reverse trajectory, determining the whole body slice of the tourist corresponding to the violation behavior belonging to the first category in the reverse frame sequence; The pre-trained tourist feature extraction model is used to extract features from the tourist's whole body slices to obtain the actor's features.

8. The method for processing video surveillance images according to claim 1, characterized in that: The reminding robot performs target tracking on the actor based on the actor's characteristics, including: The reminder robot drives to a preset position where the surveillance camera is installed based on the received characteristics of the actor; Calculate the time difference between the current time and the time node corresponding to the surveillance image with the violation, and calculate the driving route of the person relative to the preset position corresponding to the surveillance camera based on the average walking speed of the human body and the time difference, wherein the driving route at least includes the route length; The scenic spot sightseeing direction corresponding to the preset position is used as the driving direction, and the actor is tracked on the driving route according to the actor's characteristics.

9. The method for processing video surveillance images according to claim 8, characterized in that: After the reminder robot performs target tracking and violation reminder on the actor based on the actor's characteristics, the method further includes: The real-time monitoring video frame sequence corresponding to the actor is processed through the built-in edge computing module to determine whether the illegal behavior corresponding to the actor has disappeared. If so, the robot returns to the nearest stop point; Among them, the edge computing module has a built-in tourist behavior judgment model, and the tourist behavior judgment module adopts a lightweight network and can at least identify the turning behavior and smoking behavior of the actor.

10. A video surveillance image processing system, characterized in that: The system comprises: An acquisition module is used to acquire surveillance videos collected by surveillance cameras installed at preset locations in the scenic area; The acquisition module extracts frames from the surveillance video at a preset time interval to obtain a surveillance image, wherein the preset time interval is related to the flow of visitors corresponding to the preset location; An identification module, using image recognition technology to identify illegal behaviors in the monitoring image; The recognition module, if the illegal behavior belongs to the first category, performs forward recognition on the surveillance video according to the time node corresponding to the surveillance image where the illegal behavior occurs, so as to determine the characteristics of the person who acts according to the illegal behavior; The recognition module extracts the characteristics of the person in the monitoring image if the illegal behavior belongs to the second category; The sending module sends the actor's characteristics to the reminder robot, so that the reminder robot can track the actor and remind him of violations based on the actor's characteristics.

Citation Information

Patent Citations

  • Garbage throwing behavior detection method based on urban management monitoring video

    CN111611970A

  • Image recognition technology-based forest fire automatic monitoring and recognition system and method

    CN111666834A

  • Behavior detection method and device, electronic equipment and storage medium

    CN113408464A

  • Video monitoring method and device, computer equipment and storage medium

    CN114173094A

  • Office place-based illegal behavior detection system and method

    CN114913452A