Video surveillance control methods, devices, electronic equipment and computer-readable media

By using frame-by-frame processing of the camera device and detection by a pre-trained model, the problem of poor timeliness of manual identification in video surveillance has been solved, and the effect of timely identification and alarm of abnormal targets has been achieved.

CN119181039BActive Publication Date: 2025-10-31ADDX (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411204842.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-10-31
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

In existing video surveillance control systems, the timeliness of manual identification of abnormal targets is poor, making it difficult to identify abnormal targets in a timely manner.

Method used

In response to monitoring commands, the system uses video capture devices to process and segment the video into frames, determines pixel change values, adjusts the resolution, performs image preprocessing, inputs the data into a pre-trained video anomaly detection model for detection, and issues an alarm sound.

Benefits of technology

It improves the timeliness and accuracy of video surveillance, enabling timely identification and alerts for abnormal targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181039B_ABST
    Figure CN119181039B_ABST
Patent Text Reader

Abstract

This disclosure discloses video surveillance control methods, apparatuses, electronic devices, and computer-readable media. One specific implementation of the method includes: controlling a camera device to capture video of a target area according to a first video resolution; performing frame-by-frame processing on the first area surveillance video to obtain a first area surveillance frame image sequence; determining the pixel change value between every two frames in the first area surveillance frame image sequence; in response to determining that there are pixel change values ​​in the pixel change value sequence that are greater than or equal to a target change value, and the last pixel change value is greater than or equal to the target change value, controlling the camera device to capture video of the target area according to a second video resolution; and inputting the pre-processed surveillance video image set into a pre-trained video abnormal target detection model to obtain a video abnormal target detection result set. This implementation improves the timeliness of video surveillance and can promptly identify abnormal targets in the surveillance video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computers, and more specifically to video surveillance control methods, apparatus, electronic devices, and computer-readable media. Background Technology

[0002] Identifying anomalous targets within a target area (e.g., a user smoking in a non-smoking area, an animal entering the area, or a fire breaking out) and issuing alerts can improve the security of that area. Currently, the common method for video surveillance control of target areas is to manually identify anomalous targets in the video.

[0003] However, when using the above methods for video surveillance control, the following technical problems often exist: manual identification is not timely and it is difficult to identify abnormal targets in a timely manner.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide video surveillance control methods, apparatuses, electronic devices, and computer-readable media to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a video surveillance control method, the method comprising: responding to receiving a monitoring instruction for a target area, controlling an associated camera device to capture video of the target area according to a first video resolution to obtain a first area monitoring video; performing frame segmentation processing on the first area monitoring video to obtain a first area monitoring frame image sequence; determining the pixel change value between every two frames in the first area monitoring frame image sequence to obtain a pixel change value sequence; responding to determining that there are pixel change values ​​in the pixel change value sequence that are greater than or equal to a target change value, and the last pixel change value is greater than or equal to the target change value, controlling the camera device according to a second video resolution. The system acquires video of the target area to obtain a second area monitoring video; it preprocesses each frame of the monitoring video included in the second area monitoring video to generate a preprocessed monitoring video image set; it inputs the preprocessed monitoring video image set into a pre-trained video abnormal target detection model to obtain a video abnormal target detection result set, wherein the preprocessed monitoring video images in the preprocessed monitoring video image set correspond to the video abnormal target detection results in the video abnormal target detection result set; in response to determining that there are video abnormal target detection results in the video abnormal target detection result set that meet the target detection conditions, it controls the associated alarm device to issue an alarm prompt sound.

[0008] Secondly, some embodiments of this disclosure provide a video surveillance control device, comprising: a first control unit configured to, in response to receiving a monitoring instruction for a target area, control an associated camera device to capture video of the target area according to a first video resolution, thereby obtaining a first area monitoring video; a framing unit configured to perform framing processing on the first area monitoring video to obtain a first area monitoring frame image sequence; a determining unit configured to determine the pixel change value between every two frames in the first area monitoring frame image sequence, thereby obtaining a pixel change value sequence; and a second control unit configured to, in response to determining that there exists a pixel change value in the pixel change value sequence that is greater than or equal to a target change value, and that the last pixel change value is greater than or equal to the target change value, control the camera device to capture video of the target area according to the second video resolution. The aforementioned camera device acquires video of the target area to obtain a second area monitoring video; the preprocessing unit is configured to preprocess each frame of the monitoring video included in the second area monitoring video to generate a preprocessed monitoring video image, thus obtaining a preprocessed monitoring video image set; the input unit is configured to input the preprocessed monitoring video image set into a pre-trained video abnormal target detection model to obtain a video abnormal target detection result set, wherein the preprocessed monitoring video images in the preprocessed monitoring video image set correspond to the video abnormal target detection results in the video abnormal target detection result set; the third control unit is configured to control the associated alarm device to issue an alarm sound in response to determining that there is a video abnormal target detection result in the video abnormal target detection result set that meets the target detection conditions.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above-described embodiments of this disclosure have the following beneficial effects: the video surveillance control method of some embodiments of this disclosure improves the timeliness and accuracy of video surveillance, and can promptly identify abnormal targets in the surveillance video. Specifically, the reason why it is difficult to identify abnormal targets in a timely manner is that manual identification is not timely enough. Based on this, the video surveillance control method of some embodiments of this disclosure firstly, in response to receiving a monitoring instruction for a target area, controls the associated camera device to capture video of the target area according to a first video resolution, thereby obtaining a first area monitoring video. Thus, in the initial state, video capture can be performed using a lower video resolution to reduce the waste of video resources. Secondly, the first area monitoring video is processed into frames to obtain a first area monitoring frame image sequence; the pixel change value between every two frames in the first area monitoring frame image sequence is determined to obtain a pixel change value sequence. Thus, it can be confirmed whether a new target (user / animal / firelight) has appeared in the target area. Next, in response to determining that there are pixel changes in the pixel change value sequence that are greater than or equal to the target change value, and the last pixel change value is greater than or equal to the target change value, the camera device is controlled to capture video of the target area according to the second video resolution, thus obtaining a second area monitoring video. Therefore, when a new target is determined to exist in the target area, the resolution of the video capture can be increased for clearer identification. Then, each frame of the monitoring video image included in the second area monitoring video is preprocessed to generate a preprocessed monitoring video image set. This preprocessed monitoring video image set is input into a pre-trained video anomaly target detection model to obtain a video anomaly target detection result set, where the preprocessed monitoring video images in the preprocessed monitoring video image set correspond to the video anomaly target detection results in the video anomaly target detection result set. Thus, the pre-trained target detection model can perform real-time detection of video images. Finally, in response to determining that there are video anomaly target detection results in the video anomaly target detection result set that meet the target detection conditions, the associated alarm device is controlled to issue an alarm sound. This improves the timeliness and accuracy of video surveillance, enabling timely identification of abnormal targets in the surveillance video and prompting timely alarms. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1This is a flowchart of some embodiments of the video surveillance control method according to the present disclosure;

[0014] Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the video surveillance control device according to this disclosure;

[0015] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0019] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0022] Figure 1 This is a flowchart of some embodiments of the video surveillance control method according to the present disclosure. A flowchart 100 of some embodiments of the video surveillance control method according to the present disclosure is shown. The video surveillance control method includes the following steps:

[0023] Step 101: In response to receiving a monitoring instruction for the target area, control the associated camera device to capture video of the target area according to the first video resolution, and obtain the first area monitoring video.

[0024] In some embodiments, the execution entity of the video surveillance control method may, in response to receiving a monitoring instruction for a target area, control an associated camera device to capture video of the target area according to a first video resolution, thereby obtaining a first area surveillance video. The target area may refer to the area to be monitored. For example, the target area may refer to an infrastructure area, a building area, or a public place. The camera device may refer to a camera with surveillance video recording capabilities. The monitoring instruction may be an instruction to monitor the target area. The first video resolution may refer to a pre-set resolution for capturing surveillance video. For example, the associated camera device may be controlled to capture approximately 10 seconds of video of the target area in real time to obtain the first area surveillance video.

[0025] Step 102: Perform frame segmentation processing on the above-mentioned first area monitoring video to obtain the first area monitoring frame image sequence.

[0026] In some embodiments, the aforementioned executing entity may perform frame-by-frame processing on the aforementioned first area monitoring video to obtain a first area monitoring frame image sequence. The first area monitoring video can be split into individual frames to obtain the first area monitoring frame image sequence.

[0027] Step 103: Determine the pixel change value between every two frames in the first area monitoring frame image sequence to obtain the pixel change value sequence.

[0028] In some embodiments, the execution entity can determine the pixel change value between every two frames in the first region monitoring frame image sequence to obtain a pixel change value sequence. For example, the pixel change value between every two frames can be calculated using an image subtraction formula.

[0029] Step 104: In response to determining that there are pixel change values ​​in the pixel change value sequence that are greater than or equal to the target change value, and the last pixel change value is greater than or equal to the target change value, the camera device is controlled to capture video of the target area according to the second video resolution to obtain the second area monitoring video.

[0030] In some embodiments, the executing entity may, in response to determining that there are pixel changes in the pixel change value sequence that are greater than or equal to a target change value, and that the last pixel change value is greater than or equal to the target change value, control the camera device to capture video of the target area according to a second video resolution, thereby obtaining a second area monitoring video. The second video resolution is greater than the first video resolution. That is, a high resolution can be used to control the camera device to capture video of the target area to obtain a second area monitoring video.

[0031] Step 105: Perform image preprocessing on each frame of the surveillance video included in the second area surveillance video to generate preprocessed surveillance video images, thus obtaining a set of preprocessed surveillance video images.

[0032] In some embodiments, the execution entity may perform image preprocessing on each frame of the surveillance video included in the second area surveillance video to generate preprocessed surveillance video images, thereby obtaining a set of preprocessed surveillance video images.

[0033] In practice, the aforementioned implementing entity can perform image preprocessing on each frame of the surveillance video included in the second area surveillance video through the following steps:

[0034] The first step is to perform grayscale processing on the aforementioned surveillance video image information to generate a grayscale surveillance video image. The executing entity can use a preset grayscale algorithm to perform grayscale processing on the surveillance video image information to generate a grayscale surveillance video image. For example, the preset grayscale algorithm could be: component method, maximum value method, average value method, or weighted average method.

[0035] The second step involves interpolating the aforementioned grayscale surveillance video image to generate an interpolated grayscale surveillance video image. The executing entity can use a preset interpolation algorithm to perform the interpolation process on the grayscale surveillance video image to generate the interpolated grayscale surveillance video image. For example, the preset interpolation algorithm could be: nearest neighbor interpolation algorithm, bilinear interpolation algorithm, or bicubic interpolation algorithm.

[0036] The third step involves denoising the interpolated grayscale surveillance video image to generate a denoised interpolated grayscale surveillance video image, which serves as the preprocessed surveillance video image. The executing entity can use a preset denoising algorithm to denoise the interpolated grayscale surveillance video image, generating the denoised interpolated grayscale surveillance video image as the preprocessed surveillance video image. For example, the preset denoising algorithm could be: non-local mean algorithm, median filtering algorithm, or Gaussian low-pass filtering algorithm.

[0037] Step 106: Input the above preprocessed surveillance video image set into the pre-trained video abnormal target detection model to obtain the video abnormal target detection result set.

[0038] In some embodiments, the execution entity can input the preprocessed surveillance video image set into a pre-trained video anomaly target detection model to obtain a video anomaly target detection result set. The preprocessed surveillance video images in the preprocessed surveillance video image set correspond to the video anomaly target detection results in the video anomaly target detection result set. The video anomaly target detection model can be a pre-trained target detection model that takes preprocessed surveillance video images as input and outputs video anomaly target detection results. For example, the video anomaly target detection results can indicate whether user or target information exists in the preprocessed surveillance video images. Target information can represent image information such as firelight or animals. For example, the video anomaly target detection model can be a YOLOv4 model, a YOLOv5 model, an SSD (Single Shot MultiBoxDetector) model, or a DETR (Detection Transformer) model. The YOLOv4 model improves target detection performance by improving network structure, loss function, and data augmentation techniques. YOLOv4 excels in small target detection and dense target detection due to its high accuracy and good localization capabilities. YOLOv5, the latest version in the YOLO series, incorporates the concepts of EfficientDet and PANet to achieve a lightweight network structure and multi-level feature fusion, improving the efficiency and accuracy of object detection. The SSD model predicts multiple bounding boxes and class probabilities on feature maps at different scales, enabling rapid detection of objects at various scales.

[0039] The video anomaly detection model can be trained through the following steps:

[0040] The first step is to obtain a training sample set of surveillance video images. This training sample set includes both sample surveillance video images and sample labels.

[0041] The second step is to determine the initial video anomaly detection model. This initial model includes an initial feature extraction network, an initial user identification network, and an initial target object identification network. The initial feature extraction network includes an initial convolutional model sequence, and the initial convolutional networks within this sequence include initial sub-convolutional models and initial activation models. Here, the input to the initial convolutional sub-models within the initial convolutional model sequence can be the output of the initial activation model included in the previous initial convolutional model. The input to the first initial convolutional sub-model in the initial convolutional model sequence can be a sample surveillance video image. The input to the initial activation model in the initial convolutional model sequence can be the output of the initial convolutional sub-models included in the initial convolutional model sequence. For example, the number of initial convolutional models in the initial convolutional model sequence can be 22. The initial user identification network is a model that takes the features of the initial surveillance video image as input and outputs initial user identification information. The initial user identification network is used to: filter the initial surveillance video image using a preset filtering algorithm to generate initial user identification information. For example, the preset filtering algorithm could be the NMS (Non-Maximum Suppression) algorithm. The initial target object recognition network could be a model that takes the features of the initial surveillance video image as input and outputs initial object recognition information. The initial target object recognition network is used to: filter the initial surveillance video image using the preset filtering algorithm to generate initial object recognition information. For example, the preset filtering algorithm could be the NMS (Non-Maximum Suppression) algorithm.

[0042] The third step is to select a target surveillance video image training sample from the aforementioned surveillance video image training sample set. One surveillance video image training sample can be randomly selected from the aforementioned surveillance video image training sample set as the target surveillance video image training sample.

[0043] The fourth step is to input the sample surveillance video images, including the target surveillance video image training samples, into the initial feature extraction network to obtain the initial surveillance video image features.

[0044] The fifth step is to input the initial monitoring video image features into the initial user identification network to obtain the initial user identification information.

[0045] The sixth step is to input the initial monitoring video image features into the initial target object recognition network to obtain the initial target object recognition information.

[0046] Step 7: Based on a preset loss function, determine the loss value between the initial user identification information, the initial target object identification information, and the corresponding sample labels. The aforementioned loss function includes: a first user identification loss function, a second user identification loss function, a first object identification loss function, and a second object identification loss function. For example, the first user identification loss function, the second user identification loss function, the first object identification loss function, and the second object identification loss function can be the mean squared error loss function (MSE), the hinge loss function (SVM), the cross-entropy loss function, the 0-1 loss function, the absolute value loss function, the logarithmic loss function, the squared loss function, the exponential loss function, etc.

[0047] The seventh step mentioned above may include the following sub-steps:

[0048] The first sub-step involves determining the first user identification loss value between the initial user identification information and the user identification sub-labels in the sample labels, based on the aforementioned first user identification loss function.

[0049] The first user identification loss function can be:

[0050]

[0051] Where loss1 represents the loss value for the first user identification. λ coord This represents the preset first balancing parameter (e.g., 3). S represents the number of rows in all grid cells (e.g., 5). Here, the number of columns in the grid cells is the same as the number of rows. B represents the number of all prediction boxes (e.g., 2). i represents the grid cell number. j represents the prediction box number. This indicates whether the j-th prediction box in the i-th grid cell is responsible for predicting the initial user identification information. When the value is 1, it means that the j-th prediction box in the i-th grid cell is responsible for predicting the initial user identification information. When the value is 0, it indicates that the j-th prediction box in the i-th grid cell is not responsible for predicting the initial user identification information. i This represents the x-coordinate of the user-identified sub-label in the i-th grid cell. y represents the x-coordinate of the initial user identification information predicted in the i-th grid cell. i This represents the ordinate of the user-identified sub-label in the i-th grid cell. This represents the ordinate of the initial user identification information predicted in the i-th grid cell.

[0052] The second sub-step involves determining the second user identification loss value between the initial user identification information and the user identification sub-labels in the sample labels, based on the aforementioned second user identification loss function.

[0053] The second user identification loss function can be:

[0054]

[0055] Where loss3 represents the second user identification loss value. C i This represents the confidence level corresponding to the user-identified sub-label in the i-th grid cell. This represents the confidence level corresponding to the initial user identification information predicted in the i-th grid cell.

[0056] The third sub-step involves generating a user identification loss value based on the first user identification loss value and the second user identification loss value.

[0057] The fourth sub-step involves determining the first object recognition loss value between the initial target object recognition information and the object recognition sub-labels in the sample labels, based on the first object recognition loss function described above.

[0058] The fifth sub-step involves determining the second object recognition loss value between the initial target object recognition information and the object recognition sub-labels in the sample labels, based on the aforementioned second object recognition loss function.

[0059] The sixth sub-step is to generate an object recognition loss value based on the first object recognition loss value and the second object recognition loss value.

[0060] The seventh sub-step involves weighting and summing the user recognition loss value and the object recognition loss value to obtain the aforementioned loss value.

[0061] Step 8: In response to the determination that the loss value is less than or equal to the preset loss value, the above initial video anomaly detection model is determined as the trained video anomaly detection model.

[0062] Therefore, a relatively accurate video anomaly detection model can be trained by including an initial feature extraction network, an initial user recognition network, an initial target object recognition network, and four loss functions. Consequently, a relatively accurate video anomaly detection model can predict relatively accurate video anomaly detection results.

[0063] Optionally, in response to the determination that the loss value is greater than the above-mentioned preset loss value, the model parameters of the initial video abnormal target detection model are adjusted, the target monitoring video image training sample is selected from the unselected monitoring video image training samples, and the adjusted initial video abnormal target detection model is trained again.

[0064] In some embodiments, the execution entity may, in response to determining that the loss value is greater than the preset loss value, adjust the model parameters of the initial video anomaly target detection model, select target surveillance video image training samples from the unselected training samples, and retrain the adjusted initial video anomaly target detection model. For example, the difference between the aforementioned loss value and the preset loss value can be calculated. Based on this, the network parameters of the initial video anomaly target detection model can be adjusted using methods such as backpropagation and gradient descent. The setting of the preset difference value is not limited; for example, the preset difference value can be 0.1.

[0065] Step 107: In response to determining that there are video abnormal target detection results that meet the target detection conditions in the above video abnormal target detection result set, control the associated alarm device to issue an alarm prompt sound.

[0066] In some embodiments, the execution entity may, in response to determining that there is a video abnormal target detection result in the set of video abnormal target detection results that meets the target detection conditions, control an associated alarm device to issue an alarm sound. The alarm device may be a speaker for issuing the alarm sound. For example, the alarm sound may indicate that an abnormal target has appeared in the target area.

[0067] Optionally, in response to determining that there are no video abnormal target detection results that meet the above target detection conditions in the above video abnormal target detection result set, the video acquisition resolution of the above camera device is adjusted to the above first video resolution.

[0068] In some embodiments, the execution entity may adjust the video acquisition resolution of the camera device to the first video resolution in response to determining that there is no video abnormal target detection result in the video abnormal target detection result set that meets the target detection conditions.

[0069] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a video surveillance control device, which are similar to... Figure 1 Corresponding to the method embodiments shown, this video surveillance control device can be specifically applied to various electronic devices.

[0070] like Figure 2As shown, a video surveillance control device 200 in some embodiments includes: a first control unit 201, a framing unit 202, a determination unit 203, a second control unit 204, a preprocessing unit 205, an input unit 206, and a third control unit 207. The first control unit 201 is configured to, in response to receiving a monitoring command for a target area, control an associated camera device to capture video of the target area according to a first video resolution, thereby obtaining a first area monitoring video; the framing unit 202 is configured to perform framing processing on the first area monitoring video to obtain a first area monitoring frame image sequence; the determination unit 203 is configured to determine the pixel change value between every two frames in the first area monitoring frame image sequence, thereby obtaining a pixel change value sequence; the second control unit 204 is configured to, in response to determining that there are pixel change values ​​in the pixel change value sequence that are greater than or equal to a target change value, and that the last pixel change value is greater than or equal to the target change value, control the camera device to capture video of the target area according to a second video resolution. The system acquires video from a second area; a preprocessing unit 205 is configured to preprocess each frame of the video from the second area to generate a preprocessed video image set; an input unit 206 is configured to input the preprocessed video image set into a pre-trained video anomaly detection model to obtain a video anomaly detection result set, wherein the preprocessed video images in the preprocessed video image set correspond to the video anomaly detection results in the video anomaly detection result set; and a third control unit 207 is configured to control an associated alarm device to issue an alarm sound in response to determining that there is a video anomaly detection result in the video anomaly detection result set that meets the target detection conditions.

[0071] It is understandable that the units described in the video surveillance control device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the video surveillance control device 200 and the units contained therein, and will not be repeated here.

[0072] The following is for reference. Figure 3 This document illustrates a schematic diagram of an electronic device 300 (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0073] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0074] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0075] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 309, or installed from a storage device 308, or installed from a ROM 302. When the computer program is executed by the processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0076] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0077] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0078] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: in response to receiving a monitoring instruction for a target area, control an associated camera device to capture video of the target area according to a first video resolution, obtaining a first area monitoring video; perform frame segmentation processing on the first area monitoring video to obtain a first area monitoring frame image sequence; determine the pixel change value between every two frames in the first area monitoring frame image sequence, obtaining a pixel change value sequence; in response to determining that there are pixel change values ​​in the pixel change value sequence that are greater than or equal to a target change value, and the last pixel change value is greater than or equal to the target change value, execute the video capture of the target area according to a second video resolution. The system controls the aforementioned camera device to capture video of the target area, obtaining a second area monitoring video; it preprocesses each frame of the monitoring video included in the second area monitoring video to generate a preprocessed monitoring video image set; it inputs the preprocessed monitoring video image set into a pre-trained video abnormal target detection model to obtain a video abnormal target detection result set, wherein the preprocessed monitoring video images in the preprocessed monitoring video image set correspond to the video abnormal target detection results in the video abnormal target detection result set; in response to determining that there are video abnormal target detection results in the video abnormal target detection result set that meet the target detection conditions, it controls the associated alarm device to issue an alarm sound.

[0079] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0081] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor can be described as including: a first control unit, a framing unit, a determining unit, a second control unit, a preprocessing unit, an input unit, and a third control unit. The names of these units do not necessarily limit the specific unit; for example, a framing unit can also be described as "a unit that performs framing processing on the aforementioned first area monitoring video to obtain a first area monitoring frame image sequence."

[0082] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0083] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A video surveillance control method, comprising: In response to receiving a monitoring instruction for a target area, the associated camera device is controlled to capture video of the target area according to a first video resolution, thereby obtaining a first area monitoring video; The first area monitoring video is processed by frame segmentation to obtain the first area monitoring frame image sequence; Determine the pixel change value between every two frames in the first region monitoring frame image sequence to obtain a pixel change value sequence; In response to determining that there are pixel change values ​​in the pixel change value sequence that are greater than or equal to the target change value, and the last pixel change value is greater than or equal to the target change value, the camera device is controlled to capture video of the target area according to the second video resolution to obtain a second area monitoring video, wherein the second video resolution is greater than the first video resolution; Each frame of the surveillance video included in the second area is preprocessed to generate a preprocessed surveillance video image set. The preprocessed surveillance video image set is input into a pre-trained video abnormal target detection model to obtain a video abnormal target detection result set, wherein the preprocessed surveillance video images in the preprocessed surveillance video image set correspond to the video abnormal target detection results in the video abnormal target detection result set. In response to determining that there are video abnormal target detection results that meet the target detection conditions in the set of video abnormal target detection results, the associated alarm device is controlled to issue an alarm prompt sound; The loss value of the video abnormal target detection model is determined through the following steps: Based on the first user identification loss function, determine the first user identification loss value between the initial user identification information and the user identification sub-labels in the sample labels; Based on the second user identification loss function, determine the second user identification loss value between the initial user identification information and the user identification sub-labels in the sample labels; A user identification loss value is generated based on the first user identification loss value and the second user identification loss value; Based on the first object recognition loss function, determine the first object recognition loss value between the initial target object recognition information and the object recognition sub-label in the sample label; Based on the second object recognition loss function, the second object recognition loss value between the initial target object recognition information and the object recognition sub-label in the sample label is determined; An object recognition loss value is generated based on the first object recognition loss value and the second object recognition loss value; The user recognition loss value and the object recognition loss value are weighted and summed to obtain the loss value.

2. The method according to claim 1, wherein, The method further includes: In response to determining that there are no abnormal video target detection results in the set of abnormal video target detection results that meet the target detection conditions, the video acquisition resolution of the camera device is adjusted to the first video resolution.

3. The method according to claim 1, wherein, The step of preprocessing each frame of the surveillance video included in the second area surveillance video to generate a preprocessed surveillance video image includes: The surveillance video image information is processed into grayscale to generate a grayscale surveillance video image; The grayscale surveillance video image is interpolated to generate an interpolated grayscale surveillance video image; The interpolated grayscale monitoring video image is subjected to noise reduction processing to generate a noise-reduced interpolated grayscale monitoring video image, which serves as a preprocessed monitoring video image.

4. The method according to claim 1, wherein, Before inputting the preprocessed surveillance video image set into a pre-trained video anomaly detection model to obtain a video anomaly detection result set, the method further includes: Obtain a training sample set of surveillance video images, wherein the training samples of the surveillance video images in the training sample set include: sample surveillance video images and sample labels; An initial video abnormal target detection model is determined, wherein the initial video abnormal target detection model includes: an initial feature extraction network, an initial user recognition network, and an initial target object recognition network. The initial feature extraction network includes: an initial convolutional model sequence, and the initial convolutional network in the initial convolutional network sequence includes: an initial subconvolutional model and an initial activation model. Select target surveillance video image training samples from the surveillance video image training sample set; The sample surveillance video images included in the target surveillance video image training samples are input into the initial feature extraction network to obtain the initial surveillance video image features; The initial monitoring video image features are input into the initial user identification network to obtain initial user identification information; The initial monitoring video image features are input into the initial target object recognition network to obtain the initial target object recognition information; Based on a preset loss function, the loss value between the initial user identification information, the initial target object identification information, and the corresponding sample labels is determined; In response to determining that the loss value is less than or equal to a preset loss value, the initial video anomaly detection model is determined as the trained video anomaly detection model.

5. The method according to claim 4, wherein, The method further includes: In response to the determination that the loss value is greater than the preset loss value, the model parameters of the initial video abnormal target detection model are adjusted, the target monitoring video image training sample is selected from the unselected monitoring video image training samples, and the adjusted initial video abnormal target detection model is trained again.

6. A video surveillance control device, comprising: The first control unit is configured to, in response to receiving a monitoring instruction for a target area, control an associated camera device to capture video of the target area according to a first video resolution, thereby obtaining a first area monitoring video; The framing unit is configured to perform frame segmentation processing on the first area monitoring video to obtain a first area monitoring frame image sequence. The determining unit is configured to determine the pixel change value between every two frames in the first region monitoring frame image sequence, and obtain a pixel change value sequence. The second control unit is configured to, in response to determining that there are pixel changes in the pixel change value sequence that are greater than or equal to the target change value, and the last pixel change value is greater than or equal to the target change value, control the camera device to capture video of the target area according to the second video resolution, so as to obtain a second area monitoring video, wherein the second video resolution is greater than the first video resolution; The preprocessing unit is configured to perform image preprocessing on each frame of the surveillance video included in the second area surveillance video to generate preprocessed surveillance video images, thereby obtaining a set of preprocessed surveillance video images. The input unit is configured to input the preprocessed surveillance video image set into a pre-trained video abnormal target detection model to obtain a video abnormal target detection result set, wherein the preprocessed surveillance video images in the preprocessed surveillance video image set correspond to the video abnormal target detection results in the video abnormal target detection result set; The third control unit is configured to control the associated alarm device to issue an alarm sound in response to determining that there is a video abnormal target detection result that meets the target detection conditions in the set of video abnormal target detection results; The loss value of the video abnormal target detection model is determined through the following steps: Based on the first user identification loss function, determine the first user identification loss value between the initial user identification information and the user identification sub-labels in the sample labels; Based on the second user identification loss function, determine the second user identification loss value between the initial user identification information and the user identification sub-labels in the sample labels; A user identification loss value is generated based on the first user identification loss value and the second user identification loss value; Based on the first object recognition loss function, determine the first object recognition loss value between the initial target object recognition information and the object recognition sub-label in the sample label; Based on the second object recognition loss function, the second object recognition loss value between the initial target object recognition information and the object recognition sub-label in the sample label is determined; An object recognition loss value is generated based on the first object recognition loss value and the second object recognition loss value; The user recognition loss value and the object recognition loss value are weighted and summed to obtain the loss value.

7. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.

8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Target area safety monitoring method and device, electronic equipment and readable medium

    CN118334594A