Method and device for realizing detection of left objects in a monitoring scenario

By acquiring the background model pixel by pixel in the surveillance video, and filtering candidate areas with the biological historical trajectory, the accuracy of legacy items detection in complex monitoring scenarios is solved, and efficient legacy items detection is achieved.

CN110348327BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910547197.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-24
Publication Date
2025-06-27
Estimated Expiration
2040-03-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect items left in fire escapes in complex monitoring scenarios, resulting in missed inspections and missed inspections.

Method used

By obtaining the background model of the isotopic pixel points in the previous frame of the monitoring single-frame picture in the monitoring video, calculating the motion foreground pixels, and filtering the candidate areas of the legacy items according to the biological historical trajectory to obtain the accurate retention area of ​​the legacy items.

Benefits of technology

It realizes accurate detection of items left over in surveillance video, reduces the occurrence of false detection and missed detection, and is suitable for various complex surveillance scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110348327B_ABST
    Figure CN110348327B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a method and apparatus for implementing detection of left-behind items in a monitoring scenario. The method includes: obtaining, for each pixel point in a monitored single-frame picture in a monitoring video, a background model of the same-position pixel point in the previous frame picture; calculating, according to the obtained background model, moving foreground pixels in the monitored single-frame picture, where the monitored single-frame picture includes the moving foreground pixels and stationary background pixels; locating, by the moving foreground pixels, a left-behind item candidate area in the monitored single-frame picture; and filtering the left-behind item candidate area according to the historical trajectory of organisms in the monitoring video to obtain the area where the left-behind item stays. The technical solution of the embodiments of the present invention can accurately detect left-behind items in a monitoring scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video surveillance, and in particular, to a method and device for realizing the detection of left-behind items in a surveillance scenario. Background Art

[0002] A fire passage refers to a passage for firefighters to carry out rescues and evacuate trapped people, such as stairwells and aisles, which plays an inestimable role in various emergencies. Therefore, it is necessary to ensure that the fire passage is always unobstructed.

[0003] Surveillance cameras are usually installed in fire passages, and security personnel with patents are deployed to monitor the surveillance videos in real time to facilitate the timely discovery and removal of left-behind items stacked in the fire passages.

[0004] In some scenarios of automatic surveillance of fire passages, an artificial neural network model is usually used to monitor the surveillance videos corresponding to the fire passages. When left-behind items are detected in the fire passage, an alarm is automatically triggered, and relevant personnel are notified to remove the left-behind items. Therefore, there is no need to deploy surveillance personnel, saving security costs.

[0005] However, when training an artificial neural network model, a large number of surveillance videos need to be obtained and the left-behind items appearing in the surveillance videos need to be calibrated as training data, resulting in a very complicated training process. And considering that the actual surveillance scenarios are very complex, for example, the environmental light changes drastically under different fire passages, and the types of left-behind items stacked are numerous, it is difficult to train an artificial neural network model suitable for various complex scenarios, resulting in situations such as missed detection and false detection in the detection of left-behind items in the fire passage by the artificial neural network model.

[0006] Therefore, it is urgent to solve the technical problem that the left-behind items appearing in the surveillance video cannot be accurately detected in the existing implementation. Summary of the Invention

[0007] To solve the above technical problems, embodiments of the present invention provide a method and device for realizing the detection of left-behind items in a surveillance scenario, an electronic device, and a computer-readable storage medium.

[0008] Among them, the technical solution adopted by the present invention is as follows:

[0009] A method for realizing the detection of left objects in a monitoring scenario, comprising: obtaining, for each pixel point of a single monitored frame image in a monitoring video, the background model of the corresponding pixel point in the previous frame image, where the background model is used to simulate the background pixel change of the corresponding pixel point in the previous frame image in the spatial domain; calculating, according to the obtained background model, the moving foreground pixels in the single monitored frame image, where the single monitored frame image includes the moving foreground pixels and the static background pixels; positioning, by the moving foreground pixels, a candidate area for the left object in the single monitored frame image; and filtering the candidate area for the left object according to the historical trajectory of organisms in the monitoring video to obtain the area where the left object stays.

[0010] A device for realizing the detection of left objects in a monitoring scenario, comprising: a background model obtaining module, configured to obtain, for each pixel point of a single monitored frame image in a monitoring video, the background model of the corresponding pixel point in the previous frame image in the spatial domain, where the background model is used to simulate the background pixel change of the corresponding pixel point in the previous frame image; a moving foreground calculating module, configured to calculate, relative to the background model, the moving foreground pixels in the single monitored frame image, where the single monitored frame image includes the moving foreground pixels and the static background pixels; a candidate area positioning module, configured to position, by the moving foreground pixels, a candidate area for the left object in the single monitored frame image; and a candidate area filtering module, configured to filter the candidate area for the left object according to the historical trajectory of organisms in the monitoring video to obtain the area where the left object stays.

[0011] An electronic device, comprising a processor and a memory, where computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the method for realizing the detection of left objects in a monitoring scenario as described above is implemented.

[0012] A computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the method for realizing the detection of left objects in a monitoring scenario as described above.

[0013] In the above technical solution, considering that left objects are carried into the monitoring scenario by humans or other organisms, and left objects appear as moving image areas in the monitoring video, therefore, by obtaining, for each pixel point of a single monitored frame image in the monitoring video, the background model of the corresponding pixel point in the previous frame image, and then calculating the moving foreground pixels in the single monitored frame image relative to the background model, the pixel points that may represent left objects in the single monitored frame image can be obtained, and thus the candidate area for the left object is located.

[0014] Since the lost item is carried to the monitoring scene by an organism, the historical trajectory of the organism will surely appear around the area where the lost item appears. Therefore, the candidate area of the lost item is filtered according to the historical trajectory of the organism in the monitoring video, and the obtained candidate area of the lost item is the accurate staying area of the lost item.

[0015] Compared with the prior art, the above technical solution is independent of the type of the lost item detected in the monitoring video. Filtering the candidate area of the lost item according to the historical trajectory of the organism in the monitoring video can also filter out the candidate area of the lost item misdetected due to environmental light changes, so as to achieve accurate detection of the lost item.

[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0018] Figure 1 is a schematic diagram of the implementation environment related to the present invention;

[0019] Figure 2 is a block diagram of a computer device shown according to an exemplary embodiment;

[0020] Figure 3 is a flowchart of a method for detecting lost items in a monitoring scene shown according to an exemplary embodiment;

[0021] Figure 4 is a flowchart of a method for detecting lost items in a monitoring scene shown according to another exemplary embodiment;

[0022] Figure 5 is for Figure 3 the flowchart of step 210 in the corresponding embodiment in one embodiment;

[0023] Figure 6 is for Figure 3 the flowchart of step 230 in the corresponding embodiment in one embodiment;

[0024] Figure 7 is for Figure 3 the flowchart of step 230 in the corresponding embodiment in another embodiment;

[0025] Figure 8 is for Figure 3 the flowchart of step 230 in the corresponding embodiment in another embodiment;

[0026] Figure 9 is a flowchart of a method for implementing detection of left-behind items in a monitoring scenario shown according to another exemplary embodiment;

[0027] Figure 10 is a flowchart of a method for implementing detection of left-behind items in a monitoring scenario shown according to another exemplary embodiment;

[0028] Figure 11 is a flowchart of a method for implementing detection of left-behind items in a monitoring scenario shown according to another exemplary embodiment;

[0029] Figure 12 is a specific implementation diagram of a method for implementing detection of left-behind items in a monitoring scenario in an application scenario;

[0030] Figure 13 is a block diagram of a device for implementing detection of left-behind items in a monitoring scenario shown according to an exemplary embodiment. Detailed implementation manners

[0031] Here, the exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0032] Figure 1 is a schematic diagram of an implementation environment involved in a method for implementing detection of left-behind items in a monitoring scenario. The implementation environment includes a computer device 100 and a monitoring camera 200.

[0033] Among them, the monitoring camera 200 is installed in a fire passage or any other location where left-behind item monitoring is required to perform real-time monitoring and obtain a monitoring video. Usually, multiple monitoring cameras 200 ( Figure 1 4 are shown in the figure) are installed at the monitoring location for comprehensive monitoring.

[0034] A communication connection is established in advance between the monitoring camera 200 and the computer device 100 to transmit the acquired monitoring video to the computer device 100 for detection of left-behind items.

[0035] In one embodiment, when the computer device 100 detects a left-behind item in the monitoring video, it will mark the area where the left-behind item appears and display the marked area of the left-behind item, facilitating relevant personnel to accurately locate the position of the left-behind item in the monitoring location, so as to clean up the left-behind item in a timely manner.

[0036] In another embodiment, when the computer device 100 detects a left-behind item in the surveillance video, it also issues an alarm to notify relevant personnel to clean up the left-behind item in a timely manner.

[0037] Figure 2 is a block diagram of a computer device shown according to an exemplary embodiment.

[0038] It should be noted that this computer device is only an example adapted to the present invention and cannot be considered as providing any limitation to the scope of use of the present invention. This computer device cannot be interpreted as requiring dependence on or necessarily having Figure 2 one or more components in the exemplary computer device 100 shown in.

[0039] As Figure 2 shown, the computer device 100 includes a processing component 101, a memory 102, a power supply component 103, a multimedia component 104, an audio component 105, a sensor component 107, and a communication component 108. Among them, the above components are not all necessary. The computer device 100 can add other components or reduce certain components according to its own functional requirements, and this embodiment does not make any limitations.

[0040] The processing component 101 generally controls the overall operation of the computer device 100, such as operations associated with display, data communication, and log data processing. The processing component 101 may include one or more processors 109 to execute instructions to complete all or part of the above operations. In addition, the processing component 101 may include one or more modules to facilitate the interaction between the processing component 101 and other components. For example, the processing component 101 may include a multimedia module to facilitate the interaction between the multimedia component 104 and the processing component 101.

[0041] The memory 102 is configured to store various types of data to support the operation of the computer device 100. Examples of these data include instructions for any application or method operating on the computer device 100. One or more modules are stored in the memory 102, and the one or more modules are configured to be executed by the one or more processors 109 to complete all or part of the steps in the method for detecting left-behind items in a surveillance scenario described in the following embodiments.

[0042] The power supply component 103 provides power for various components of the computer device 100. The power supply component 103 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the computer device 100.

[0043] The multimedia component 104 includes a screen that provides an output interface between the computer device 100 and the user. In some embodiments, the screen may include a TP (Touch Panel) and an LCD (Liquid Crystal Display). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.

[0044] The audio component 105 is configured to output and / or input audio signals. For example, the audio component 105 includes a microphone that is configured to receive external audio signals when the computer device 100 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 102 or transmitted via the communication component 108. In some embodiments, the audio component 105 further includes a speaker for outputting audio signals.

[0045] The sensor component 107 includes one or more sensors for providing a status assessment of various aspects of the computer device 10. For example, the sensor component 107 can detect the on / off state of the computer device 100 and can also detect changes in the temperature of the computer device 100.

[0046] The communication component 108 is configured to facilitate communication between the computer device 100 and other devices in a wired or wireless manner. The computer device 100 can access a wireless network based on a communication standard, such as WiFi (WIreless - Fidelity). In an exemplary embodiment, the communication component 108 receives broadcast signals or broadcast - related information from an external broadcast management system via a broadcast channel.

[0047] It can be understood that Figure 2 the structure shown is only schematic, and the computer device 100 may include more or fewer components than Figure 2 shown therein, or have components different from those Figure 2 shown. Figure 2 Each of the components shown therein can be implemented using hardware, software, or a combination thereof.

[0048] Please refer to Figure 3 , in an exemplary embodiment, a method for implementing the detection of left - behind items in a monitoring scenario is applicable to Figure 1 the computer device in the implementation environment shown, and the structure of the computer device can be as Figure 2 shown.

[0049] The method for detecting left-behind items in a monitoring scenario can be executed by a computer device and may include the following steps:

[0050] Step 210: Obtain the background model of the corresponding pixel points in the previous frame picture for each pixel point of the monitored single-frame picture in the monitoring video.

[0051] As mentioned above, for fire corridors or some other locations where stacking left-behind items is not allowed, real-time monitoring is usually required to avoid serious problems such as endangering human safety due to the illegal stacking of left-behind items.

[0052] In the existing implementation, a trained artificial neural network model is usually used to automatically detect left-behind items in the monitoring video to save labor costs. However, due to factors such as the complexity of the actual monitoring scenario, the variety of left-behind item types, and the frequent change of environmental illumination, it is difficult to train an artificial neural network model that adapts to different monitoring scenarios, resulting in the inability to accurately detect left-behind items.

[0053] Based on the property that left-behind items are brought into the monitoring location by humans or other organisms, this embodiment can conclude that left-behind items appear as moving image regions in the monitoring video. Therefore, by detecting the moving foreground of the monitoring video, the detected moving foreground can be obtained as the left-behind items appearing in the monitoring video, thus eliminating the need to train an artificial neural network model and making the detection of left-behind items more convenient.

[0054] In this embodiment, the detection of left-behind items (i.e., moving foreground) in the monitoring video is performed on the monitored single-frame picture of the monitoring video. If a picture area corresponding to a left-behind item is detected in the monitored single-frame picture, it indicates that a left-behind item appears in the monitoring video.

[0055] The detection of left-behind items performed on the monitored single-frame picture is carried out pixel by pixel and depends on the background model of the corresponding pixel points in the previous frame picture in the spatial domain. It should be understood that this spatial domain refers to the spatial domain corresponding to the pixel points, that is, the pixel domain, and pixel-level image superposition and other processes can be performed in the spatial domain. The corresponding pixel points refer to the pixel points with the same position in different monitored single-frame pictures. For example, the pixel points located in the nth row and the mth column in different monitored single-frame pictures are corresponding pixel points to each other, and the corresponding pixel points correspond to the same spatial domain.

[0056] For the background model of the same-position pixel points in the spatial domain in the previous frame of the picture, it is obtained by sampling the neighboring pixels of the same-position pixel points in the first frame of the monitoring video. After obtaining the background model of the same-position pixel points in the spatial domain in the first frame of the picture, in the subsequent detection of the left-behind items in the subsequent frames of the picture, if it is detected that the same-position pixel point is the pixel point corresponding to the static background, then it is obtained by updating the background model of the same-position pixel point and its adjacent pixel points in the previous frame of the picture according to this same-position pixel point.

[0057] That is to say, except for the background model of the same-position pixel points in the spatial domain in the first frame of the picture, the background models corresponding to the same-position pixel points in other frames of the picture are all updated based on the background model corresponding to the previous frame of the picture. Therefore, for the background model of the same-position pixel points in the previous frame of the picture obtained in this embodiment, it should contain some same-position pixel points and the adjacent pixel points of the same-position pixel points in the historical frame pictures, and these pixel points constitute the background pixel points of the same-position pixel points in the previous frame of the picture.

[0058] Thus, compared with the background model of the same-position pixel points in the spatial domain in the first frame of the picture, the background model of the same-position pixel points in the previous frame of the picture obtained in this embodiment can simulate the background pixel changes of the same-position pixel points in the previous frame of the picture in the spatial domain.

[0059] Step 230, calculate the moving foreground pixels in the monitored single-frame picture according to the obtained background model.

[0060] Among them, the moving foreground pixels in the monitored single-frame picture refer to the pixel points that appear as the moving foreground in the monitored single-frame picture, that is, the pixel points that appear as the left-behind items.

[0061] The monitored single-frame picture also correspondingly includes static background pixels, which refer to the pixel points that appear as the static background in the monitored single-frame picture, such as the walls and floors of the fire passage, etc., which appear as static or slowly moving image areas in the monitoring video.

[0062] As mentioned above, for the background model of the same-position pixel points in the previous frame of the picture obtained in step 210, it should contain the pixel values of some same-position pixel points and the pixel values of adjacent pixel points in the historical frame pictures as background pixel points. Therefore, for the pixel points in the monitored single-frame picture currently undergoing the detection of left-behind items, if its distribution characteristics are the same as those of the background pixel points included in the background model of the same-position pixel points in the previous frame of the picture, it means that this pixel point is the same as or similar to the image content represented by the same-position pixel points in the historical frame pictures. Therefore, it is determined that this pixel point is a static background pixel.

[0063] If the distribution characteristics of the pixel points in a single monitored frame image are different from those of the background pixel points included in the corresponding background model, it indicates that the image content represented by the pixel points is different from that of the corresponding pixel points in the historical frame images, and these pixel points are manifested as left-behind items.

[0064] Therefore, relative to the background model of the corresponding pixel points in the previous frame image in the spatial domain, by calculating the similarity between the pixel points in the single monitored frame image and the distribution characteristics of the background pixel points included in the corresponding background model pixel by pixel, the pixel points with a similarity reaching a certain threshold are determined as the moving foreground pixels in the single monitored frame image, while the pixel points with a similarity less than a certain threshold are determined as the static background pixels in the single monitored frame image.

[0065] Step 250: Locate the candidate areas of left-behind items from the moving foreground pixels in the single monitored frame image.

[0066] As mentioned above, the moving foreground pixels are the pixel points in the single monitored frame image that are manifested as left-behind items, and the pixel area connected by several moving foreground pixels in the single monitored frame image correspondingly represents the picture area corresponding to the left-behind items.

[0067] Therefore, according to the connection relationship between the moving foreground pixels, the candidate areas of left-behind items in the single monitored frame image can be located.

[0068] In an exemplary embodiment, considering that the left-behind items affecting the smoothness of the fire passage should have a certain size, by judging whether the area of the connected moving foreground pixels in the single monitored frame image reaches a set area threshold, the connected area reaching the area threshold is located as the candidate area of the left-behind item, thereby improving the positioning accuracy of the candidate area of the left-behind item.

[0069] Since the sizes of all pixel points in the single monitored frame image are the same, it is also possible to locate the connected area where the number of moving foreground pixels is greater than the set threshold by judging whether the number of connected moving foreground pixels in the single monitored frame image is greater than the set threshold as the candidate area of the left-behind item.

[0070] It can be concluded that compared with the prior art, in this embodiment, by calculating the moving foreground pixels pixel by pixel for the single monitored frame image, the candidate areas of left-behind items in the single monitored frame image can be located according to the obtained moving foreground pixels, without the need for complex artificial neural network model training and without being affected by the variety of left-behind items.

[0071] Step 270: Filter the candidate areas of left-behind items according to the historical trajectories of organisms in the monitored video to obtain the areas where the left-behind items stay.

[0072] Among them, the historical trajectories of organisms in the monitored video refer to the movement trajectories where organisms appear and continuously move in the monitored video.

[0073] Considering that the moving foreground pixel points constituting the candidate area of the left-behind item are obtained through the aforementioned moving foreground detection, these moving foreground pixels may also appear as moving organisms or other moving objects in the surveillance video, and may also be misjudged as moving foreground pixels due to the influence of environmental light on the stationary background pixels. Therefore, there is still a possibility of false detection in the candidate area of the left-behind item located according to step 250.

[0074] Still based on the property that the left-behind item is carried to the surveillance scene by an organism, there should be an organism historical trajectory around the area where the left-behind item appears. Thus, a strong association between the organism and the left-behind item is established. Filtering the candidate area of the left-behind item located in step 250 according to the organism historical trajectory in the surveillance video can largely remove the above-mentioned falsely detected candidate areas of the left-behind item, and the finally obtained candidate area of the left-behind item is the accurate staying area of the left-behind item.

[0075] Therefore, compared with the prior art, the method provided in this embodiment is independent of the type of the left-behind item detected in the surveillance video, and can also filter out the falsely detected candidate areas of the left-behind item due to environmental light changes according to the organism historical trajectory in the surveillance video, which can eliminate the influence of the type of the left-behind item and environmental light changes on the detection accuracy of the left-behind item, and achieve accurate detection of the left-behind item.

[0076] Please refer to Figure 4 , in an exemplary embodiment, before step 210, the method for implementing the detection of the left-behind item in the surveillance scene further includes the following steps:

[0077] Step 310, eliminating the influence of environmental light on each pixel point of the surveillance single-frame picture to obtain an update of the pixel value in the RGB color space where it is located.

[0078] Among them, due to different environmental lights, there will be a certain deviation between the color of the image collected by the surveillance camera and the real color, which is likely to lead to false detection of the moving foreground pixels and stationary background pixels in the surveillance single-frame picture, and will also affect the detection accuracy of the left-behind item to a certain extent.

[0079] In order to further eliminate the influence of environmental light on the detection of the left-behind item in the surveillance single-frame picture, before detecting the left-behind item, the influence of environmental light on each pixel point of each surveillance single-frame picture in the surveillance video can be eliminated in advance to restore the original scene of the surveillance single-frame picture. Based on the original scene of the surveillance single-frame picture, the recognition accuracy of each pixel point in the surveillance single-frame picture can be improved.

[0080] The process of eliminating the influence of ambient light on each pixel of a monitored single-frame image is essentially a process of updating the pixel values of each pixel in the monitored single-frame image in the RGB color space. The updated pixel value is the original pixel value of the corresponding pixel after removing the influence of ambient light. It should be understood that in existing implementations, image acquisition by a monitoring camera is all based on the RGB color space. Therefore, the monitored single-frame images in a monitoring video should also correspond to the RGB color space.

[0081] In an exemplary embodiment, the gray world algorithm can be used to eliminate the influence of ambient light on each pixel of a monitored single-frame image. The gray world algorithm assumes that the average value of the average reflection of natural scenes for ambient light is a fixed value overall, and this fixed value is approximately "gray". By forcibly applying this assumption to the monitored single-frame image, the influence of ambient light can be eliminated from the monitored single-frame image, and the original scene of the monitored single-frame image can be restored.

[0082] The process of using the gray world algorithm to eliminate the influence of ambient light on each pixel of a monitored single-frame image is as follows:

[0083] First, determine the average reflection mean value of ambient light Exemplarily, the average channel values of the monitored single-frame image in the R, G, and B channels can be calculated, and the average value of these three average channel values is taken as the average reflection mean value of ambient light, that is, It is also possible to directly obtain half of the maximum gray value that the monitored single-frame image can display as the average reflection mean value of ambient light, and there is no limitation here.

[0084] Then, calculate the gain coefficients of each pixel in the monitored single-frame image in the R, G, and B channels respectively. These gain coefficients are the ratios between the average reflection mean value of ambient light and the channel values of the current pixel. Among them, the gain coefficients are correspondingly expressed as

[0085] Finally, calculate the product of the channel values of each pixel in the monitored single-frame image and the gain coefficients of the corresponding channels, so as to obtain the original channel values of each pixel after removing the influence of ambient light. For each pixel in the monitored single-frame image, by updating the current channel value with the respective original channel values, the original pixel value after removing the influence of ambient light can be obtained.

[0086] In another exemplary embodiment, to accelerate the speed of eliminating the influence of ambient light on a monitored single-frame image and save computing resources, before eliminating the influence of ambient light on each pixel of the monitored single-frame image, a downsampling process is also performed on the monitored single-frame image to reduce the resolution of the monitored single-frame image. Exemplarily, the resolution of the monitored single-frame image can be reduced to 400×320.

[0087] Step 330: Convert the monitored single-frame image from the RGB color space to the HSV color space according to the updated pixel values.

[0088] Among them, the HSV color space is a representation method of points in the RGB color space in an inverted cone, including three channels: hue, saturation, and value. Its expression of colors is more similar to human color perception. Compared with the RGB color space, the HSV color space is more intuitive.

[0089] Thus, after eliminating the influence of ambient light on the monitored single-frame image, the monitored single-frame image is also converted from the RGB color space to the HSV color space, and the converted monitored single-frame image is used to perform the detection of left-behind items, so as to detect the left-behind items by simulating the real vision of humans, further improving the accuracy of detecting left-behind items in the monitored single-frame image.

[0090] In an exemplary embodiment, the process of converting the monitored single-frame image from the RGB color space to the HSV color space is as follows:

[0091] For each pixel point in the monitored single-frame image, first calculate the ratio between the color components in the RGB color space and the maximum gray value that the monitored single-frame image can display, that is, calculate R' = R / 255, G' = G / 255, B' = B / 255 respectively, and obtain the maximum ratio C max = max(R', G', B'), the minimum ratio C min = min(R', G', B'), and the difference Δ between the maximum ratio and the minimum ratio = C max - C min .

[0092] When the difference Δ is zero, the hue component of the pixel point in the HSV color space is 0°; when C max = R', the hue component of the pixel point is When C max = G', the hue component of the pixel point is When C max = B', the hue component of the pixel point is

[0093] When C max = 0, the saturation component of the pixel point in the HSV color space is zero; when C max ≠ 0, the saturation component of the pixel point is

[0094] The value component of the pixel point in the HSV color space is C max .

[0095] Therefore, according to the hue component, saturation component, and value component of the obtained pixel points in the HSV color space, replacing the R, G, and B channel values of the RGB color space where the pixel points are located can achieve the conversion of the pixel points from the RGB color space to the HSV color space.

[0096] Please refer to Figure 5 , in an exemplary embodiment, when the current single-frame monitoring picture for detecting left-behind items is the first frame picture in the monitoring video, before step 210, the above method further includes the following steps:

[0097] Step 211, for each pixel point in the single-frame monitoring picture, by randomly sampling the neighboring pixels of the pixel point, obtain the set of background pixel points of the pixel point in the spatial domain;

[0098] Step 213, obtain the set of background pixel points of the pixel point in the spatial domain to form a background model.

[0099] Among them, the neighborhood of a pixel point in the single-frame monitoring picture refers to a certain area around the current pixel point. For example, for a certain pixel point in the single-frame monitoring picture, its neighborhood can be the four sides of the pixel point or the upper area of the pixel point, and no limitation is imposed here.

[0100] The neighboring pixels of a pixel point in the single-frame monitoring picture refer to the pixel points within the range of the neighborhood where the current pixel point is located.

[0101] Assume that the pixel values of each pixel point and its neighboring pixels in the single-frame monitoring picture have a similar distribution in the spatial domain. Based on this assumption, each pixel point can be represented as a corresponding background model according to the pixel points in its neighborhood. However, it should be noted that in order to ensure that the background model conforms to statistical laws, the range of the neighborhood of the pixel points in the single-frame monitoring picture should be large enough to ensure that a certain scale of neighboring pixels can be sampled, that is, a certain scale of background pixel points can be sampled.

[0102] By randomly sampling the neighboring pixels of each pixel point in the first frame picture several times, the set of background pixel points of each pixel point in the spatial domain is obtained accordingly, and the background model of each pixel point in the spatial domain is formed by the sampled set of background pixel points.

[0103] In an exemplary embodiment, by randomly sampling the neighboring pixels of the pixel point P(x) in the first frame picture 8 times, the background pixel points P(1) to P(8) are obtained in sequence, and then the background model of the pixel point P(x) in the spatial domain is constituted by the background pixel points P(1) to P(8).

[0104] Thus, after obtaining the background model of each pixel point in the spatial domain of the first-frame image, the identification of left-behind objects can be performed pixel by pixel on the second-frame image. Exemplarily, for a certain pixel point in the second-frame image, the calculation of moving foreground pixels is performed on the background model of the corresponding pixel point in the spatial domain of the first-frame image. If it is obtained that the pixel point is a static background pixel, the static background pixel is used as a new background pixel point to update the background model of the corresponding pixel point in the first-frame image, and thus the background model corresponding to the second-frame image can be obtained. Repeat this process, and for each monitoring single-frame image for left-behind object detection, the left-behind object detection is performed according to the background model corresponding to the corresponding pixel point in the previous-frame image.

[0105] It can be concluded therefrom that in this embodiment, only the initialization process of the background model needs to be performed on the first-frame image in the monitoring video, that is, the random sampling process of neighborhood pixels is performed, and there is no need to spend a lot of time generating corresponding background models for each pixel point in each monitoring single-frame image, and the real-time performance of left-behind object detection for the monitoring video can be achieved.

[0106] Please refer to Figure 6 , in an exemplary embodiment, step 230 may include the following steps:

[0107] Step 231, corresponding to the color space where the monitoring single-frame image is located, by calculating the distances between the pixel points in the monitoring single-frame image and each background pixel point in the background model, obtain the background pixel points in the background model whose pixel values are similar to that of the pixel point.

[0108] Among them, for the background pixel points in the background model, the fact that their pixel values are similar to the pixel points in the monitoring single-frame image means that the picture content represented by the background pixel points is similar to that of the pixel points in the monitoring single-frame image, their pixel values should also be similar, and the distance between them in the color space should also be less than a certain threshold.

[0109] Therefore, after calculating the distances between each background pixel point in the background model and the pixel points in the monitoring single-frame image respectively, obtain the background pixel points with distances less than the set distance threshold as the similar background pixel points in the background model.

[0110] In an exemplary embodiment, in the HSV color space, since the hue channel is insensitive to environmental illumination, by calculating the distance between the background pixel point and the pixel point in the monitoring single-frame image in the hue channel to obtain the similarity, the influence of environmental illumination can be further eliminated to further improve the accuracy of left-behind object detection in the monitoring single-frame image.

[0111] In other embodiments, the spatial distance between the background pixel point and the pixel point in the monitoring single-frame image in the HSV color space can also be calculated, and this is not limited herein.

[0112] Step 233: When the number of similar background pixel points is less than the quantity threshold, obtain the pixel points in the monitored single-frame image as foreground moving pixels.

[0113] As described above, if the distribution characteristics of the current pixel point in the monitored single-frame image are different from those of the background pixel points included in the corresponding background model, it indicates that the image content represented by the current pixel point is different from that of the co-located pixel point in the historical frame image, and this pixel point is then manifested as a left-behind item.

[0114] When the number of background pixel points similar to the corresponding pixel point in the monitored single-frame image in the background model is less than the quantity threshold, it indicates that the distribution characteristics of the background pixel points included in the background model are different from those of the corresponding pixel point in the monitored single-frame image, and thus the pixel points in the monitored single-frame image are obtained as foreground moving pixels.

[0115] In this embodiment, by calculating the distances between the pixel points in the monitored single-frame image and each background pixel point in the corresponding background model, the relationship between the distribution characteristics of the pixel points in the monitored single-frame image and each background pixel point in the background model can be accurately obtained, and thus the foreground moving pixels in the monitored single-frame image can be accurately detected.

[0116] Please refer to Figure 7 , in an exemplary embodiment, step 230 further includes the following steps:

[0117] Step 232: When the number of similar background pixel points reaches the quantity threshold, or when it is detected that the pixel points in the monitored single-frame image belong to a special area, obtain these pixel points as static background pixels.

[0118] Among them, when the number of similar background pixel points in the background model reaches the quantity threshold, it indicates that the distribution characteristics of the background pixel points included in the background model are the same as those of the corresponding pixel point in the monitored single-frame image, and thus the pixel points in the monitored single-frame image are obtained as static background pixels.

[0119] Moreover, when the pixel points in the monitored single-frame image are detected to belong to special areas such as a highlight area or a shadow area caused by environmental light influence, these pixel points are also obtained as static background pixels. It should be noted that the detection of the highlight or shadow area of the pixel points in the monitored single-frame image is implemented according to corresponding algorithms such as a highlight discrimination algorithm and a shadow discrimination algorithm, and will not be described in detail here.

[0120] Step 234: Use the static background pixels to randomly update the background pixel points in the background model to obtain an updated background model.

[0121] Among them, after detecting the static background pixels in the monitored single-frame image, it is necessary to update the static background pixels as new background pixel points to the background model.

[0122] The random update of the background pixel points means randomly replacing the pixel value of the static background pixel with a background pixel point in the background model.

[0123] By updating the background model corresponding to the static background pixels, the background model can contain the co-located pixel points in the historical frame images, enabling the background model to accurately simulate the changes of the background pixels, which is conducive to the recognition of the remaining items in the subsequent frame images.

[0124] Step 236: Randomly update the background model corresponding to the neighborhood pixels of the static background pixels according to the set probability.

[0125] Among them, the set probability means that when a pixel point is detected as a static background, it has a 1 / rate probability of updating the background model corresponding to the neighborhood pixels, where rate represents the time sampling factor, and generally takes a value of 16.

[0126] The random update of the background model corresponding to the neighborhood pixels of the static background pixels means that according to the set probability, first randomly select an adjacent pixel point from the neighborhood of the static background pixel, and then still use the pixel value of the static background pixel as the new background pixel point to randomly select and replace a background pixel point in the background model corresponding to the adjacent pixel point.

[0127] In this embodiment, not only the background model corresponding to the static background pixels is randomly updated, but also the background model corresponding to its neighborhood pixels is randomly updated, making full use of the spatial propagation characteristics of the pixel values, enabling the background model to gradually spread to the neighborhood pixels, which is conducive to the recognition of the ghost regions.

[0128] Please refer to Figure 8 , in an exemplary embodiment, after step 233, the following steps may further be included:

[0129] Step 410: Identify the misjudgment of the moving foreground pixels according to whether the moving foreground pixels belong to the ghost region, and correct the misjudged moving foreground pixels to static background pixels.

[0130] Among the moving foreground pixels detected in step 233, there may be ghost regions caused by environmental light interference and be misdetected as moving foreground pixels, which will also affect the accuracy of the detection of the remaining items. Therefore, it is necessary to further eliminate the ghost regions in the monitored single-frame image.

[0131] For moving foreground pixels that are considered as moving foreground pixels for multiple consecutive frames, it can be determined whether they belong to the ghost area based on the information of the HSV color space and the set channel threshold.

[0132] As mentioned above, since the hue channel is not sensitive to ambient lighting, the channel value distribution range of the ghost on the hue channel can be set in advance. If the hue channel value of the moving foreground pixel is within the set hue channel value distribution range, it means that the moving foreground pixel belongs to the ghost area, and it is corrected to a static background pixel.

[0133] Step 430 , updating the corresponding background model for the corrected static background pixels.

[0134] As mentioned above, a random update of the corresponding background model should be performed for each static background pixel, and the specific update process will not be described in detail here.

[0135] Therefore, this embodiment can further eliminate the ghost area caused by ambient lighting, and eliminate the influence of ambient lighting on the detection of left-behind items to the greatest extent possible, so that this embodiment is robust to changes in ambient lighting and further improves the accuracy of left-behind item detection.

[0136] See also Figure 9 In an exemplary embodiment, the method for detecting left-behind items in a monitoring scene further includes the following steps:

[0137] Step 510 , by performing organism detection on a single-frame surveillance image in the surveillance video, an image region corresponding to the organism in the single-frame surveillance image is obtained.

[0138] As mentioned above, since the abandoned items must be carried to the monitoring scene by organisms, there must be historical tracks of organisms around the area where the abandoned items appear. By filtering the candidate area of ​​abandoned items located in step 250 according to the historical tracks of organisms in the monitoring video, the accurate area where the abandoned items stay can be obtained.

[0139] Therefore, when starting to perform left-behind item detection on a single-frame surveillance image in a surveillance video, it is necessary to synchronously perform organism detection and maintenance on the single-frame surveillance image, so that in the process of performing left-behind item detection on the single-frame surveillance image, the candidate areas for left-behind items in the single-frame surveillance image can be filtered based on the synchronously maintained historical trajectory of the organism.

[0140] Exemplarily, organism detection in a single-frame monitoring image is also achieved through an artificial neural network model that has been pre-trained and used to perform organism detection, such as a yolov3 model, a fastercnn model, etc., which can correspondingly obtain the image area corresponding to the organism in the single-frame monitoring image.

[0141] Step 530: Maintain the continuously detected picture areas to obtain the historical trajectory of organisms in the surveillance video.

[0142] Among them, the continuously detected picture areas refer to the picture areas corresponding to organisms detected in each frame of a number of consecutive surveillance single-frame pictures during the organism detection. Therefore, by maintaining these picture areas, the historical trajectory of organisms can be obtained.

[0143] In an exemplary embodiment, when the picture area corresponding to an organism is first detected in a surveillance single-frame picture, the accumulation of the picture areas starts, and the historical trajectory of the organism is obtained by accumulating the continuously detected picture areas.

[0144] When no organism is detected in the surveillance single-frame picture, counting starts. When the count reaches a certain threshold, it means that no organism has been detected within a certain period of time. At this time, the maintained historical trajectory of the organism is cleared. When an organism is detected again, the accumulation of the picture area corresponding to the organism is performed again.

[0145] Thus, in the filtering of the candidate areas of left-behind items performed in step 270, for the current frame picture for left-behind item detection, the picture area corresponding to the organism in the current frame picture can be obtained according to the maintained historical trajectory of the organism. Then, the candidate areas of left-behind items included in or adjacent to this picture area are obtained as the areas where the left-behind items stay, and the accurate areas where the left-behind items stay can be obtained.

[0146] For the candidate areas of left-behind items not filtered in the surveillance single-frame picture, that is, the candidate areas of left-behind items are not included in and not adjacent to the picture area corresponding to the organism, it means that these candidate areas of left-behind items are the picture areas corresponding to non-left-behind items. Therefore, all the moving foreground pixels included in these candidate areas of left-behind items are corrected to static background pixels, and the corresponding background model is updated for these corrected static background pixels.

[0147] Please refer to Figure 10 , in an exemplary embodiment, the method for implementing the detection of left-behind items in a surveillance scenario further includes the following steps:

[0148] Step 610: In the detection of left-behind items performed on a surveillance single-frame picture, if the area where the left-behind item stays is first detected in the surveillance single-frame picture, start counting;

[0149] Step 630: Execute an alarm when the counted value exceeds the alarm threshold.

[0150] Among them, for the application scenario of automatically alarming for the area where the detected left-behind item stays, when the area where the left-behind item stays is first detected in a single monitored frame image, counting starts. When the counted value exceeds the alarm threshold, it means that the left-behind item has stayed in the monitored scenario for a period of time, so automatic alarm is executed.

[0151] Compared with the way of immediately executing alarm after the left-behind item appears in the monitored scenario, this embodiment can avoid unnecessary alarms caused by the short stay of the left-behind item in the monitored scenario.

[0152] Please refer to Figure 11 , in another exemplary embodiment, still considering the situation that the short stay of the left-behind item in the monitored scenario may cause unnecessary alarms, before step 630, the method further includes the following steps:

[0153] Step 710, when it is detected that the area where the left-behind item stays is not included in the single monitored frame image, stop counting;

[0154] Step 730, if the time when the counting stops reaches the set number of frames, clear the count.

[0155] Among them, when it is detected that the area where the left-behind item stays is not included in the single monitored frame image, it means that the left-behind item has moved out of the monitored scenario. By stopping counting, it is avoided that the value continues to increase and triggers a false alarm.

[0156] When the time when the counting stops reaches the set number of frames, it means that the left-behind item is determined to have moved out of the monitored scenario, and automatic alarm is no longer executed. Therefore, the current count is cleared. When it is detected again that the area where the left-behind item stays is included in the single monitored frame image, counting starts from zero again to conform to the actual monitored scenario.

[0157] In another embodiment, considering that in the actual monitored scenario, the long stay of organisms in the monitored scenario may cause false alarms, and the left-behind item may be taken away by organisms after a short stay in the monitored scenario, etc. If there is a picture area corresponding to the organism around the area where the left-behind item stays, even if the value of the counter exceeds the alarm threshold, no alarm will be executed.

[0158] Figure 12 is a specific implementation schematic diagram of a method for detecting left-behind items in a monitored scenario in an application scenario. This application scenario is specifically applied to the detection of left-behind items in a fire passage.

[0159] Such as Figure 12As shown, for the current frame image in the surveillance video for detecting left-behind items, on the one hand, after performing resolution adjustment, eliminating the illumination influence by the gray world algorithm, and converting from the RGB color space to the HSV color space, the current frame image obtained is judged pixel by pixel for static background pixels or moving foreground pixels, and the connected region of moving foreground pixels with an area exceeding the threshold is obtained as the left-behind item candidate region.

[0160] On the other hand, it is necessary to perform organism detection on the current frame image, and when the image region corresponding to the organism is detected in the image, maintain the organism historical trajectory based on the image region corresponding to the organism detected in the historical image; if the image region corresponding to the organism is not detected in the image, perform counting through a counter, and clear the maintained organism historical trajectory when the count exceeds the threshold.

[0161] After obtaining the left-behind item candidate region in the current frame image and the maintained organism historical trajectory through the above two aspects, filter the left-behind item candidate region according to the maintained organism historical trajectory to obtain the region where the left-behind item stays, and count the images detected to contain the region where the left-behind item stays through another counter, and perform automatic alarm when the count exceeds the threshold.

[0162] Thus, when a left-behind item appears in the fire passage, an automatic alarm will be triggered to notify the relevant personnel to clean up the left-behind item in time, thereby effectively ensuring the unobstructedness of the fire passage and avoiding serious accidents caused by the blockage of the fire passage in case of various emergencies.

[0163] The following is an embodiment of the device of the present invention, which can be used to implement the method for detecting left-behind items in the surveillance scenario involved in the present invention. For the details not disclosed in the embodiment of the device of the present invention, please refer to the embodiment of the method for detecting left-behind items in the surveillance scenario involved in the present invention.

[0164] Please refer to Figure 13 , in an exemplary embodiment, a device for implementing the detection of left-behind items in a surveillance scenario includes a background model acquisition module 810, a moving foreground calculation module 830, a candidate region positioning module 850, and a candidate region filtering module 870.

[0165] The background model acquisition module 810 is used to obtain the background model of the corresponding pixel points in the previous frame image pixel by pixel for the surveillance single-frame image in the surveillance video, and this background model is used to simulate the background pixel change of the corresponding pixel points in the previous frame image in the spatial domain.

[0166] The moving foreground calculation module 830 is used to calculate the moving foreground pixels in the surveillance single-frame image relative to the background model, and this surveillance single-frame image includes moving foreground pixels and static background pixels.

[0167] The candidate region positioning module 850 is used to locate the candidate regions of the left-behind items in the monitored single-frame image by the foreground motion pixels.

[0168] The candidate region filtering module 870 is used to filter the candidate regions of the left-behind items according to the historical trajectories of organisms in the monitored video, and obtain the regions where the left-behind items stay.

[0169] In an exemplary embodiment, the device for detecting left-behind items in the monitoring scenario further includes an environmental light elimination module and a color space conversion module.

[0170] The environmental light elimination module is used to eliminate the influence of environmental light on each pixel point of the monitored single-frame image, and obtain the updated pixel values in the RGB color space.

[0171] The color space conversion module is used to convert the monitored single-frame image from the RGB color space to the HSV color space according to the updated pixel values.

[0172] In an exemplary embodiment, when the monitored single-frame image is the first-frame image in the monitored video, the device for detecting left-behind items in the monitoring scenario further includes a neighborhood pixel sampling module and a background pixel acquisition module.

[0173] The neighborhood pixel sampling module is used to perform several random samplings on the neighborhood pixels of each pixel point in the monitored single-frame image, and obtain the set of background pixel points of the pixel point in the spatial domain.

[0174] The background pixel acquisition module is used to obtain the set of background pixel points of the pixel point in the spatial domain to form a background model.

[0175] In an exemplary embodiment, the foreground motion calculation module 830 includes a spatial distance calculation unit and a foreground motion pixel acquisition unit.

[0176] The spatial distance calculation unit is used to calculate the distances between the pixel points in the monitored single-frame image and each background pixel point in the background model corresponding to the color space where the monitored single-frame image is located, and obtain the background pixel points whose pixel values are similar to the pixel points in the monitored single-frame image in the background model.

[0177] The foreground motion pixel acquisition unit is used to obtain the pixel point as a foreground motion pixel when the number of similar background pixel points is less than the number threshold.

[0178] In an exemplary embodiment, the foreground motion calculation module 830 further includes a static background pixel acquisition unit, a background model update unit, and a neighborhood pixel update unit.

[0179] The static background pixel acquisition unit is used to acquire a pixel as a static background pixel when the number of similar background pixels reaches a threshold or the pixel is detected to belong to a special area, and the special area includes a high-light area and a shadow area.

[0180] The background model update unit is used to randomly update the background pixels in the background model with the static background pixels as new background pixels to obtain an updated background model.

[0181] The neighborhood pixel update unit is used to randomly update the background model corresponding to the neighborhood pixels of the static background pixels according to a set probability.

[0182] In an exemplary embodiment, the moving foreground calculation module 830 further includes a ghost misjudgment recognition unit and a correction update unit.

[0183] The ghost misjudgment recognition unit is used to recognize the misjudgment of the moving foreground pixels according to whether the moving foreground pixels belong to the ghost area, and correct the misjudged moving foreground pixels to static background pixels.

[0184] The correction update unit is used to perform the update of the corresponding background model for the static background pixels obtained by correction.

[0185] In an exemplary embodiment, the above device for detecting left-behind items in a monitoring scenario further includes a biological body detection module and a historical track maintenance module.

[0186] The biological body detection module is used to perform biological body detection on a single monitoring frame picture in the monitoring video to obtain a picture area corresponding to the biological body in the single monitoring frame picture.

[0187] The historical track maintenance module is used to maintain the continuously detected picture areas to obtain the historical track of the biological body in the monitoring video.

[0188] In an exemplary embodiment, the candidate area filtering module 870 includes a picture area acquisition unit and a target acquisition unit.

[0189] The picture area acquisition unit is used to obtain the picture area corresponding to the biological body according to the maintained historical track of the biological body for a single monitoring frame picture for left-behind item detection.

[0190] The target acquisition unit is used to acquire the candidate area of the left-behind item included in or adjacent to the picture area as the area where the left-behind item stays.

[0191] In an exemplary embodiment, the above device for detecting left-behind items in a monitoring scenario further includes a counting module and an alarm module.

[0192] The counting module is used to start counting when, in the detection of left-behind items performed on a monitored single-frame image, the area where a left-behind item stays is first detected in the monitored single-frame image.

[0193] The alarm module is used to execute an alarm when the counted value exceeds the alarm threshold.

[0194] In an exemplary embodiment, the device further includes a counting stop control module and a counting reset module.

[0195] The counting stop control module is used to control the stop of counting when it is detected that the monitored single-frame image does not contain the area where a left-behind item stays.

[0196] The counting reset module is used to reset the count when the time of counting stop reaches the set number of frames.

[0197] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment and will not be elaborated here.

[0198] The present invention also provides an electronic device, including a processor and a memory. Among them, computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the method for detecting left-behind items in a monitoring scenario as described above is implemented.

[0199] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, the method for detecting left-behind items in a monitoring scenario as described above is implemented.

[0200] The above content is only a preferred exemplary embodiment of the present application and is not used to limit the implementation of the present application. Those of ordinary skill in the art can make corresponding adaptations or modifications very conveniently according to the main concept and spirit of the present application. Therefore, the protection scope of the present application should be subject to the protection scope required by the claims.

Claims

1. A method for implementing the detection of left-behind items in a monitoring scenario, characterized in that, The method includes: Eliminating the influence of ambient light pixel by pixel on the monitored single-frame image in the monitoring video corresponding to the fire passage to obtain the pixel value update in the RGB color space where it is located; Converting the monitored single-frame image from the RGB color space to the HSV color space according to the updated pixel values; Obtaining the background model of the corresponding pixel in the previous frame image for each pixel of the monitored single-frame image, and the background model is used to simulate the background pixel change of the corresponding pixel in the previous frame image in the spatial domain; Calculating the moving foreground pixels in the monitored single-frame image according to the obtained background model, and the monitored single-frame image includes the moving foreground pixels and the stationary background pixels; Locating the candidate area of the left-behind item from the moving foreground pixels in the monitored single-frame image; when starting to detect the left-behind item in the monitored single-frame image of the monitoring video, synchronously performing organism detection on the monitored single-frame image in the monitoring video to obtain the image area corresponding to the organism in the monitored single-frame image; Maintaining the continuously detected image areas to obtain the organism historical trajectory in the monitoring video; For the monitored single-frame image for which the left-behind item detection is performed, obtaining the image area corresponding to the organism according to the maintained organism historical trajectory, and obtaining the area where the left-behind item stays as the candidate area of the left-behind item included in or adjacent to the image area; For the candidate areas of the left-behind items that are not filtered in the monitored single-frame image, correcting all the included moving foreground pixels to stationary background pixels, and updating the corresponding background model for the stationary background pixels.

2. The method according to claim 1, wherein Before obtaining the background model of the corresponding pixel in the previous frame image for each pixel of the monitored single-frame image, the method further includes: For each pixel in the first-frame image of the monitoring video, obtaining the set of background pixel points of the pixel in the spatial domain by randomly sampling the neighboring pixels of the pixel; Obtaining the set of background pixel points of the pixel in the spatial domain to form the background model.

3. The method according to claim 1, characterized in that, The calculating the moving foreground pixels in the monitored single-frame image according to the obtained background model includes: Corresponding to the color space where the monitored single-frame image is located, obtaining the background pixel points with pixel values similar to the pixel by calculating the distance between the pixel and each background pixel point in the background model; When the number of the similar background pixel points is less than the quantity threshold, obtaining the pixel as a moving foreground pixel.

4. The method according to claim 3, characterized in that, The method further includes: When the number of the similar background pixel points reaches the quantity threshold or it is detected that the pixel belongs to a special area, obtaining the pixel as a stationary background pixel, and the special area includes a highlight area and a shadow area; Randomly updating the background pixel points in the background model with the stationary background pixel as a new background pixel point to obtain the updated background model; Randomly updating the background model corresponding to the neighboring pixels of the stationary background pixel according to a set probability.

5. The method according to claim 4, characterized in that, After calculating the moving foreground pixels in the monitored single-frame image according to the obtained background model, the method further includes: Identifying misjudgments of the moving foreground pixels according to whether the moving foreground pixels belong to the ghost region, and correcting the misjudged moving foreground pixels to static background pixels; Performing an update of the corresponding background model for the corrected static background pixels.

6. The method according to claim 1, characterized in that, The method further includes: In the detection of the left-behind item in the monitored single-frame image, if the area where the left-behind item stays in the monitored single-frame image is first detected, start counting; Execute an alarm when the value of the count exceeds the alarm threshold.

7. The method according to claim 6, wherein Before executing the alarm when the value of the count exceeds the alarm threshold, the method further includes: Stop the counting when it is detected that the area where the left-behind item stays is not included in the monitored single-frame image; If the time when the counting stops reaches the set number of frames, clear the count.

8. A device for implementing the detection of left items in a monitoring scenario, characterized in that, The device includes: An ambient light elimination module, configured to eliminate the influence of ambient light on each pixel point of the monitored single-frame image in the monitoring video corresponding to the fire passage, and obtain an update of the pixel value in the RGB color space where it is located; A color space conversion module, configured to convert the monitored single-frame image from the RGB color space to the HSV color space according to the updated pixel value; A background model acquisition module, configured to acquire the background model of the same-position pixel points in the previous frame image for each pixel point of the monitored single-frame image, and the background model is used to simulate the change of the background pixels of the same-position pixel points in the previous frame image in the spatial domain; A moving foreground calculation module, configured to calculate the moving foreground pixels in the monitored single-frame image according to the obtained background model, and the monitored single-frame image includes the moving foreground pixels and the static background pixels; A candidate region positioning module, configured to locate the left-behind item candidate region in the monitored single-frame image from the moving foreground pixels; When starting to detect the left-behind item in the monitored single-frame image of the monitoring video, synchronously perform organism detection on the monitored single-frame image in the monitoring video to obtain the image region corresponding to the organism in the monitored single-frame image; Maintaining the continuously detected image regions to obtain the organism historical trajectory in the monitoring video; A candidate region filtering module, configured to, for the monitored single-frame image for which the left-behind item is detected, obtain the image region corresponding to the organism according to the maintained organism historical trajectory, and obtain the region where the left-behind item stays as the region that includes or is adjacent to the left-behind item candidate region in the image region; and for the unfiltered left-behind item candidate regions in the monitored single-frame image, correct all the included moving foreground pixels to static background pixels, and perform an update of the corresponding background model for the static background pixels.

9. An electronic device, characterized in that, Includes: A memory, storing computer-readable instructions; A processor, reading the computer-readable instructions stored in the memory to execute the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored thereon, which, when executed by a processor of a computer, cause the computer to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Video-based detection method for drowning event in swimming pool

    CN106022230A

  • Foreground detection method based on background model and frame difference

    CN106548488A