Camera shielding detection method, electronic equipment and computer program product

By detecting humanoids in the elevator and replacing humanoid areas in the image frame, combined with the similarity calculation of the air-environment frame, the problem of low occlusion detection accuracy in complex scenarios in the prior art is solved, and higher detection accuracy and lower false alarm rate are achieved.

CN120147648APending Publication Date: 2025-06-13TP-LINK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312066.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The occlusion detection algorithm of existing elevator cameras has low detection accuracy in complex scenarios, and the risk of false alarms or missed alarms is high, especially in the case of multiple objects, large changes in light or reflections.

Method used

By detecting whether there is someone in the elevator, and when someone is detected, the humanoid area in the image frame is replaced with the target area at the same position in the empty frame, a new image frame is obtained, and the similarity between the image frame and the empty frame is calculated. If the similarity is lower than the preset threshold, it is determined that the camera is blocked.

Benefits of technology

It effectively reduces the false alarm rate, improves the accuracy of camera occlusion detection, is suitable for more complex application scenarios, reduces detection costs, and reduces the risk of false alarms and underreports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147648A_ABST
    Figure CN120147648A_ABST
Patent Text Reader

Abstract

The invention provides a camera shielding detection method, electronic equipment and a computer program product. Relates to the technical field of computer vision. The camera shielding detection method comprises the steps that for a first image frame collected by a camera, under the condition that it is detected that no person exists in an elevator, the first image frame is stored as an air environment frame; for a second image frame collected by the camera, under the condition that it is detected that a person exists in the elevator, a human shape area in the second image frame is replaced with a target area in the air environment frame, and a third image frame is obtained; the similarity between the third image frame and the air environment frame is calculated, and under the condition that the similarity is lower than a preset similarity threshold value, it is determined that a camera in the elevator is shielded; wherein the second image frame is acquired after the first image frame, and the target area is an area which is the same as the human-shaped area in position in the air frame. According to the embodiment of the invention, the camera shielding detection accuracy can be improved, and the risk of false alarm or missing alarm can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a method for detecting camera occlusion, an electronic device, and a computer program product. Background Art

[0002] In the elevator application scenario, users may bring dangerous items such as electric vehicles upstairs after blocking the camera, posing a safety hazard. To solve this problem, the camera in the elevator is configured with a detection algorithm to detect whether there is a situation of blocking the camera; when it is detected that the camera is blocked, an alarm is issued to prevent users from bringing in dangerous items.

[0003] Currently, the detection algorithms configured for cameras in elevators calculate the similarity of each group of image blocks corresponding to adjacent frame positions based on the pixel values of pixel points, and judge whether the lens is blocked based on the similarity. Or, a depth map is collected, and occlusion detection is performed according to the average depth of the depth map. Or, the features of the current video frame are extracted, and the features are sent to a classification module for occlusion detection. For some complex scenarios, such as when there are multiple objects, changes in lighting, or reflections, when there is occlusion, the texture of the occluder is rich, the depth is difficult to distinguish, and the image features are diverse, resulting in low accuracy of the detection results and a high risk of false alarms or missed alarms. Summary of the Invention

[0004] According to various embodiments of the present application, there is provided a method for detecting camera occlusion, an electronic device, and a computer program product; which can improve the accuracy of camera occlusion detection and reduce the risk of false alarms or missed alarms.

[0005] In a first aspect, the present application provides a method for detecting camera occlusion, the method comprising:

[0006] For the first image frame collected by the camera, when it is detected that there is no one in the elevator, the first image frame is saved as an empty scene frame; for the second image frame collected by the camera, when it is detected that there is someone in the elevator, the human-shaped area in the second image frame is replaced with the target area in the empty scene frame to obtain a third image frame; calculate the similarity between the third image frame and the empty scene frame, and when the similarity is lower than a preset similarity threshold, determine that the camera in the elevator is blocked; wherein, the second image frame is collected after the first image frame, and the target area is the area in the empty scene frame that has the same position as the human-shaped area.

[0007] In the above manner, when it is detected that there is someone in the elevator, the human-shaped area detected in the image frame is replaced with the target area at the same position in the empty-scene frame to obtain a new third image frame. By calculating the similarity between the third image frame and the empty-scene frame, the false alarm rate can be effectively reduced; compared with the traditional detection based on image texture, the occlusion situation of the camera image can be detected more effectively, applicable to more complex application scenarios, reducing the detection cost; and it has strong usability and practicality.

[0008] In a possible implementation manner of the first aspect, for the first image frame collected by the camera, when it is detected that there is no one in the elevator, saving the first image frame as the empty-scene frame includes:

[0009] When it is detected that there is no one in the elevator for a consecutive first number of the first image frames, saving the last first image frame in the first number as the empty-scene frame.

[0010] In a possible implementation manner of the first aspect, the method further includes:

[0011] When it is detected that there is no one in the elevator for the first number of the first image frames, adjusting the first number to a second number; the second number is greater than the first number; when it is detected that there is no one in the elevator for a consecutive second number of the first image frames, updating the empty-scene frame based on the last first image frame in the second number.

[0012] In a possible implementation manner of the first aspect, calculating the similarity between the third image frame and the empty-scene frame, and when the similarity is lower than a preset similarity threshold, determining that the camera in the elevator is occluded, includes:

[0013] Calculating the similarity between each of a consecutive third number of the third image frames and the empty-scene frame, and when the similarities are all lower than the similarity threshold, determining that the camera in the elevator is occluded.

[0014] In a possible implementation manner of the first aspect, after it is detected that there is someone in the elevator for the second image frame collected by the camera, the method further includes:

[0015] For the fourth image frame collected by the camera, when it is detected that there is no one in the elevator, calculating the similarity between the fourth image frame and the empty-scene frame; the fourth image frame is collected after the second image frame; when it is detected that there is no one in the elevator for a consecutive fourth number of the fourth image frames and the similarities between them and the empty-scene frame are all lower than a preset similarity threshold, determining that the camera in the elevator is occluded.

[0016] In a possible implementation of the first aspect, the step of replacing the humanoid region in the second image frame with the target region in the empty scene frame to obtain a third image frame includes:

[0017] Replacing the pixel values of the humanoid region in the second image frame with the pixel values of the target region in the empty scene frame to obtain the third image frame.

[0018] In a possible implementation of the first aspect, an image frame captured by a camera is input into a humanoid detection model, and after being processed by the humanoid detection model, a detection result is output;

[0019] In the case where the detection result does not include the coordinates of the humanoid frame or the confidence level corresponding to the included coordinates of the humanoid frame is lower than a preset confidence threshold, it is determined that there is no one in the elevator; in the case where the detection result includes the coordinates of the humanoid frame and the confidence level corresponding to the coordinates of the humanoid frame is greater than the preset confidence threshold, it is determined that there is someone in the elevator.

[0020] In a possible implementation of the first aspect, the calculation of the similarity between the third image frame and the empty scene frame includes:

[0021] Inputting the third image frame and the empty scene frame into a siamese neural network model, and after being calculated by the siamese neural network model, outputting the similarity.

[0022] In a second aspect, the present application provides a camera occlusion detection device, which includes:

[0023] A detection module, configured to save the first image frame captured by the camera as an empty scene frame when it is detected that there is no one in the elevator;

[0024] A replacement module, configured to replace the humanoid region in the second image frame with the target region in the empty scene frame to obtain a third image frame when it is detected that there is someone in the elevator for the second image frame captured by the camera;

[0025] A calculation module, configured to calculate the similarity between the third image frame and the empty scene frame, and determine that there is an occlusion in the camera in the elevator when the similarity is lower than a preset similarity threshold.

[0026] In a third aspect, the present application provides an electronic device, including a memory and a processor, where the memory stores a computer program, and the processor implements the method according to any one of the first aspect when executing the computer program.

[0027] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the method according to any one of the first aspect when being executed by a processor.

[0028] In a fifth aspect, the present application provides a computer program product that, when running on a device, causes the device to execute the method described in any one of the above first aspects.

[0029] It can be understood that for the beneficial effects of the above second to fifth aspects, reference can be made to the relevant descriptions in the above first aspect, and details are not elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0031] Figure 1 Schematic diagram of the implementation process of the camera occlusion detection method provided by an embodiment of the present application;

[0032] Figure 2 Schematic diagram of the image processing architecture of the camera occlusion detection method provided by an embodiment of the present application;

[0033] Figure 3 Schematic diagram of the image processing architecture of the camera occlusion detection method provided by an embodiment of the present application;

[0034] Figure 4 Schematic diagram of the overall implementation process of the camera occlusion detection method provided by an embodiment of the present application;

[0035] Figure 5 Schematic diagram of the structure of the camera occlusion detection device provided by an embodiment of the present application;

[0036] Figure 6 Schematic diagram of the structure of the electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following will describe in detail the embodiments of the technical solutions of the present application with reference to the drawings. The following embodiments are only used to illustrate the technical solutions of the present application more clearly, so they are only examples and cannot be used to limit the protection scope of the present application.

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above description of the drawings are intended to cover non-exclusive inclusion.

[0039] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity, specific order, or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "a plurality of" is more than two, unless otherwise specifically defined.

[0040] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase does not necessarily refer to the same embodiment at every occurrence in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0041] In the description of the embodiments of this application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0042] The camera in the elevator can emit an alarm sound and control the elevator not to close the door when it detects that the user is blocking the camera by configuring an occlusion detection algorithm, so as to prevent the user from bringing dangerous items in.

[0043] However, traditional occlusion detection algorithms include detection methods based on the pixel values of pixel points in the collected images, which are easily affected by changes in light and texture. When the light changes greatly or the occluder has rich texture, false detections are likely to occur, resulting in missed reports; or, the occlusion situation is judged based on the obtained depth map. In complex application scenarios (such as multiple objects, large light changes, and reflections), it is difficult to distinguish the depth of the image, resulting in a high risk of both false reports and missed reports; or, features are extracted from the current video frame for occlusion classification, ignoring the context information, which is likely to lead to false reports or missed reports. Moreover, due to the rich variety of elevator types and occluder types, the training of this model requires collecting a large amount of diverse data to cover various possible scenarios, resulting in high costs.

[0044] In view of the above technical problems, an embodiment of the present application provides a method for detecting camera occlusion. By detecting whether there is anyone in the elevator, and when it is detected that there is someone in the elevator, replacing the human-shaped area detected in the image frame with the target area at the same position in the empty-scene frame to obtain a third image frame, and calculating the similarity between the third image frame and the empty-scene frame, the false alarm rate can be effectively reduced; there is no need to deploy additional depth cameras or binocular cameras to obtain depth images, reducing equipment costs; based on the similarity comparison between the new third image frame and the latest empty-scene frame, fully combining context information, the risks of false alarms and missed detections are reduced.

[0045] The following introduces the specific implementation process of the method for detecting camera occlusion through embodiments.

[0046] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the implementation of the method for detecting camera occlusion provided by an embodiment of the present application. As Figure 1 shown, the execution subject of this method for detecting camera occlusion can be an electronic device, which can be a camera device in the elevator, or a background server communicatively connected to the camera device in the elevator, etc.; this method can include the following steps:

[0047] S201, for the first image frame collected by the camera, when it is detected that there is no one in the elevator, save the first image frame as the empty-scene frame.

[0048] In an embodiment of the present application, a camera is installed in the elevator. After the camera is started, it can collect continuous images inside the elevator in real time, such as continuous first image frames. For the collected first image frame, based on the configured detection model, each image frame is recognized to detect whether there is anyone in the elevator at the corresponding moment of each image frame. Among them, the detection model can be a human-shaped detection model. Based on this human-shaped detection model, by analyzing features such as shape, color, and movement in the image frame, it is determined whether there is anyone in the elevator.

[0049] Exemplarily, the empty-scene frame is an image frame in a state where there is no one in the elevator. The electronic device recognizes the continuous first image frames. When it is detected that there is no one in the elevator, it saves the latest collected first image frame and uses it as the empty-scene frame; continues to recognize the subsequently collected image frames to detect whether there is anyone in the elevator in the subsequent time period.

[0050] Exemplarily, a camera device configured with a detection model may be provided inside the elevator. The camera device is configured with the detection model. After the camera device captures an image frame inside the elevator through a camera, it identifies the image frame based on the configured detection model to detect whether there is anyone inside the elevator. Alternatively, the camera installed inside the elevator is communicatively connected to a backend server, and the backend server is configured with a detection model. After the camera captures a plurality of consecutive first image frames, it sends the plurality of consecutive first image frames to the backend server based on the communication connection. The backend server identifies the received plurality of first image frames based on the configured detection model to detect whether there is anyone inside the elevator during the period corresponding to the plurality of consecutive first image frames.

[0051] In some embodiments, for the first image frame captured by the camera, when it is detected that there is no one inside the elevator, saving the first image frame as an empty scene frame includes:

[0052] When it is detected that there is no one inside the elevator for each of a consecutive first number of first image frames, saving the last first image frame among the first number as an empty scene frame.

[0053] Exemplarily, in order to save storage space, when it is detected that there is no one inside the elevator for each of a consecutive first number of captured first image frames, saving the latest first image frame as an empty scene frame.

[0054] Exemplarily, a parameter of the number of consecutive detection frames, such as the first number, may also be set in the electronic device. As Figure 2 shown, the first number may be N (N is an integer greater than or equal to 1). When it is detected that there is no one inside the elevator for each of N consecutive image frames through the identification of the detection model, the Nth image frame in the consecutive image frames satisfies the condition that no human form has been detected in N consecutive frames, and the Nth image frame is saved as an empty scene frame.

[0055] Correspondingly, when it is detected that there is no one inside the elevator for each of the next consecutive N first image frames, the saved empty scene frame can be updated based on the Nth image frame in the next N frames. For example, if no human form is detected in frames 1 to 10, then the 10th image frame is saved as an empty scene frame. If no human form is still detected in frames 11 to 20, then the empty scene frame is updated based on the 20th image frame, that is, the 20th image frame is saved as an empty scene frame.

[0056] Through the above method, when the subsequent scene inside the elevator changes, the saved empty scene frame is made more referenceable relative to the subsequently captured image frames, reducing the interference of continuously changing factors inside the elevator on the subsequent judgment of camera occlusion.

[0057] In some embodiments, the method further includes:

[0058] When it is detected that there is no one in the elevator in the first image frames of the first quantity, adjust the first quantity to the second quantity; the second quantity is greater than the first quantity; when it is detected that there is no one in the elevator in the consecutive first image frames of the second quantity, update the empty scene frame based on the last first image frame in the second quantity.

[0059] Exemplarily, the parameter of the consecutive detection frame number set in the electronic device can be dynamically adjusted; for example, the first quantity set by N is 10 frames. When no human form is detected in 10 consecutive frames, save the 10th frame image as the empty scene frame; adjust N to a larger second quantity, such as 20 frames (it can also be other frame numbers, such as 50 frames, etc.), and continue to identify the consecutive first image frames of the second quantity.

[0060] For example, N = 10. After the electronic device is started, for the 1st to 10th frames, the human form detection model does not detect a human form. Meeting the condition "when no human form is detected in consecutive N frames", it is determined that there is no one in the elevator, and save the image of the 10th frame as the empty mirror frame; assume that for the 11th to 20th frames, the human form detection model does not detect a human form, then the condition "when no human form is detected in consecutive N frames" is met again, it is determined that there is no one in the elevator, update the empty mirror frame, and save the 20th frame image as the empty mirror frame.

[0061] Correspondingly, when it is detected that there is no one in the elevator in the consecutive first image frames of the second quantity, update the empty scene frame to the last first image frame in the second quantity. As Figure 2 shown, when it is detected that there is no one in the elevator from the (N + 1)th frame image to the 3Nth frame image, save the 3Nth frame image as the empty scene frame, that is, update the empty scene frame to the 3Nth frame image.

[0062] Among them, the initially set first quantity can be smaller, so that the empty scene frame can be obtained quickly; subsequently, the consecutive detection frame number can be set larger to reduce the resource consumption caused by replacing the empty scene frame; by continuously updating the empty scene frame later, the influence of changes in the monitoring angle, billboards, etc. in the elevator on the subsequent area replacement process is avoided.

[0063] In some embodiments, the method further includes:

[0064] Input the image frame collected by the camera into the human form detection model, and after being processed by the human form detection model, output the detection result; when the detection result does not include the coordinates of the human form box, or the confidence level corresponding to the coordinates of the included human form box is lower than the preset confidence level threshold, determine that there is no one in the elevator; when the detection result includes the coordinates of the human form box and the confidence level corresponding to the coordinates of the human form box is greater than the preset confidence level threshold, determine that there is someone in the elevator.

[0065] Exemplarily, the above detection model can be a humanoid detection model, and this humanoid detection model can be a real-time object detection algorithm, the YOLO (You Only Look Once) v11 neural network model. Each collected image frame is input into this humanoid detection model, and after being processed by the humanoid detection model, a detection result is output; the detection result can include the coordinates of the humanoid box, the corresponding confidence level, and the detection result category.

[0066] Among them, the YOLOv11 neural network model can include a backbone network, a neck network, and a head network; features in the image frame are extracted through the backbone network, the extracted features are input into the neck network, feature fusion is performed through the neck network, and the head network performs object detection based on the fused features and outputs a detection result; the YOLOv11 neural network model also uses anchor boxes to identify objects of different sizes and shapes in the image frame, which can better adapt to the complex and changeable scenarios inside the elevator; the YOLOv11 neural network model can be trained based on the loss function EIoU (Extended IoU), and this loss function can combine the overlapping area between the predicted box and the ground truth box, the aspect ratio, and the deviation of the center point. During the training process, the parameters are optimized to improve the model detection accuracy, achieve the detection ability for objects of different sizes and shapes in the image frame, and also maintain the real-time inference speed, which is applicable to the application scenario of real-time detection of continuous image frames inside the elevator.

[0067] Correspondingly, if the coordinates of the humanoid box are not output, it is determined that the humanoid area is not detected; or if the coordinates of the humanoid box are output and the corresponding confidence level does not meet the confidence threshold, then this humanoid box is not credible and it is determined that there is no one in the elevator; when the coordinates of the humanoid box are output and the corresponding confidence level is greater than the confidence threshold, then this humanoid box is credible; in the case where the number of credible humanoid box coordinates output is greater than or equal to 1, it is determined that there is someone in the elevator of the current image frame.

[0068] Among them, when the electronic device detects a humanoid based on the humanoid detection model, it can output the normalized coordinates of the humanoid box. For example, the humanoid box coordinates include the coordinates of the upper left corner and the lower right corner (x1, y1, x2, y2), representing the position of the detected person. For example, if the output coordinates are [0.125, 0.0833, 0.25, 0.5], multiplying them by the image resolution can obtain the specific coordinates of the humanoid in this image frame; for example, if the obtained image size is 800*600, the upper left corner coordinates of the rectangular box of the humanoid area can be calculated as (100, 50), and the lower right corner coordinates can be calculated as (200, 300), so as to determine the position of the humanoid area in the image frame.

[0069] Based on the above detection model, the detection of whether there is anyone in the elevator is performed. The training of the detection model is relatively simple, and it is only trained to identify whether there is anyone in the elevator. The implementation of the algorithm architecture is also relatively simple, reducing the implementation complexity of the detection algorithm.

[0070] S202. For the second image frame collected by the camera, when it is detected that there is someone in the elevator, replace the human-shaped area in the second image frame with the target area in the empty scene frame to obtain the third image frame.

[0071] Among them, the second image frame is collected after the first image frame, and the target area is the area in the empty scene frame that has the same position as the human-shaped area.

[0072] In the embodiment of the present application, after saving the empty scene frame, the electronic device continues to identify the subsequent continuously collected image frames, and detects whether there is a human figure in each frame of the image. Based on the detection model, when at least one human-shaped box is detected in the image frame, based on the coordinates of the human-shaped box and the size of the image frame, determine the human-shaped area in the image frame.

[0073] Exemplarily, as Figure 2 shown, after saving the Nth frame of the image as the empty scene frame, detect the human-shaped box in the (N + 2)th frame of the image collected subsequently; based on the coordinates of the human-shaped box and the size of the image frame, calculate the position coordinates of the human-shaped box in the image frame, and these position coordinates can be represented based on the image pixel coordinates, so as to determine the position of the human-shaped area in the image frame.

[0074] Correspondingly, after determining the position of the human-shaped area in the second image frame, determine the target area with the same position coordinates in the empty scene frame, such as Figure 2 the target area at the same position in the empty scene frame shown in, and replace the image feature value of this area with the human-shaped area in the second image frame to obtain a new third image frame.

[0075] Among them, the electronic device uses the same detection model as that for the first image frame above to identify the human-shaped box in the second image frame, and based on the size of the second image frame, calculate the position coordinates of this human-shaped box in the second image frame.

[0076] In some embodiments, replacing the human-shaped area in the second image frame with the target area in the empty scene frame to obtain the third image frame includes:

[0077] Replace the pixel values of the human-shaped area in the second image frame with the pixel values of the target area in the empty scene frame to obtain the third image frame.

[0078] Exemplarily, as Figure 2As shown, replace the pixel value of each pixel in the target area of the empty scene frame with the corresponding pixel value of each pixel in the human-shaped area of the second image frame; for a color image frame, the pixel value is used to represent the color and brightness information of the pixel at a specific position, and for a grayscale image frame, the pixel value can also be the grayscale value of the pixel, representing the brightness information of the pixel.

[0079] Among them, the above Figure 2 The area replacement shown is only for illustrative purposes and does not limit the actual data processing process. There can also be multiple human-shaped areas detected in the second image frame; for example, judge the number of people detected currently according to the confidence threshold. For example, if there are 3 human-shaped frames with a confidence greater than 0.8, it is considered that 3 people are detected. Correspondingly, 3 human-shaped areas can be determined, or a connected human-shaped area can be determined according to the 3 human-shaped frames.

[0080] It should be noted that when there are many people in the elevator, the elevator background only accounts for a small part of the image frame. When performing similarity comparison later, only a small part of the elevator background in the empty scene frame is similar to the image frame with people in the elevator, so it is considered that the difference is large, resulting in false alarms for occlusion situations; and when there are people moving in the elevator, most of the image frames collected are human objects, resulting in a low similarity after comparison; by replacing the detected human-shaped area with the target area of the empty scene frame, it is possible to reduce misjudgments caused by full occupancy or movement of people, thereby improving the judgment accuracy of the subsequent model for camera occlusion situations.

[0081] S203, calculate the similarity between the third image frame and the empty scene frame, and determine that the camera in the elevator is occluded when the similarity is lower than the preset similarity threshold.

[0082] In the embodiments of the present application, the third image frame and the empty scene frame are input into the neural network model. The neural network model analyzes and calculates the feature differences between the two frames of images and outputs the similarity. For example, the similarity between the two frames of images is represented by a similarity score.

[0083] Exemplarily, the above neural network model can be a Siamese neural network model. As Figure 2 shown, input the empty scene frame and the third image frame into the Siamese neural network model, analyze the feature differences between the two frames of images through the Siamese neural network model, and output a similarity score (a value between 0 and 1, the larger the value, the higher the similarity). Correspondingly, when the similarity score is less than the similarity threshold, it can be determined that the camera in the elevator is occluded.

[0084] Exemplarily, in the process of calculating the similarity between two image frames, the similarity can also be represented by calculating the mean square error or the structural similarity index; for example, the mean square error is used to measure the average squared difference of the corresponding pixel values between two image frames, and a lower mean square error indicates a higher similarity between the two image frames; correspondingly, the structural similarity index is used to measure the differences in structural information, brightness, and contrast between two image frames. The closer the structural similarity index is to 1, the higher the similarity between the two image frames, and the closer it is to 0, the greater the difference. Additionally, a convolutional neural network can also be used to extract key features in the image frames and compare the similarity between the features.

[0085] As Figure 2 shown, there may be Region 1 and Region 2 in the third image frame. After comparison, Region 1 is a similar region to the empty scene frame, and Region 2 with differences makes the similarity of the entire image frame lower than the similarity threshold, thereby determining that there is an occlusion in the captured image. Figure 2 The occlusion situation in only exemplarily illustrates the situation where the similarity is lower than the similarity threshold, that is, the scene where there are significant differences between the empty scene frame and the third image frame, and does not limit the presentation of the image frames in the actual application scenario and the similarity calculation process.

[0086] In some embodiments, calculating the similarity between the third image frame and the empty scene frame includes:

[0087] Inputting the third image frame and the empty scene frame into a siamese neural network model, and after the calculation of the siamese neural network model, outputting the similarity.

[0088] Exemplarily, inputting the third image frame and the empty scene frame into a siamese neural network model, analyzing the differences in high-dimensional features between the two image frames through this siamese neural network model, and calculating the similarity between the two features; among them, high-dimensional features refer to features with multiple dimensions, which can include the elevator monitoring perspective, elevator color and material, billboard shape, etc.

[0089] Exemplarily, the structure of the siamese neural network model can consist of two sub-networks sharing parameters (such as shared weights) and can process a pair of inputs in parallel. The structure of the siamese neural network model can include an input layer, a shared network layer, and an output layer; the input layer receives two input images: the third image frame and the empty scene frame; the shared network layer propagates forward through the same convolutional neural network to extract the features of the two input images respectively. The features can be edge and texture features, and can also be high-level semantic information, such as the elevator monitoring perspective, elevator color and material, billboard shape, etc.; since the two sub-networks share parameters, similar input images will use similar feature representations; each sub-network outputs a feature vector representing the features of the input image. The output layer compares the feature vectors generated by the two sub-networks; for example, using the Euclidean distance: calculating the distance between the two feature vectors to determine the similarity of the two input images; or, using the cosine similarity to evaluate the similarity of the two feature vectors; or, using a classifier (such as a fully connected layer) to predict whether the two inputs belong to the same category.

[0090] Among them, the convolutional neural network structure of the shared network layer can adopt a deep convolutional network structure such as the Visual Geometry Group (VGG) network.

[0091] In some embodiments, calculating the similarity between the third image frame and the empty scene frame, and determining that the camera in the elevator is blocked when the similarity is lower than a preset similarity threshold includes:

[0092] Calculating the similarity between each of the consecutive third number of third image frames and the empty scene frame, and determining that the camera in the elevator is blocked when the similarities are all lower than the similarity threshold.

[0093] Exemplarily, in order to improve the accuracy of occlusion detection, after detecting the human-shaped area in a continuous plurality of second image frames, replacing the human-shaped area with the corresponding target area in the empty scene frame to obtain a continuous plurality of third image frames, calculating the similarity between each third image frame and the empty scene frame, and determining that the camera in the elevator is blocked when the similarities are all less than the similarity threshold.

[0094] For example, such as Figure 2As shown, the similarity threshold is 0.6 and N = 10. After the electronic device is started, for the 1st to 10th frames, the human detection model does not detect a human figure. When the condition "no human figure is detected for consecutive N frames" is met, it is determined that there is no one in the elevator, and the image of the 10th frame is saved as an empty frame. Assume that no human figure is detected in the 11th frame (the (N + 1)th frame), and a human figure is detected in the 12th frame (the (N + 2)th frame). Then, the corresponding target area of the empty frame is used to replace the human figure area in the 12th frame to obtain a new image frame. The new image frame and the empty frame (the 10th frame) are input into the neural network model to predict the similarity, and a similarity score is output. Assume that the output similarity score is 0.8, which is greater than the threshold, then it is confirmed that there is no occlusion of the camera in the elevator.

[0095] In one case, if a human figure is detected in the 12th frame (the (N + 2)th frame), the corresponding target area of the empty frame is used to replace the human figure area in the 12th frame to obtain a new image frame. The new image frame and the empty frame (the 10th frame) are input into the neural network model to predict the similarity, and a similarity score is output. Assume that the output similarity score is 0.4, which is lower than the threshold. The situations of the 13th to 16th frames are the same as that of the 12th frame. Then, the 16th frame meets the condition "the similarity scores of consecutive K frames (the third quantity K = 4) are lower than the set similarity threshold". At this time, since a human figure can still be detected and the similarity is lower than the similarity threshold, it indicates that there may be partial occlusion of the camera, and an alarm is issued.

[0096] In some embodiments, the method further includes:

[0097] For the fourth image frame collected by the camera, when it is detected that there is no one in the elevator, the similarity between the fourth image frame and the empty frame is calculated. When no one is detected in consecutive fourth quantities of the fourth image frames, and the similarities with the empty frame are all lower than the preset similarity threshold, it is determined that there is occlusion of the camera in the elevator.

[0098] Exemplarily, based on the above example, as Figure 3 shown, in another case, assume that a human figure is detected in the 12th frame. After replacing the human figure area in the 12th frame, the calculated similarity is greater than the similarity threshold, and it is determined that there is no occlusion at the 12th frame. No one is detected in the 13th frame (the (N + 3)th frame), and the 13th frame and the empty frame (the 10th frame) are directly input into the siamese neural network to predict the similarity. Assume that the output similarity score is 0.3. The situations of the 14th to 17th frames are the same as that of the 13th frame. Then, at the 17th frame, the condition "the similarity scores of consecutive K frames (the fourth quantity K = 4) are lower than the set similarity threshold" is met, and it is determined that there is an occlusion event of the camera in the current elevator, and an alarm is issued.

[0099] As Figure 4As shown in the figure, the following is a schematic diagram of the overall implementation process of the camera occlusion detection method provided by the embodiments of the present application. Based on the same implementation principle as the above embodiments, it will not be elaborated here; the method may include the following steps:

[0100] S401, Detect the number of people in the elevator based on the human detection algorithm.

[0101] S402, Detect whether the number of people is 0. If so, execute S403; if not, execute S405.

[0102] S403, Continuously detect N frames. Whether the number of people detected in all frames is 0. If so, execute S404.

[0103] S404, Update the empty scene frame.

[0104] S405, Obtain the human-shaped area.

[0105] S406, Region replacement: Replace the human-shaped area with the target area of the empty scene frame.

[0106] S407, Similarity detection: Detect the similarity between the empty scene frame and the image frame after region replacement based on the Siamese neural network model.

[0107] S408, Whether the similarity is less than the similarity threshold. If so, execute S409; if not, execute S411.

[0108] S409, Whether the similarity of continuously K frames is less than the similarity threshold. If so, execute S410; if not, execute S411.

[0109] S410, The camera in the elevator is blocked.

[0110] S411, The camera in the elevator is not blocked.

[0111] In the embodiments of the present application, by detecting whether there are people in the elevator, and when it is detected that there are people in the elevator, replacing the detected human-shaped area in the image frame with the target area at the same position in the empty scene frame to obtain a new third image frame, and by calculating the similarity between the third image frame and the empty scene frame, the false alarm rate can be effectively reduced; there is no need to deploy additional depth cameras or binocular cameras to obtain depth images, reducing equipment costs; based on the similarity comparison between the new third image frame and the latest empty scene frame, fully combining context information, the risks of false alarms and missed detections are reduced.

[0112] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0113] Corresponding to the camera occlusion detection method provided in the above embodiments, the camera occlusion detection device provided in the embodiments of the present application, as Figure 5 shown, is a schematic structural diagram of the camera occlusion detection device provided in the embodiments of the present application. For the sake of illustration, only the parts related to the embodiments of the present application are shown.

[0114] The camera occlusion detection device includes:

[0115] A detection module 51, configured to save the first image frame collected by the camera as an empty scene frame when it is detected that there is no one in the elevator.

[0116] A replacement module 52, configured to replace the human figure area in the second image frame with the target area in the empty scene frame to obtain a third image frame when it is detected that there is someone in the elevator for the second image frame collected by the camera.

[0117] A calculation module 53, configured to calculate the similarity between the third image frame and the empty scene frame, and determine that the camera in the elevator is occluded when the similarity is lower than a preset similarity threshold.

[0118] In a possible implementation manner, the detection module is further configured to save the last first image frame in the first quantity as the empty scene frame when it is detected that there is no one in the elevator for each of the first quantity of consecutive first image frames.

[0119] In a possible implementation manner, the detection module is further configured to adjust the first quantity to a second quantity when it is detected that there is no one in the elevator for each of the first quantity of first image frames; the second quantity is greater than the first quantity; and when it is detected that there is no one in the elevator for each of the second quantity of consecutive first image frames, update the empty scene frame based on the last first image frame in the second quantity.

[0120] In a possible implementation manner, the calculation module is further configured to calculate the similarity between each of the third quantity of consecutive third image frames and the empty scene frame, and determine that the camera in the elevator is occluded when the similarities are all lower than the similarity threshold.

[0121] In a possible implementation manner, the calculation module is further configured to calculate the similarity between the fourth image frame and the empty scene frame when it is detected that there is no one in the elevator for the fourth image frame collected by the camera; the fourth image frame is collected after the second image frame; and when it is detected that there is no one in the elevator for each of the fourth quantity of consecutive fourth image frames and the similarities between them and the empty scene frame are all lower than a preset similarity threshold, determine that the camera in the elevator is occluded.

[0122] In a possible implementation, the replacement module is further configured to replace the pixel values of the humanoid region in the second image frame with the pixel values of the target region in the empty scene frame to obtain the third image frame.

[0123] In a possible implementation, the detection module is further configured to input the image frame collected by the camera into a humanoid detection model, process it through the humanoid detection model, and output a detection result; when the detection result does not include the coordinates of the humanoid frame or the confidence level corresponding to the included coordinates of the humanoid frame is lower than a preset confidence level threshold, it is determined that there is no one in the elevator; when the detection result includes the coordinates of the humanoid frame and the confidence level corresponding to the coordinates of the humanoid frame is greater than the preset confidence level threshold, it is determined that there is someone in the elevator.

[0124] In a possible implementation, the calculation module is further configured to input the third image frame and the empty scene frame into a siamese neural network model, and output the similarity through the calculation of the siamese neural network model.

[0125] Through the embodiments of the present application, by detecting whether there is someone in the elevator, and when it is detected that there is someone in the elevator, replacing the detected humanoid region in the image frame with the target region at the same position in the empty scene frame to obtain a new third image frame, and by calculating the similarity between the third image frame and the empty scene frame, the false alarm rate can be effectively reduced; there is no need to deploy additional depth cameras or binocular cameras to obtain depth images, reducing equipment costs; based on the similarity comparison between the new third image frame and the latest empty scene frame, fully combining context information, the risks of false alarms and missed detections are reduced.

[0126] Figure 6 Fig. shows a schematic hardware structure diagram of the electronic device 6.

[0127] As Figure 6 shown, the electronic device 6 of this embodiment includes: at least one processor 61 ( Figure 6 only one is shown in the figure), a memory 62, and a computer program 63 stored in the memory 62 that can run on the processor 61. When the processor 61 executes the computer program 63, it implements the steps in the above method embodiments, such as Figure 1 shown in S101 to S103. Alternatively, when the processor 61 executes the computer program 63, it implements the functions of each module / unit in the above device embodiments.

[0128] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 6. In other embodiments of the present application, the electronic device 6 may include more or fewer components than those shown in the figures, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0129] The electronic device 6 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 6 merely examples of the electronic device 6 do not constitute a limitation on the electronic device 6, and it may include more or fewer components than those shown in the figures, or combine certain components, or have different components. For example, the server may further include an input and transmission device, a network access device, a bus, etc.

[0130] The above-mentioned processor 61 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0131] A memory may also be provided in the processor 61 for storing instructions and data. In some embodiments, the memory in the processor 61 is a cache memory. This memory can store the instructions or data that the processor 61 has just used or recycled. If the processor 61 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 61, and thus improves the efficiency of the system.

[0132] In some embodiments, the memory 62 may be an internal storage unit of the electronic device 6, such as a hard disk or memory of the electronic device 6. The memory 62 may also be an external storage device of the electronic device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 6. Further, the memory 62 may include both an internal storage unit and an external storage device of the electronic device 6. The memory 62 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as program codes of a computer program. The memory 62 may also be used to temporarily store data that has been sent or is to be sent.

[0133] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0134] It should be noted that the structure of the above-mentioned electronic device is only illustrative, and based on different application scenarios, it may also include other physical structures, and the physical structure of the electronic device is not limited here.

[0135] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0136] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0137] An embodiment of the present application provides a computer program product. When the computer program product runs on a server, the server can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0138] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0139] The electronic device, computer storage medium, and computer program product provided in the above embodiments of this application are all used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the method provided above, and will not be elaborated here.

[0140] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should all be considered as the scope described in this specification.

[0141] It should be understood that the above is only to help those skilled in the art better understand the embodiments of this application, rather than to limit the scope of the embodiments of this application. Those skilled in the art can obviously make various equivalent modifications or changes according to the above examples. For example, some steps in the above-described embodiments of the detection method may not be necessary, or some steps may be newly added, etc. Or any combination of any two or any number of the above embodiments. Such modified, changed, or combined solutions also fall within the scope of the embodiments of this application.

[0142] It should also be understood that the division of the manners, situations, categories, and embodiments in the embodiments of this application is only for the convenience of description and should not constitute a special limitation. The features in various manners, categories, situations, and embodiments can be combined without conflict.

[0143] It should also be understood that in various embodiments of the present application, if there is no special description and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.

[0144] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0145] In the embodiments provided in the present application, it should be understood that the disclosed devices / network devices and methods can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.

[0146] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0147] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

[0148] Finally, it should be noted that the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A camera occlusion detection method, characterized in that: The method comprises: For the first image frame captured by the camera, when it is detected that there is no one in the elevator, the first image frame is saved as an empty frame; For the second image frame captured by the camera, when a person is detected in the elevator, the human-shaped area in the second image frame is replaced with the target area in the empty frame to obtain a third image frame; Calculating the similarity between the third image frame and the empty frame, and determining that the camera in the elevator is blocked when the similarity is lower than a preset similarity threshold; The second image frame is acquired after the first image frame, and the target area is an area in the empty space frame that has the same position as the human-shaped area.

2. The method according to claim 1, characterized in that The first image frame captured by the camera, when detecting that there is no one in the elevator, saving the first image frame as an empty frame, comprises: For a first number of consecutive first image frames, when it is detected that there is no one in the elevator, the last first image frame in the first number is saved as the empty frame.

3. The method according to claim 2, characterized in that The method further comprises: When the first number of first image frames all detect that there is no one in the elevator, adjusting the first number to a second number; the second number is greater than the first number; When the second number of consecutive first image frames all detect that there is no one in the elevator, the empty frame is updated based on the last first image frame in the second number.

4. The method according to claim 1, characterized in that: The calculating the similarity between the third image frame and the empty frame, and determining that the camera in the elevator is blocked when the similarity is lower than a preset similarity threshold, includes: The similarities between a third number of consecutive third image frames and the empty space frame are calculated respectively, and when the similarities are all lower than the similarity threshold, it is determined that there is occlusion on the camera in the elevator.

5. The method according to claim 1, characterized in that After detecting that there is someone in the elevator in the second image frame captured by the camera, the method further includes: For a fourth image frame captured by the camera, when it is detected that no one is in the elevator, the similarity between the fourth image frame and the empty frame is calculated; the fourth image frame is captured after the second image frame; When a fourth number of consecutive fourth image frames all detect that there is no one in the elevator and their similarities with the empty frame are lower than a preset similarity threshold, it is determined that there is occlusion on the camera in the elevator.

6. The method according to claim 1, characterized in that The step of replacing the human-shaped region in the second image frame with the target region in the empty frame to obtain a third image frame includes: The pixel values ​​of the human-shaped area in the second image frame are replaced with the pixel values ​​of the target area in the empty frame to obtain the third image frame.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Input the image frames captured by the camera into the human figure detection model, process them through the human figure detection model, and output the detection results; When the detection result does not include the human frame coordinates, or the confidence level corresponding to the included human frame coordinates is lower than a preset confidence level threshold, determining that no one is in the elevator; When the detection result includes the human frame coordinates and the confidence corresponding to the human frame coordinates is greater than a preset confidence threshold, it is determined that there is a person in the elevator.

8. The method according to any one of claims 1 to 6, characterized in that: The calculating the similarity between the third image frame and the empty space frame comprises: The third image frame and the empty space frame are input into the twin neural network model, and the similarity is output after calculation by the twin neural network model.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 8 when executing the computer program.

10. A computer program product, characterized in that When the computer program product is executed on a device, the device is caused to execute the method according to any one of claims 1 to 8.