Access control method and device based on video recognition

By dividing the access control system into a dangerous area prone to pinching and a safe buffer zone, and using video recognition technology to detect target objects in real time, the problem of the existing system's inability to promptly identify turning behavior is solved, the security and intelligence of the access control device are improved, and user safety and system efficiency are ensured.

CN118823912BActive Publication Date: 2025-10-03WONLY SECURITY & PROTECTION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411099724.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-12
Publication Date
2025-10-03
Estimated Expiration
2044-08-12

AI Technical Summary

Technical Problem

Existing access control systems are limited in data processing speed and recognition accuracy and are unable to identify the return behavior in a timely manner, causing the door to trap the user during the closing process.

Method used

Through the access control method based on video recognition, the door area is accurately divided into the easy-to-prevent pinching danger zone and the safety buffer zone, and the target objects in the door environment image are detected in real time. The trained target detection model is used for automatic identification and positioning, and control information is output according to the matching results to control the operating status of the access control device.

Benefits of technology

It effectively prevents the access control device from opening incorrectly when people or objects are in dangerous areas, significantly reduces the risk of pinching accidents, improves the security and intelligence level of the access control system, ensures that the access control device is opened or closed at the right time and in the right way, and improves operational efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118823912B_ABST
    Figure CN118823912B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and more specifically to a method and device for access control based on video recognition, the method comprising: obtaining a door area division result and a real-time door environment image; inputting the real-time door environment image into a trained target detection model to identify a target object corresponding to the real-time door environment image and the position information of the target object; matching the real-time door environment image with the door area division result based on the position information of the target object; and outputting control information based on the matching result to control the operating state of an access control device in the access control system. The present invention improves data processing speed and recognition accuracy by accurately dividing door areas and using target objects in the target detection model image, effectively preventing the access control device from mistakenly closing the door when a person or object is in a dangerous area, thereby significantly reducing the risk of pinching accidents and improving the overall security of the access control system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an access control method and device based on video recognition. Background Art

[0002] Access control system is an important part of modern security technology. Its goal is to improve the efficiency and security of access control management by integrating advanced sensor technology and intelligent algorithms.

[0003] The response speed of an access control system directly impacts its real-time performance and reliability. Existing access control systems integrate multiple sensing technologies, such as millimeter-wave radar, infrared laser, and lidar, for real-time target detection and response. While these sensors provide highly accurate environmental perception, their data processing speed and recognition accuracy can be limited. For example, millimeter-wave radar and lidar typically require complex signal processing and data parsing, resulting in slow recognition speeds. Infrared lasers, on the other hand, have limited detection capabilities for small, dim targets with small imaging areas and low reflectivity, and are susceptible to interference from background noise and clutter, resulting in low recognition accuracy.

[0004] In the above scheme, when a user turns back when passing through the access control (for example, the user returns after entering the door because he forgets something), the existing access control system may not be able to identify the turning back behavior in time due to limited data processing speed and recognition accuracy, causing the door to pinch the user during the closing process. Summary of the Invention

[0005] In view of this, the present invention provides an access control method and device based on video recognition to solve the problem that the existing access control system may not be able to identify the return behavior in time due to limited data processing speed and recognition accuracy, causing the door to clamp the user during the closing process.

[0006] In a first aspect, the present invention provides an access control method based on video recognition, which is applied to a controller of an access control system, and comprises:

[0007] Obtaining door area division results and real-time door environment images; the door area division results include easy-to-prevent-pinch danger zones and safety buffer zones;

[0008] Inputting the real-time door environment image into the trained target detection model to identify the target object corresponding to the real-time door environment image and the position information of the target object;

[0009] Matching the real-time door environment image with the door area division result according to the position information of the target object;

[0010] Control information is output according to the matching result to control the operating state of the access control device in the access control system.

[0011] The above solution can effectively prevent the access control device from opening incorrectly when people or objects are in the danger zone by accurately dividing the door area into an anti-pinch danger zone and a safety buffer zone, and detecting target objects in the door environment image in real time, thereby significantly reducing the risk of pinching accidents and improving the overall security of the access control system. The trained target detection model is used to process real-time images to achieve automatic recognition and positioning of target objects, significantly enhancing the intelligence level of the access control system, enabling it to respond more flexibly and accurately to the needs of different scenarios. It also ensures that the access control device is opened or closed at the appropriate time and in the appropriate manner through rapid response and precise control, avoiding user inconvenience caused by misjudgment or delay. At the same time, the intelligent area division and detection mechanism improves the overall user experience. In addition, the above solution reduces the need for manual intervention and improves the operating efficiency of the access control system through automated and intelligent processing. The precise target detection and area matching also reduce system downtime caused by misoperation or failure, further improving the overall effectiveness of the access control system.

[0012] In an optional embodiment, matching the real-time door environment image with the door area division result according to the position information of the target object includes:

[0013] According to the position information of the target object, position matching processing is performed on the target object in the real-time door environment image to obtain whether the target object is located in the easy-to-prevent-pinch danger zone or the safety buffer zone.

[0014] Through precise position matching, the above solution can instantly determine whether the target object (such as pedestrians, pets, etc.) is in a high-risk area or a safe buffer zone, ensuring that the access control device will not malfunction in potentially dangerous situations, thereby effectively preventing the occurrence of safety accidents such as pinching, and greatly improving the security of the access control system.

[0015] In an optional embodiment, outputting control information according to the matching result to control the operating state of the access control device in the access control system includes:

[0016] If the matching result indicates that the target object is located in the easy-to-anti-pinch danger zone, a first level signal is sent to the access control device to control the access control device in the access control system to be in an open door opener state;

[0017] If the matching result indicates that the target object is located in the safety buffer zone, a second level signal is sent to the access control device to control the access control device in the access control system to a safety protection state; the safety protection state includes an instruction to pause the door closing action and an instruction to issue an alarm.

[0018] The above scheme can output corresponding control information according to the matching results and adjust the operating status of the access control device in real time, ensuring that the access control device can make correct decisions in the shortest time and improving the overall security protection effect.

[0019] In an optional embodiment, the target detection model is trained by the following steps:

[0020] Get historical door environment images;

[0021] Annotating the historical door environment image to determine a target object in the historical door environment image and location information of the target object;

[0022] Constructing a training data set based on the historical door environment image and the annotation information corresponding to the historical door environment image;

[0023] The target detection model to be trained is trained using the training data set, and the trained target detection model is obtained.

[0024] The above solution constructs a training dataset by annotating historical door environment images, which helps the target detection model learn accurate feature representation, thereby improving the recognition accuracy and robustness of the model in practical applications.

[0025] In an optional embodiment, the structure of the target detection model includes an input end, a backbone network, a neck structure, and a head structure;

[0026] The input end is used to preprocess the training data set and the real-time door environment image input into the target detection model;

[0027] The backbone network is used to perform feature extraction processing on the training data set and the real-time door environment image input to the target detection model;

[0028] The neck structure is used to perform feature fusion processing on the features extracted by the backbone network;

[0029] The head structure is used to output the target object identified by the target detection model and the position information of the target object.

[0030] In an optional embodiment, training the target detection model to be trained using the training data set and obtaining the trained target detection model includes:

[0031] The images in the training data set are sequentially subjected to adaptive anchor frame calculation processing, adaptive image scaling processing, mosaic data enhancement processing, and normalization processing;

[0032] Performing feature extraction on the normalized training data set and obtaining the extracted features;

[0033] Perform multi-scale feature fusion processing on the extracted features and output feature maps of different scales;

[0034] Based on the feature maps of different scales, the identified target object and the location information of the target object are output, and a trained target detection model is obtained.

[0035] The above scheme uses adaptive anchor box calculation processing to enable the model to dynamically adjust the size and proportion of the anchor box according to the actual size and shape of the target object, thereby more accurately predicting the bounding box of the target object and enhancing the model's detection ability for targets of different sizes and shapes; adaptive image scaling and normalization processing ensure that the images input to the model have consistent size and distribution, which helps the model better learn features; at the same time, mosaic data augmentation processing increases the diversity and complexity of training data by mixing multiple images, which helps to improve the generalization and robustness of the model.

[0036] In an optional embodiment, after outputting control information according to the matching result to control the operating state of the access control device in the access control system, the method further includes:

[0037] The motion trajectory of the target object is obtained from a plurality of continuous real-time door environment images, and the operating state of the access control device is dynamically adjusted according to the motion trajectory.

[0038] By tracking the target object's trajectory in real time, the above solution can more accurately predict the target object's future position, thereby making corresponding safety responses in advance and improving the safety of the system.

[0039] In a second aspect, the present invention provides an access control system based on video recognition, the system comprising: a camera, an access control device, and a controller; the camera is arranged on a door frame for collecting real-time images of the door environment; the controller is used to:

[0040] Obtaining door area division results and real-time door environment images; the door area division results include easy-to-prevent-pinch danger zones and safety buffer zones;

[0041] Inputting the real-time door environment image into the trained target detection model to identify the target object corresponding to the real-time door environment image and the position information of the target object;

[0042] Matching the real-time door environment image with the door area division result according to the position information of the target object;

[0043] Control information is output according to the matching result to control the operating state of the access control device in the access control system.

[0044] In a third aspect, the present invention provides an access control device based on video recognition, which is applied to a controller of an access control system, and includes:

[0045] An acquisition module is used to obtain door area division results and real-time door environment images; the door area division results include easy-to-prevent-pinch danger zones and safety buffer zones;

[0046] A target recognition module is used to input the real-time door environment image into the trained target detection model to identify the target object corresponding to the real-time door environment image and the position information of the target object;

[0047] a matching module, configured to match the real-time door environment image with the door area division result according to the position information of the target object;

[0048] The access control module is used to output control information according to the matching result to control the operating state of the access control device in the access control system.

[0049] In a fourth aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the computer instructions to thereby execute a video recognition-based access control method according to the first aspect or any corresponding embodiment thereof.

[0050] In a fifth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute a video recognition-based access control method according to the first aspect or any corresponding embodiment thereof.

[0051] In a sixth aspect, the present invention provides a computer program product comprising computer instructions, which are used to enable a computer to execute a video recognition-based access control method according to the first aspect or any corresponding embodiment thereof.

[0052] The technical solution provided by the present invention can have the following beneficial effects:

[0053] The present invention can effectively prevent the access control device from opening incorrectly when a person or object is in a dangerous area by accurately dividing the door area into an easy-to-prevent pinching danger zone and a safe buffer zone, and detecting target objects in the door environment image in real time, thereby significantly reducing the risk of pinching accidents and improving the overall safety of the access control system; the trained target detection model is used to process the real-time image to achieve automatic recognition and positioning of the target object, significantly enhancing the intelligence level of the access control system, enabling it to respond more flexibly and accurately to the needs of different scenarios; it also ensures that the access control device is opened or closed at the right time and in the right manner through rapid response and precise control, avoiding user inconvenience caused by misjudgment or delay. At the same time, the intelligent area division and detection mechanism improves the overall user experience. In addition, the present invention reduces the need for manual intervention and improves the operating efficiency of the access control system through automated and intelligent processing procedures, and the precise target detection and area matching also reduce system downtime caused by misoperation or failure, further improving the overall effectiveness of the access control system. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0055] Figure 1 1 is a schematic structural diagram of an access control system based on video recognition according to an embodiment of the present invention;

[0056] Figure 2 is a flow chart of a method for access control based on video recognition according to an embodiment of the present invention;

[0057] Figure 3 is a flow chart of another access control method based on video recognition according to an embodiment of the present invention;

[0058] Figure 4 is a flow chart of another access control method based on video recognition according to an embodiment of the present invention;

[0059] Figure 5 is a structural block diagram of an access control device based on video recognition according to an embodiment of the present invention;

[0060] Figure 6 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0061] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0062] Figure 1 FIG. 1 is a schematic structural diagram of an access control system based on video recognition according to an embodiment of the present invention. Figure 1 As shown, the system includes a camera 110, an access control device 120 and a controller 130;

[0063] The camera 110 can be set on the door frame to collect real-time door environment images; the real-time door environment images are used to indicate the surrounding environment of the target door, such as whether there are target objects such as people and pets;

[0064] The access control device 120 is used to control the door opening and closing states;

[0065] The controller 130 may be a monitoring device based on edge computing, such as a tablet and a personal computer;

[0066] The controller 130 is used to:

[0067] Obtain door area division results and real-time door environment images; the door area division results include easy-to-prevent pinching danger zones and safety buffer zones;

[0068] Inputting the real-time door environment image into the trained target detection model to identify the target object corresponding to the real-time door environment image and the position information of the target object;

[0069] According to the position information of the target object, the real-time door environment image is matched with the door area division result;

[0070] Control information is output according to the matching result to control the operating state of the access control device 120 in the access control system.

[0071] Furthermore, this embodiment combines image recognition, feature extraction, scene matching, and control response, such as Figure 1As shown, the area around the target door is first divided into a high-risk area for easy pinching and a safety buffer zone (including a safety buffer zone in front of the door and a safety buffer zone behind the door). The camera 110 captures a video stream of a specific area (such as the door area) to obtain a real-time door environment image and a historical door environment image. The operator uses the controller 130 (such as a tablet) to select the spatial range to be identified, the target object, and the location of the target object in the historical door environment image, and annotates it, such as by inputting information such as the type of the target object and the area where it is located. After the annotation is completed, this embodiment uploads the annotation information together with the corresponding historical door environment image to the image processing module in the controller 130. After receiving the uploaded historical door environment image and the annotation information, the image processing module uses an image processing algorithm (such as the YOLOv5 target detection model) to extract features from the historical door environment image. The extracted features may include information such as shape, texture, and color, thereby completing the training of the YOLOv5 target detection model. The extracted features and the annotation information are stored in a database for subsequent matching. Afterwards, when camera 110 captures a new video stream, the system processes the video frames in real time to obtain a real-time door environment image. The trained YOLOv5 object detection model is used to identify the target object in the real-time door environment image and extract features. The extracted features are then compared with features stored in the database to find the best match.

[0072] If a dangerous area prone to pinching is identified in the real-time door environment image and a target object matching the object stored in the database exists in the safety buffer, the corresponding control logic is triggered and instructions are sent to the access control device 120. The instructions may include controlling the opening speed and time of the door opener, or the anti-pinch function of the access control device 120.

[0073] Furthermore, the controller 130 in this embodiment uses a monitoring device with edge computing technology to delegate data processing to the monitoring device, realize real-time data processing and analysis, improve response speed and system efficiency, and have low system power consumption, suitable for long-term continuous operation, simple structure, and low installation and maintenance costs.

[0074] In summary, the present invention can effectively prevent the access control device from opening incorrectly when a person or object is in a dangerous area by accurately dividing the door area into an easy-to-prevent pinching danger zone and a safety buffer zone, and detecting target objects in the door environment image in real time, thereby significantly reducing the risk of pinching accidents and improving the overall safety of the access control system; the trained target detection model is used to process the real-time image to achieve automatic recognition and positioning of the target object, significantly enhancing the intelligence level of the access control system, enabling it to respond more flexibly and accurately to the needs of different scenarios; it also ensures that the access control device is opened or closed at the right time and in the right manner through rapid response and precise control, avoiding user inconvenience caused by misjudgment or delay. At the same time, the intelligent area division and detection mechanism improves the overall user experience. In addition, the present invention reduces the need for manual intervention and improves the operating efficiency of the access control system through automated and intelligent processing procedures, and the precise target detection and area matching also reduce system downtime caused by misoperation or failure, further improving the overall effectiveness of the access control system.

[0075] According to an embodiment of the present invention, an embodiment of an access control method based on video recognition is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0076] This embodiment provides an access control method based on video recognition, which can be used to Figure 1 In the controller of the access control system, Figure 2 FIG. 1 is a flow chart of a method for access control based on video recognition according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0077] Step S201 , obtaining a door area division result and a real-time door environment image; the door area division result includes an easy-to-prevent-pinch danger zone and a safety buffer zone.

[0078] Furthermore, in this embodiment, the door area division result is predefined and is used to distinguish areas of different security levels of the door. The door area division result in this embodiment includes an easy-to-prevent-pinch danger zone (that is, the area where the human body or object is most likely to be pinched when the door is about to close or has been closed, which is the area closest to the door frame) and a safety buffer zone (a relatively safe area, including a safety buffer zone in front of the door and a safety buffer zone behind the door). The door area division result can be accurately divided through physical markings (such as lines on the ground), software configuration, or in combination with sensors. The real-time door environment image is an image captured in real time by a camera installed on the door frame, so that the subsequent target detection model can accurately identify objects in the image.

[0079] Step S202 : inputting the real-time door environment image into a trained target detection model to identify a target object corresponding to the real-time door environment image and position information of the target object.

[0080] Furthermore, the target detection model in this embodiment can be a trained YOLOv5 target detection model or other deep learning model. This embodiment processes the input image through the target detection model to identify the target objects in the image (such as people, animals, objects, etc.) and the location information of these objects (such as whether they are located in the door area division result). Before the real-time door environment image is input into the target detection model, the target detection model has been trained using a large amount of labeled image data to learn how to identify different target objects and the location information of the target objects.

[0081] Step S203 : matching the real-time door environment image with the door area division result according to the position information of the target object.

[0082] Furthermore, this embodiment compares the position information (such as the bounding box) of the target object identified by the target detection model with the door area division result to determine which area each target object is currently in (the easy-to-prevent-pinch danger zone or the safety buffer zone). Based on the position information, the system can determine whether the target object is in the danger zone to decide on subsequent control actions.

[0083] Step S204: outputting control information according to the matching result to control the operating state of the access control device in the access control system.

[0084] Furthermore, this embodiment generates corresponding control information based on the matching result between the target object and the door area division result, and the control information is sent to the access control device to adjust its operating status, for example, to adjust the opening and closing speed of the door, delay the closing time, or completely prevent the door from closing.

[0085] In summary, the present invention can effectively prevent the access control device from opening incorrectly when a person or object is in a dangerous area by accurately dividing the door area into an easy-to-prevent pinching danger zone and a safety buffer zone, and detecting target objects in the door environment image in real time, thereby significantly reducing the risk of pinching accidents and improving the overall safety of the access control system; the trained target detection model is used to process the real-time image to achieve automatic recognition and positioning of the target object, significantly enhancing the intelligence level of the access control system, enabling it to respond more flexibly and accurately to the needs of different scenarios; it also ensures that the access control device is opened or closed at the right time and in the right manner through rapid response and precise control, avoiding user inconvenience caused by misjudgment or delay. At the same time, the intelligent area division and detection mechanism improves the overall user experience. In addition, the present invention reduces the need for manual intervention and improves the operating efficiency of the access control system through automated and intelligent processing procedures, and the precise target detection and area matching also reduce system downtime caused by misoperation or failure, further improving the overall effectiveness of the access control system.

[0086] In this embodiment, another access control method based on video recognition is provided, which can be used Figure 1 In the controller of the access control system, Figure 3 FIG. 1 is a flow chart of another access control method based on video recognition according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0087] Step S301, obtaining a door area division result; the door area division result includes an easy-to-prevent-pinch danger zone and a safety buffer zone.

[0088] Further, Figure 4 This is a flow chart of another access control method based on video recognition according to an embodiment of the present invention. This embodiment first determines which parts of the area around the door are easy-to-prevent-pinch danger zones (i.e. Figure 4 Dangerous areas refer to areas that are prone to causing harm to people or objects), and which are safe buffer zones (i.e. Figure 4 The safety zone in this example refers to a relatively safe area that will not cause immediate harm. In this embodiment, the easy-to-prevent-pinch danger zone can be the door frame, and the safety buffer zone can be one meter in front and behind the door. The operator can use the tablet to select and mark the easy-to-prevent-pinch danger zone and the safety buffer zone.

[0089] Step S302 : Acquire a historical door environment image and annotate the historical door environment image to determine a target object in the historical door environment image and position information of the target object.

[0090] Furthermore, this embodiment collects environmental images of the door area in the past, i.e., historical door environmental images, and annotates target objects (such as people, pets, goods, etc.) in these images, marks their location information, uses cameras and other devices to capture historical images, and accurately annotates the target objects and location information in these images manually or automatically (such as using image annotation tools).

[0091] Step S303 : constructing a training data set based on the historical door environment image and the annotation information corresponding to the historical door environment image.

[0092] Furthermore, this embodiment organizes the annotated historical door environment images into a data set that can be used to train the target detection model, that is, organizes the annotated images and their corresponding annotation information into a certain format to form a training data set, which will be used to train the model to identify target objects and their locations in different environments.

[0093] Step S304: train the target detection model to be trained using the training data set, and obtain the trained target detection model.

[0094] In an optional embodiment, the structure of the target detection model includes an input end, a backbone network, a neck structure, and a head structure;

[0095] The input end is used to preprocess the training data set and the real-time door environment image input into the target detection model;

[0096] The backbone network is used to perform feature extraction processing on the training data set and the real-time door environment image input to the target detection model;

[0097] The neck structure is used to perform feature fusion processing on the features extracted by the backbone network;

[0098] The header structure is used to output the target object identified by the target detection model and the location information of the target object.

[0099] In an optional implementation, step S304 includes:

[0100] The images in the training dataset are sequentially processed with adaptive anchor box calculation, adaptive image scaling, mosaic data enhancement, and normalization.

[0101] Performing feature extraction on the normalized training data set and obtaining the extracted features;

[0102] Perform multi-scale feature fusion processing on the extracted features and output feature maps of different scales;

[0103] Based on the feature maps of different scales, the identified target object and the location information of the target object are output, and a trained target detection model is obtained.

[0104] Furthermore, this embodiment uses a training dataset to train a target detection model that can accurately identify target objects and their locations. First, the images in the training dataset undergo a series of preprocessing steps, such as adaptive anchor box calculation, image scaling, data augmentation, and normalization, to improve the model's generalization and detection accuracy. Next, a backbone network is used to extract features from the images, perform feature fusion based on the neck structure, and output detection results based on the head structure. Finally, through multiple iterations of training, the model parameters are continuously adjusted until the model achieves satisfactory performance on the validation set.

[0105] Furthermore, the adaptive anchor box calculation process is to better match the shape and size of the target object in the image. The anchor box is a predefined set of rectangular boxes of different sizes and proportions used as the initial estimate of the target object's bounding box during the training process. The steps of the adaptive anchor box calculation process include:

[0106] Analyze the shape and size distribution of the target objects in the training dataset; based on the analysis results, dynamically adjust the size and scale of the anchor box according to the actual bounding box of the target object, so that the anchor box is closer to the actual bounding box of the target object (during the training process, the model will use adaptive anchor boxes to predict the bounding box of the target object).

[0107] Since the images in the training dataset may have different sizes and aspect ratios, these images need to be scaled in order to be uniformly input into the model. Therefore, this embodiment adopts adaptive image scaling. The steps of adaptive image scaling include:

[0108] According to the input requirements of the model, set the target size (such as 640x640, 800x800, etc.); use a scaling algorithm (such as bilinear interpolation, nearest neighbor interpolation, etc.) to scale the historical door environment image to the target size; when scaling, fill the extra space by adding padding (which can be black or gray) on the edge of the image to maintain the aspect ratio of the image.

[0109] Mosaic data augmentation combines four training images to create a new training image, thereby increasing the generalization and robustness of the model. The steps of Mosaic data augmentation include:

[0110] Four training images (historical door environment images) are randomly selected from the training dataset and stitched together to form a new stitched image that contains more information and context; the stitched image is annotated to ensure the accuracy of the bounding box of the target object in the new image.

[0111] Normalization involves scaling the pixel values ​​of image data to a specific range (such as 0 to 1 or -1 to 1) to facilitate model learning and convergence. Normalization involves transforming each pixel value in the image so that it falls within the specified range. Normalization helps eliminate differences in brightness, contrast, and other factors between images, allowing the model to focus more on the image's content and structure.

[0112] In addition, this embodiment uses a backbone network (such as ResNet and VGG) to extract features from the normalized image to capture key information in the image. This includes: inputting the normalized image into the backbone network, which then passes through a series of convolutional layers, pooling layers, and other structures to gradually extract low-level to high-level features in the image; ultimately, it outputs a series of feature maps, which contain abstract representations of the image at different scales.

[0113] In this embodiment, the multi-scale feature fusion processing of the extracted features is to improve the model's detection ability for target objects of different sizes. It is necessary to fuse feature maps of different scales. The steps include: fusing feature maps of different scales output by the backbone network through the neck structure (such as FPN, PANet, etc.). The fusion method may include upsampling, downsampling, feature splicing, etc. The fused feature map contains both high-resolution detail information and low-resolution semantic information, which helps the model detect target objects more accurately.

[0114] Finally, based on the fused feature map, the model of this embodiment outputs the category and bounding box position of the target object. This includes: a head structure (such as a detection head) receives the fused feature map as input; predicts the position on each feature map, and outputs whether there is a target object at that position, the category of the target object, and the position and size of the bounding box; removes redundant bounding boxes through post-processing techniques such as non-maximum suppression (NMS) to obtain the final detection result, which includes the target object and the location information of the target object. In this embodiment, the model parameters are continuously adjusted to minimize the loss function during repeated iterative training. During the training process, the performance of the model is evaluated through the validation set, and the training strategy is adjusted as needed. When the performance of the model on the validation set reaches the predetermined standard, the training is stopped and the model is saved, and finally a model that can accurately detect the target object is obtained.

[0115] That is, during the training process, this embodiment first collects the historical door environment image through the camera, annotates the historical door environment image, and then obtains the target object contained in the image and its corresponding position in the image. Then, the image file and the corresponding annotation file are put into the Yolov5 target detection model. The Yolov5 target detection model mainly includes four modules: Input input terminal, Backbone backbone network, Neck structure and Head structure (Prediction). They are respectively responsible for input image preprocessing, feature extraction, feature fusion, and output detection information.

[0116] Regarding the Input input end: The Input input end mainly preprocesses the input image. The preprocessing mainly involves scaling the input image to the network input size and performing normalization and other operations. In addition, this embodiment also proposes an adaptive anchor frame calculation and adaptive image scaling method. Using mosaic data enhancement processing, four images are randomly read from the training data set, and the four images are flipped (the original image is flipped left and right), scaled (the original image is scaled in size), and the color gamut is changed (the brightness, saturation, and hue of the original image are changed). Then, the four images are combined into one image, and some useless frames are filtered out. Traditional scaling methods scale images to their original proportions and fill them with black to the target size. However, in practice, many images have different aspect ratios. Therefore, after scaling and filling, the black borders at both ends may not be the same size. Excessive filling can lead to a large amount of information redundancy, which can affect the inference speed of the entire algorithm. This embodiment uses letterbox image processing technology, which adds gray, white, or other colored borders to the top, bottom, left, and right sides of the original historical door environment image to adapt the image's aspect ratio to a specific ratio. This allows for the adaptive addition of minimal black borders to the scaled image, reducing the computational effort during inference and improving target detection speed. For example, if an image is 800*600 and the original scaled size is 416x416, dividing both by the original image size yields two scaling factors, 0.52 and 0.69. The smaller factor, 0.52, is selected. To calculate the scaled size, the original image's length and width are multiplied by the minimum scaling factor, 0.52, resulting in a width of 416 and a height of 312. Calculate the black edge padding value and subtract 416 - 312 = 104 to get the original height required. Use the numpy method np.mod to take the remainder, which gives 8 pixels. Divide by 2 to get the required padding value at both ends of the image.

[0117] For the Backbone backbone network: Backbone (BottleNeckCSP structure) is constructed in series by Focus structure, three groups of CBL+CSP1_x and CBL+SPP. Among them, CBL: consists of convolution Conv+batch normalization BN+activation function Leaky Relu. Focus: Slice the image and then Concat. CSP1_x: consists of CBL module, Res uint module and convolution layer Concat, where x means there are x CSP1 modules. CSP2_x: No longer uses the Res unit residual component, and consists of convolution layer CBL module Concat. SPP: Uses 1×1, 5×5, 9×9, 13×13 maximum pooling method for multi-scale fusion. Res unit: residual component, after the input passes through two CBLs, it is added with the original input. It is a conventional residual unit, so that the network can extract deeper features while avoiding gradient disappearance or explosion. Upsampling: Increases the size of a feature map by replicating elements, such as linear interpolation. Concat: Concatenates tensors, expanding the dimensions of two tensors to achieve multi-scale feature fusion. For example, concatenating two tensors of 26×26×256 and 26×26×512 results in 26×26×768. Add: Adds tensors directly, without increasing their dimensions. For example, adding 104×104×128 and 104×104×128 still results in 104×104×128. Y1, Y2, and Y3: Represent the outputs of Yolov5 at three scales.

[0118] It can be seen that the structure of the Yolov5 target detection model in this embodiment deletes two convolutional layers, changes the activation function, and performs model quantization operations, which greatly improves the recognition frame rate of the rknn chip during operation.

[0119] Step S305: Acquire a real-time door environment image.

[0120] Furthermore, this embodiment uses a camera to capture the current environment image of the door area in real time to perform target detection.

[0121] Step S306 : inputting the real-time door environment image into the trained target detection model to identify the target object corresponding to the real-time door environment image and the position information of the target object.

[0122] Furthermore, this embodiment utilizes a trained target detection model to identify target objects and their positions in real-time door environment images, that is, the real-time image is input into the target detection model, and the model outputs the identified target objects and their position information.

[0123] Step S307 : performing position matching processing on the target object in the real-time door environment image according to the position information of the target object to determine whether the target object is located in the easy-to-prevent-pinch danger zone or the safety buffer zone.

[0124] Furthermore, this embodiment performs position matching based on the position information of the target object and the door area division result to determine whether the identified target object is located in the easy-to-prevent-pinch danger zone or the safety buffer zone.

[0125] Step S308: output control information according to the matching result to control the operating state of the access control device in the access control system.

[0126] In an optional implementation, step S308 includes:

[0127] If the matching result indicates that the target object is located in the easy-to-prevent-pinch danger zone, a first level signal is sent to the access control device to control the access control device in the access control system to be in an open door opener state;

[0128] If the matching result indicates that the target object is located in the safety buffer zone, a second level signal is sent to the access control device to control the access control device in the access control system to a safety protection state; the safety protection state includes an instruction to pause the door closing action and an instruction to issue an alarm.

[0129] Furthermore, this embodiment controls the operating state of the access control device based on the location of the target object to ensure the safety of people or objects. If the target object is located in a high-risk area, a first level signal is sent to the access control device to activate the door opener to prevent pinching. If the target object is located in a safety buffer zone, a second level signal is sent to the access control device to activate a safety protection state, which may include pausing door closing and sounding an alarm. The first level signal and the second level signal may differ, such as when the first level signal is high and the second level signal is low, or alternatively, when the first level signal is low and the second level signal is high.

[0130] Step S309: Acquire the motion trajectory of the target object from a plurality of continuous real-time door environment images, and dynamically adjust the operating state of the access control device according to the motion trajectory.

[0131] Furthermore, this embodiment tracks the motion trajectory of the target object from multiple continuous real-time door environment images, predicts its possible future movement path, and adjusts the operating status of the access control device in a timely manner based on the prediction results to ensure the safe passage of people or objects.

[0132] In summary, the present invention can effectively prevent the access control device from opening incorrectly when a person or object is in a dangerous area by accurately dividing the door area into an easy-to-prevent pinching danger zone and a safety buffer zone, and detecting target objects in the door environment image in real time, thereby significantly reducing the risk of pinching accidents and improving the overall safety of the access control system; the trained target detection model is used to process the real-time image to achieve automatic recognition and positioning of the target object, significantly enhancing the intelligence level of the access control system, enabling it to respond more flexibly and accurately to the needs of different scenarios; it also ensures that the access control device is opened or closed at the right time and in the right manner through rapid response and precise control, avoiding user inconvenience caused by misjudgment or delay. At the same time, the intelligent area division and detection mechanism improves the overall user experience. In addition, the present invention reduces the need for manual intervention and improves the operating efficiency of the access control system through automated and intelligent processing procedures, and the precise target detection and area matching also reduce system downtime caused by misoperation or failure, further improving the overall effectiveness of the access control system.

[0133] This embodiment also provides a video recognition-based access control device, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0134] This embodiment provides an access control device based on video recognition, which is applied to Figure 1 The controller of the access control system shown is as follows Figure 5 Shown, including:

[0135] The acquisition module 501 is used to obtain the door area division result and the real-time door environment image; the door area division result includes the easy-to-prevent-pinch danger zone and the safety buffer zone;

[0136] The target recognition module 502 is used to input the real-time door environment image into the trained target detection model to identify the target object corresponding to the real-time door environment image and the location information of the target object;

[0137] A matching module 503 is configured to match the real-time door environment image with the door area division result according to the location information of the target object;

[0138] The access control module 504 is configured to output control information according to the matching result to control the operating state of the access control device in the access control system.

[0139] In some optional implementations, the matching module 503 is further configured to:

[0140] According to the position information of the target object, position matching processing is performed on the target object in the real-time door environment image to obtain whether the target object is located in the easy-to-prevent-pinch danger zone or the safety buffer zone.

[0141] In some optional implementations, the access control module 504 is further configured to:

[0142] If the matching result indicates that the target object is located in the easy-to-prevent-pinch danger zone, a first level signal is sent to the access control device to control the access control device in the access control system to be in an open door opener state;

[0143] If the matching result indicates that the target object is located in the safety buffer zone, a second level signal is sent to the access control device to control the access control device in the access control system to a safety protection state; the safety protection state includes an instruction to pause the door closing action and an instruction to issue an alarm.

[0144] In some optional embodiments, the device is further used to:

[0145] The object detection model is trained by the following steps:

[0146] Get historical door environment images;

[0147] Annotating the historical door environment image to determine a target object in the historical door environment image and location information of the target object;

[0148] Constructing a training data set based on the historical door environment image and the annotation information corresponding to the historical door environment image;

[0149] The target detection model to be trained is trained using the training data set, and a trained target detection model is obtained.

[0150] In some optional embodiments, the structure of the target detection model includes an input end, a backbone network, a neck structure, and a head structure;

[0151] The input end is used to preprocess the training data set and the real-time door environment image input into the target detection model;

[0152] The backbone network is used to perform feature extraction processing on the training data set and the real-time door environment image input to the target detection model;

[0153] The neck structure is used to perform feature fusion processing on the features extracted by the backbone network;

[0154] The header structure is used to output the target object identified by the target detection model and the location information of the target object.

[0155] In some optional embodiments, the device is further used to:

[0156] The images in the training dataset are sequentially processed with adaptive anchor box calculation, adaptive image scaling, mosaic data enhancement, and normalization.

[0157] Performing feature extraction on the normalized training data set and obtaining the extracted features;

[0158] Perform multi-scale feature fusion processing on the extracted features and output feature maps of different scales;

[0159] Based on the feature maps of different scales, the identified target object and the location information of the target object are output, and a trained target detection model is obtained.

[0160] In some optional embodiments, the device is further used to:

[0161] The motion trajectory of the target object is obtained from a plurality of continuous real-time door environment images, and the operating state of the access control device is dynamically adjusted according to the motion trajectory.

[0162] The further functional description of each of the above modules is the same as that of the above corresponding embodiments and will not be repeated here.

[0163] In this embodiment, a video recognition-based access control device is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0164] In summary, the present invention can effectively prevent the access control device from opening incorrectly when a person or object is in a dangerous area by accurately dividing the door area into an easy-to-prevent pinching danger zone and a safety buffer zone, and detecting target objects in the door environment image in real time, thereby significantly reducing the risk of pinching accidents and improving the overall safety of the access control system; the trained target detection model is used to process the real-time image to achieve automatic recognition and positioning of the target object, significantly enhancing the intelligence level of the access control system, enabling it to respond more flexibly and accurately to the needs of different scenarios; it also ensures that the access control device is opened or closed at the right time and in the right manner through rapid response and precise control, avoiding user inconvenience caused by misjudgment or delay. At the same time, the intelligent area division and detection mechanism improves the overall user experience. In addition, the present invention reduces the need for manual intervention and improves the operating efficiency of the access control system through automated and intelligent processing procedures, and the precise target detection and area matching also reduce system downtime caused by misoperation or failure, further improving the overall effectiveness of the access control system.

[0165] The present invention also provides a computer device. Figure 6 , Figure 6 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 610, memory 620, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 610 is taken as an example.

[0166] Processor 610 may be a central processing unit, a network processor, or a combination thereof. Processor 610 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0167] The memory 620 stores instructions that can be executed by at least one processor 610, so that the at least one processor 610 executes the method shown in the above embodiment.

[0168] The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 620 may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 620 may optionally include a memory remotely located relative to the processor 610, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0169] The memory 620 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 620 may also include a combination of the above types of memory.

[0170] The computer device further includes a communication interface 630 for the computer device to communicate with other devices or a communication network.

[0171] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0172] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0173] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the defined scope.

Claims

1. A video recognition-based access control method, characterized in that: The method is applied to a controller of an access control system, and the method includes: Obtaining door area division results and real-time door environment images; the door area division results include easy-to-prevent-pinch danger zones and safety buffer zones; Inputting the real-time door environment image into the trained target detection model to identify the target object corresponding to the real-time door environment image and the position information of the target object; Matching the real-time door environment image with the door area division result according to the position information of the target object; Outputting control information based on the matching result to control the operating state of the access control device in the access control system, including: if the matching result indicates that the target object is located in the easy-to-prevent-pinch danger zone, sending a first level signal to the access control device to control the access control device in the access control system to an open door opener state; if the matching result indicates that the target object is located in the safety buffer zone, sending a second level signal to the access control device to control the access control device in the access control system to a safety protection state; the safety protection state includes an indication to pause the door closing action and an indication to issue an alarm.

2. The method according to claim 1, characterized in that The matching of the real-time door environment image with the door area division result according to the position information of the target object includes: According to the position information of the target object, position matching processing is performed on the target object in the real-time door environment image to obtain whether the target object is located in the easy-to-prevent-pinch danger zone or the safety buffer zone.

3. The method according to claim 1, characterized in that The target detection model is trained by the following steps: Get historical door environment images; Annotating the historical door environment image to determine a target object in the historical door environment image and location information of the target object; Constructing a training data set based on the historical door environment image and the annotation information corresponding to the historical door environment image; The target detection model to be trained is trained using the training data set, and the trained target detection model is obtained.

4. The method according to claim 3, characterized in that The structure of the target detection model includes an input end, a backbone network, a neck structure and a head structure; The input end is used to preprocess the training data set and the real-time door environment image input into the target detection model; The backbone network is used to perform feature extraction processing on the training data set and the real-time door environment image input to the target detection model; The neck structure is used to perform feature fusion processing on the features extracted by the backbone network; The head structure is used to output the target object identified by the target detection model and the position information of the target object.

5. The method according to claim 4, characterized in that The step of training the target detection model to be trained using the training data set and obtaining the trained target detection model includes: The images in the training data set are sequentially subjected to adaptive anchor frame calculation processing, adaptive image scaling processing, mosaic data enhancement processing, and normalization processing; Performing feature extraction on the normalized training data set and obtaining the extracted features; Perform multi-scale feature fusion processing on the extracted features and output feature maps of different scales; Based on the feature maps of different scales, the identified target object and the location information of the target object are output, and a trained target detection model is obtained.

6. The method according to any one of claims 1 to 5, characterized in that After outputting control information according to the matching result to control the operating state of the access control device in the access control system, the method further includes: The motion trajectory of the target object is obtained from a plurality of continuous real-time door environment images, and the operating state of the access control device is dynamically adjusted according to the motion trajectory.

7. An access control system based on video recognition, characterized in that: The system includes: a camera, an access control device, and a controller; the camera is set on the door frame for collecting real-time door environment images; the controller is used to: Obtaining door area division results and real-time door environment images; the door area division results include easy-to-prevent-pinch danger zones and safety buffer zones; Inputting the real-time door environment image into the trained target detection model to identify the target object corresponding to the real-time door environment image and the position information of the target object; Matching the real-time door environment image with the door area division result according to the position information of the target object; Outputting control information based on the matching result to control the operating state of the access control device in the access control system, including: if the matching result indicates that the target object is located in the easy-to-prevent-pinch danger zone, sending a first level signal to the access control device to control the access control device in the access control system to an open door opener state; if the matching result indicates that the target object is located in the safety buffer zone, sending a second level signal to the access control device to control the access control device in the access control system to a safety protection state; the safety protection state includes an indication to pause the door closing action and an indication to issue an alarm.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the access control method based on video recognition according to any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the access control method based on video recognition according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent door control system on basis of machine learning and control method implemented by intelligent door control system

    CN107909687A

  • Vehicle door control method and device, electronic equipment and storage medium

    CN118148461A