Intelligent access control identification system
By combining video recognition, motion recognition, and voice monitoring, the intelligent access control system solves the problems of multimodal authentication and dynamic behavior monitoring in traditional access control systems, achieving high security and flexible authentication, and adapting to the needs of complex environments and international venues.
Patent Information
- Application Number
- CN202510164221.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-02-14
AI Technical Summary
Traditional access control systems lack multimodal authentication and dynamic behavior monitoring, leading to security risks and false alarms, and are difficult to adapt to the needs of complex environments and international venues.
By employing image acquisition, object recognition, action recognition, and voice recognition modules, combined with a well-trained access control recognition model, multi-dimensional identity verification and dynamic behavior monitoring are achieved. The access control system is controlled by video recognition of object categories, action recognition, and voice analysis.
It improves the system's security and resistance to attacks, can identify abnormal behavior and issue timely alerts, adapts to different environmental conditions, and enhances the user experience.
Smart Images

Figure CN120071494B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent access control, and in particular to an intelligent access control recognition system and method. BACKGROUND
[0002] With the rapid development of social economy and the continuous progress of technology, people's demand for safety management is increasing, especially in places with high security requirements such as residential communities, office buildings, financial institutions, government agencies, etc. The traditional access control system has been difficult to meet the high requirements of modern security.
[0003] Traditional access control systems usually rely on a single identity verification method, such as password input, access card (RFID card), fingerprint recognition, or facial recognition, etc. Although these systems have improved security to some extent, they still have many problems and potential risks. Specifically, traditional access control systems lack dynamic monitoring of user behavior (such as actions), and cannot timely identify abnormal behavior (such as tailing, loitering, violent destruction, etc.), leading to security risks. Traditional access control systems are prone to false positives in complex environments (such as insufficient light, environmental noise interference, obstruction influence, etc.). For example, facial recognition systems may not accurately identify users in dark or low-light environments, and voice recognition systems may misjudge legitimate instructions in noisy environments. These false positives not only affect user experience, but also increase the workload of security personnel. In complex usage scenarios (such as multi-language environments, multi-cultural backgrounds, different lighting conditions, etc.), traditional access control systems often perform poorly. For example, a single voice recognition function cannot meet the needs of multi-language users, and facial recognition may fail in environments with insufficient light or direct sunlight. This limits the widespread application of access control systems, especially in internationalized places or complex environments. SUMMARY
[0004] Therefore, it is necessary to provide an intelligent access control recognition system and method to solve the technical problems that the existing access control system cannot perform multi-modal identity verification and lacks dynamic behavior monitoring.
[0005] To solve the above problems, the present application provides an intelligent access control monitoring system, comprising:
[0006] An image acquisition module is configured to acquire a monitoring video and a voice recording identified by an access control system, and obtain an image frame sequence from the monitoring video;
[0007] An object recognition module is configured to identify an object category in an image based on a trained complete access control recognition model and the image frame sequence;
[0008] An action recognition module is configured to identify an action of an object based on a trained complete access control recognition model and the object category and the image frame sequence.
[0009] a voice recognition module, configured to analyze the voice record based on the trained access control recognition model to obtain an analysis result, and control opening and closing of the access control system according to the analysis result.
[0010] In a possible implementation, before the object category recognition and the action recognition are performed, the method further includes:
[0011] obtaining a multi-type sample set; the multi-type sample set includes a face sample set, an object sample set, and a scene sample set; the face sample set includes human sample data and a face label, the object sample set includes object data and an object label, and the scene sample set includes scene data and a scene label;
[0012] training an initial access control recognition model based on the multi-type sample set to obtain the trained access control recognition model.
[0013] In a possible implementation, the access control recognition model includes an object recognition module, an action recognition module, and a voice recognition module.
[0014] The object recognition module includes a first convolutional layer, a first pooling layer, a second convolutional layer, a third convolutional layer, a second pooling layer, a fourth convolutional layer, a fifth convolutional layer, and a third pooling layer; and the object category in the image is recognized based on the image frame sequence, and includes:
[0015] first scale feature extraction is performed on the image frame sequence based on the first convolutional layer and the first pooling layer to obtain a first feature map;
[0016] second scale feature extraction is performed on the first feature map based on the second convolutional layer, the third convolutional layer, and the second pooling layer to obtain a second feature map;
[0017] third scale feature extraction is performed on the second feature map based on the fourth convolutional layer, the fifth convolutional layer, and the third pooling layer to obtain the object category.
[0018] In a possible implementation, the action recognition module includes a convolutional layer, an adjustment layer, and an action connection layer; and the action of the object is recognized based on the object category and the image frame sequence, and includes:
[0019] sampling feature extraction is performed on corresponding features of the object category based on the convolutional layer to obtain sampling features;
[0020] adjustment features are obtained by adjusting the sampling features based on the features of the object category based on the adjustment layer;
[0021] The action connection layer performs supervised learning on the adjustment features to obtain an action recognition result of the object.
[0022] In a possible implementation, the speech recognition module comprises a first full connection layer, a second full connection layer, and an activation function layer; and the analyzing the speech record according to the action of the object to obtain an analysis result comprises:
[0023] The first full connection layer is configured to perform classification processing on the features corresponding to the action of the object to obtain an action classification result.
[0024] The second full connection layer is configured to perform extraction on the content of the speech record according to the classification result of the action of the object to obtain speech features.
[0025] The activation function layer is configured to perform activation processing on the speech features to obtain a recognition result of speech recognition.
[0026] In a possible implementation, the activation function layer is formed according to an activation function established based on a first image frame of the image frame sequence.
[0027] In a possible implementation, the analyzing the speech record to obtain an analysis result and controlling opening and closing of an access control system according to the analysis result comprises:
[0028] extracting acoustic features from the speech record;
[0029] decoding the acoustic features according to the speech recognition module to obtain a hidden state label sequence;
[0030] matching each hidden state label in the hidden state label sequence with a plurality of preset state labels to obtain access opening information corresponding to the speech data;
[0031] controlling opening and closing of the access control system according to the access opening information.
[0032] Each preset state label is used to represent an access opening information.
[0033] In a possible implementation, the obtaining a monitoring video recognized by the access control system and obtaining an image frame sequence according to the monitoring video comprises:
[0034] obtaining the monitoring video, and extracting a video segment of a preset time length from the monitoring video to obtain an initial image frame sequence;
[0035] selecting image frames at odd positions in the initial image frame sequence to obtain the image frame sequence.
[0036] The beneficial effects of the present application are: first, obtaining the monitoring video and voice record recognized by the access control system, and obtaining an image frame sequence according to the monitoring video; then, based on the trained access control recognition model, recognizing the object category in the image according to the image frame sequence; and based on the trained access control recognition model, recognizing the action of the object in the image frame sequence according to the object category; finally, combining the action of the object, analyzing the voice record to obtain an analysis result, and controlling the opening and closing of the access control system according to the analysis result. By combining video recognition, action recognition and voice monitoring recognition, multi-dimensional identity verification is realized, thereby improving the security and attack resistance of the system, and at the same time, the user's action such as walking speed, gesture, posture, etc. is analyzed in real time to identify potential abnormal behaviors such as loitering, violent destruction, etc., and timely alarm or protective measures are taken. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 The structure schematic diagram of an embodiment of the intelligent access control system provided by the present application is shown in the figure.
[0038] Figure 2 In the intelligent access control system provided by the present application, the structure schematic diagram of an embodiment of the object recognition module is shown in the figure.
[0039] Figure 3 In the intelligent access control system provided by the present application, the structure schematic diagram of an embodiment of the action recognition module is shown in the figure.
[0040] Figure 4 The running environment schematic diagram of an embodiment of the electronic device provided by the present application is shown in the figure. DETAILED DESCRIPTION
[0041] The preferred embodiments of the present application will be specifically described below in combination with the drawings, wherein the drawings constitute a part of the present application, and are used to illustrate the principles of the embodiments of the present application, and are not used to limit the scope of the present application.
[0042] One specific embodiment of the present application discloses an intelligent access control system 10, please refer to Figure 1 , comprising:
[0043] The image acquisition module 110 is used to acquire the monitoring video and voice record recognized by the access control system, and obtain an image frame sequence according to the monitoring video;
[0044] The object recognition module 120 is used to recognize the object category in the image based on the trained access control recognition model according to the image frame sequence;
[0045] The action recognition module 130 is used to recognize the action of the object based on the trained access control recognition model according to the object category and the image frame sequence;
[0046] The voice recognition module 140 is configured to analyze the voice record based on the trained access control recognition model and the action of the object to obtain an analysis result, and control the opening and closing of the access control system according to the analysis result.
[0047] In the embodiment, first, the monitoring video and the voice record recognized by the access control system are obtained, and an image frame sequence is obtained according to the monitoring video. Then, the object category in the image is recognized based on the trained access control recognition model and the image frame sequence. The action of the object in the image frame sequence is recognized based on the trained access control recognition model and the object category. Finally, the voice record is analyzed based on the action of the object to obtain an analysis result, and the opening and closing of the access control system is controlled according to the analysis result. Through the combination of video recognition, action recognition and voice monitoring recognition, multi-dimensional identity verification is realized, thereby improving the security and attack resistance of the system. At the same time, potential abnormal behaviors such as loitering and violent destruction are identified by analyzing the user's action such as walking speed, gesture and posture in real time, and an alarm or protective measures are sent in time.
[0048] In a specific implementation, before the object category recognition and the action recognition, the following steps are further included.
[0049] A multi-type sample set is obtained. The multi-type sample set includes a face sample set, an object sample set and a scene sample set. The face sample set includes face data and face labels, the object sample set includes object data and object labels, and the scene sample set includes scene data and scene labels.
[0050] The initial access control recognition model is trained based on the multi-type sample set to obtain the trained access control recognition model.
[0051] In the embodiment, after the training of the multi-type sample set, the trained access control recognition model can accurately recognize the face and perform identity verification, can recognize the objects related to the access control such as the door handle and the access card, can adapt to different environmental conditions such as light and background changes, and can further recognize the action of the user (such as card swiping and door handle pressing) based on the recognition results of the face, the object and the scene.
[0052] In an embodiment, the access control recognition model includes an object recognition module, an action recognition module and a voice recognition module.
[0053] Please refer to Figure 2 The object recognition module includes a first convolutional layer 210, a first pooling layer 220, a second convolutional layer 230, a third convolutional layer 240, a second pooling layer 250, a fourth convolutional layer 260, a fifth convolutional layer 270 and a third pooling layer 280. The object category in the image is recognized based on the image frame sequence, which includes:
[0054] perform first scale feature extraction on the image frame sequence based on the first convolutional layer and the first pooling layer to obtain a first feature map;
[0055] perform second scale feature extraction on the first feature map based on the second convolutional layer, the third convolutional layer and the second pooling layer to obtain a second feature map;
[0056] perform third scale feature extraction on the second feature map based on the fourth convolutional layer, the fifth convolutional layer and the third pooling layer to obtain the object category.
[0057] In the embodiment, specifically, the convolution kernel scale of the first convolutional layer is 3x3, the step is 1, the convolution kernel scale of the second convolutional layer is 11x11, the step is 2, the convolution kernel scale of the third convolutional layer is 3x3, the step is 1, and the convolution kernel scale of the fourth convolutional layer and the fifth convolutional layer is 3x3, the step is 1.
[0058] After the image frame sequence is input into the object recognition module, the first convolutional layer 210 with a size of 3x3 and a step of 1 is directly used to perform convolution operation, i.e., feature extraction operation, on the image frame sequence, the obtained feature data is input into the first pooling layer 220 for processing to reduce the calculation amount and prevent overfitting, after the first feature map is obtained, the second convolutional layer 230 with a size of 11x11 and a step of 2 is sequentially used to perform convolution operation on the first feature map, the obtained feature map is continuously input into the third convolutional layer 240 with a size of 3x3 and a step of 1 to perform convolution operation, different scale feature extraction is realized, and then the second feature map with rich feature data of different levels is obtained through the second pooling layer 250, and then the fourth convolutional layer 260 with a size of 3x3 and a step of 1 and the fifth convolutional layer 270 with a size of 3x3 and a step of 1 are sequentially used to perform convolution operation on the second feature map, and the third pooling layer 280 is used to process the extracted feature data to obtain the object category.
[0059] Compared with the feature data obtained by the feature extraction module of the single task network, the object category obtained by the method of the embodiment contains richer domain features and image features, and better meets the requirement of recognition accuracy of rich object categories.
[0060] Further, the image sharing feature of the to-be-recognized image is extracted through the object recognition module, hidden public information and correlation between features between different attribute person recognitions are mined, and the recognition performance is improved.
[0061] In some embodiments, the action recognition module includes a convolutional layer, an adjustment layer and an action connection layer; and the action of the object is recognized according to the object category and the image frame sequence, including:
[0062] The convolutional layer samples and extracts features of the object category to obtain sample features;
[0063] The adjustment layer adjusts the sample features based on the features of the object category to obtain adjustment features;
[0064] The action connection layer performs supervised learning on the adjustment features to obtain an action recognition result of the object.
[0065] In this embodiment, the adjustment features obtained by combining the convolutional layer and the adjustment layer contain more action characteristics, and the adjustment features have stronger resolution and can more accurately distinguish action characteristics, thereby improving the accuracy of the action recognition result.
[0066] In some embodiments, referring to Figure 3 , the speech recognition module includes a first fully connected layer 310, a second fully connected layer 320, and an activation function layer 330; the analysis of the speech record based on the action of the object to obtain an analysis result includes:
[0067] The first fully connected layer classifies and processes features corresponding to the action of the object to obtain an action classification result;
[0068] The second fully connected layer extracts content of the speech record based on the classification result of the action of the object to obtain speech features;
[0069] The activation function layer activates the speech features to obtain a recognition result of speech recognition.
[0070] and analyzing the speech record to obtain an analysis result, and controlling the opening and closing of the access control system based on the analysis result, includes:
[0071] Extracting acoustic features from the speech record;
[0072] Decoding the acoustic features based on the speech recognition module to obtain a hidden state label sequence;
[0073] Matching each hidden state label in the hidden state label sequence with a plurality of preset state labels to obtain access opening information corresponding to the speech data;
[0074] Controlling the opening and closing of the access control system based on the access opening information;
[0075] Each of the preset state labels is used to represent a type of access opening information.
[0076] In the present embodiment, a Hidden Markov Model (HMM) is used to extract acoustic features from the voice recording.
[0077] An HMM consists of two key components: hidden states and observation states. Hidden states are states that cannot be directly observed and represent some internal state of the system. Observation states are states that can be directly observed and represent the data that we can observe. The basic principle of an HMM is that the hidden states form a probability transition matrix that defines the transition probabilities between states. It should be understood that other existing speech recognition models can also be used as the preset speech recognition model.
[0078] In the above process, the hidden state sequence is a sequence of labels of hidden states obtained by a decoding algorithm in a speech recognition task, used to represent the evolution process of the input speech and the corresponding speech units. It is a sequence composed of a series of labels of hidden states. Each hidden state corresponds to a specific label, representing the speech unit (such as phonemes, words, syllables, etc.) corresponding to the speech signal at a certain time point.
[0079] The generation of the hidden state sequence is based on the decoding process of the speech signal, which infers the observation probability and state transition probability in the HMM model. In the decoding process, according to the feature extraction of the audio signal and the parameters of the HMM model, the Viterbi algorithm and other decoding algorithms are used to find the most likely hidden state sequence, thereby realizing the modeling and recognition of the input speech.
[0080] The hidden state sequence is a discrete representation of the speech signal in time, and each time point in a piece of speech corresponds to a hidden state label. By obtaining the hidden state sequence, we can understand the changes of the speech signal in time and the composition of the speech units according to the order and conversion relationship of the labels, and then perform tasks such as speech recognition, speech understanding, and speech synthesis.
[0081] Further, the activation layer can take the first frame of image in the sequence of image frames as a reference to perform denoising processing on each feature map, and eliminate pixels with small changes, i.e. eliminate pixels containing less action information, so that the neural network can identify the pixels containing action information, and improve the identification speed.
[0082] Specifically, in a preferred embodiment, the activation function established is:
[0083]
[0084]
[0085] wherein, is In the grayscale feature map under the channel, located at the first OK, The grayscale value of the pixels in the column. for In the first frame grayscale image under the channel, located at the... line, number The grayscale value of the pixels in the column. for In the background grayscale image under the channel, the one located at the 1st line, number The grayscale value of the pixels in the column. These are the color channels for red, green, and blue, respectively.
[0086] The above formula distinguishes differences by directly comparing the grayscale feature map, the first frame grayscale image, and the background grayscale image in each color channel. Pixels whose grayscale feature map does not match the grayscale values of both the first frame and the background grayscale image are retained, while pixels whose grayscale values are the same in all three images are considered to contain no action information and are set to 0 to filter out pixels containing action information. It is understandable that in practice, other computational methods can be flexibly used to establish the activation function depending on the specific situation.
[0087] The timing of extracting video segments of a preset duration can be flexibly set according to actual conditions. For example, segments can be extracted at random times, and all frames in the segment can be used as the initial image sequence. Furthermore, this embodiment further filters the initial image frame sequence by discarding one frame from every two adjacent frames. This ensures that the image frame sequence retains the action information of the person opening the access control system while effectively reducing its size, decreasing subsequent data processing volume, and improving processing speed.
[0088] like Figure 4 As shown, based on the aforementioned intelligent access control system, the present invention also provides an electronic device, which can be a mobile terminal, desktop computer, laptop, handheld computer, server, or other computing electronic device. The electronic device includes a processor 410, a memory 420, and a display 430. Figure 4 Only some components of the electronic device are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0089] The memory 420 can be an internal storage unit of the electronic device in some embodiments, such as a hard disk or a memory of the electronic device. The memory 420 can also be an external storage device of the electronic device in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory 420 can include both an internal storage unit and an external storage device of the electronic device. The memory 420 is used to store application software and various data installed on the electronic device, such as program codes installed on the electronic device. The memory 420 can also be used to temporarily store data that has been output or will be output. In an embodiment, the memory 420 stores a smart access control system program 440, which can be executed by the processor 410 to implement the smart access control system of the embodiments of the present application.
[0090] The processor 410 can be a central processing unit (CPU), a microprocessor, or other data processing chip in some embodiments, used to run program codes or process data stored in the memory 420, such as executing the smart access control system, etc.
[0091] The display 430 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, etc. in some embodiments. The display 430 is used to display information of the smart access control system electronic device and to display a visual user interface. The components 410-430 of the electronic device communicate with each other through a system bus.
[0092] Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer readable storage medium. The computer readable storage medium is a disk, an optical disk, a read-only memory, a random access memory, etc.
[0093] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed by the present application can be easily thought of by those skilled in the art, which should be covered within the protection scope of the present application.
Claims
1. A smart access control system, characterized in that, The method comprises the following steps: An image acquisition module is configured to acquire a monitoring video and a voice record recognized by an access control system, and to obtain an image frame sequence according to the monitoring video An object recognition module is configured to recognize an object category in an image according to the image frame sequence based on a trained access control recognition model An action recognition module is configured to recognize an action of the object according to the object category and the image frame sequence based on the trained access control recognition model A voice recognition module is configured to analyze the voice record to obtain an analysis result according to the action of the object based on the trained access control recognition model, and to control the opening and closing of the access control system according to the analysis result Before the object category recognition and the action recognition, the method further comprises the following steps: A plurality of sample sets are acquired; the plurality of sample sets comprise a face sample set, an object sample set and a scene sample set; the face sample set comprises face data and face labels, the object sample set comprises object data and object labels, and the scene sample set comprises scene data and scene labels The initial access control recognition model is trained based on the plurality of sample sets to obtain the trained access control recognition model; the access control recognition model comprises the object recognition module, the action recognition module and the voice recognition module; the object recognition module comprises a first convolutional layer, a first pooling layer, a second convolutional layer, a third convolutional layer, a second pooling layer, a fourth convolutional layer, a fifth convolutional layer and a third pooling layer; the object category in the image is recognized according to the image frame sequence by performing first scale feature extraction on the image frame sequence based on the first convolutional layer and the first pooling layer to obtain a first feature map, performing second scale feature extraction on the first feature map based on the second convolutional layer, the third convolutional layer and the second pooling layer to obtain a second feature map, and performing third scale feature extraction on the second feature map based on the fourth convolutional layer, the fifth convolutional layer and the third pooling layer to obtain the object category The action recognition module comprises a convolutional layer, an adjustment layer and an action connection layer; the action of the object is recognized according to the object category and the image frame sequence by performing sampling feature extraction on the corresponding features of the object category based on the convolutional layer to obtain sampling features, adjusting the sampling features based on the features of the object category by using the adjustment layer to obtain adjusted features, and performing supervised learning on the adjusted features based on the action connection layer to obtain an action recognition result of the object The voice recognition module comprises a first fully connected layer, a second fully connected layer and an activation function layer; the voice record is analyzed to obtain an analysis result according to the action of the object by performing classification processing on the corresponding features of the action of the object based on the first fully connected layer to obtain an action classification result, extracting the content of the voice record based on the classification result of the action of the object by using the second fully connected layer to obtain voice features, and performing activation processing on the voice features based on the activation function layer to obtain a voice recognition result.
2. The intelligent access control system of claim 1, wherein, 3. The intelligent access control system of claim 2, wherein, The activation function layer is formed according to the establishment of an activation function of a first frame image of the image frame sequence.
4. The intelligent access control system of claim 1, wherein, The analysis of the voice record obtains an analysis result, and the opening and closing of the access control system is controlled according to the analysis result, including: Acoustic features are extracted from the voice record; The acoustic features are decoded according to the voice recognition module to obtain a hidden state label sequence; Each hidden state label in the hidden state label sequence is matched with a plurality of preset state labels to obtain access opening information corresponding to the voice data; The opening and closing of the access control system is controlled according to the access opening information. Each preset state label is used to represent one kind of access opening information.
5. The intelligent access control system of claim 1, wherein, The monitoring video recognized by the access control system is obtained, and an image frame sequence is obtained according to the monitoring video, including: The monitoring video is obtained, and a video segment with a preset time length is intercepted from the monitoring video to obtain an initial image frame sequence; Image frames at odd positions in the initial image frame sequence are selected to obtain the image frame sequence.
Citation Information
Patent Citations
Action recognition method based on Tensorflow target detection
CN111860103A
Elevator video call system with voice recognition function
CN117278708A
Image recognition method and device, equipment and storage medium
CN117830699A
Multi-mode biological recognition access control system
CN119251945A