Campus safety intelligent monitoring system based on computer vision
By adopting a multi-angle panoramic stereo optimized face recognition model and a multimodal feature extraction and fusion method in the campus security monitoring system, the problems of low face recognition accuracy and low efficiency in dangerous behavior recognition in the existing system are solved, higher recognition accuracy and efficiency are achieved, and the reliability of the system is enhanced.
Patent Information
- Application Number
- CN202510781572.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-12
AI Technical Summary
In existing campus security monitoring systems, video face recognition accuracy is low, and dangerous behavior recognition relies on single-modality action recognition, which has a high error rate and low efficiency.
A face recognition model based on multi-angle panoramic stereo optimization integrates feature information from different angles and depths to improve face recognition accuracy. Furthermore, a multimodal feature extraction and fusion method based on channel convolutional attention module optimization, combined with a high-resolution network and the Lite-HRNet algorithm, identifies key points on the human body and constructs a skeletal heat map, achieving highly accurate and efficient behavior recognition.
It improves the accuracy of face recognition, reduces false alarms and false positives, achieves high-accuracy and high-efficiency recognition of dangerous behaviors, and enhances the reliability of the campus security monitoring system.
Smart Images

Figure CN120673459A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent monitoring, and in particular relates to a campus safety intelligent monitoring system based on computer vision. Background Art
[0002] The background of campus security monitoring systems stems from the importance and demand for campus safety issues. With the development of society and the advancement of technology, campus safety issues have received increasing attention. School administrators need to ensure the personal and property safety of teachers, students and staff on campus, while also maintaining campus order and managing various activities on campus. Therefore, higher requirements are placed on campus safety management. However, the accuracy of existing surveillance for video face recognition is low, and the identification of dangerous behaviors on campus relies on single-mode motion recognition, which has a high recognition error rate and low efficiency. Summary of the Invention
[0003] In view of the above situation, in order to overcome the defects of the existing technology and the problem of low accuracy of video face recognition in existing monitoring, the present invention creatively adopts a face recognition model based on multi-angle panoramic stereo optimization. By integrating feature information of different angles and depths, the model obtains a more comprehensive and accurate face representation, reducing the degree of influence by factors such as lighting and posture. At the same time, through feature simplification, scaling and selection, redundant information is removed, which can be more accurately recognized, reduce false alarms and false alarms, and improve the reliability of the system. For existing monitoring, the recognition of dangerous behaviors on campus relies on single-mode action recognition, and the recognition error rate is very low. To solve the problem of high accuracy and low efficiency, the present invention adopts a multimodal feature extraction fusion method based on channel convolution attention module optimization, adopts a high-resolution network to extract spatial position information of images, uses the top-down posture estimation algorithm Lite-HRNet to identify key points of the human body, introduces Gaussian distribution to construct bone heat map, and adopts an action recognition model based on channel attention optimization to extract bone modal features from bone heat map, and adopts an RGB feature extraction model based on channel attention optimization to extract RGB features from pedestrian images, and finally recognizes actions with fused features to achieve highly accurate and efficient behavior recognition.
[0004] The technical solution provided by the present invention is as follows: a campus security intelligent monitoring system based on computer vision, comprising a video monitoring module, an image enhancement module, an intelligent face recognition module, a skeleton-based behavior recognition module, an alarm module and a central control module;
[0005] The video monitoring module collects pedestrian images on campus through a high-definition camera and transmits the pedestrian images to the image enhancement module;
[0006] The image enhancement module performs spatial domain-based enhancement and frequency domain-based enhancement processing on the pedestrian image to obtain an enhanced image, and sends the enhanced image to the intelligent face recognition module and the skeleton-based behavior recognition module;
[0007] The spatial domain-based enhancement includes gamma correction and histogram equalization;
[0008] The frequency domain-based enhancement includes Gaussian low-pass filtering;
[0009] The intelligent face recognition module intelligently identifies the identity of pedestrians in the enhanced image through a face recognition model based on multi-angle panoramic stereo optimization;
[0010] The skeleton-based behavior recognition module performs behavior recognition on pedestrians in the enhanced image, and includes a human target detection unit, a key point recognition unit, and an action recognition unit;
[0011] The human target detection unit detects human targets in the enhanced image using the YOLOX target detection algorithm;
[0012] The key point recognition unit performs human key point recognition on a human target using a human key point recognition method of a lightweight high-resolution network Lite-HRNet;
[0013] The action recognition unit recognizes pedestrian behavior and includes a skeleton modal feature extraction unit, an RGB modal feature extraction unit and a feature fusion unit;
[0014] The skeleton modal feature extraction unit extracts the skeleton modal features by using a channel attention optimization feature extraction method based on Gaussian distribution;
[0015] The RGB modal feature extraction unit extracts RGB modal features by using a channel attention optimization feature extraction method;
[0016] The feature fusion unit combines the skeleton modal features with the RGB modal features and activates and outputs them to obtain pedestrian behavior recognition results;
[0017] The alarm module classifies the pedestrian behavior recognition results and issues an alarm when there is a dangerous behavior classification;
[0018] The central control module controls all modules and stores information.
[0019] Furthermore, in the intelligent face recognition module, the data processing of the face recognition model based on multi-angle panoramic stereo optimization specifically includes the following steps:
[0020] Step S1: Multi-angle convolution: collect the face images in the enhanced image from a angles, and use a completely identical graph convolutional network to convolve a face images at different angles to obtain a initial feature maps of dimensions H×W×C, where the face images at a angles fully cover the face;
[0021] Step S2: Feature splicing: splice the a initial feature maps, keep the height and width unchanged, and increase the number of channels to a times the original number, i.e., aC, to obtain the full feature map;
[0022] Step S3: Redundant information is removed, and 1×1 point convolution is performed on the full feature map to obtain a simplified feature map: ; ;
[0023] Where, is a simplified feature map, For the picture The coordinates in the nth channel are Elements, is the weight of the convolution kernel of the nth simplified feature map channel of the kth channel of the full feature map, The coordinates representing the kth channel of the full feature map are Elements of
[0024] Step S4: Feature downsampling, use global depth convolution to downsample the simplified feature map to obtain a simplified feature vector: ; ;
[0025] Where, It is the value of the b-th channel after the global depth convolution of the simplified feature map, representing the information of the b-th channel. Represents the depth convolution kernel parameter matrix of the b-th channel, Represents the features of the b-th channel of the simplified feature map, Represents the simplified feature vector obtained by global depth convolution;
[0026] Step S5: Feature vector compression, compressing the simplified feature vector to obtain a compressed feature vector: ;
[0027] Where, represents the activation function, represents the parameter matrix of the fully connected layer, represents the compressed feature vector;
[0028] Step S6: Weight vector calculation: ;
[0029] Where, represents the parameter matrix of the fully connected layer, is the weight vector with dimension C of b, where m ∈ [1,a] , is the activation function;
[0030] Step S7: Feature selection and fusion to obtain facial features: ; ;
[0031] Where, Representative An initial feature map, is the b-th weight vector with dimension C, Represents facial features that filter out redundant information;
[0032] Step S8: Face recognition, by matching facial features with pedestrian identities to obtain face recognition results.
[0033] Furthermore, in the skeletal modal feature extraction unit, the channel attention optimization feature extraction method based on Gaussian distribution specifically includes the following steps:
[0034] Step Q1: Construction of key point heatmap based on Gaussian distribution: heatmap (x,y) = e - D((x,y),seg[ a k , b k ] ) 2 2* s 2 *min( ca k ,c b k ) ;
[0035] Where, is the standard deviation of the Gaussian distribution, are the key point coordinates, From the key point To line segment seg[ a k , b k ] distance, For key points and The confidence score of represents the key point heat map;
[0036] Step Q2: 3D heatmap construction, by stacking key point heatmaps along the time dimension to obtain a 3D heatmap;
[0037] Step Q3: Model building, including The model includes multiple ResNet layers, and each ResNet layer includes multiple bottleneck blocks;
[0038] Step Q4: Model optimization, using channel attention mechanism The bottleneck block of the ResNet layer in the model is replaced, and the channel attention mechanism processing steps are as follows: ; M S (F)= M C (F)⊗ s ( f 7 × 7 ([MaxPool( M C (F));AvgPool( M C (F))])) ;
[0039] Where, Represents the input of the ResNet layer, stands for multi-layer vector machine processing, represents maximum pooling, represents average pooling, represents element-wise multiplication, represents the Sigmoid activation function, is an intermediate variable, represents 7×7 convolution, Represents the calculation result of the channel attention mechanism;
[0040] Step Q5: Feature extraction, through optimized Model,feature extraction and activation are performed on the 3D heat map to obtain the bone modal features;
[0041] Furthermore, in the RGB modality feature extraction unit, the channel attention optimization feature extraction method specifically includes the following steps:
[0042] Step P1: Slowonly model construction, where the Slowonly model includes multiple ResNet layers, and each ResNet layer includes multiple bottleneck blocks;
[0043] Step P2: Slowonly model optimization, using the channel attention mechanism to replace the bottleneck block of the ResNet layer in the Slowonly model;
[0044] Step P3: Feature extraction: Through the optimized Slowonly model, feature extraction and activation are performed on the enhanced image to obtain RGB modal features.
[0045] The beneficial effects achieved by the present invention using the above scheme are as follows:
[0046] (1) In response to the low accuracy of video face recognition in existing surveillance, the present invention creatively adopts a face recognition model based on multi-angle panoramic stereo optimization. By integrating feature information from different angles and depths, the model obtains a more comprehensive and accurate face representation, reducing the degree of influence of factors such as lighting and posture. At the same time, through feature simplification, scaling and selection, redundant information is removed, which enables more accurate recognition, reduces false alarms and false alarms, and improves the reliability of the system.
[0047] (2) In response to the problem that the existing monitoring system relies on single-mode action recognition to identify dangerous behaviors on campus, with high recognition error rate and low efficiency, the present invention adopts a multimodal feature extraction and fusion method based on channel convolution attention module optimization, uses a high-resolution network to extract spatial position information from images, uses the top-down posture estimation algorithm Lite-HRNet to identify key points of the human body, introduces Gaussian distribution to construct bone heat map, and uses an action recognition model based on channel attention optimization to extract bone modal features from bone heat map, and uses an RGB feature extraction model based on channel attention optimization to extract RGB features from pedestrian images. Finally, the fused features are used to identify actions, achieving highly accurate and efficient behavior recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A module diagram of the computer vision-based campus security intelligent monitoring system provided by the present invention;
[0049] Figure 2 The figure is a flowchart of data processing in a face recognition model based on multi-angle panoramic stereo optimization.
[0050] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0052] Example 1, see Figure 1 The present invention provides a campus security intelligent monitoring system based on computer vision, which is characterized by:
[0053] It includes video surveillance module, image enhancement module, intelligent face recognition module, skeleton-based behavior recognition module, alarm module and central control module;
[0054] The video monitoring module collects pedestrian images on campus through a high-definition camera and transmits the pedestrian images to the image enhancement module;
[0055] The image enhancement module performs spatial domain-based enhancement and frequency domain-based enhancement processing on the pedestrian image to obtain an enhanced image, and sends the enhanced image to the intelligent face recognition module and the skeleton-based behavior recognition module;
[0056] The spatial domain-based enhancement includes gamma correction and histogram equalization;
[0057] The frequency domain-based enhancement includes Gaussian low-pass filtering;
[0058] The intelligent face recognition module intelligently identifies the identity of pedestrians in the enhanced image through a face recognition model based on multi-angle panoramic stereo optimization to obtain a pedestrian identification result;
[0059] The skeleton-based behavior recognition module performs behavior recognition on pedestrians in the enhanced image, and includes a human target detection unit, a key point recognition unit, and an action recognition unit;
[0060] The alarm module classifies the pedestrian behavior recognition results and issues an alarm when there is a dangerous behavior classification;
[0061] The central control module controls all modules and stores information.
[0062] Example 2: This example is based on the above example. In the skeleton-based behavior recognition module,
[0063] The human target detection unit detects human targets in the enhanced image using the YOLOX target detection algorithm;
[0064] The key point recognition unit performs human key point recognition on a human target using a human key point recognition method of a lightweight high-resolution network Lite-HRNet;
[0065] The action recognition unit recognizes pedestrian behavior and includes a skeleton modal feature extraction unit, an RGB modal feature extraction unit and a feature fusion unit;
[0066] The skeleton modal feature extraction unit extracts the skeleton modal features by using a channel attention optimization feature extraction method based on Gaussian distribution;
[0067] The RGB modal feature extraction unit extracts RGB modal features by using a channel attention optimization feature extraction method;
[0068] The feature fusion unit combines the skeleton modal features with the RGB modal features and activates and outputs them to obtain pedestrian behavior recognition results.
[0069] Example 3, see Figure 2 This embodiment is based on the above embodiment. In the intelligent face recognition module, the data processing of the face recognition model based on multi-angle panoramic stereo optimization specifically includes the following steps:
[0070] Step S1: Multi-angle convolution: collect the face images in the enhanced image from a angles, and use a completely identical graph convolutional network to convolve a face images at different angles to obtain a initial feature maps of dimensions H×W×C, where the face images at a angles fully cover the face;
[0071] Step S2: Feature splicing: splice the a initial feature maps, keep the height and width unchanged, and increase the number of channels to a times the original number, i.e., aC, to obtain the full feature map;
[0072] Step S3: Redundant information is removed, and 1×1 point convolution is performed on the full feature map to obtain a simplified feature map: ; ;
[0073] Where, is a simplified feature map, For the picture The coordinates in the nth channel are Elements, is the weight of the convolution kernel of the nth simplified feature map channel of the kth channel of the full feature map, The coordinates representing the kth channel of the full feature map are Elements of
[0074] Step S4: Feature downsampling, use global depth convolution to downsample the simplified feature map to obtain a simplified feature vector: ; ;
[0075] Where, It is the value of the b-th channel after the global depth convolution of the simplified feature map, representing the information of the b-th channel. Represents the depth convolution kernel parameter matrix of the b-th channel, Represents the features of the b-th channel of the simplified feature map, Represents the simplified feature vector obtained by global depth convolution;
[0076] Step S5: Feature vector compression, compressing the simplified feature vector to obtain a compressed feature vector: ;
[0077] Where, represents the activation function, represents the parameter matrix of the fully connected layer, represents the compressed feature vector;
[0078] Step S6: Weight vector calculation: ;
[0079] Where, represents the parameter matrix of the fully connected layer, is the weight vector with dimension C of b, where m ∈ [1,a] , is the activation function;
[0080] Step S7: Feature selection and fusion to obtain facial features: ; ;
[0081] Where, Representative An initial feature map, is the b-th weight vector with dimension C, Represents facial features that filter out redundant information;
[0082] Step S8: Face recognition, by matching facial features with pedestrian identities, obtaining pedestrian identity recognition results.
[0083] By performing the above operations, in order to address the problem of low accuracy of video face recognition in existing monitoring, the present invention creatively adopts a face recognition model based on multi-angle panoramic stereo optimization. By integrating feature information from different angles and depths, the model obtains a more comprehensive and accurate face representation, reducing the degree of influence by factors such as lighting and posture. At the same time, through feature simplification, scaling and selection, redundant information is removed, which enables more accurate recognition, reduces false alarms and false alarms, and improves the reliability of the system.
[0084] Embodiment 4: This embodiment is based on the above embodiment.
[0085] In the skeletal modality feature extraction unit, the channel attention optimization feature extraction method based on Gaussian distribution specifically includes the following steps:
[0086] Step Q1: Construction of key point heatmap based on Gaussian distribution: heatmap (x,y) = e - D((x,y),seg[ a k , b k ] ) 2 2* s 2 *min( ca k ,c b k ) ;
[0087] Where, is the standard deviation of the Gaussian distribution, are the key point coordinates, From the key point To line segment seg[ a k , b k ] distance, For key points and The confidence score of represents the key point heat map;
[0088] Step Q2: 3D heatmap construction, by stacking key point heatmaps along the time dimension to obtain a 3D heatmap;
[0089] Step Q3: Model building, including The model includes multiple ResNet layers, and each ResNet layer includes multiple bottleneck blocks;
[0090] Step Q4: Model optimization, using channel attention mechanism The bottleneck block of the ResNet layer in the model is replaced, and the channel attention mechanism processing steps are as follows: ; M S (F)= M C (F)⊗ s ( f 7 × 7 ([MaxPool( M C (F));AvgPool( M C (F))])) ;
[0091] Where, Represents the input of the ResNet layer, stands for multi-layer vector machine processing, represents maximum pooling, represents average pooling, represents element-wise multiplication, represents the Sigmoid activation function, is an intermediate variable, represents 7×7 convolution, Represents the calculation result of the channel attention mechanism;
[0092] Step Q5: Feature extraction, through optimized Model,feature extraction and activation are performed on the 3D heat map to obtain the bone modal features.
[0093] Embodiment 5: Based on the above embodiment, in the RGB modality feature extraction unit, the channel attention optimization feature extraction method specifically includes the following steps:
[0094] Step P1: Slowonly model construction, where the Slowonly model includes multiple ResNet layers, and each ResNet layer includes multiple bottleneck blocks;
[0095] Step P2: Slowonly model optimization, using the channel attention mechanism to replace the bottleneck block of the ResNet layer in the Slowonly model;
[0096] Step P3: Feature extraction: Through the optimized Slowonly model, feature extraction and activation are performed on the enhanced image to obtain RGB modal features.
[0097] By performing the above operations, in order to address the problem that the existing monitoring relies on single-modal action recognition to identify dangerous behaviors on campus, with high recognition error rate and low efficiency, the present invention adopts a multimodal feature extraction and fusion method based on channel convolutional attention module optimization, adopts a high-resolution network to extract spatial position information from images, uses the top-down posture estimation algorithm Lite-HRNet to identify key points of the human body, introduces Gaussian distribution to construct bone heat map, and adopts an action recognition model based on channel attention optimization to extract bone modal features from the bone heat map, and adopts an RGB feature extraction model based on channel attention optimization to extract RGB features from pedestrian images. Finally, the fused features are used to identify actions, achieving highly accurate and efficient behavior recognition.
[0098] Example 6: This example is based on the above example. The pedestrian behavior recognition results and classification details are as follows:
[0099] Risky behavior categories:
[0100] Violent behavior: such as fighting, pushing, and hurting others;
[0101] Illegal climbing: climbing over campus walls and entering restricted areas;
[0102] Abnormal gathering: students gather in a certain area for a long time, causing congestion;
[0103] Carrying dangerous items: knives, sticks, etc.;
[0104] Fall / Unconscious: If the student falls or does not move for a long time.
[0105] Example 7: This example is based on the above example and is an actual application scenario of the present invention:
[0106] Entrance and exit monitoring:
[0107] Facial recognition cameras are deployed at the campus gate and dormitory entrances to record all people entering and leaving;
[0108] When unauthorized persons attempt to enter, an audible and visual alarm is triggered and the door is prohibited from opening;
[0109] Teaching area monitoring:
[0110] Deploy dangerous behavior detection cameras in crowded areas such as corridors and stairways of teaching buildings;
[0111] When the system detects abnormal gathering, pushing, and other behaviors, it will issue early warnings to prevent stampedes;
[0112] Playground and public area surveillance:
[0113] The playground deploys 360° panoramic cameras to monitor large activity areas;
[0114] When a violent act or fall is detected, the incident location is automatically located and security personnel are notified;
[0115] Restricted area monitoring:
[0116] In restricted areas such as laboratories and computer rooms, the system combines access control and facial recognition to prevent unauthorized entry;
[0117] In case of unauthorized access, the system will automatically record and alarm.
[0118] Example 8: This example is based on the above example and is optimized to further expand the functions of face recognition:
[0119] Identity authentication process:
[0120] Students, faculty, and staff must enter their facial information into the database in advance, and visitors must register their identities and have their faces photographed at the entrance;
[0121] The system conducts real-time comparison of people entering and leaving the campus to determine whether they are authorized personnel;
[0122] Abnormal identity detection:
[0123] Unauthorized person intrusion: The system identifies unregistered persons and triggers an alarm;
[0124] Blacklist comparison: matching with the blacklist database (including long-term residents, potential threat personnel, etc.);
[0125] Attendance function:
[0126] Automatically record students’ attendance data through facial recognition to avoid omissions in manual registration.
[0127] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0128] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
[0129] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. Campus security intelligent monitoring system based on computer vision, characterized by: It includes video surveillance module, image enhancement module, intelligent face recognition module, skeleton-based behavior recognition module, alarm module and central control module; The video monitoring module collects pedestrian images on campus through a high-definition camera and transmits the pedestrian images to the image enhancement module; The image enhancement module performs spatial domain-based enhancement and frequency domain-based enhancement processing on the pedestrian image to obtain an enhanced image, and sends the enhanced image to the intelligent face recognition module and the skeleton-based behavior recognition module; The spatial domain-based enhancement includes gamma correction and histogram equalization; The frequency domain-based enhancement includes Gaussian low-pass filtering; The intelligent face recognition module intelligently identifies the identity of pedestrians in the enhanced image through a face recognition model based on multi-angle panoramic stereo optimization to obtain a pedestrian identification result; The skeleton-based behavior recognition module performs behavior recognition on pedestrians in the enhanced image, and includes a human target detection unit, a key point recognition unit, and an action recognition unit; The alarm module classifies the pedestrian behavior recognition results and issues an alarm when there is a dangerous behavior classification; The central control module controls all modules and stores information.
2. The computer vision-based campus security intelligent monitoring system according to claim 1 is characterized in that: The human target detection unit detects human targets in the enhanced image using the YOLOX target detection algorithm; The key point recognition unit recognizes the key points of a human target using a human key point recognition method based on a lightweight high-resolution network Lite-HRNet; The action recognition unit recognizes pedestrian behavior and includes a skeleton modal feature extraction unit, an RGB modal feature extraction unit and a feature fusion unit; The skeleton modal feature extraction unit extracts the skeleton modal features by using a channel attention optimization feature extraction method based on Gaussian distribution; The RGB modal feature extraction unit extracts RGB modal features by using a channel attention optimization feature extraction method; The feature fusion unit combines the skeleton modal features with the RGB modal features and activates and outputs them to obtain pedestrian behavior recognition results.
3. The computer vision-based campus security intelligent monitoring system according to claim 1 is characterized in that: In the intelligent face recognition module, the data processing of the face recognition model based on multi-angle panoramic stereo optimization specifically includes the following steps: Step S1: Multi-angle convolution: collect the face images in the enhanced image from a angles, and use a completely identical graph convolutional network to convolve a face images at different angles to obtain a initial feature maps of dimensions H×W×C, where the face images at a angles fully cover the face; Step S2: Feature splicing, splicing a initial feature maps, keeping the height and width unchanged, and changing the number of channels to a times the original a, that is, aC, to obtain the full feature map; Step S3: Redundant information is removed, and 1×1 point convolution is performed on the full feature map to obtain a simplified feature map; Step S4: feature downsampling, using global depth convolution to downsample the simplified feature map to obtain a simplified feature vector; Step S5: feature vector compression, compressing the simplified feature vector to obtain a compressed feature vector; Step S6: weight vector calculation; Step S7: Feature selection and fusion to obtain facial features: ; ; Where, Representative An initial feature map, is the b-th weight vector with dimension C, Represents facial features that filter out redundant information; Step S8: Face recognition, by matching facial features with pedestrian identities, obtaining pedestrian identity recognition results.
4. The computer vision-based campus security intelligent monitoring system according to claim 2, characterized in that: In the skeletal modality feature extraction unit, the channel attention optimization feature extraction method based on Gaussian distribution specifically includes the following steps: Step Q1: Construction of key point heatmap based on Gaussian distribution: ; Where, is the standard deviation of the Gaussian distribution, are the key point coordinates, From the key point To line segment distance, For key points and The confidence score of represents the key point heat map; Step Q2: 3D heatmap construction, by stacking key point heatmaps along the time dimension to obtain a 3D heatmap; Step Q3: Model building, including The model includes multiple ResNet layers, and each ResNet layer includes multiple bottleneck blocks; Step Q4: Model optimization, using channel attention mechanism Replace the bottleneck block of the ResNet layer in the model; Step Q5: Feature extraction, through optimized Model,feature extraction and activation are performed on the 3D heat map to obtain the bone modal features.
Citation Information
Patent Citations
Multi-angle human face detecting method based on weighting of deformable components
CN102622604A
Personnel identification method and device, terminal equipment and storage medium
CN115311706A
Student expression recognition method, teaching state evaluation method and related equipment
CN115457619A
Coal mine belt full life cycle monitoring method based on visual large model
CN118396971A
Face recognition method based on artificial intelligence
CN118711228A