A security monitoring camera detection method and system based on face recognition
Patent Information
- Application Number
- CN202511268153.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-09-05
AI Technical Summary
[0005]本申请提供一种基于人脸识别的安防监控摄像头检测方法及系统,用以解决现有技术中低光照环境下人脸识别的准确率低和鲁棒性差的问题
[0051]本申请实现全天候人脸数据采集,确保在低光照环境下仍能获取有效的人脸信息。建立温度与纹理的精确对应关系,为后续融合处理提供数据基础。综合红外和可见光数据优势,提升图像信息的完整性和可用性。构建具有空间结构信息的人脸模型,克服二维图像的角度和遮挡限制。利用三维结构特征提高识别准确性,增强对复杂情况的适应能力。
Smart Images

Figure CN121121822B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of facial recognition technology, and in particular to a method and system for detecting security monitoring cameras based on facial recognition. Background Technology
[0002] In low-light security monitoring scenarios such as bank night shifts and underground parking garages, traditional visible light cameras are limited by lighting conditions and struggle to clearly capture facial details, leading to decreased recognition accuracy. Therefore, there is an urgent need for a facial recognition technology that can operate stably in low-light environments, preserving facial texture features while overcoming problems such as insufficient lighting, angular deviations, and partial occlusion, ensuring the reliability and real-time performance of security monitoring.
[0003] Currently, some security systems employ near-infrared imaging-based liveness detection combined with 2D face recognition technology. This method acquires facial images using a near-infrared camera, extracts features using a deep learning model, and then performs matching. This approach can operate in low-light environments and possesses a certain degree of anti-spoofing capability, making it suitable for nighttime surveillance scenarios.
[0004] This scheme relies on two-dimensional image features, and key facial texture information may still be lost under extremely low light or strong backlight conditions, affecting recognition accuracy. Furthermore, because it only uses planar image data, it is poorly adaptable to complex situations such as side profiles and occlusions, easily leading to false positives or false negatives, making it difficult to meet the needs of high-security scenarios. Summary of the Invention
[0005] This application provides a method and system for detecting security surveillance cameras based on face recognition, in order to solve the problems of low accuracy and poor robustness of face recognition in low light environments in the prior art.
[0006] In a first aspect, this application provides a method for detecting security surveillance cameras based on face recognition, including:
[0007] Receive infrared temperature data and visible light texture data of the target face collected by security monitoring cameras;
[0008] Spatiotemporal registration is performed on the infrared temperature data and the visible light texture data to generate temperature texture association information;
[0009] Based on the temperature texture association information, an enhanced multi-channel image is generated using multispectral fusion technology;
[0010] Three-dimensional face reconstruction is performed on the enhanced multi-channel image to generate a three-dimensional face mesh model corresponding to the target face;
[0011] A convolutional graph neural network is used to extract three-dimensional topological features from the three-dimensional mesh model of the face. The three-dimensional topological features are then matched and compared with three-dimensional feature templates in a pre-stored registered face database. Based on the matching and comparison results, the person identifier corresponding to the target face is determined.
[0012] Optionally, the step of performing three-dimensional face reconstruction on the enhanced multi-channel image to generate a three-dimensional face mesh model corresponding to the target face includes:
[0013] Facial landmark localization is performed on the enhanced multi-channel image;
[0014] Generate a spatial point cloud of the face based on the located facial key points;
[0015] Connect adjacent points in the face spatial point cloud to form a set of triangular patches;
[0016] The set of triangular facets is optimized, and the optimized set of triangular facets is reconstructed into a three-dimensional mesh model of a human face.
[0017] Optionally, generating a facial spatial point cloud based on the located facial key points includes:
[0018] Based on the two-dimensional image coordinates of the facial key points, and combined with the parameters of the security monitoring camera, the initial three-dimensional coordinates of each facial key point in the spatial coordinate system are calculated.
[0019] Using the initial three-dimensional coordinates as a reference point, dense sampling points are generated by expanding along the normal direction of the face surface;
[0020] Arrange the dense sampling points corresponding to all facial key points into a three-dimensional coordinate set according to a preset topological order;
[0021] A normalization operation is performed on the three-dimensional coordinate set to generate a normalized three-dimensional coordinate set, which is then used as the face spatial point cloud.
[0022] Optionally, optimizing the set of triangular facets and reconstructing the optimized set of triangular facets into a 3D face mesh model includes:
[0023] Traverse each vertex of the triangle facet set, calculate the geometric position deviation between the vertex of the triangle facet and the vertex of the adjacent triangle facet, and when the geometric position deviation is greater than a preset deviation threshold, adjust the vertex spatial coordinates to obtain an optimized triangle facet set.
[0024] Detect abnormal faces in the optimized triangular face set that meet preset conditions, remove all abnormal faces, and perform patching operation in the area where abnormal faces were removed to obtain the optimized triangular face set. The preset conditions include the maximum value of the face's interior angle being greater than a preset angle threshold and the minimum value of the face's side length being less than a preset length threshold.
[0025] The triangular faces in the optimized triangular facet set are reconnected to generate a 3D mesh model of the face.
[0026] Optionally, the step of extracting three-dimensional topological features from the three-dimensional face mesh model using a convolutional graph neural network, matching and comparing the three-dimensional topological features with three-dimensional feature templates in a pre-stored registered face database, and determining the person identifier corresponding to the target face based on the matching and comparison results includes:
[0027] The vertex connection relationships of the three-dimensional face mesh model are input into a convolutional graph neural network;
[0028] The convolutional graph neural network aggregates the geometric features of each vertex and its neighborhood.
[0029] Based on the geometric features of all aggregated vertices, generate 3D topological features;
[0030] Calculate the similarity score between the three-dimensional topological feature and the three-dimensional feature template, where the similarity score is the matching comparison result;
[0031] When the matching result is greater than the preset similarity threshold, the personnel identifier associated with the three-dimensional feature template is output.
[0032] Optionally, the step of performing spatiotemporal registration of the infrared temperature data and the visible light texture data to generate temperature texture association information includes:
[0033] A positional mapping relationship is established based on the pixel coordinates of the infrared temperature data and the pixel coordinates of the visible light texture data;
[0034] Based on the location mapping relationship, infrared temperature data and visible light texture data at the same time stamp are superimposed on the same spatial coordinate system, and a corresponding associated information unit is generated for each spatial coordinate point.
[0035] The associated information units corresponding to each spatial coordinate point are integrated into temperature texture associated information.
[0036] Optionally, generating an enhanced multi-channel image based on the temperature texture association information using multispectral fusion technology includes:
[0037] The infrared radiation intensity value in the temperature texture association information is converted into first channel data;
[0038] Convert the visible light texture detail values in the temperature texture association information into second channel data;
[0039] The first channel data and the second channel data are fused into an initial multi-channel image using multispectral fusion technology.
[0040] By adjusting the radiation intensity and texture detail ratio of the initial multi-channel image through the inter-channel weight allocation relationship, an adjusted multi-channel image is generated, and the adjusted multi-channel image is used as an enhanced multi-channel image.
[0041] Secondly, this application provides a security monitoring camera detection system based on face recognition, comprising:
[0042] The receiving module is used to receive infrared temperature data and visible light texture data of the target face captured by the security monitoring camera;
[0043] The first generation module is used to perform spatiotemporal registration of the infrared temperature data and the visible light texture data to generate temperature texture association information.
[0044] The second generation module is used to generate an enhanced multi-channel image based on the temperature texture association information using multispectral fusion technology;
[0045] The reconstruction module is used to perform three-dimensional face reconstruction on the enhanced multi-channel image and generate a three-dimensional face mesh model corresponding to the target face.
[0046] The extraction module is used to extract three-dimensional topological features from the three-dimensional mesh model of the face using a convolutional graph neural network, match and compare the three-dimensional topological features with three-dimensional feature templates in a pre-stored registered face database, and determine the person identifier corresponding to the target face based on the matching and comparison results.
[0047] Thirdly, this application provides a computing device, including a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute a security monitoring camera detection method based on face recognition as described in any of the first aspects.
[0048] Fourthly, this application provides a computer storage medium storing computer program instructions thereon, which, when executed by a processor, implement the face recognition-based security monitoring camera detection method described in any one of the first aspects.
[0049] This application provides a method for detecting a security surveillance camera based on face recognition. The method includes: receiving infrared temperature data and visible light texture data of a target face collected by the security surveillance camera; performing spatiotemporal registration on the infrared temperature data and the visible light texture data to generate temperature-texture association information; generating an enhanced multi-channel image based on the temperature-texture association information using multispectral fusion technology; performing three-dimensional face reconstruction on the enhanced multi-channel image to generate a three-dimensional mesh model of the target face; extracting three-dimensional topological features from the three-dimensional mesh model of the face using a convolutional graph neural network; matching and comparing the three-dimensional topological features with three-dimensional feature templates in a pre-stored registered face database; and determining the personnel identifier corresponding to the target face based on the matching and comparison results.
[0050] The technical solution provided in this application has the following beneficial effects:
[0051] This application enables all-weather facial data acquisition, ensuring effective facial information can be obtained even in low-light environments. It establishes a precise correspondence between temperature and texture, providing a data foundation for subsequent fusion processing. By combining the advantages of infrared and visible light data, the integrity and usability of image information are improved. A facial model with spatial structural information is constructed to overcome the limitations of angle and occlusion in two-dimensional images. Three-dimensional structural features are utilized to improve recognition accuracy and enhance adaptability to complex situations.
[0052] Furthermore, this application also performs facial key point localization on enhanced multi-channel images, generates spatial point clouds, constructs a set of triangular facets, and then optimizes and reconstructs them to form an accurate three-dimensional facial mesh model.
[0053] Furthermore, the 3D face model constructed in this step fully preserves the spatial structural features of the face, effectively solving the recognition problem of 2D images under conditions of angle change and partial occlusion, and significantly improving the reliability of face recognition in complex scenes.
[0054] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 A flowchart illustrating a security surveillance camera detection method based on face recognition, provided for an embodiment of this application;
[0057] Figure 2 A schematic diagram of a security monitoring camera detection system based on face recognition provided in an embodiment of this application;
[0058] Figure 3 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0059] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0060] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0061] In low-light security monitoring scenarios, existing near-infrared imaging-based facial recognition technology has significant shortcomings: although it can operate in low-light conditions, important facial details are still lost when the light is too dim or there is strong backlight, affecting recognition accuracy. More importantly, this method relies solely on planar image information; when the face is turned to the side or partially obscured, the recognition effect drops drastically, making it difficult to meet the actual needs of high-security locations such as banks and underground parking garages.
[0062] To address these issues, this application proposes a face recognition-based detection method for security surveillance cameras. This method utilizes both infrared temperature data and visible light image data to construct more complete facial information. Specifically, the method first precisely aligns and fuses the two types of data to generate an enhanced image containing temperature distribution and texture details. Then, based on this information, a precise 3D face model is reconstructed. Finally, recognition is performed by analyzing the 3D structural features of the face. This method not only functions normally in completely dark environments, but more importantly, it effectively solves problems that traditional techniques struggle with, such as side profiles and occlusions, through 3D modeling. This improves the reliability of recognition in complex scenes and provides more reliable security for high-security locations.
[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0064] Figure 1 A flowchart illustrating a security surveillance camera detection method based on face recognition, as provided in this application embodiment, is shown below. Figure 1 As shown, the method includes:
[0065] Step 101: Receive infrared temperature data and visible light texture data of the target face collected by the security monitoring camera.
[0066] In step 101, infrared temperature data refers to the surface temperature distribution information of the face obtained through a thermal imaging sensor, with each data point corresponding to the temperature value of a specific location on the face. Visible light texture data refers to detailed images of the face surface obtained through a regular camera, including visual features such as skin tone and facial contours.
[0067] In this embodiment, the security monitoring camera simultaneously activates both an infrared thermal imaging module and a visible light imaging module. The infrared module collects the thermal radiation signal emitted by the face and converts it into a temperature value matrix, while the visible light module collects the optical image reflected from the face. The two modules are synchronized via hardware to ensure consistent acquisition time. The temperature data matrix and the visible light image are transmitted to the processing system as raw input data for subsequent processing.
[0068] For example, in a nighttime surveillance scenario at Bank A, when someone enters the monitored area, the camera simultaneously activates two modules: the infrared module outputs temperature data at a resolution of 256×192 pixels at 30 frames per second, with each data point representing the temperature value of an area of approximately 1 square centimeter; the visible light module outputs a grayscale image at a resolution of 1280×720 pixels. The system automatically records a timestamp to ensure synchronization between the two sets of data, with the temperature data ranging from 32 to 42 degrees Celsius, corresponding to the normal human body temperature range.
[0069] Step 102: Perform spatiotemporal registration on the infrared temperature data and the visible light texture data to generate temperature texture association information.
[0070] In step 102, spatiotemporal registration refers to the process of aligning data collected by different sensors in terms of spatial location and time series. Temperature texture association information is a data structure that contains the correspondence between registered temperature values and texture values.
[0071] In this embodiment, the coordinate transformation relationship between infrared and visible light images is first established based on the camera calibration parameters, and the corresponding visible light image region is found for each temperature data point. Then, the position mapping is optimized through a feature point matching algorithm to eliminate errors caused by lens distortion. Finally, the registered temperature values and texture values are packaged and stored to form associated data units with spatial correspondence.
[0072] For example, the system reads pre-stored camera calibration parameters, where each pixel in the infrared image corresponds to a 2×2 pixel area in the visible light image. Feature point detection identifies 50 matching points in each of the two images, and a precise coordinate transformation matrix is calculated. After registering the 256×192 temperature data with the 1280×720 visible light image, 256×192 associated units are generated, each containing one temperature value and four visible light pixel values.
[0073] Step 103: Based on the temperature texture association information, generate an enhanced multi-channel image using multispectral fusion technology.
[0074] In step 103, multispectral fusion technology refers to a method of integrating optical information from different bands. Enhanced multichannel images are image data containing multiple feature channels after fusion.
[0075] In this embodiment, temperature data is normalized and used as the first channel, while the grayscale value of the visible light image is used as the second channel. A weighted fusion algorithm is used to calculate the fusion value for each pixel, where the weight of the temperature channel is dynamically adjusted according to ambient lighting conditions. During the fusion process, a pyramid decomposition method is employed to preserve the features of each frequency band, ultimately generating an enhanced image containing rich information.
[0076] For example, in the low-light environment of underground parking garage B, the system linearly transforms the temperature data to the range of 0-255 as the first channel, and uses the grayscale values of the visible light image as the second channel. Based on the ambient light sensor readings, the fusion weights are set to 0.6 for the temperature channel and 0.4 for the visible light channel. A three-layer pyramid fusion algorithm is used to generate an enhanced image with a resolution of 640×480, which preserves temperature characteristics while enhancing texture details.
[0077] Step 104: Perform 3D face reconstruction on the enhanced multi-channel image to generate a 3D face mesh model corresponding to the target face.
[0078] In step 104, 3D face reconstruction refers to the process of recovering the 3D geometric structure from a 2D image. The 3D face mesh model is a 3D face surface model represented by triangular facets.
[0079] In this embodiment, key facial feature points are first located on the enhanced image, and the three-dimensional coordinates of each point are calculated based on the principle of perspective. Then, dense sampling points are supplemented between the feature points using a point cloud generation algorithm. Next, the sampling points are connected using the Delaunay triangulation method to form an initial mesh. Finally, the surface is smoothed and topological errors are corrected using a mesh optimization algorithm.
[0080] For example, in nighttime surveillance of Community C, the system detected 68 facial key points from the enhanced image, calculated their 3D coordinates, and generated a point cloud containing approximately 10,000 points. After forming an initial mesh through triangulation, 15 mesh anomalies were detected and repaired, and finally, a smooth 3D face model composed of approximately 20,000 triangular faces was output, accurately restoring the three-dimensional shape of the face.
[0081] Step 105: Extract three-dimensional topological features from the face three-dimensional mesh model using a convolutional graph neural network, match and compare the three-dimensional topological features with the three-dimensional feature templates in the pre-stored registered face database, and determine the person identifier corresponding to the target face based on the matching and comparison results.
[0082] In step 105, the three-dimensional topological feature is a feature vector describing the three-dimensional geometric structure of a face. The three-dimensional feature template is a standard sample of three-dimensional facial features pre-stored in the database. The target face is the face of the person associated with the three-dimensional feature template. The personnel identifier refers to the identity information of the target personnel determined through face recognition matching and comparison. Specifically, it includes the unique identification code, name, employee number, and other textual or numerical identifiers that can clearly distinguish the individual's identity, which are pre-stored in the registered face database. These identifiers correspond one-to-one with the three-dimensional feature templates. When a match is successful, the system automatically retrieves and returns the corresponding identifier content for personnel identity verification in security monitoring.
[0083] In this embodiment, the vertices and edges of the 3D mesh model are input into a convolutional network. Through feature aggregation calculations at each layer of the network, local and global geometric features are extracted step by step. Finally, the output feature vector is compared with templates in the database to calculate similarity and identify the personnel information corresponding to the template with the highest matching degree.
[0084] For example, in the access control system of Company D, the reconstructed 3D mesh is input into a pre-trained network containing five graph convolutional layers, each aggregating information from three neighboring layers around a vertex. The final output 512-dimensional feature vector is compared with 500 employee templates in the database, and the template with the highest similarity is found, successfully identifying the current employee.
[0085] This method effectively addresses the performance degradation issue of traditional face recognition under low-light conditions through multimodal data fusion and 3D reconstruction technology. The introduction of temperature data ensures information acquisition in low-light environments, 3D modeling overcomes the influence of angle and occlusion, and graph neural network feature extraction improves recognition accuracy. The entire solution demonstrates stable and reliable recognition capabilities in complex monitoring scenarios.
[0086] To address the issue of insufficient accuracy in 3D face reconstruction under low-light conditions, in some embodiments, step 104: performing 3D face reconstruction on the enhanced multi-channel image to generate a 3D mesh model of the target face includes:
[0087] Step 201: Perform facial landmark localization on the enhanced multi-channel image.
[0088] In step 201, facial landmark localization refers to the process of identifying facial feature contour points from enhanced multi-channel images. These landmarks include the location information of features such as the corners of the eyes, the tip of the nose, and the corners of the mouth.
[0089] In this embodiment, the system first loads a pre-trained keypoint detection model. This model analyzes the facial region features in the image to gradually locate the precise positions of each keypoint. Multi-scale feature fusion technology is employed during the detection process to ensure accurate identification of keypoints under different lighting conditions. After localization, a data structure containing the coordinates of all keypoints is output, providing reference points for subsequent 3D reconstruction.
[0090] Step 202: Generate a facial spatial point cloud based on the located facial key points.
[0091] In step 202, the face spatial point cloud refers to a set of spatial points consisting of the 3D coordinates of key points and their interpolated supplementary points, with each point containing 3D positional information. The point cloud density determines the level of detail in the subsequent mesh model.
[0092] In this embodiment, initial three-dimensional coordinates are calculated based on the two-dimensional coordinates of key points and camera parameters. A point diffusion algorithm based on the normal direction is used to generate supplementary points between the key points. During the supplementation process, the curvature characteristics of the face are considered, and the point density is increased in areas with large curvature changes, ultimately forming a uniformly distributed set of spatial points.
[0093] Step 203: Connect adjacent points in the face spatial point cloud to form a set of triangular patches.
[0094] In step 203, the triangular patch set refers to the combination of triangular patches formed by connecting three adjacent points in the spatial point cloud, with each patch representing a small area of the human face surface.
[0095] In this embodiment, an improved Delaunay triangulation algorithm is used to connect spatial points. First, an initial tetrahedral bounding box is constructed, and then points in the point cloud are gradually inserted and the connection relationships are optimized. Curvature constraints are introduced during the triangulation process to ensure that the orientation of the face patches conforms to the curved features of a human face.
[0096] Step 204: Optimize the set of triangular facets and reconstruct the optimized set of triangular facets into a 3D face mesh model.
[0097] In this embodiment, the optimal position of each vertex is calculated iteratively to ensure a smooth transition between adjacent facets; elongated or excessively large abnormal facets are detected and removed; and interpolation patches are applied to empty regions based on the geometric features of surrounding facets. The final output is a three-dimensional mesh model that conforms to the anatomical structure of a human face.
[0098] Here is a specific example:
[0099] In the nighttime surveillance scenario at Bank A, based on the aforementioned 640×480 enhanced multi-channel image, the system first uses an improved feature point detection algorithm to locate 72 facial key points. The three-dimensional coordinates of each key point are calculated using the perspective projection formula z=f×B / d, where f is the camera focal length (8 mm), B is the binocular viewing distance (60 mm), and d is the parallax of the feature points in the left and right images. The depth information error of the key points calculated in this way is controlled within 0.3 mm. Then, using these key points as centers, according to the curvature of the face, 4 sampling points are generated per square centimeter in high-curvature areas such as the bridge of the nose and eye sockets, and 1 sampling point is generated per square centimeter in flat areas such as the cheeks. A point cloud containing approximately 9,500 spatial points was ultimately formed. Then, an improved triangulation algorithm was used to connect adjacent points to generate an initial mesh. This algorithm prioritizes ensuring that the side length of each triangular facet is between 3 and 8 millimeters, and the interior angle is between 30 and 120 degrees, generating approximately 18,500 initial triangular facets. Subsequently, the initial mesh was optimized, detecting and deleting 142 unacceptable facets. Simultaneously, 130 new facets were added to the deleted areas based on the orientation characteristics of surrounding facets. Finally, a 3D face mesh model composed of 18,488 high-quality triangular facets was generated. This model accurately reproduces the three-dimensional features of a human face, including a nose bridge height with 0.5 millimeters of precision and a lip contour with 1 millimeter of precision.
[0100] In this embodiment of the application, the three-dimensional reconstruction method effectively overcomes the deficiency of insufficient two-dimensional image information under low light conditions by accurately locating key points and generating intelligent point clouds, combined with mesh construction and optimization that takes into account the characteristics of human faces. The reconstructed three-dimensional model completely preserves the spatial structural features of the human face, providing an accurate geometric data basis for subsequent recognition.
[0101] To further improve the accuracy and completeness of 3D face reconstruction, in some embodiments, step 202: generating a face spatial point cloud based on the located facial key points includes:
[0102] Step 301: Based on the two-dimensional image coordinates of the facial key points and combined with the parameters of the security monitoring camera, calculate the initial three-dimensional coordinates of each facial key point in the spatial coordinate system.
[0103] In step 301, the camera parameters include internal and external parameters that affect 3D calculation, such as focal length and sensor size. The initial 3D coordinates refer to the spatial positions of key points calculated using 2D image coordinates and camera parameters, including positional information in the horizontal, vertical, and depth directions.
[0104] In this embodiment, the system first reads the pre-stored camera calibration parameters and converts the two-dimensional image coordinates of key points into three-dimensional spatial coordinates based on the principle of perspective projection. Lens distortion correction is considered during the conversion process to ensure the accuracy of the coordinate calculation. After the calculation is completed, a data structure containing the three-dimensional coordinates of all key points is output.
[0105] Step 302: Using the initial three-dimensional coordinates as the reference point, generate dense sampling points by expanding along the normal direction of the face surface.
[0106] In step 302, the normal direction of the face surface refers to the direction vector perpendicular to the face skin surface. This direction is calculated from the surface geometry features of the area surrounding the key points and is used to guide the expansion of dense sampling points along the extension direction of the natural curve of the face, ensuring that the newly added sampling points conform to the actual three-dimensional shape of the face. Dense sampling points refer to additional spatial points generated around the key points according to certain rules, used to fill the blank areas between the key points.
[0107] In this embodiment, sampling points are generated along the pre-calculated vertical direction of the face surface, centered on each key point. During the generation process, the sampling density is automatically adjusted according to the curvature changes of surrounding key points, and the number of sampling points is increased in complex areas such as facial features to ensure that the point cloud can accurately reflect the detailed features of the face.
[0108] Step 303: Arrange the dense sampling points corresponding to all facial key points into a three-dimensional coordinate set according to a preset topological order.
[0109] In step 303, the preset topological order refers to the point arrangement rules designed according to the anatomical characteristics of the human face, typically organizing the point data in the order from forehead to chin and from left ear to right ear. The three-dimensional coordinate set refers to the data structure that organizes all sampling points according to their spatial positional relationships, where each point contains coordinate values in the X, Y, and Z directions. These points are arranged in an anatomical order from forehead to chin and from left ear to right ear, forming an ordered set of points that can completely describe the three-dimensional shape of the human face.
[0110] In this embodiment, the system establishes point sorting rules based on a standard face model, and arranges the generated sampling points in an orderly manner according to region division and spatial position relationship. During the arrangement process, the spatial continuity of adjacent points is maintained, forming a structured three-dimensional coordinate set.
[0111] Step 304: Perform a normalization operation on the three-dimensional coordinate set to generate a normalized three-dimensional coordinate set, and use the normalized three-dimensional coordinate set as the face spatial point cloud.
[0112] In step 304, the normalization operation refers to the process of optimizing point cloud data, including operations such as removing redundant points, supplementing missing points, and homogenizing point distribution.
[0113] In this embodiment, redundant points that are too close together are first detected and merged. Then, new sampling points are inserted in areas with insufficient point density. During the insertion process, the distribution characteristics of surrounding points are referenced to ensure that the new points connect naturally with the existing points. Finally, the overall point cloud is smoothed to output a uniformly distributed set of three-dimensional coordinates.
[0114] Here is a specific example:
[0115] In the nighttime surveillance scenario at Bank A, based on the aforementioned 72 facial key points and their 3D coordinates, the system first calculates the normal direction according to the curvature of the surface in the area where each key point is located. The normal direction is determined by the local surface fitting plane formed by the key point and its eight adjacent points, ensuring that the direction vector accurately reflects the orientation of the face surface. Then, along the normal direction, 2 sampling points are generated per millimeter in the bridge of the nose area, 1.5 sampling points per millimeter in the eye socket area, and 0.8 sampling points per millimeter in the cheek area. The sampling point spacing is determined by the formula Δs=k / R, where Δs represents the sampling spacing and k is an adjustment coefficient of 0.6. R represents the local surface radius. This calculation yields high-density sampling with a spacing of 0.5 mm in the bridge of the nose region and 1.2 mm in the flat region. The generated sampling points are then arranged in order from the center of the forehead outwards and from top to bottom, forming an initial three-dimensional coordinate set containing approximately 9800 points. Finally, this set is normalized, and 150 redundant points with a spacing of less than 0.3 mm are detected and merged. 120 new sampling points are added in areas with obvious features such as the corners of the eyes, ultimately generating a normalized point cloud of 9670 points. The point cloud density reaches 4 points per square centimeter in the bridge of the nose region and 1 point per square centimeter in the cheek region.
[0116] In this embodiment, the point cloud generation method constructs a spatial point cloud that fully reflects the three-dimensional structure of a face by accurately calculating three-dimensional coordinates and intelligently distributing sampling points, combined with normalization processing that takes into account facial features. This effectively solves the problem of insufficient sampling in complex areas by traditional methods and lays a solid foundation for high-quality three-dimensional face reconstruction.
[0117] To further improve the quality and accuracy of the 3D mesh model, in some embodiments, step 204: optimizing the triangular facet set and reconstructing the optimized triangular facet set into a human face 3D mesh model includes:
[0118] Step 401: Traverse each vertex of the triangle facet set, calculate the geometric position deviation between the vertex of the triangle facet and the vertex of the adjacent triangle facet, and when the geometric position deviation is greater than a preset deviation threshold, adjust the vertex spatial coordinates to obtain the optimized triangle facet set.
[0119] In step 401, geometric position deviation refers to the degree of difference in spatial position between the current vertex and its neighboring vertices, which is measured by calculating the average offset of the coordinates of neighboring vertices. The preset deviation threshold is the maximum allowable deviation value set according to the smoothness requirements of the face surface. Adjusting the vertex spatial coordinates adjusts the vertex of the triangle currently being traversed, not the vertices of adjacent triangles. This adjustment operation corrects the spatial position of the current vertex based on the geometric position deviation between the current vertex and its surrounding neighboring vertices.
[0120] In this embodiment, the system sequentially checks the vertex position of each triangular facet and calculates the average positional difference between that vertex and all its directly connected adjacent vertices. When the difference exceeds a set threshold, the vertex coordinates are gradually adjusted towards the average position of its adjacent vertices to make the surface transition smoother. During the adjustment process, important geometric features of the original vertices are preserved to avoid over-smoothing that could lead to loss of detail.
[0121] Step 402: Detect abnormal faces in the optimized triangular face set that meet preset conditions, remove all abnormal faces, and perform patching operation in the area where abnormal faces were removed to obtain the optimized triangular face set. The preset conditions include the maximum value of the face's interior angle being greater than a preset angle threshold and the minimum value of the face's side length being less than a preset length threshold.
[0122] In step 402, abnormal facets refer to triangular facets that do not conform to the geometric characteristics of the human face, including facets that are too long or too narrow or too large. The preset angle threshold and length threshold are facet quality evaluation criteria set based on the anatomical features of the human face.
[0123] In this embodiment, the system scans all triangular facets, detects facets with interior angles exceeding a threshold or side lengths shorter than a threshold, and marks them. After removing abnormal facets, feature points are searched at the edges of the resulting holes, and these points are connected according to the orientation of the surrounding facets to generate new compliant facets. During the patching process, the geometric continuity between the new facets and the surrounding facets is ensured.
[0124] Step 403: Reconnect the triangular faces in the optimized triangular face set to generate a 3D face mesh model.
[0125] In step 403, reconnection refers to constructing a new mesh topology based on the optimized vertex and face relationships.
[0126] In this embodiment, the system re-establishes the vertex connection table based on the optimized patch set, checking and correcting all unreasonable connections. Then, it organizes the patch data according to the face region, generating a structurally complete and topologically correct 3D mesh model, ensuring that the model has no holes or overlapping patches.
[0127] Here is a specific example:
[0128] In the nighttime surveillance scenario of Bank A, based on the 18,500 initial triangular faces generated above, the system first calculates the average positional deviation between each vertex and its six adjacent vertices. When the deviation exceeds 0.4 mm, a weighted adjustment algorithm is used to move the vertex coordinates towards the average position of the adjacent vertices. The moving distance Δx = α × δ, where α is a smoothing coefficient of 0.3 and δ is the original deviation value. A total of 345 vertices are adjusted. Then, the face quality is detected, and faces with interior angles greater than 105 degrees or side lengths less than 2.8 mm are marked as abnormal faces. A total of 155 abnormal faces are removed. In the removed area, 148 new faces are added based on the average side length and angle characteristics of the surrounding faces. The side length of the new faces is determined by the formula L = (L1 + L2 + L3) / 3, where L1, L2, and L3 are the side lengths of the three surrounding reference faces. Finally, all faces are reconnected to generate an optimized mesh composed of 18,493 triangular faces.
[0129] In this embodiment, the mesh optimization method effectively eliminates geometric defects in the initial mesh by adjusting vertex positions and processing abnormal patches, combined with patching operations that take into account facial features. The generated mesh model has good smoothness and integrity, providing an accurate three-dimensional data foundation for face recognition.
[0130] To further improve the accuracy and reliability of face recognition, in some embodiments, step 105: extracting three-dimensional topological features from the face 3D mesh model using a convolutional graph neural network, matching and comparing the three-dimensional topological features with three-dimensional feature templates in a pre-stored registered face database, and determining the person identifier corresponding to the target face based on the matching and comparison results, includes:
[0131] Step 501: Input the vertex connection relationship of the three-dimensional face mesh model into the convolutional graph neural network.
[0132] In step 501, vertex connectivity refers to the connection structure information between vertices in the 3D mesh model, including the adjacency relationship and geometric distance between vertices, which is used to describe the topological structure of the face.
[0133] In this embodiment, the system first converts the 3D mesh model into a graph data structure, where each vertex is a graph node and face edges are graph edges. During the conversion, the spatial coordinates and curvature of vertices are preserved as node features, and the length and connection angle of edges are preserved as edge features, thus constructing a complete graph representation input network.
[0134] Step 502: Aggregate the geometric features of each vertex and its neighborhood through the convolutional graph neural network.
[0135] In step 502, geometric feature aggregation refers to the process of gradually fusing the features of a vertex itself and the features of its surrounding vertices through network layers, capturing local and global geometric characteristics through multi-level feature transfer.
[0136] In this embodiment, the network employs a hierarchical processing approach. The first layer aggregates features from the directly connected neighborhoods of each vertex, and each subsequent layer expands the aggregation range. Each layer updates vertex features through feature transformation and neighborhood information fusion, gradually extracting more abstract geometric feature representations.
[0137] Step 503: Generate 3D topological features based on the geometric features of all aggregated vertices.
[0138] In this embodiment, global pooling is performed on the features of all vertices in the last layer of the network, and the features of each vertex are fused by weighted summation to generate a fixed-dimensional feature vector. This vector preserves both the global shape features of the face and contains detailed information of key regions.
[0139] Step 504: Calculate the similarity score between the three-dimensional topological feature and the three-dimensional feature template. The similarity score is the matching comparison result.
[0140] In step 504, the similarity score is a matching metric calculated by comparing the degree of difference between the current facial features and the database template features.
[0141] In this embodiment, the cosine similarity method is used to calculate the cosine of the angle between the feature vectors in space as the similarity score. During calculation, the feature vectors are normalized to eliminate the influence of dimensions and ensure that the score ranges from 0 to 1.
[0142] Step 505: When the matching comparison result is greater than the preset similarity threshold, output the personnel identifier associated with the three-dimensional feature template.
[0143] In this embodiment, the system maintains a database organized according to feature vectors, and quickly retrieves the most similar templates during matching. When the highest similarity exceeds a threshold, the system returns the identity information associated with that template; otherwise, it marks it as an unknown person.
[0144] Here is a specific example:
[0145] In the nighttime surveillance scenario of Bank A, based on the aforementioned 3D face mesh model composed of approximately 20,000 triangular facets, the system first converts the mesh into graph structure data. Each vertex serves as a graph node, totaling approximately 10,000 nodes, and each triangular facet's edge serves as a graph edge, totaling approximately 30,000 edges. Node features include vertex coordinates, curvature, and normal vector information. This graph data is then input into a 5-layer graph convolutional network for processing. The first layer aggregates features from a one-layer neighborhood around each vertex, and the fifth layer expands to a five-layer neighborhood. The network uses the feature transformation formula H(l+1)=σ(D~-1 / 2A~D~-1 / 2H(l)W (l)), where A~ is the adjacency matrix with self-connection, D~ is the degree matrix, H(l) is the feature of the l-th layer, W(l) is the learnable parameter, and σ is the activation function; the network finally outputs a 512-dimensional feature vector, which is compared with 1000 three-dimensional feature templates stored in the bank employee database. An improved cosine similarity is used to calculate S=1-||F1-F2|| / max(||F1||,||F2||), where F1 and F2 represent the input feature and the template feature, respectively; when the similarity with a certain template reaches the threshold of 0.94, the system automatically returns the employee ID and name information associated with that template.
[0146] In this embodiment of the application, the method effectively overcomes the limitations of traditional two-dimensional face recognition under angle changes and occlusion conditions by learning multi-level features of graph convolutional networks and accurately matching three-dimensional topological features, significantly improving the recognition success rate in complex scenarios and providing a reliable identity verification means for high-security locations.
[0147] To further improve the accuracy and efficiency of multimodal data registration, in some embodiments, step 102: performing spatiotemporal registration of the infrared temperature data and the visible light texture data to generate temperature texture association information includes:
[0148] Step 601: Establish a position mapping relationship based on the pixel coordinates of the infrared temperature data and the pixel coordinates of the visible light texture data.
[0149] In step 601, the position mapping relationship refers to the coordinate correspondence between infrared image pixels and visible light image pixels. Data alignment is achieved by establishing transformation rules between the two image coordinate systems.
[0150] In this embodiment, the system first reads pre-calibrated camera intrinsic and extrinsic parameters, including focal length and principal point position. Then, it uses a feature point matching algorithm to find corresponding feature point pairs in the two images and calculates a precise coordinate transformation matrix using these matching points. Finally, the transformation parameters are saved for subsequent data processing.
[0151] Step 602: Based on the location mapping relationship, infrared temperature data and visible light texture data at the same timestamp are superimposed on the same spatial coordinate system, and a corresponding associated information unit is generated for each spatial coordinate point.
[0152] In step 602, the associated information unit refers to a composite data structure formed by packaging infrared temperature values and visible light texture values at the same spatial location under a unified coordinate system.
[0153] In this embodiment, the system locates the corresponding visible light image region for each infrared pixel based on the calculated position mapping relationship. Then, the infrared temperature value at that location is combined with the visible light pixel value to form an associated unit containing multimodal information. Time synchronization is considered during processing to ensure that the data originates from the same moment.
[0154] Step 603: Integrate the associated information units corresponding to each spatial coordinate point into temperature texture associated information.
[0155] In this embodiment, the system arranges the associated units at each location into a matrix structure according to the image scanning order. Each element in the matrix contains a temperature value and a corresponding texture value, forming structured data that can be directly used for subsequent processing.
[0156] Here is a specific example:
[0157] In a nighttime surveillance scenario at Bank A, after the system simultaneously acquires infrared temperature data at a resolution of 256×192 and visible light images at a resolution of 1280×720, it first establishes a coordinate mapping relationship based on pre-stored camera calibration parameters. Each pixel in the infrared image corresponds to a 2×2 pixel region in the visible light image. A feature point matching algorithm is used to select 50 corresponding feature points from each of the two images. The precise affine transformation matrix M = [a11 a12 a13;a21 a22 a23;0 0] is then calculated using the least squares method. [1], where parameters a11, a12, etc. represent scaling, rotation, and translation transformations; then, the temperature data and visible light images acquired at the same time are aligned according to this matrix, and for each temperature data point in the infrared image, four corresponding visible light pixels are found, generating 256×192 associated units. Each unit contains one temperature value and four visible light pixel values, where the temperature data range remains unchanged from 32 to 42 degrees Celsius, and the visible light pixel values are grayscale values from 0 to 255; finally, these associated units are integrated into a temperature texture association information matrix in row and column order, where the position of each element in the matrix corresponds one-to-one with the pixel position of the original infrared image, ensuring that subsequent processing can accurately obtain multimodal data at each spatial position.
[0158] In this embodiment of the application, the registration method effectively solves the alignment problem of different sensor data in space and time through precise coordinate mapping and multimodal data combination. The generated correlation information completely preserves the correspondence between temperature distribution and texture features, providing a high-quality data foundation for subsequent fusion processing.
[0159] To further improve the quality and practicality of multispectral image fusion, in some embodiments, step 103: generating an enhanced multichannel image based on the temperature texture association information using multispectral fusion technology includes:
[0160] Step 701: Convert the infrared radiation intensity value in the temperature texture association information into first channel data.
[0161] In step 701, the infrared radiation intensity value specifically refers to the infrared thermal radiation energy value after quantization, while the infrared temperature data is the raw temperature measurement value directly obtained from the sensor. The former is the latter's representation after a specific conversion. The first channel data is the image data after converting the raw temperature value to a standard range.
[0162] In this embodiment, the system reads the original temperature values from the temperature texture association information and maps them to a preset numerical range through a linear transformation. The transformation process preserves the relative relationships of the temperature distribution, ensuring that thermal feature information is not lost, and generates standard format data that can be used for image processing.
[0163] Step 702: Convert the visible light texture detail values in the temperature texture association information into second channel data.
[0164] In step 702, the visible light texture detail value originates from the visible light texture data, which represents the surface microstructure features of the target face in the visible light band, acquired by the visible light sensor of the security monitoring camera. The second channel data is the preprocessed visible light image data.
[0165] In this embodiment, the system performs noise reduction and enhancement processing on the visible light pixel values in the associated information to highlight the texture features of key areas such as facial features. The processed data retains its original resolution and forms a correspondence with the data in the first channel.
[0166] Step 703: Use multispectral fusion technology to fuse the first channel data and the second channel data into an initial multichannel image.
[0167] In step 703, the initial multi-channel image refers to the preliminary fusion result generated by simply superimposing different spectral data, which contains the original feature information from different sensors.
[0168] In this embodiment, the system combines the data from two channels according to their spatial correspondence to generate a dual-channel image. Data alignment is maintained during the combination process to avoid information misalignment, providing a foundation for subsequent optimization processing.
[0169] Step 704: Adjust the radiation intensity and texture detail ratio of the initial multi-channel image through the inter-channel weight allocation relationship to generate an adjusted multi-channel image, and use the adjusted multi-channel image as an enhanced multi-channel image.
[0170] In step 704, the channel weight allocation relationship refers to dynamically adjusting the contribution ratio of each channel's data in the final image according to different environmental conditions and application requirements. Radiation intensity and infrared radiation intensity value are different expressions of the same meaning, both referring to the converted infrared channel data value. Texture detail and visible light texture detail value are different expressions of the same meaning, both referring to the processed visible light channel image feature data.
[0171] In this embodiment, the system automatically calculates the optimal weight combination based on ambient light intensity and target detection requirements. A weighted fusion algorithm is used to adjust the display intensity of each channel, generating an enhanced image with balanced visual effects and feature extraction results.
[0172] Here is a specific example:
[0173] In the nighttime surveillance scenario of Bank A, based on the aforementioned 256×192 temperature texture association units, the system first linearly converts the temperature data of 32-42 degrees Celsius into the first channel data in the range of 0-255 using the formula N=(T-32) / (42-32)×255, where N is the converted grayscale value and T is the original temperature value; then, the average value of the four visible light pixels in the association unit is taken as the second channel data; next, a wavelet transform-based fusion algorithm is used to combine the data of the two channels into an initial multi-channel image, where the low-frequency part is taken as 70% of the temperature data and 30% of the visible light data, and the high-frequency part is taken as 30% of the temperature data and 70% of the visible light data; finally, based on the light intensity value L detected by the ambient light sensor, the temperature channel weight W_t is calculated using the dynamic weight formula W_t=0.5+0.3×(1-L / L_max), and the visible light channel weight W_v=1-W_t, where L_max is the maximum range of the sensor, generating an enhanced multi-channel image with a resolution of 640×480.
[0174] In this embodiment, the multispectral fusion method effectively integrates the advantages of infrared and visible light data through intelligent weight allocation and hierarchical fusion strategy. The generated enhanced image retains temperature distribution characteristics and highlights key texture details, providing reliable data support for face recognition under different lighting conditions.
[0175] Figure 2 A schematic diagram of a security monitoring camera detection system based on face recognition is provided in an embodiment of this application, as shown below. Figure 2 As shown, the system includes:
[0176] The receiving module 21 is used to receive infrared temperature data and visible light texture data of the target face collected by the security monitoring camera.
[0177] The first generation module 22 is used to perform spatiotemporal registration of the infrared temperature data and the visible light texture data to generate temperature texture association information.
[0178] The second generation module 23 is used to generate an enhanced multi-channel image based on the temperature texture association information using multispectral fusion technology.
[0179] The reconstruction module 24 is used to perform three-dimensional face reconstruction on the enhanced multi-channel image and generate a three-dimensional face mesh model corresponding to the target face.
[0180] Extraction module 25 is used to extract three-dimensional topological features from the three-dimensional mesh model of the face using a convolutional graph neural network, match and compare the three-dimensional topological features with the three-dimensional feature templates in the pre-stored registered face database, and determine the person identifier corresponding to the target face based on the matching and comparison results.
[0181] Figure 2 The aforementioned security surveillance camera detection system based on facial recognition can perform... Figure 1 The implementation principle and technical effects of the face recognition-based security camera detection method described in the illustrated embodiment will not be repeated here. The specific methods by which each module and unit of the face recognition-based security camera detection system in the above embodiments perform their operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0182] In one possible design, Figure 2 The security surveillance camera detection system based on face recognition in the illustrated embodiment can be implemented as a computing device, such as... Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;
[0183] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 32.
[0184] The processing component 32 is used to perform the above. Figure 1 The embodiment describes a security monitoring camera detection method based on face recognition.
[0185] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-described method.
[0186] Storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0187] Of course, computing devices may also include other components, such as input / output interfaces, display components, communication components, etc.
[0188] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.
[0189] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.
[0190] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.
[0191] This application also provides a computer storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 1 The illustrated embodiment presents a security surveillance camera detection method based on face recognition.
[0192] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0193] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0194] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for detecting security surveillance cameras based on face recognition, characterized in that, include: Receive infrared temperature data and visible light texture data of the target face collected by security monitoring cameras; Spatiotemporal registration is performed on the infrared temperature data and the visible light texture data to generate temperature texture association information; Based on the temperature texture association information, an enhanced multi-channel image is generated using multispectral fusion technology; Three-dimensional face reconstruction is performed on the enhanced multi-channel image to generate a three-dimensional face mesh model corresponding to the target face; A convolutional graph neural network is used to extract three-dimensional topological features from the three-dimensional mesh model of the face. The three-dimensional topological features are then matched and compared with three-dimensional feature templates in a pre-stored registered face database. Based on the matching and comparison results, the person identifier corresponding to the target face is determined. The step of performing 3D face reconstruction on the enhanced multi-channel image to generate a 3D face mesh model corresponding to the target face includes: Facial landmark localization is performed on the enhanced multi-channel image; Generate a spatial point cloud of the face based on the located facial key points; Connect adjacent points in the face spatial point cloud to form a set of triangular patches; The set of triangular facets is optimized, and the optimized set of triangular facets is reconstructed into a three-dimensional mesh model of a human face. The step of generating a facial spatial point cloud based on the located facial key points includes: Based on the two-dimensional image coordinates of the facial key points, and combined with the parameters of the security monitoring camera, the initial three-dimensional coordinates of each facial key point in the spatial coordinate system are calculated. Using the initial three-dimensional coordinates as a reference point, dense sampling points are generated by expanding along the normal direction of the face surface; Arrange the dense sampling points corresponding to all facial key points into a three-dimensional coordinate set according to a preset topological order; Perform a normalization operation on the three-dimensional coordinate set to generate a normalized three-dimensional coordinate set, and use the normalized three-dimensional coordinate set as the face spatial point cloud; The optimization of the triangular facet set, and the reconstruction of the optimized triangular facet set into a 3D face mesh model, includes: Traverse each vertex of the triangle facet set, calculate the geometric position deviation between the vertex of the triangle facet and the vertex of the adjacent triangle facet, and when the geometric position deviation is greater than a preset deviation threshold, adjust the vertex spatial coordinates to obtain an optimized triangle facet set. Detect abnormal faces in the optimized triangular face set that meet preset conditions, remove all abnormal faces, and perform patching operation in the area where abnormal faces were removed to obtain the optimized triangular face set. The preset conditions include the maximum value of the face's interior angle being greater than a preset angle threshold and the minimum value of the face's side length being less than a preset length threshold. The triangular faces in the optimized triangular facet set are reconnected to generate a 3D mesh model of the face.
2. The method according to claim 1, characterized in that, The step of extracting 3D topological features from the 3D face mesh model using a convolutional graph neural network, matching and comparing the 3D topological features with 3D feature templates in a pre-stored registered face database, and determining the person identifier corresponding to the target face based on the matching and comparison results includes: The vertex connection relationships of the three-dimensional face mesh model are input into a convolutional graph neural network; The convolutional graph neural network aggregates the geometric features of each vertex and its neighborhood. Based on the geometric features of all aggregated vertices, generate 3D topological features; Calculate the similarity score between the three-dimensional topological feature and the three-dimensional feature template, where the similarity score is the matching comparison result; When the matching result is greater than the preset similarity threshold, the personnel identifier associated with the three-dimensional feature template is output.
3. The method according to claim 1, characterized in that, The step of performing spatiotemporal registration of the infrared temperature data and the visible light texture data to generate temperature texture association information includes: A positional mapping relationship is established based on the pixel coordinates of the infrared temperature data and the pixel coordinates of the visible light texture data; Based on the location mapping relationship, infrared temperature data and visible light texture data at the same time stamp are superimposed on the same spatial coordinate system, and a corresponding associated information unit is generated for each spatial coordinate point. The associated information units corresponding to each spatial coordinate point are integrated into temperature texture associated information.
4. The method according to claim 1, characterized in that, The step of generating an enhanced multi-channel image based on the temperature texture association information using multispectral fusion technology includes: The infrared radiation intensity value in the temperature texture association information is converted into first channel data; Convert the visible light texture detail values in the temperature texture association information into second channel data; The first channel data and the second channel data are fused into an initial multi-channel image using multispectral fusion technology. By adjusting the radiation intensity and texture detail ratio of the initial multi-channel image through the inter-channel weight allocation relationship, an adjusted multi-channel image is generated, and the adjusted multi-channel image is used as the enhanced multi-channel image.
5. A security monitoring camera detection system based on face recognition, characterized in that, include: The receiving module is used to receive infrared temperature data and visible light texture data of the target face captured by the security monitoring camera; The first generation module is used to perform spatiotemporal registration of the infrared temperature data and the visible light texture data to generate temperature texture association information. The second generation module is used to generate an enhanced multi-channel image based on the temperature texture association information using multispectral fusion technology; The reconstruction module is used to perform three-dimensional face reconstruction on the enhanced multi-channel image and generate a three-dimensional face mesh model corresponding to the target face. The extraction module is used to extract three-dimensional topological features from the three-dimensional mesh model of the face using a convolutional graph neural network, match and compare the three-dimensional topological features with the three-dimensional feature templates in the pre-stored registered face database, and determine the personnel identifier corresponding to the target face based on the matching and comparison results. The step of performing 3D face reconstruction on the enhanced multi-channel image to generate a 3D face mesh model corresponding to the target face includes: Facial landmark localization is performed on the enhanced multi-channel image; Generate a spatial point cloud of the face based on the located facial key points; Connect adjacent points in the face spatial point cloud to form a set of triangular patches; The set of triangular facets is optimized, and the optimized set of triangular facets is reconstructed into a three-dimensional mesh model of a human face. The step of generating a facial spatial point cloud based on the located facial key points includes: Based on the two-dimensional image coordinates of the facial key points, and combined with the parameters of the security monitoring camera, the initial three-dimensional coordinates of each facial key point in the spatial coordinate system are calculated. Using the initial three-dimensional coordinates as a reference point, dense sampling points are generated by expanding along the normal direction of the face surface; Arrange the dense sampling points corresponding to all facial key points into a three-dimensional coordinate set according to a preset topological order; Perform a normalization operation on the three-dimensional coordinate set to generate a normalized three-dimensional coordinate set, and use the normalized three-dimensional coordinate set as the face spatial point cloud; The optimization of the triangular facet set, and the reconstruction of the optimized triangular facet set into a 3D face mesh model, includes: Traverse each vertex of the triangle facet set, calculate the geometric position deviation between the vertex of the triangle facet and the vertex of the adjacent triangle facet, and when the geometric position deviation is greater than a preset deviation threshold, adjust the vertex spatial coordinates to obtain an optimized triangle facet set. Detect abnormal faces in the optimized triangular face set that meet preset conditions, remove all abnormal faces, and perform patching operation in the area where abnormal faces were removed to obtain the optimized triangular face set. The preset conditions include the maximum value of the face's interior angle being greater than a preset angle threshold and the minimum value of the face's side length being less than a preset length threshold. The triangular faces in the optimized triangular facet set are reconnected to generate a 3D mesh model of the face.
6. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a security monitoring camera detection method based on face recognition as described in any one of claims 1-4.
7. A computer storage medium, characterized in that, The device contains a computer program that, when executed by a computer, implements a security monitoring camera detection method based on face recognition as described in any one of claims 1-4.
Citation Information
Patent Citations
Face recognition method and system
CN111985348A
Three-dimensional face reconstruction method for removing invalid points based on coding structured light
CN118918288A
Access control face recognition system and method based on three-dimensional technology
CN119418434A