A zero-interaction permission recognition method and system based on face recognition and ERP / MES integration
By integrating a multi-camera network with an ERP/MES-integrated identity recognition system, the problems of low efficiency and poor security of traditional identity recognition methods in industrial environments have been solved, achieving efficient access control that is seamless and imperceptible.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN NEW INSPIRATION TECH CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional identity verification methods suffer from problems such as easy loss and duplication, low recognition rate, inability to be deeply integrated with ERP/MES systems, and the need for manual interaction in complex industrial environments, resulting in low efficiency and insufficient security of access control.
The system uses a multi-camera network to capture facial images in real time, and combines computer vision and artificial intelligence technologies for preprocessing and liveness detection. It then compares these images with the employee feature database in real time through the ERP/MES system, generates permission binding tags, automatically triggers device authorization, and records audit logs.
It achieves seamless and efficient access control, eliminates the loopholes of proxy access and fraudulent claims, improves identification accuracy and response speed, and enhances the security and flexibility of the system.
Smart Images

Figure CN122133122A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information security identification technology, and in particular to a zero-interaction permission identification method and system based on face recognition and ERP / MES integration. Background Technology
[0002] With the development of intelligent manufacturing, traditional identification methods relying on physical cards, password input, or manual verification are no longer sufficient to meet the demands of efficient, secure, and traceable modern production management. Especially in factory settings with frequent personnel movement, complex work areas, and strict hierarchical access control, achieving accurate personnel identification without increasing operational burden, and integrating with existing information management systems to complete real-time access control, has become a critical issue that urgently needs to be addressed.
[0003] Currently, in most manufacturing enterprises, employee access to specific functional areas (such as workshops, warehouses, and laboratories) typically relies on card swiping, fingerprint scanning, or other manual interaction methods for access control. While these methods meet the basic requirements of access management to some extent, they have revealed many shortcomings in practical applications. For example, cards are easily lost, copied, or lent to others, creating security management vulnerabilities; fingerprint recognition is greatly affected by skin condition, resulting in a high failure rate; furthermore, these traditional authentication methods mostly operate independently within security systems, lacking effective integration with business systems, and cannot dynamically obtain user job information, task assignments, etc., leading to a static and rigid permission granting process that is difficult to adapt to rapidly changing production rhythms. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides a zero-interaction permission recognition method and system based on the integration of face recognition with ERP / MES.
[0005] Firstly, this application provides a zero-interaction permission recognition method based on face recognition and ERP / MES integration, employing the following technical solution: A zero-interaction permission recognition method based on face recognition and ERP / MES integration, the recognition method comprising: Real-time facial images are captured through a multi-camera network deployed in different scenarios; The face image is preprocessed to output a standardized face image; A facial feature relationship network is constructed based on the standardized face images, facial feature vectors are extracted, liveness detection is performed, and liveness detection confidence is generated. When the confidence level of the liveness detection exceeds a preset confidence threshold, the identity feature with the liveness detection confidence level is output. The identity features are compared in real time with the pre-stored employee feature database in the ERP / MES system. Based on the comparison results, the corresponding employee identity identifier and real-time permission data are bound together to generate permission binding tags. Based on the permission binding tag, match the current scenario and automatically trigger the corresponding device automatic authorization command; Record the execution events of automatic authorization commands for equipment and the associated ERP / MES task documents to generate structured audit logs.
[0006] By adopting the above-mentioned technical solution, a series of cutting-edge technologies such as computer vision, artificial intelligence, and industrial IoT are fully integrated. Based on the three design concepts of "unmanned operation, seamless access, and seamless integration," a highly integrated, dynamically responsive, and fully traceable security management ecosystem has been created. Compared with the traditional manual card-swiping and paper-based approval management model, this technical solution significantly shortens the time required for identity authentication, eliminates the risk of fraudulent card swiping and misappropriation, enhances the emergency response speed for sudden events, and effectively supports the implementation of the strategic goal of intelligent manufacturing transformation and upgrading.
[0007] Optionally, the step of preprocessing the face image to output a standardized face image includes: Receives facial images from different scenes captured by a multi-camera network; Perform grayscale conversion on the face image to generate a grayscale face image; Detect the locations of facial key points in the grayscale face image; Based on the positions of the facial key points, a rotation correction matrix is calculated, facial alignment is performed, and a corrected image is generated. Environmental interference compensation processing is performed on the corrected image to generate an anti-interference image; The anti-interference image is scaled to a preset size, and a standardized face image block is output.
[0008] By adopting the above technical solution, a complete preprocessing pipeline was constructed through optimized design for complex shooting conditions in industrial scenarios. Compared with traditional solutions that rely solely on simple filtering and scaling, the technical solution of this application deeply integrates two major categories of methods: geometric correction and physical modeling. It exhibits stronger robustness and generalization ability when facing harsh working conditions such as strong light glare and smoke and dust, thereby significantly improving the accuracy and timeliness of the entire face recognition system.
[0009] Optionally, the steps of constructing a facial feature relationship network based on the standardized face image, extracting facial feature vectors, performing liveness detection, and generating liveness detection confidence include: The standardized face image is divided into a set of local regions and the initial feature vector of each local region is extracted; Construct the topological relationships between the initial feature vectors of each local region; Based on the aforementioned topological relationships, local region features are fused to generate a global face feature vector; Extracting physiological dynamic features from a continuous sequence of video frames; The global facial feature vector and physiological dynamic features are combined to make predictions and generate a liveness detection confidence score.
[0010] By adopting the above technical solutions, a topological graph model between local facial regions is constructed, and a graph neural network propagation mechanism is used to extract deep semantic associations, effectively solving the problem of false detection under complex posture changes and occlusion. At the same time, with the addition of a non-invasive physiological feature acquisition path, the ideal application scenario requirements of non-interactive and all-weather operation are truly realized.
[0011] Optionally, the step of comparing the identity features with the pre-stored employee feature database in the ERP / MES system in real time, and generating permission binding tags by binding the corresponding employee identity identifier and real-time permission data according to the comparison results includes: Access the pre-stored employee feature database stored in the ERP / MES system; wherein, the pre-stored employee feature database contains employee identification and associated basic feature data; Calculate the similarity value between the identity features and each basic feature data in the pre-stored employee feature database in real time; Based on a preset similarity threshold, the matching employee identity identifier is determined according to the similarity value; Query the ERP / MES system to obtain the real-time permission data corresponding to the employee's identity identifier; The employee identification is bound to the real-time permission data to generate permission binding tags.
[0012] By adopting the above technical solutions, facial recognition technology is deeply integrated with ERP / MES systems to build a fully automated access control system. This system not only achieves real-time mapping between identity and permissions but also possesses excellent scalability and security. It can effectively address the access control needs in complex and ever-changing production environments, significantly improving the intelligence level and response speed of enterprise information systems in terms of access control.
[0013] Optionally, the step of automatically triggering the corresponding device automatic authorization instruction based on the permission binding tag matching the current scenario includes: Obtain the physical location identifier and device type identifier of the current scene; Match the real-time permission data in the permission binding tag with the device type identifier of the current scene; When a match is successful, the corresponding automatic device authorization command is generated; The device automatic authorization command is sent to the target device execution terminal.
[0014] By adopting the above technical solutions, not only has the scenario-based, dynamic, and secure nature of access control been achieved, but the direct execution of access commands has also been realized through the Industrial Internet of Things protocol. This has significantly improved the system's response efficiency and ease of operation, breaking through the technical bottlenecks of traditional access control systems, such as the need for manual intervention, coarse access control, and delayed response. It has truly realized a "zero-interaction, high-security, and fine-grained" intelligent device authorization mechanism, which has broad application prospects and significant technological progress.
[0015] Optionally, when the confidence level of the liveness detection exceeds a preset confidence threshold, a multimodal verification step is initiated, specifically including: Trigger the infrared camera to collect thermal imaging data and generate a temperature distribution matrix; Detecting live biological feature patterns in the temperature distribution matrix; The live biological feature patterns are compared with a pre-existing live biological spectral library to output an auxiliary confidence level. The auxiliary confidence score and the liveness detection confidence score are combined to generate a composite verification result.
[0016] By adopting the above technical solution, the traditional visible light liveness detection mechanism has been effectively enhanced. This mechanism relies on the unique thermodynamic properties of living tissue and, through intelligent fusion with the original detection results, not only improves the recognition accuracy but also enhances the system's ability to resist various deception attacks. It is especially suitable for application scenarios with extremely high security requirements, such as industrial control and security access control.
[0017] Optionally, after the step of generating structured audit logs, the following may also be included: Parse the ERP / MES task document operation records in the structured audit log; Based on the ERP / MES task document operation record, retrieve the camera video stream data with the corresponding timestamp and extract the operation behavior trajectory; The operation behavior trajectory is matched with the standard operation process template of the task document to obtain the matching deviation. When the matching deviation exceeds the limit, an permission exception event is generated and the automatic authorization command of the device is frozen.
[0018] By adopting the above technical solution, the log data analysis capabilities within the information system are cleverly integrated with visual perception methods from the external physical world, enabling two-way verification and deviation early warning throughout the entire process from business instruction issuance to on-site operation execution. This technical solution not only enhances the post-event traceability depth of the original access control system but also endows it with forward-looking risk control and adjustment functions, making the access control system in the entire intelligent manufacturing environment more intelligent, refined, and reliable.
[0019] Optionally, a permission conflict detection step is added before the step of executing the device automatic authorization instruction, specifically including: Retrieve the real-time status identifier and current occupant identifier of the target device; When the target device is occupied, compare the permission priority parameter of the current occupant with the current requester corresponding to the device's automatic authorization instruction. A conflict resolution strategy is generated based on the permission priority parameter, including an instruction queuing sequence or preemptive execution of instructions; Write the conflict resolution strategy into an additional field of the structured audit log.
[0020] By adopting the above technical solution, starting from the monitoring of underlying equipment status and combining it with a refined permission priority modeling method, a flexible and efficient instruction scheduling function is achieved. Furthermore, a comprehensive log recording system enhances the overall risk control level. Compared to the traditional static authorization management model, this application significantly improves the security adaptability and operational management flexibility in the face of complex and ever-changing industrial environments, effectively avoiding the probability of work delays or even safety accidents caused by multiple people vying for critical equipment.
[0021] Secondly, this application provides a zero-interaction permission recognition system based on face recognition and ERP / MES integration, employing the following technical solution: A zero-interaction permission recognition system based on facial recognition and ERP / MES integration, the recognition system comprising: The image acquisition module is used to acquire facial images in real time through a network of multiple cameras deployed in different scenarios; The image preprocessing module is used to preprocess the face image and output a standardized face image; The liveness detection module is used to construct a facial feature relationship network based on the standardized face image, extract facial feature vectors, perform liveness detection, and generate liveness detection confidence. The identity feature output module is used to output identity features with liveness detection confidence when the liveness detection confidence exceeds a preset confidence threshold. The permission binding module is used to compare the identity features with the pre-stored employee feature database in the ERP / MES system in real time, bind the corresponding employee identity identifier and real-time permission data according to the comparison results, and generate permission binding tags. The device authorization module is used to match the current scenario based on the permission binding tag and automatically trigger the corresponding device automatic authorization instruction; The log generation module is used to record the execution events of automatic authorization commands for devices and the associated ERP / MES task documents, generating structured audit logs.
[0022] Thirdly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect. Attached Figure Description
[0023] Figure 1 This is a first flowchart illustrating a zero-interaction permission recognition method based on face recognition and ERP / MES integration, which is one embodiment of this application.
[0024] Figure 2 This is a second flowchart illustrating a zero-interaction permission recognition method based on face recognition and ERP / MES integration, which is one embodiment of this application.
[0025] Figure 3 This is a schematic diagram of the third process of a zero-interaction permission recognition method based on face recognition and ERP / MES integration, which is one embodiment of this application.
[0026] Figure 4 This is a schematic diagram of the fourth process of a zero-interaction permission recognition method based on face recognition and ERP / MES integration, which is one embodiment of this application.
[0027] Figure 5 This is a schematic diagram of the fifth process of a zero-interaction permission recognition method based on face recognition and ERP / MES integration, which is one embodiment of this application.
[0028] Figure 6 This is a schematic diagram of the sixth process of a zero-interaction permission recognition method based on face recognition and ERP / MES integration, which is one embodiment of this application.
[0029] Figure 7 This is a schematic diagram of the seventh process of a zero-interaction permission recognition method based on face recognition and ERP / MES integration, which is one embodiment of this application.
[0030] Figure 8 This is the eighth flowchart of a zero-interaction permission recognition method based on face recognition and ERP / MES integration, which is one embodiment of this application. Detailed Implementation
[0031] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figure 1-8 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.
[0032] Currently, in manufacturing enterprises, traditional access control methods mostly rely on physical keys, access cards, and passwords. These methods have drawbacks such as being easily lost, easily copied, and complex to manage. With the development of facial recognition technology, some enterprises have begun to apply facial recognition access control systems. However, existing systems often suffer from the following problems: they require personnel to actively cooperate with the identification process, affecting work efficiency; the access control system operates independently from the ERP / MES system, leading to data inconsistencies and delays in access updates; single-camera recognition is affected by factors such as lighting and angle in complex industrial environments, resulting in low accuracy; they lack liveness detection capabilities, posing a risk of being fooled by photos or videos; and they cannot achieve seamless, zero-interaction access control, requiring employees to undergo multiple verifications when accessing different areas and devices.
[0033] While existing technologies have improved recognition efficiency through setting authentication security levels and system linkage, they still require personnel to cooperate in recognition at certain locations and lack deep integration with ERP / MES systems. Some existing solutions for facial recognition and verification devices used in production process monitoring systems, while solving the recognition problem at densely populated entrances and exits, are mainly applied to specific areas and have not been extended to access management across the entire production environment. Therefore, existing access management systems still have many shortcomings in practical applications, particularly in terms of recognition efficiency, system integration, environmental adaptability, and security, requiring further improvement and refinement.
[0034] Based on this, this application discloses a zero-interaction permission recognition method based on face recognition and ERP / MES integration.
[0035] Reference Figure 1 A zero-interaction permission recognition method based on face recognition and ERP / MES integration, the recognition method includes: Step S101: Real-time acquisition of facial images through a multi-camera network deployed in different scenarios; Cameras, serving as front-end sensing units, are widely distributed in key locations such as factory entrances, workstations alongside production lines, and warehouse management areas. This distributed deployment not only ensures coverage of all personnel movement paths but also, with the support of edge computing architecture, enables preliminary image screening to be completed locally, thereby reducing the load on the central server.
[0036] Furthermore, using a combination of high-definition CMOS sensors and wide-angle lenses can effectively expand the monitoring range and improve capture efficiency. It should be noted that a multi-camera network is not merely a physical network of multiple devices connected in a mesh topology; it also embodies a deeper level of redundancy and backup mechanism in the visual information acquisition system. Even if one video stream is interrupted or blocked, the remaining nodes can continue to provide stable service, ensuring the robustness and availability of the overall system.
[0037] Step S102: Preprocess the face image to output a standardized face image; The preprocessing steps include grayscale conversion, face alignment, and environmental interference compensation. Specifically, the main purpose of this stage is to eliminate various noise sources in the original image, laying a good foundation for subsequent feature extraction. Grayscale conversion refers to converting a color RGB format image into a grayscale image represented by a single intensity value. This reduces storage space requirements and speeds up computation, while also avoiding the risk of misjudgment caused by color differences. Face alignment uses face detection algorithms such as HOG+SVM or MTCNN to accurately locate the coordinates of the eyes, nose tip, and corners of the mouth. Then, affine transformations are performed on these key parts to ensure that all face samples are in a uniform orientation and planar scale, facilitating subsequent modeling and training. As for environmental interference compensation, it covers two main categories: illumination correction and haze removal. The former usually uses CLAHE (Contrast Limited Adaptive Histogram Equalization) to alleviate the shadow blurring problem caused by backlighting, while the latter introduces dark channel prior theory to restore the image distortion caused by scattering of airborne particles. The combined effect of both significantly improves the model's generalization ability in complex environments.
[0038] Step S103: Construct a facial feature relationship network based on standardized face images, extract facial feature vectors and perform liveness detection, and generate liveness detection confidence. In this process, the CNN (Convolutional Neural Network) model under the deep learning framework plays a core role. It gradually abstracts a high-level representation with semantic meaning by performing convolutional pooling operations on the input image layer by layer, and then forms a set of fixed-dimensional (e.g., 128-dimensional) floating-point arrays to uniquely identify individual identity attributes.
[0039] At the same time, in order to prevent photo attacks, mask forgery and other non-real person impersonation incidents, it is also necessary to carry out Liveness Detection simultaneously. This involves using multiple sensor signals, such as infrared thermal imaging and near-infrared reflectivity measurement, to jointly evaluate the probability score of the target's authenticity. Only when the score is higher than a set threshold is the next step of the verification process allowed.
[0040] Step S104: When the confidence level of liveness detection exceeds the preset confidence threshold, output the identity features with the confidence level of liveness detection. This can be achieved through deep learning models, such as convolutional neural networks (CNNs), by extracting and encoding key facial points to generate a fixed-dimensional floating-point vector, i.e., the identity feature. This identity feature vector not only retains the main visual features of the face but also possesses a certain degree of robustness, maintaining a high recognition accuracy under interference factors such as changes in lighting and pose differences. The preset confidence threshold can be flexibly adjusted according to the actual application scenario. For example, it can be set to above 0.95 for extremely high-security locations, while for ordinary office areas, a threshold of 0.8 may be sufficient to meet basic protection requirements. Once a user is identified as legitimate, their corresponding vector encoding can be mapped to their registration file in the ERP / MES database, forming a complete chain of digital identity credentials.
[0041] Step S105: Compare the identity features with the pre-stored employee feature database in the ERP / MES system in real time, and bind the corresponding employee identity identifier and real-time permission data according to the comparison results to generate permission binding tags. The core idea of this step is to establish a mechanism for information sharing and coordinated response between cross-domain heterogeneous information systems. ERP (Enterprise Resource Planning) and MES (Manufacturing Execution System) respectively carry functional modules such as maintaining organizational personnel data and approving production orders. They often store basic files and job descriptions for each employee.
[0042] Through API calls or middleware bridging technology, the feature codes extracted from the visual side can be quickly retrieved into the corresponding entry records, and relevant context parameters such as job level, work group, and current to-do list can be read from them. Then, they are encapsulated into a structured Tag object and passed to the next stage for use.
[0043] Step S106: Match the current scenario based on the permission binding tag and automatically trigger the corresponding device automatic authorization command; Scene matching is essentially a state machine transition process driven by a rule engine. It requires comprehensive consideration of multiple dimensions of variables, such as geographical location fence division, time window constraints, and task priority ranking, in order to make an accurate response.
[0044] For example, if the identification result shows that Zhang San, the foreman of Class A, is near the assembly station on Line B, an unlock command should be immediately sent to the PLC controller to turn on the power switch of the relevant robotic arm; conversely, if an outside visitor attempts to approach the area of the confidential document cabinet, the security department must be notified immediately and any attempt to operate the device should be prohibited. It is clear that this "automatic triggering" is not a blind response, but rather an intelligent scheduling behavior guided by a detailed configuration strategy.
[0045] Step S107: Record the execution events of the automatic authorization instructions for the equipment and the associated ERP / MES task documents to generate a structured audit log.
[0046] Structured audit logs should include at least the following fields: trigger time timestamp, operator's name and employee ID, affected terminal number, and relevant work order serial number, for future review and statistical analysis. More importantly, this metadata can be further imported into a big data platform for long-term archiving and used in conjunction with BI tools to generate visual reports and charts to assist management in making informed decisions.
[0047] The above implementation fully integrates cutting-edge technologies such as computer vision, artificial intelligence, and industrial IoT, creating a highly integrated, dynamically responsive, and fully traceable security management ecosystem based on the three design concepts of "unmanned operation, seamless access, and seamless integration." Compared to the traditional manual card-swiping and paper-based approval management model, this technical solution significantly reduces the time required for identity authentication, eliminates the risk of fraudulent claims, enhances the emergency response speed for sudden events, and effectively supports the implementation of the strategic goal of intelligent manufacturing transformation and upgrading.
[0048] Reference Figure 2 As one implementation of step S102, the step of preprocessing the face image and outputting a standardized face image includes: Step S201: Receive face images of different scenes captured by a multi-camera network; The system first acquires video frames or still images captured by camera devices deployed in different scenarios, and then extracts the portion containing the face area as the object to be processed. Because industrial environments typically have complex lighting conditions, dust pollution, mechanical vibrations, and other factors, which may lead to unstable or even blurry image quality, this step, although seemingly simple, is actually a fundamental prerequisite for ensuring the effectiveness of all subsequent processing steps.
[0049] To ensure the validity and consistency of image data, practical applications often require the use of front-end hardware configurations (such as high-sensitivity CMOS sensors) and preliminary screening mechanisms (such as quickly determining whether a face region is present based on skin color models or Haar-like features) to avoid invalid data entering subsequent processes and causing resource waste.
[0050] Step S202: Perform grayscale conversion on the face image to generate a grayscale face image; The core purpose of this step is to reduce image dimensionality and simplify computational complexity, while retaining enough visual information for subsequent processing.
[0051] Specifically, color images typically have three color channels (red R, green G, and blue B), while converting them into single-channel grayscale images can significantly reduce memory usage and computational overhead, which is especially beneficial for keypoint detection algorithms that rely on edge response or texture characteristics.
[0052] The standard formula used in this process is Gray = 0.299 × R + 0.587 × G + 0.114 × B, where different weighting coefficients reflect the differences in human eye sensitivity to different wavelengths of light—green contributes the most, followed by red, and blue the least. This non-uniform weighting method can compress the color space while maintaining the consistency of subjective brightness in the image as much as possible, thus providing a more stable and reliable input signal for the next step of keypoint detection.
[0053] Step S203: Detect the location of facial key points in the grayscale face image; This step is a typical computer vision task, which aims to accurately locate several semantically meaningful facial landmarks from a two-dimensional planar image. Typical examples include the corners of the eyes, the tip of the nose, and the corners of the mouth.
[0054] In some embodiments, regression prediction tasks can be performed using pre-trained deep neural network models, which may be based on convolutional neural network architectures (CNNs), cascade regression frameworks, or more advanced Transformer variants. Once a sufficient number of well-distributed keypoint coordinates have been successfully extracted, it means that we have obtained important clues about the current facial pose, expression state, and even some individual differences, which is crucial for subsequent pose correction.
[0055] Step S204: Calculate the rotation correction matrix based on the facial key point positions, perform face alignment processing, and generate a corrected image; In this step, an affine transformation matrix is established using the facial keypoint information obtained in the previous step to eliminate spatial distortion caused by head rotation and tilt. This restores the facial image, which might otherwise exhibit skewness or tilting, to an ideal frontal orientation, allowing for better matching with reference samples in the standard template library.
[0056] Specifically, affine transformation is essentially a linear mapping that preserves the parallelism and proportionality invariance between images. In practice, a set of reference control points (such as a triangle formed by the line connecting the centers of the eyes and the center of the mouth) are often selected as the target reference frame, and the optimal transformation parameters are solved using the least squares method or other optimization strategies. After such geometric normalization, not only is the similarity consistency between images improved, but a good foundation is also laid for subsequent work such as feature extraction and classification.
[0057] Step S205: Perform environmental interference compensation processing on the corrected image to generate an anti-interference image; In particular, considering the common problems in industrial environments such as strong backlighting, local shadow occlusion, and scattering of suspended particles, even after pose normalization, it is still difficult to guarantee that the image quality reaches the ideal level. Therefore, it is especially necessary to introduce a compensation mechanism specifically designed for such noise sources.
[0058] In this embodiment, on the one hand, a contrast-limited adaptive histogram equalization technique is used to address the excessive contrast caused by uneven light source distribution. The basic idea is to redistribute the pixel intensity distribution curve in a local area to make it flatter, thereby enhancing the visibility of details without causing over-sharpening artifacts. On the other hand, the dark channel prior theory is used to carry out dehazing and restoration work. This theory points out that most fog-free outdoor natural images have at least one color channel with a very small value in their neighborhood window. Based on this, the energy function expression describing the haze degradation process can be derived as I(x)=J(x)t(x)+A(1−t(x)), and then the transmittance map t(x) is inverted and combined with the estimated atmospheric illumination value A to reconstruct the image, thereby removing the blurring effect caused by air particles and improving the overall visual clarity.
[0059] Step S206: Scale the anti-interference image to a preset size and output a standardized face image block.
[0060] The task of this step is to standardize image scale specifications so that they conform to the fixed input format requirements of downstream machine learning models, especially convolutional neural networks.
[0061] In this embodiment, a bicubic interpolation algorithm can be used for high-quality resampling. Compared with nearest neighbor interpolation or bilinear interpolation, this method can more effectively suppress jagged edge distortion and maintain good smooth transition characteristics even at high magnification.
[0062] It should be noted that although the image size is adjusted to a fixed 128×128 pixels, the output results have good standardization and rich discriminative features because the previous preprocessing steps have fully considered and mitigated various potential disturbances. They can be directly fed into backbone networks such as ResNet and MobileNet for further identity authentication or behavior analysis tasks.
[0063] In the above embodiments, targeted optimization designs were carried out to address the complex and ever-changing actual shooting conditions in industrial application scenarios, constructing a complete preprocessing pipeline that includes grayscale conversion, key point-driven alignment, dual environmental compensation, and standardized cropping. Compared to traditional solutions that rely solely on simple filtering and scaling, the technical solution of this application deeply integrates two major categories of methods: geometric correction and physical modeling. This results in stronger robustness and generalization ability when facing harsh conditions such as strong light glare and dense smoke, thereby significantly improving the accuracy and timeliness of the entire face recognition system.
[0064] Reference Figure 3 As one implementation of step S103, the steps of constructing a facial feature relationship network based on standardized face images, extracting facial feature vectors, performing liveness detection, and generating liveness detection confidence include: Step S301: Divide the standardized face image into a set of local regions and extract the initial feature vector of each local region; The core idea of this step is to introduce a fine-grained feature representation mechanism, that is, instead of using a single global feature to characterize the entire face, it decomposes it into several sub-regions and models them separately.
[0065] Specifically, the image is divided into an M×N grid (such as 8×8 or 16×16), with each small grid forming an independent local receptive field. A convolutional neural network (CNN) is then used to extract low-level visual features from each local region, such as edges, textures, and color distribution. It's important to note that to preserve the positional context, corresponding positional encoding information is appended to the output of these local features, thereby enhancing local perception capabilities and providing a foundation for establishing cross-regional semantic connections in the next step.
[0066] Step S302: Construct the topological relationships between the initial feature vectors of each local region; During this process, the system calculates the similarity between various local features. One common method is to use the cosine similarity function to measure the angle between two vectors, thus reflecting their directional consistency. When two sets of features are highly matched, it is considered that their corresponding spatial locations have a strong intrinsic connection. Next, an appropriate threshold parameter is set to filter out significantly related node combinations, and an adjacency matrix is constructed accordingly, thus forming an abstract graph structure. This graph is essentially a weighted undirected graph, where vertices represent local feature nodes, and edge weights are determined by the aforementioned similarity. The topological relationships established in this way can effectively capture structural knowledge such as the relative layout of facial features and the trend of facial contours, overcoming the problem of traditional CNNs neglecting long-distance dependencies.
[0067] Step S303: Based on the topological relationship, fuse local region features to generate a global face feature vector; One approach is to introduce a Graph Neural Network (GNN) architecture, which leverages its unique message passing mechanism to achieve state sharing and iterative optimization among neighbors. In other words, GNN allows each node to continuously absorb the latest state information from its surrounding nodes and update its own expression accordingly, eventually reaching a stable equilibrium state.
[0068] This approach is particularly well-suited for processing non-Euclidean structural data (such as social networks, molecular structures, and facial topology maps in this paper), as it better preserves the complex geometric constraints of the original image. After training, the entire map is mapped into a compact vector of fixed length (e.g., 128 dimensions), which comprehensively reflects the overall appearance characteristics of the face and the coordination between its internal parts, making it more discriminative and robust compared to traditional hand-designed features.
[0069] Step S304: Extract physiological dynamic features from the continuous video frame sequence; Unlike traditional liveness verification methods that require explicit behavioral responses such as blinking, shaking the head, and opening the mouth, this application adopts a passive monitoring strategy that focuses on identifying signs of life activities unique to humans.
[0070] Specifically, it mainly includes two aspects: first, tracking the displacement trajectory changes of reflective spots on the iris surface caused by subtle eye movements; and second, analyzing the subtle deformation and fluctuation patterns of facial skin tissue caused by factors such as respiration and heartbeat. The former relies on the principle of optical reflection, while the latter involves spectral analysis techniques. By performing differential operations on multiple consecutive frames of images and applying signal processing tools such as Fast Fourier Transform (FFT), the periodic perturbation components existing within a specific frequency range can be accurately estimated. These physiological characteristics naturally possess unimitable and real-time response characteristics, greatly enhancing the ability to resist various deception attacks.
[0071] Step S305: Combine global facial feature vectors with physiological dynamic features to make predictions and generate liveness detection confidence scores.
[0072] In this process, two types of heterogeneous features (static appearance + dynamic physiology) are first concatenated into a joint feature vector, which is then fed into a fully connected classifier for binary classification prediction—that is, determining whether the target belongs to a real human individual. The output is a real probability value in the range [0,1], with values closer to 1 indicating a higher likelihood of a live face sample. This confidence assessment mechanism not only provides clear result labels but also feeds back more credibility level reference indicators to the access control system, facilitating the development of flexible security policy configurations.
[0073] In the above implementation, a topological graph model between local facial regions is constructed, and a graph neural network propagation mechanism is used to extract deep semantic associations, which effectively solves the problem of false detection under complex pose changes and occlusion. At the same time, it is supplemented by a non-invasive physiological feature acquisition path, which truly realizes the ideal application scenario requirements of non-interactive and all-weather operation.
[0074] Reference Figure 4 As one implementation of step S105, the step of comparing the identity features with the pre-stored employee feature database in the ERP / MES system in real time, and generating an permission binding tag by binding the corresponding employee identity identifier and real-time permission data according to the comparison result includes: Step S401: Access the pre-stored employee feature database stored in the ERP / MES system; wherein, the pre-stored employee feature database contains employee identification and associated basic feature data; Specifically, the pre-stored employee feature database is essentially a structured database containing each employee's identity identifier (such as employee number, name, etc.) and the basic feature data associated with them.
[0075] It should be noted that these basic feature data can be standard facial feature vectors collected when employees register in the system, or average feature vectors generated by aggregating multiple collections to improve recognition stability. By accessing this feature library, the system obtains a target set that can be used for comparison, providing data support for the next step of similarity calculation.
[0076] Step S402: Calculate the similarity value between the identity features and each basic feature data in the pre-stored employee feature database in real time; Specifically, this step uses mathematical methods to quantify the spatial distance or directional relationship between two vectors, typically employing the cosine similarity algorithm. Cosine similarity measures the degree of similarity between two vectors by calculating the cosine of the angle between them, with a value ranging from [-1, 1]. The closer the value is to 1, the more similar the two vectors are.
[0077] In this embodiment, the system compares the currently collected identity feature vector with the basic feature vectors of all employees in the feature database one by one to obtain a set of similarity values. This calculation method is not only highly efficient, but also insensitive to the length of the feature vector, and can effectively cope with feature differences caused by different collection conditions.
[0078] Step S403: Based on a preset similarity threshold, determine the matching employee identity identifier according to the similarity value; At this stage, the system sets a reasonable similarity threshold. This threshold is usually an empirical value set based on training data or experience, but can also be dynamically adjusted through an adaptive algorithm. When the highest similarity value exceeds this threshold, the current identity feature vector is considered to have successfully matched the basic feature vector of an employee, thus confirming the employee's identity. If all similarity values are below the threshold, an anomaly handling mechanism may be triggered, such as requiring manual review or re-collecting images. This step ensures the accuracy and security of identity recognition, preventing incorrect permission allocation due to misidentification.
[0079] Step S404: Query the ERP / MES system to obtain the real-time permission data corresponding to the employee's identity identifier; Real-time access data includes job changes, task assignments, and device access permissions; Specifically, traditional access control systems often rely on fixed permission configuration tables, making it difficult to handle changes in employee roles and task switching. This solution, however, uses an API interface to call the real-time data service of the ERP / MES system to obtain the latest permission information related to the employee's identity. This permission data includes not only job responsibilities and equipment operation permissions, but may also cover dynamic attributes such as current task assignment status and work area restrictions. This real-time query mechanism gives access control high flexibility and responsiveness, automatically adjusting permission configurations as business processes change, effectively avoiding security risks caused by outdated permissions.
[0080] Step S405: Bind the employee's identity identifier with the real-time permission data to generate a permission binding tag.
[0081] The permission binding tag includes the employee's identity identifier and its corresponding real-time permission data; Specifically, this step packages the identity identifier and permission data into a structured permission binding tag. The tag can also include a timestamp and session identifier for subsequent auditing and permission revoke processes. The generated permission binding tag is then transmitted to the permission decision engine, which determines whether to grant the employee the corresponding system access permissions. This approach not only improves the automation level of permission allocation but also enhances the traceability and controllability of the system, meeting the security and compliance requirements of permission management in modern smart manufacturing environments.
[0082] In this embodiment, permission data is typically dynamically generated by upper-level business systems such as Enterprise Resource Planning (ERP), Manufacturing Execution System (MES), or Human Resource Management System (HRMS), and parsed and encapsulated by a permission decision engine. The purpose of this tag is to abstract complex permission logic into a transmissible and parsable data structure, enabling downstream modules to quickly obtain and utilize the permission context information. For example, when an employee enters a specific area, the system can quickly identify whether they have permission to operate devices within that area using this tag, without needing to call multiple business interfaces for permission verification again. This design is particularly important in an Industrial Internet of Things (IIoT) environment because it significantly reduces the latency of permission queries and improves the system's response speed and operational efficiency.
[0083] In the above implementation, facial recognition technology is deeply integrated with the ERP / MES system to build a fully automated access control system. This system not only achieves real-time mapping between identity and permissions but also possesses excellent scalability and security. It can effectively address the access control needs in complex and ever-changing production environments, significantly improving the intelligence level and response speed of enterprise information systems in terms of access control.
[0084] In this embodiment, employee permissions can be dynamically adjusted based on task allocation and job changes in the ERP / MES system. For example, employee A can work at both the polishing and assembly stations. When employee A arrives at the polishing station, the monitoring system automatically identifies him as the person on duty and can automatically switch his system permission identity. If he is not the person on duty, no operation is performed. When employee A arrives at the assembly station, the system can also automatically identify his system permission identity, meaning the backend system pre-assigns two different identity permissions.
[0085] Reference Figure 5 As one implementation of step S106, the step of automatically triggering the corresponding device automatic authorization instruction based on the permission binding tag matching the current scenario includes: Step S501: Obtain the physical location identifier and device type identifier of the current scene; The purpose of this step is to provide spatial and device-level contextual support for subsequent permission matching. Specifically, physical location identifiers are typically determined through location codes collected by cameras or other sensors deployed in a specific area. For example, the deployment coordinates of a camera can be obtained through GPS, indoor positioning systems (such as Bluetooth beacons, UWB positioning), or image recognition-based spatial mapping techniques, thereby generating a precise location identifier.
[0086] The equipment type identification is achieved by parsing the equipment code (such as a QR code, barcode, or RFID tag) captured by the equipment's external camera, thereby identifying the type information of the current equipment. For example, the equipment code may correspond to a equipment type table in a database, thus mapping "DVC-001" to "polishing machine" and "DVC-002" to "assembly robotic arm," etc.
[0087] Understandably, this dual-resolution mechanism ensures that the system can accurately identify devices and locations in the current scene, thereby providing a reliable data foundation for permission matching. This step reflects the high importance that this application embodiment places on "scene awareness," that is, achieving dynamic identification and modeling of the operating environment through the deep integration of the physical world and digital information.
[0088] Step S502: Match the real-time permission data in the permission binding tag with the device type identifier of the current scene; Specifically, the system extracts a list of device types that the employee can currently operate from the permission data (i.e., the operation whitelist), and then determines whether the device type in the current scenario is within the whitelist. For example, if employee A's permission data only contains operation permissions for "polishing machine" and "inspection instrument", and the device type in the current scenario is "welding robot", the match will fail; conversely, if the device type is "polishing machine", the match will succeed.
[0089] This matching mechanism not only enables dynamic verification of permissions but also ensures the principle of least privilege in access control, meaning that employees can only perform operations on authorized devices, thus effectively preventing abuse of permissions. Furthermore, this matching logic can be extended to include a time dimension, such as introducing a "time window" mechanism, which allows employees to have operating permissions only during specific time periods, further enhancing the flexibility and security of access control.
[0090] Step S503: When the matching is successful, generate the corresponding automatic device authorization command; The system first selects a suitable template from a pre-set instruction template library based on the device type identifier. For example, for an access control system, the template might include the instruction to "open the electromagnetic lock"; for a workstation terminal, the template might be the instruction to "automatically log in to the system".
[0091] Subsequently, the system injects the employee's identity and the current scenario's timestamp into the instruction template, generating an authorization instruction with contextual information. The introduction of the timestamp helps prevent instruction replay attacks, enhancing the timeliness and security of the instruction. Finally, the instruction is encapsulated using encryption algorithms (such as AES and RSA) to form a secure data packet, ensuring it cannot be tampered with or stolen during transmission. This design not only improves system security but also makes the authorization instruction traceable, facilitating subsequent auditing and log management.
[0092] Step S504: Send the device automatic authorization instruction to the target device execution terminal.
[0093] This step relies on Industrial Internet of Things (IIoT) communication protocols such as OPC UA, Modbus TCP, and MQTT. These protocols offer high reliability, low latency, and cross-platform compatibility, making them particularly suitable for complex industrial control environments. Through these protocols, authorization commands are accurately transmitted to the target device controller, triggering corresponding access control actions, such as opening access gates, unlocking user interfaces, and starting equipment. Simultaneously, the system records the execution status of the commands and feeds the results back to the log module for subsequent auditing, anomaly analysis, and system optimization. This closed-loop feedback mechanism not only enhances system controllability but also provides crucial data support for maintenance personnel.
[0094] The above implementation not only realizes the scenario-based, dynamic, and secure management of permissions, but also enables direct execution of permission commands through the Industrial Internet of Things protocol, significantly improving the system's response efficiency and ease of operation. It breaks through the technical bottlenecks of traditional permission control systems, such as the need for manual intervention, coarse permissions, and delayed response, and truly realizes a "zero-interaction, high-security, and fine-grained" intelligent device authorization mechanism, which has broad application prospects and significant technological progress.
[0095] Reference Figure 6 As a further implementation of the zero-interaction permission recognition method, when the liveness detection confidence exceeds a preset confidence threshold, a multimodal verification step is initiated, specifically including: Step S601: Trigger the infrared camera to collect thermal imaging data and generate a temperature distribution matrix; The infrared camera constructs a two-dimensional grayscale or pseudo-color image, or thermal image, reflecting the temperature distribution based on the differences in infrared energy emitted or reflected by the object itself. By performing pixel-level analysis and numerical mapping on this image, the system can extract the temperature value corresponding to each pixel and then organize it into a temperature distribution matrix with spatial dimensions.
[0096] This matrix not only contains information on temperature gradient changes in local areas, but also implicitly contains the unique thermal conduction characteristics and physiological metabolic activity features of human tissues. For example, in the facial region, the tip of the nose, lips, and the area around the eyes often exhibit different thermal response patterns, which constitute an important basis for subsequent liveness detection.
[0097] Step S602: Detect the live biological feature patterns in the temperature distribution matrix; Because living tissues possess specific blood circulation mechanisms and metabolic rates, they exhibit dynamic temperature change behavior and steady-state heat distribution patterns in the infrared band that differ from those of non-living materials (such as photographs and masks). Therefore, the system employs a method combining statistics and image processing to locate and encode feature vectors key regions (such as facial contours, blood vessel patterns, and respiratory hotspots) within the temperature distribution matrix.
[0098] For example, facial boundaries can be identified using edge detection algorithms, and then core hot zones can be extracted using Gaussian filtering and region growing methods. Principal component analysis (PCA), linear discriminant analysis (LDA), or deep neural network models can be further used to abstract and model these thermal features, ultimately forming a set of discriminative feature patterns that can effectively distinguish between live and fake samples. This process essentially transforms physical thermodynamic phenomena into a data representation that can be used as input to machine learning classifiers.
[0099] Step S603: Compare the live biological feature patterns with the pre-existing live biological spectral library and output the auxiliary confidence level; The pre-vivor spectral library refers to a set of typical infrared thermal imaging datasets of real human bodies under different environmental conditions, collected and labeled in advance during the system training phase. After normalization, dimensionality reduction, and clustering, these data form a representative reference template database.
[0100] After the thermal features of the sample to be tested are extracted, the system uses similarity measurement methods (such as Euclidean distance, cosine similarity, Mahalanobis distance, etc.) to calculate the degree of matching between it and the templates of each category in the library. The matching results are presented in the form of a probability distribution or membership function, where the highest score item corresponds to the most likely true liveness category, and this score value is the "auxiliary confidence score". This confidence score reflects the degree to which the current sample conforms to the known liveness features and serves as a supplementary judgment criterion to the output results of the original visible light liveness detection module.
[0101] Step S604: Combine the auxiliary confidence score and the liveness detection confidence score to generate a composite verification result.
[0102] Understandably, in facial recognition systems, liveness detection using a single modality (such as an RGB image) is easily affected by factors such as changes in lighting and material camouflage, leading to an increased false positive rate.
[0103] To address this, the system employs weighted fusion, Bayesian inference, or DS evidence theory to jointly evaluate the original liveness confidence score from the visible light channel and the auxiliary confidence score provided by the infrared channel. The fusion process considers not only the absolute values of the two confidence scores but also incorporates historical data feedback, environmental context information, and model uncertainty factors to optimize the decision boundary and improve the overall system's robustness and generalization ability. For example, if the confidence score in the visible light channel decreases due to strong light exposure, but the infrared channel maintains a high auxiliary confidence score, the fusion result can still maintain a positive determination that the target is alive, thereby effectively reducing the risk of missed detections and false alarms.
[0104] The above implementation method effectively enhances the traditional visible light liveness detection mechanism. This mechanism relies on the unique thermodynamic properties of living tissue and intelligently integrates with the original detection results, which not only improves the recognition accuracy but also enhances the system's ability to resist various deception attacks. It is especially suitable for application scenarios with extremely high security requirements, such as industrial control and security access control.
[0105] Reference Figure 7 As a further implementation of the zero-interaction permission identification method, after the step of generating structured audit logs, the method further includes: Step S701: Parse the ERP / MES task document operation records in the structured audit log; Specifically, task documents generated by ERP (Enterprise Resource Planning) or MES (Manufacturing Execution System) typically contain key fields such as clear work instruction numbers, operator IDs, target equipment identifiers, and process node descriptions. By structurally parsing these fields, the complete business intent and expected behavioral path behind each authorized operation can be reconstructed. This parsing process relies on an understanding of communication protocols between heterogeneous systems, such as XML / JSON data encapsulation methods, API call specifications, and database table relationship modeling. Based on this, a transaction-driven behavior mapping framework can be established, providing initial anchor points for subsequent behavior trajectory tracking.
[0106] Step S702: Retrieve the camera video stream data with the corresponding timestamp based on the ERP / MES task document operation record, and extract the operation behavior trajectory; Since the timestamps of the aforementioned task orders are accurate to the second or even millisecond level, they can be used to locate high-definition surveillance video clips within a specific time period and reconstruct the operator's actual action sequence from them using computer vision technology.
[0107] In the embodiments of this application, the process of behavior trajectory extraction not only includes basic human pose estimation (such as the application of open-source algorithms like OpenPose), but also involves higher-level action semantic segmentation and spatiotemporal feature encoding. For example, a 3D convolutional neural network (CNN) combined with a long short-term memory network (LSTM) is used to jointly train continuous frame images to obtain a more stable behavior vector representation; or the YOLO series object detectors are used in conjunction with the SORT / Kalman filtering algorithm to complete individual tracking and separation in multi-person scenes.
[0108] In addition, challenges such as changes in lighting and occlusion interference must be considered in industrial environments. Therefore, it is often necessary to introduce adaptive background modeling methods or cross-camera collaborative positioning strategies to improve trajectory reconstruction accuracy.
[0109] Step S703: Match the operation behavior trajectory with the standard operation process template of the task document to obtain the matching deviation. Standard operating procedure templates are typically a set of standardized actions pre-defined based on industry safety regulations, process manuals, and expert experience. They abstractly express the logical flow rules between various processes using finite state machines (FSMs) or Petri nets. The current operational behavior trajectory, on the other hand, is a time-ordered list of specific action instances derived from preceding video analysis. There is a semantic hierarchy between the two: the former is a conceptual, idealized model, while the latter is a real-world projection.
[0110] To bridge this gap, various similarity metrics can be used for quantitative evaluation, such as edit distance, longest common subsequence (LCS), or hidden Markov models (HMM). Edit distance is suitable for measuring the minimum transformation cost between two strings; LCS can be used to discover co-evolutionary trends; and HMM can capture more complex temporal dependencies. Furthermore, attention mechanisms can be introduced to enhance the model's focus on key operational nodes, making the final matching results more discriminative.
[0111] Step S704: When the matching deviation exceeds the limit, generate a permission exception event and freeze the device automatic authorization command.
[0112] The matching deviation is a comprehensive indicator that may encompass information aggregation results from multiple dimensions, including the proportion of missed actions, the number of unauthorized interventions, and the frequency of unauthorized access. Once it exceeds the system's set threshold, it indicates that the current operation is highly likely to violate established procedures and poses a potential security risk. At this point, the system will automatically trigger an alarm mechanism, generate a permission anomaly event log with a severity level marker, and simultaneously notify relevant administrators to intervene and investigate.
[0113] More importantly, this mechanism also features an immediate authorization circuit breaker function, suspending the issuance of any new authorization commands for the affected device until manually verified. This effectively prevents the spread of cascading failures caused by misoperation or malicious tampering, improving the overall system's fault tolerance and emergency response efficiency.
[0114] The above implementation cleverly integrates the log data analysis capabilities within the information system with visual perception methods from the external physical world, achieving two-way verification and deviation warning throughout the entire process from business instruction issuance to on-site operation execution. This technical solution not only enhances the post-event traceability depth of the original access control system but also endows it with forward-looking risk control and adjustment functions, making the access control system in the entire intelligent manufacturing environment more intelligent, refined, and reliable.
[0115] Reference Figure 8 As a further implementation of the zero-interaction permission identification method, a permission conflict detection step is added before the execution of the device's automatic authorization command. Specifically, this includes: Step S801: Retrieve the real-time status identifier and the current occupant's identity identifier of the target device; Among them, the real-time status identifier of the equipment is usually a logical tag maintained by the Manufacturing Execution System (MES) or Enterprise Resource Planning System (ERP) to indicate whether a specific piece of equipment is idle, running, faulty, or in other usage states; while the "current occupant identity identifier" refers to the unique digital credentials of the user operating the equipment, such as employee number, role classification information, etc.
[0116] These two parameters together form the key basis for determining whether there are potential operational conflicts. Based on this, the system can obtain information about who is currently using the device and what its operating status is, providing basic support for subsequent permission assessment.
[0117] Step S802: When the target device is in an occupied state, compare the current occupant with the permission priority parameter of the current requester corresponding to the device automatic authorization instruction; The permission priority parameter is not a simple hierarchical classification, but a quantitative indicator that comprehensively considers multiple factors, including but not limited to job responsibilities, security certification level, and historical operation records. These parameters can be pre-stored in the permission database and dynamically updated based on organizational structure changes, personnel promotions, and other factors.
[0118] Understandably, by comparing and weighting the two sets of priority parameters, the system can determine which party has higher priority. If the current requester has a higher priority than the existing user, they are allowed to take over control of the device; otherwise, they are placed in a waiting queue. This process embodies the rule-driven automated decision-making approach, which helps improve response efficiency on the production floor and reduce the uncertainty risks caused by human intervention.
[0119] Step S803: Generate a conflict resolution strategy based on the permission priority parameter, including an instruction queuing sequence or preemptive execution instruction; Instruction queuing involves sequentially arranging multiple pending device operation requests into an ordered list according to a certain sorting principle, so that they can be responded to one by one at the appropriate time. "Preemptive execution of instructions," on the other hand, means that a high-priority user can directly interrupt the execution of a low-priority task, thereby gaining immediate control of the device. The choice between these two strategies depends on the specific business scenario requirements and the enterprise's security management policies.
[0120] For example, in situations where emergency repair tasks need to be inserted into regular production line operations, equipment control can be quickly switched through preemption. In non-emergency situations, a fair queuing system is used to ensure that the rights of all legitimate users are not infringed. In addition, to prevent frequent switching from causing excessive system load, a minimum holding time threshold can be set as a buffer condition, so that only tasks that meet the continuous usage duration requirements can be preempted.
[0121] Step S804: Write the conflict resolution strategy into an additional field of the structured audit log.
[0122] Specifically, regardless of whether a queuing or preemptive measure is ultimately adopted, the decision-making basis and relevant contextual information should be fully preserved for future reference and analysis. This supplementary log content not only includes basic elements such as the identity information of the participating parties, application time points, and final rulings, but may also include relevant corporate management system clauses and approval opinions from relevant responsible persons on which this scheduling was based.
[0123] In this embodiment, the structured audit log is not an isolated static collection of documents. In fact, it has a close linkage with other subsystems such as ERP / MES. For example, it can trigger status change notifications of corresponding modules or reassign work task list items through reverse interface calls, thereby promoting the evolution of the overall operation process towards a more efficient and collaborative direction.
[0124] In the above embodiments, starting from the monitoring of underlying equipment status and combining it with a refined permission priority modeling method, a flexible and efficient instruction scheduling function is achieved, and the risk control level of the entire process is strengthened by a comprehensive log recording system. Compared with the traditional static authorization management mode, this application significantly improves the security adaptability and operational management flexibility in the face of complex and ever-changing industrial environments, and effectively avoids the probability of work delays or even safety accidents caused by multiple people competing for critical equipment.
[0125] This application integrates technologies such as complex environment adaptability, multimodal data fusion, real-time communication, and dynamic permission adjustment to achieve a comprehensive upgrade of permission management for manufacturing enterprises, yielding significant benefits. Firstly, this application achieves true zero-interaction permission recognition. Through AI image recognition technology, it automatically binds system users and automatically assigns them corresponding permissions upon successful authentication. This effectively reduces data errors in system documents caused by employees failing to switch identities in a timely manner, significantly improving work efficiency and production continuity, and reducing increased labor costs due to human error.
[0126] Secondly, through deep integration of AI facial recognition technology with the ERP / MES system, the accuracy and real-time nature of access control are achieved, ensuring the consistency of enterprise data and the timeliness of access control updates. Thirdly, this application employs a multi-camera network with full coverage, effectively eliminating blind spots in recognition. Combined with liveness detection, it effectively prevents security risks such as photo or video spoofing, significantly improving the system's reliability and security.
[0127] Furthermore, by incorporating variables related to aerosols and illumination to optimize the facial recognition algorithm, the system's recognition accuracy in complex industrial environments has been significantly improved, enhancing its adaptability and robustness under harsh conditions such as dust and water vapor. Simultaneously, based on dynamic permission adjustment technology, the system can automatically adjust employee permissions according to task assignments and job changes in ERP / MES, achieving intelligent and automated permission management. Finally, the system provides detailed operation log recording capabilities, offering strong support for comprehensive auditing and compliance management, ensuring traceability and standardized management of the production process.
[0128] This application also discloses a zero-interaction permission recognition system based on the integration of face recognition and ERP / MES.
[0129] A zero-interaction access control system based on facial recognition and ERP / MES integration includes: The image acquisition module is used to acquire facial images in real time through a network of multiple cameras deployed in different scenarios; The image preprocessing module is used to preprocess face images and output standardized face images; The liveness detection module is used to construct a facial feature relationship network based on standardized face images, extract facial feature vectors, perform liveness detection, and generate liveness detection confidence scores. The identity feature output module is used to output identity features with liveness detection confidence when the confidence level of liveness detection exceeds a preset confidence threshold. The permission binding module is used to compare identity features with the pre-stored employee feature database in the ERP / MES system in real time, and bind the corresponding employee identity identifier and real-time permission data according to the comparison results to generate permission binding tags. The device authorization module is used to match the current scenario based on the permission binding tags and automatically trigger the corresponding device automatic authorization command; The log generation module is used to record the execution events of automatic authorization commands for devices and the associated ERP / MES task documents, generating structured audit logs.
[0130] The zero-interaction permission recognition system of this application embodiment can implement any of the above-described zero-interaction permission recognition methods, and the specific working process of each module in the zero-interaction permission recognition system can be referred to the corresponding process in the above-described method embodiments.
[0131] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0132] This application also discloses a computer-readable storage medium.
[0133] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the zero-interaction permission recognition methods based on face recognition and ERP / MES integration.
[0134] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0135] In this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0136] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.
[0137] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A zero-interaction permission recognition method based on face recognition and ERP / MES integration, characterized in that, The identification method includes: Real-time facial images are captured through a multi-camera network deployed in different scenarios; The face image is preprocessed to output a standardized face image; A facial feature relationship network is constructed based on the standardized face images, facial feature vectors are extracted, liveness detection is performed, and liveness detection confidence is generated. When the confidence level of the liveness detection exceeds a preset confidence threshold, the identity feature with the liveness detection confidence level is output. The identity features are compared in real time with the pre-stored employee feature database in the ERP / MES system. Based on the comparison results, the corresponding employee identity identifier and real-time permission data are bound together to generate permission binding tags. Based on the permission binding tag, match the current scenario and automatically trigger the corresponding device automatic authorization command; Record the execution events of automatic authorization commands for equipment and the associated ERP / MES task documents to generate structured audit logs.
2. The zero-interaction permission recognition method based on face recognition and ERP / MES integration according to claim 1, characterized in that, The steps for preprocessing the face image to output a standardized face image include: Receives facial images from different scenes captured by a multi-camera network; Perform grayscale conversion on the face image to generate a grayscale face image; Detect the locations of facial key points in the grayscale face image; Based on the positions of the facial key points, a rotation correction matrix is calculated, facial alignment is performed, and a corrected image is generated. Environmental interference compensation processing is performed on the corrected image to generate an anti-interference image; The anti-interference image is scaled to a preset size, and a standardized face image block is output.
3. The zero-interaction permission recognition method based on face recognition and ERP / MES integration according to claim 2, characterized in that, The steps of constructing a facial feature relationship network based on the standardized face image, extracting facial feature vectors, performing liveness detection, and generating liveness detection confidence scores include: The standardized face image is divided into a set of local regions and the initial feature vector of each local region is extracted; Construct the topological relationships between the initial feature vectors of each local region; Based on the aforementioned topological relationships, local region features are fused to generate a global face feature vector; Extracting physiological dynamic features from a continuous sequence of video frames; The global facial feature vector and physiological dynamic features are combined to make predictions and generate a liveness detection confidence score.
4. The zero-interaction permission recognition method based on face recognition and ERP / MES integration according to claim 1, characterized in that, The steps of comparing the identity features with the pre-stored employee feature database in the ERP / MES system in real time, and generating permission binding tags by binding the corresponding employee identity identifier and real-time permission data based on the comparison results include: Access the pre-stored employee feature database stored in the ERP / MES system; wherein, the pre-stored employee feature database contains employee identification and associated basic feature data; Calculate the similarity value between the identity features and each basic feature data in the pre-stored employee feature database in real time; Based on a preset similarity threshold, the matching employee identity identifier is determined according to the similarity value; Query the ERP / MES system to obtain the real-time permission data corresponding to the employee's identity identifier; The employee identification is bound to the real-time permission data to generate permission binding tags.
5. The zero-interaction permission recognition method based on face recognition and ERP / MES integration according to claim 4, characterized in that, The steps for automatically triggering the corresponding device automatic authorization command based on the permission binding tag matching the current scenario include: Obtain the physical location identifier and device type identifier of the current scene; Match the real-time permission data in the permission binding tag with the device type identifier of the current scene; When a match is successful, the corresponding automatic device authorization command is generated; The device automatic authorization command is sent to the target device execution terminal.
6. The zero-interaction permission recognition method based on face recognition and ERP / MES integration according to claim 3, characterized in that, When the confidence level of the liveness detection exceeds a preset confidence threshold, a multimodal verification step is initiated, which specifically includes: Trigger the infrared camera to collect thermal imaging data and generate a temperature distribution matrix; Detecting live biological feature patterns in the temperature distribution matrix; The live biological feature patterns are compared with a pre-existing live biological spectral library to output an auxiliary confidence level. The auxiliary confidence score and the liveness detection confidence score are combined to generate a composite verification result.
7. The zero-interaction permission recognition method based on face recognition and ERP / MES integration according to claim 1, characterized in that, Following the step of generating structured audit logs, the following also includes: Parse the ERP / MES task document operation records in the structured audit log; Based on the ERP / MES task document operation record, retrieve the camera video stream data with the corresponding timestamp and extract the operation behavior trajectory; The operation behavior trajectory is matched with the standard operation process template of the task document to obtain the matching deviation. When the matching deviation exceeds the limit, an permission exception event is generated and the automatic authorization command of the device is frozen.
8. A zero-interaction permission recognition method based on face recognition and ERP / MES integration according to any one of claims 1 to 7, characterized in that, Before the step of executing the device automatic authorization instruction, a permission conflict detection step is added, specifically including: Retrieve the real-time status identifier and current occupant identifier of the target device; When the target device is occupied, compare the permission priority parameter of the current occupant with the current requester corresponding to the device's automatic authorization instruction. A conflict resolution strategy is generated based on the permission priority parameter, including an instruction queuing sequence or preemptive execution of instructions; Write the conflict resolution strategy into an additional field of the structured audit log.
9. A zero-interaction permission recognition system based on facial recognition and ERP / MES integration, characterized in that, The identification system includes: The image acquisition module is used to acquire facial images in real time through a network of multiple cameras deployed in different scenarios; The image preprocessing module is used to preprocess the face image and output a standardized face image; The liveness detection module is used to construct a facial feature relationship network based on the standardized face image, extract facial feature vectors, perform liveness detection, and generate liveness detection confidence. The identity feature output module is used to output identity features with liveness detection confidence when the liveness detection confidence exceeds a preset confidence threshold. The permission binding module is used to compare the identity features with the pre-stored employee feature database in the ERP / MES system in real time, bind the corresponding employee identity identifier and real-time permission data according to the comparison results, and generate permission binding tags. The device authorization module is used to match the current scenario based on the permission binding tag and automatically trigger the corresponding device automatic authorization instruction; The log generation module is used to record the execution events of automatic authorization commands for devices and the associated ERP / MES task documents, generating structured audit logs.
10. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 8.