A real-time risk management and control method and system based on edge computing
By implementing an NPU-based real-time risk management method on edge computing devices, combined with SCRFD and ArcFace algorithms, high-precision, low-latency risk identification of bank self-service terminal devices is achieved, solving the problems of real-time performance and high cost in existing technologies, and providing a stable and reliable risk management solution.
Patent Information
- Application Number
- CN202511293784.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-11
AI Technical Summary
In existing technologies, the risk control system of bank self-service terminal equipment has problems such as poor real-time performance, high deployment cost, dependence on network, and difficulty in balancing performance and accuracy at the edge. In particular, the inference speed of high-precision face recognition models on edge computing devices is insufficient, which cannot meet the real-time requirements.
A real-time risk management method based on edge computing is adopted. By using an edge computing device equipped with a neural network processing unit (NPU), face detection and quality assessment are performed by capturing video frame sequences in real time. Face tracking and feature extraction are performed by combining SCRFD and ArcFace algorithm models. A dual screening mechanism is designed to achieve high-precision and low-latency risk decision-making.
It achieves high-precision, low-latency risk management at the edge, reduces operating costs, improves system stability and data security, can respond to risk events instantly, avoids false alarms, and ensures robustness and accuracy in complex environments.
Smart Images

Figure CN120805118B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and specifically to a real-time risk management method and system based on edge computing. Background Technology
[0002] Currently, banks, human resources and social security departments, and other government and enterprise institutions have widely deployed self-service terminals and remote video teller systems to provide efficient services to customers and citizens, supplemented by on-site guides and online customer service personnel. However, due to the inherent limitations of the equipment's interaction logic, coupled with the high turnover rate of on-site guides and their lack of experience in risk identification, there are risks such as unauthorized personnel replacement, unauthorized personnel entering the premises, and authorized personnel leaving their posts without authorization. Traditional monitoring systems mostly rely on manual review of recordings after the fact, but this manual review is extremely labor-intensive, cannot provide real-time alerts, and is inefficient.
[0003] Existing automation solutions often employ a "cloud-terminal" architecture, uploading video streams to cloud servers for analysis. This architecture has the following drawbacks:
[0004] 1. High latency: There is network latency in uploading video data and distributing analysis results, which cannot meet the real-time requirements.
[0005] 2. High bandwidth cost: Continuous transmission of high-definition video streams consumes a large amount of network bandwidth, resulting in high operating costs.
[0006] 3. Data privacy risks: Sensitive biometric information such as facial recognition data is transmitted over the public internet, posing a risk of leakage.
[0007] 4. Poor reliability: Once the network is interrupted, the entire risk management system will immediately fail.
[0008] While edge computing devices can address the aforementioned issues, conventional edge computing devices have limited computing power, making it difficult to efficiently run high-precision deep learning models locally. This results in low risk identification accuracy and high false positive rates. Therefore, there is an urgent need for a technical solution that can achieve high-precision, low-latency real-time risk management on miniaturized, low-power edge computing devices. Summary of the Invention
[0009] The purpose of this application is to provide a real-time risk management method and system based on edge computing, which can achieve high-precision, low-latency real-time risk management.
[0010] Firstly, this application provides a real-time risk management method based on edge computing, applied to an edge computing device equipped with a neural network processing unit (NPU), comprising:
[0011] Real-time capture of video frame sequences;
[0012] Face detection is performed on real-time captured video frames using an NPU to obtain face bounding boxes.
[0013] The quality of the detected face bounding boxes is evaluated, and a comprehensive quality score is calculated. ; retain the overall quality score Face bounding boxes exceeding a preset quality threshold; subsequent processing only applies to the retained faces, i.e., only to the comprehensive quality score. Face bounding boxes exceeding a preset quality threshold are tracked and their features extracted for subsequent risk decisions.
[0014] Based on the NPU, faces within the preserved face bounding boxes are tracked and their features extracted (recognized) to obtain the number of faces and the actual face feature vectors. Compared with the standard feature vectors of pre-stored authorized personnel The relationship between faces and the temporal information of face appearance;
[0015] Based on the number of faces and actual facial feature vectors Compared with the standard feature vectors of pre-stored authorized personnel Based on the relationship between the faces and the timing information of their appearance, we can determine whether there are pre-set risk scenarios.
[0016] An alert is triggered when the assessment indicates that a risk exists.
[0017] In one possible implementation, the step of performing face detection on real-time captured video frames based on the NPU to obtain face bounding boxes includes: performing face detection on real-time captured video frames based on the NPU, and outputting the detected face bounding boxes and their corresponding confidence scores. The confidence level This refers to the probability value that the face bounding box corresponds to a real face; after outputting the face bounding box, it is calculated based on the set confidence threshold. Filter the test results: retain only Face bounding box, discard The bounding box of the face.
[0018] In one possible implementation, the quality of the detected face bounding boxes is evaluated, and a comprehensive quality score is calculated. The process includes: first, calculating the size score, sharpness score, and pose score of the face bounding box separately; then, weighting the size score, sharpness score, and pose score to obtain the overall quality score. .
[0019] In one possible implementation, faces are detected using a face detection model; the face detection model employs the SCRFD algorithm.
[0020] Extract actual facial feature vectors using a facial feature extraction model. The facial feature extraction model uses the ArcFace algorithm.
[0021] Face tracking is performed using a centroid tracking algorithm.
[0022] In one possible implementation, the edge computing device is a Rockchip device, and both the SCRFD algorithm model and the ArcFace algorithm model are converted and quantized from ONNX format to RKNN format to adapt for accelerated inference on the NPU.
[0023] To convert ONNX format to RKNN format, you can use the RKNN Toolkit. The RKNN Toolkit is a tool provided by Rockchip for converting deep learning models to RKNN format for deployment and operation on Rockchip devices.
[0024] In one possible implementation, the determination of whether a preset risk scenario exists includes at least one of the following:
[0025] Determining the risk of replacement: Based on actual facial feature vectors Compared with the standard feature vectors of pre-stored authorized personnel Calculate the actual facial feature vector Compared with the standard feature vectors of pre-stored authorized personnel Cosine similarity between If cosine similarity If the similarity is less than the preset similarity threshold, it is determined that there is a risk of replacement.
[0026] This risk scenario indicates that after the authorized person's face leaves, an unauthorized person's face appears with a similarity to the authorized person's face feature that is lower than a preset similarity threshold;
[0027] Determine if there is a risk of multiple people appearing in the same frame: Based on the number of faces tracked, if the number of faces detected in the current video frame is greater than 1, it is determined that there is a risk of multiple people appearing in the same frame.
[0028] Determine if there is a risk of leaving midway: Based on the timing information of the appearance of the face, maintain a timer to record the duration during which the authorized personnel are not detected, that is, the time when the authorized personnel's face leaves the site. If the duration exceeds the preset departure time threshold, it is determined that there is a risk of leaving midway.
[0029] Secondly, this application provides a real-time risk management system based on edge computing. This system is deployed on an edge computing device equipped with a neural network processing unit (NPU), and includes:
[0030] The image acquisition module is used to capture video frame sequences in real time.
[0031] The face detection unit is used to perform face detection on real-time captured video frames based on the NPU to obtain face bounding boxes;
[0032] The face quality assessment unit is used to evaluate the quality of detected face bounding boxes and calculate a comprehensive quality score. ; retain the overall quality score Face bounding boxes that exceed a preset quality threshold;
[0033] The face tracking and feature extraction unit is used to track and extract features from faces within the preserved face bounding boxes based on the NPU, obtaining the number of faces and the actual face feature vectors. Compared with the standard feature vectors of pre-stored authorized personnel The relationship between faces and the temporal information of face appearance;
[0034] Risk decision-making unit, used to determine the number of faces and actual facial feature vectors. Compared with the standard feature vectors of pre-stored authorized personnel Based on the relationship between faces and the timing of their appearance, we can determine whether there are any pre-set risk scenarios, and thus make risk decisions.
[0035] The early warning output module is used to trigger an early warning when the risk decision-making unit determines that there is a risk.
[0036] Among them, the face detection unit, face quality assessment unit, and face tracking and feature extraction unit are the core processing units of the system.
[0037] The risk scenarios identified by the risk decision-making unit include at least one of the following: risk of personnel replacement, risk of multiple people appearing in the same frame, and risk of leaving midway.
[0038] Thirdly, this application provides an electronic device, including: a memory and a processor;
[0039] The memory is used to store computer programs;
[0040] The processor is used to invoke the computer program to execute the method described above.
[0041] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed on an electronic device, causes the electronic device to perform the method described above.
[0042] Fifthly, this application provides a computer program product, including a computer program that, when run on an electronic device, causes the electronic device to perform the method described above.
[0043] The specific implementation methods of the second to fifth aspects of this application can refer to the implementation methods of the first aspect, and will not be elaborated here.
[0044] Beneficial effects:
[0045] 1. Fully edge-driven and highly real-time: All computations are performed on local miniaturized devices without network transmission, resulting in extremely low system response latency and the ability to react instantly to risk events.
[0046] 2. Low cost and high reliability: No expensive cloud servers and high bandwidth are required, reducing deployment and operation costs. The system operates stably and reliably, unaffected by network fluctuations.
[0047] 3. High data security: Facial data does not leave the local device, effectively protecting user privacy and data security.
[0048] 4. High performance and high precision: This application is based on a dual screening mechanism formed by confidence level and comprehensive quality score. The first layer of confidence level screening ensures that the face is real; the second layer of quality assessment screening ensures that the face is high quality (suitable for feature extraction). The combination of the two avoids non-face interference in the tracking and feature extraction process, and avoids low-quality real faces such as blurry or extremely small faces from consuming computing power, ultimately achieving the goal of high precision and high efficiency risk management at the edge.
[0049] By selecting the SCRFD-2.5G face detection model and ArcFace-R50 face feature extraction model, which perform excellently on the RK3588 NPU, and performing quantization conversion from ONNX to RKNN, fast inference at the edge was achieved (detection <16ms, recognition <19ms) while maintaining extremely high accuracy (quantization accuracy retention rate >99%), ensuring the accuracy of risk identification. Attached Figure Description
[0050] Figure 1 This is a flowchart of a real-time risk management method based on edge computing in one embodiment of this application;
[0051] Figure 2 This is a diagram illustrating the architecture of a real-time risk management system based on edge computing in one embodiment of this application.
[0052] Figure 3 This is an example diagram illustrating the risk decision-making process for mid-process personnel replacement in one embodiment of this application. Detailed Implementation
[0053] To enable those skilled in the art to better understand the present application, the technical solution of the present application will be further described in detail below with reference to the embodiments and accompanying drawings.
[0054] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification and claims of this application are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.
[0055] This application aims to address the technical problems of existing risk management systems, such as poor real-time performance, high deployment costs, reliance on networks, and difficulty in balancing performance and accuracy at the edge. Specifically, this application seeks to provide a system and method capable of efficiently and accurately performing real-time detection and early warning of risk scenarios such as personnel swapping, multiple people in the same frame, and mid-journey departure on edge computing devices (such as miniaturized devices based on the RK3588 chip). This application faces the following technical challenges:
[0056] Technical Challenge 1: The challenge of efficiently deploying high-precision AI models on resource-constrained edge computing devices. Current state-of-the-art (SOTA) face detection and recognition models (such as SCRFD and ArcFace) offer high accuracy but require massive computation, typically relying on cloud computing or high-performance GPU servers. Directly porting them to edge computing devices like the RK3588, which have limited computing power (NPU computing power is limited, such as 6 TOPS), memory, and power consumption, results in severely insufficient inference speed and an inability to meet real-time requirements. The primary technical challenge is how to enable these complex models to "run fast" on edge NPUs without significantly sacrificing accuracy.
[0057] Technical Challenge 2: Ensuring Real-Time Performance and Controlling Latency in End-to-End Processing. A complete risk management process includes multiple sequential steps such as image acquisition, detection, quality assessment, tracking, recognition, decision-making, and early warning. Even if a single step (such as NPU inference including detection and recognition) is fast, the accumulated overhead of data copying, CPU computation, and logical judgment between steps can easily cause the end-to-end latency of the entire processing pipeline to exceed the real-time threshold (for example, for 25fps video, the total latency must be less than 40ms). Designing an efficient system architecture and data flow to minimize unnecessary computation and waiting time is crucial to ensuring the overall real-time performance of the system.
[0058] Technical Challenge 3: Resource Scheduling and Performance Balancing Under Multi-Task Concurrency. The system needs to run two core AI tasks simultaneously: face detection and face recognition, as well as multiple CPU-intensive tasks such as face quality assessment and centroid tracking. On heterogeneous computing platforms like the RK3588, CPU, NPU, and memory resources are shared. Improper scheduling, such as performing recognition on all detected faces in every frame, can lead to excessive preemption of NPU resources, causing detection tasks to lag or lose frames, which in turn affects the stability of tracking and decision-making. How to intelligently schedule and allocate computing tasks to achieve a balance between performance and power consumption is a complex multi-task concurrency management challenge.
[0059] Technical Challenge 4: Ensuring Robustness and Accuracy in Complex Real-World Scenarios. Real-world monitoring environments are characterized by variable lighting, diverse human postures, and potential for partial occlusion or motion blur. These factors severely interfere with the accuracy of face detection and recognition. A low-quality, unclear, or extremely angled face image, even when fed into a face recognition model, may generate erroneous feature vectors, leading to misidentification (e.g., misidentifying authorized individuals as unauthorized individuals), resulting in numerous false alarms. Effectively filtering out invalid information and ensuring high-quality data input to the recognition process is the core challenge in improving the system's robustness in real-world environments.
[0060] Technical Challenge 5: The challenge of accurately identifying complex risk scenarios based on temporal information. Risk behaviors such as personnel changes and mid-course departures are not isolated single-frame events, but rather continuous state transition processes. For example, personnel changes require first determining if the original authorized person has left, and then determining if a new unauthorized person has entered; mid-course departures require timing the authorized person's exit (disappearance). Designing a stable and reliable state machine and decision logic that can correctly interpret the temporal results of face tracking and recognition, and accurately distinguish between similar but different situations such as temporary occlusion and genuine departure, or visitors and unauthorized personnel changes, is crucial to avoiding false alarms and missed detections.
[0061] To better understand the technical solution of this application, the relevant technical terms in this application will be explained first.
[0062] (1) NPU: The full name of NPU is Neural-network Processing Unit. NPU is a dedicated processor designed for neural network computing. It is a type of AI (Artificial Intelligence) accelerator and is mainly used to efficiently perform inference or training tasks of deep learning models, replacing the inefficiency of traditional CPU / GPU in AI computing scenarios.
[0063] (2) SCRFD algorithm: The full English name is Sample and Computation Redistribution for Efficient Face Detection. The SCRFD algorithm is a lightweight face detection algorithm. This algorithm improves the overall detection accuracy by re-evaluating the difficulty and value of faces at different scales in the training data (samples) and allocating more computing resources to those more difficult and important samples (such as small-scale faces).
[0064] (3) ArcFace algorithm model: The full English name is Additive Angular Margin Loss for DeepFace Recognition, which is translated into Chinese as Additive Angular Margin Loss Function for Deep Face Recognition. The ArcFace algorithm model is one of the most influential loss functions in the field of deep face recognition. This loss function significantly improves the discrimination ability of face recognition models by introducing additive angular margins in the angular space.
[0065] (4) ONNX format: The full English name is Open Neural Network Exchange. The Chinese translation is Open Neural Network Exchange Format. The ONNX format is an open source neural network model format. Its design purpose is to provide an intermediate representation so that models trained by different deep learning frameworks can interact and be used in a unified format.
[0066] (5) RKNN format: The full English name is Rockchip Neural Network, and the Chinese translation is Rockchip Neural Network Model Format. It is a dedicated neural network model format launched by the domestic chip manufacturer Rockchip. It is designed for the NPU of Rockchip series chips (such as RK3568, RK3588, RK3599) and is a core component of Rockchip AI development toolchain (RKNN Toolkit). It is responsible for converting general models (such as ONNX, TensorFlow) into executable models adapted to Rockchip NPU.
[0067] The following will refer to Figure 1 A specific implementation method according to this application is described.
[0068] Example 1:
[0069] like Figure 1 As shown, this application provides a real-time risk management method based on edge computing, which is applied to an edge computing device equipped with a neural network processing unit (NPU).
[0070] When the system starts up, system initialization and authorized personnel registration can be performed first. For example, facial images of authorized operators can be collected, and their standard facial feature vectors can be extracted and stored. and record its tracking ID as .
[0071] Real-time capture of video frame sequences; face detection is performed on the real-time captured video frames using an NPU to obtain face bounding boxes; the quality of the detected face bounding boxes is evaluated, and a comprehensive quality score is calculated. ; retain the overall quality score Face bounding boxes larger than a preset quality threshold are identified. Using an NPU, faces within these bounding boxes are tracked and their features extracted to obtain the number of faces and the actual face feature vectors. Compared with the standard feature vectors of pre-stored authorized personnel The relationship between faces and the temporal information of face appearance; based on the number of faces and the actual facial feature vectors Compared with the standard feature vectors of pre-stored authorized personnel The system uses the relationship between faces and the timing of their appearance to determine if a pre-set risk scenario exists; when the determination result indicates that a risk exists, an alert is triggered.
[0072] This step involves real-time video stream processing, processing each video frame in a loop.
[0073] The quality of the detected face bounding boxes is evaluated, and their overall quality score is calculated. Only for the overall quality score Face bounding boxes exceeding a preset quality threshold are tracked and their features extracted for subsequent risk decisions.
[0074] Face quality assessment plays a crucial role as a smart filter after face detection but before recognition. This is based on the detected face bounding box. It can quickly calculate multiple quality indicators and integrate them into a comprehensive quality score. Several quality metrics include: face bounding box size, sharpness (e.g., using the Laplacian operator to evaluate blur), and pose (e.g., head tilt angle). Only when... Higher than the preset quality threshold Only then is the face considered credible and of high quality, and is sent to the subsequent identification process.
[0075] In some embodiments, the face bounding box size score, sharpness score, and pose score can be calculated separately; then, a comprehensive quality score is calculated by weighting the face bounding box size score, sharpness score, and pose score. .
[0076] Among them, the face bounding box size score The calculation method is as follows:
[0077] First, calculate the area of the face bounding box. , and respectively The width and height of the face bounding box; The unit is pixels²;
[0078] Then, to Normalization is performed to obtain the face bounding box size score. :
[0079] ;
[0080] In the formula, The defined effective area range; and These are the minimum effective area and the ideal area, respectively; for example, (48×48) (100×100, calculations outside this range are based on 10000). To normalize the area, the larger the area, The closer the score is to 1, the less information the face contains, resulting in a score of 0.
[0081] Sharpness score The calculation method is as follows:
[0082] First, calculate the Laplacian operator response value for the face region image (the image within the face bounding box). , The larger the value, the clearer the image (the sharper the edges);
[0083] Then, to Normalization is performed to obtain the sharpness score. :
[0084] ;
[0085] In the formula, For the set resolution threshold range, This is the minimum sharpness threshold; anything below this value is considered blurry. For ideal sharpness; the higher the Laplacian value, the better. The closer the score is to 1, the less likely the face is to be blurred and score 0.
[0086] Posture score The calculation method is as follows:
[0087] First, calculate the head rotation angle using facial landmarks (such as the eyes, nose tip, and corners of the mouth), including the horizontal rotation angle (left and right head turns, range). ) and vertical deflection angle (tilt / descent, range) The maximum absolute value of the two values is taken as the head deflection angle. ;
[0088] Then, to Normalization is performed to obtain the pose score. :
[0089] ;
[0090] in, To set the attitude tolerance range, This is the maximum acceptable deflection angle; exceeding this range is considered an attitude failure. The smaller the deflection angle, the better. The closer to 1, the more points you get for turning your face at a large angle.
[0091] Overall quality score The weighted summation of the scores from the three dimensions yields the following expression:
[0092] ;
[0093] in: , , and These are the weights for the face bounding box size score, sharpness score, and pose score, respectively. These weights are determined empirically; in some embodiments, , , ;
[0094] This step calculates a comprehensive quality score for each face. The system filters out faces that meet the requirements.
[0095] In some embodiments, faces are detected by a face detection model; the face detection model adopts the SCRFD algorithm model.
[0096] The SCRFD algorithm model is used to detect all faces in the current video frame. For example, a set of face bounding boxes is obtained. , of which The personal face bounding box is denoted as , , This represents the total number of faces detected.
[0097] In some embodiments, a centroid tracking algorithm is used for face tracking.
[0098] The centroid tracking algorithm requires only simple coordinate calculations from the CPU, making it extremely fast. Once a face is assigned a tracking ID, in subsequent consecutive frames, as long as its position doesn't change significantly, the system assumes its identity remains unchanged, eliminating the need to call the recognition model again. This solves the huge performance overhead of recognizing all faces in every frame. A unique tracking ID (e.g., ...) is assigned to each face appearing consecutively in the video sequence. ), and associate it with its feature vector (such as ).
[0099] The centroid tracking algorithm is used to update the tracking IDs of all faces. For newly emerging high-quality faces, the ArcFace model is used to extract their feature vectors. For existing faces, their feature vectors can be selectively updated according to a strategy.
[0100] In some embodiments, the actual facial feature vector is extracted using a facial feature extraction model (facial recognition model). The facial feature extraction model can employ the ArcFace algorithm.
[0101] In some embodiments, the edge computing device is a Rockchip device (such as RK3588), and both the SCRFD algorithm model and the ArcFace algorithm model are converted and quantized from ONNX format to RKNN format to adapt for accelerated inference on the NPU.
[0102] The face detection model employs an optimized SCRFD algorithm model that has undergone format conversion and quantization. Quantizing the SCRFD algorithm model, converting its weights from 32-bit floating-point (FP32) to 8-bit integer (INT8), significantly reduces model size and memory usage, fully utilizing the NPU's INT8 computing power and improving inference speed by several times. The model performs efficient and high-precision face detection on the input video frames on the NPU, outputting the face bounding box coordinates and confidence scores. .
[0103] The confidence level This refers to the probability that the bounding box of a face corresponds to a real face; the value typically ranges from 100 to 100. After the model outputs the face bounding box, it is then analyzed according to the set confidence threshold. (e.g., 0.8); only retain The detection results (high-reliability face) are discarded directly. The result (low reliability candidate region). Therefore, invalid detection results can be filtered out, ensuring the accuracy of subsequent processes and reducing the computing power consumption of edge devices.
[0104] The face feature extraction model uses an optimized ArcFace algorithm, optimized using the same method as the SCRFD algorithm, which can extract 512-dimensional feature vectors from high-quality face images on an NPU. .
[0105] To convert ONNX format to RKNN format, you can use the RKNN Toolkit. The RKNN Toolkit is a tool provided by Rockchip for converting deep learning models to RKNN format for deployment and operation on Rockchip devices.
[0106] The determination of whether a preset risk scenario exists can be made by making a risk scenario decision every frame or every few frames.
[0107] In some embodiments, determining whether a preset risk scenario exists includes at least one of the following: determining whether there is a risk of replacement, determining whether there is a risk of multiple people appearing in the same frame, or determining whether there is a risk of leaving midway.
[0108] Determining the risk of replacement: Based on actual facial feature vectors Compared with the standard feature vectors of pre-stored authorized personnel Calculate the actual facial feature vector Compared with the standard feature vectors of pre-stored authorized personnel Cosine similarity between If cosine similarity If the similarity is less than the preset similarity threshold, it is determined that there is a risk of replacement.
[0109] Specifically, determining whether there is a risk of personnel replacement, i.e., making a personnel replacement risk decision, may include: 1) Checking if there is an authorized person with the tracking ID. 2) If not, or if the person's associated face is not detected in consecutive frames, then the authorized person is deemed to have left the premises. 3) If the authorized person has left the premises, checking if there are other unauthorized personnel (tracking ID). 4) Feature vectors associated with other unauthorized personnel. Standard feature vectors of authorized personnel Calculate cosine similarity : ,in This represents the function for calculating cosine similarity. Indicates the calculation of vector length; 5) If Less than the preset similarity threshold If the value is 0.6, it is considered that there is a risk of replacement.
[0110] Risk level of player replacement It can be defined as: ;in To determine the risk factor for replacing personnel, a value can be determined based on experience. It represents the maximum similarity between all unauthorized personnel and authorized personnel in the current scenario.
[0111] The risk scenario of personnel replacement indicates that after the authorized person's face leaves, an unauthorized person's face appears with a similarity to the authorized person's face feature that is lower than a preset similarity threshold.
[0112] Determine if there is a risk of multiple people appearing in the same frame: Based on the number of faces tracked, if the number of faces detected in the current video frame is greater than 1, it is determined that there is a risk of multiple people appearing in the same frame.
[0113] For example, calculate the number of faces that are stably tracked in the current video frame. .like This would trigger the risk of multiple people appearing in the same frame.
[0114] The risk level of multiple people appearing in the same frame can be defined by the following formula: ,in The risk factor is calculated based on the number of people in the same frame.
[0115] The criteria for determining a stably tracked face can be: in the current video frame, a SCRFD model detects the face, and the overall quality score is higher than a preset quality threshold. The face is then assigned a tracking ID by the centroid tracking algorithm. The centroid tracking algorithm can be configured to calculate the number of consecutive frames (e.g., 5 frames). If the same face is detected within a consecutive number of frames, a tracking ID is assigned to that face.
[0116] Determine if there is a risk of leaving midway: Based on the timing information of face appearance, maintain a timer to record the duration for which authorized personnel are not detected (the time for authorized personnel to leave the site). If the duration exceeds the preset departure time threshold, it is determined that there is a risk of leaving midway.
[0117] For example, the system maintains a timer. Used to record authorized personnel Duration during which it was not detected; when detected, reset. When not detected, accumulate. ;like Exceeding the preset threshold for absence from work And there are no other faces in the current scene ( If this occurs, it triggers the risk of leaving midway. Risk level for leaving midway. It can be defined as: ,in This represents the risk factor for leaving midway.
[0118] In some embodiments, if , and If any risk level exceeds the preset trigger threshold, an alarm will be issued through the early warning output module.
[0119] This solution provides an innovative risk quantification model: a risk level calculation formula based on the number of faces, feature vector similarity, and absence time, which transforms risk assessment from a simple binary judgment (yes / no) into a continuous quantitative indicator, making it easier to set different warning thresholds according to different scenarios, and making the system more flexible and intelligent.
[0120] Secondly, this application provides a real-time risk management system based on edge computing. This system is deployed on an edge computing device equipped with a neural network processing unit (NPU), and includes:
[0121] The image acquisition module is used to capture video frame sequences in real time.
[0122] The face detection unit is used to perform face detection on real-time captured video frames based on the NPU to obtain face bounding boxes;
[0123] The face quality assessment unit is used to evaluate the quality of detected face bounding boxes and calculate a comprehensive quality score. ; retain the overall quality score Face bounding boxes that exceed a preset quality threshold;
[0124] The face tracking and feature extraction unit is used to track and extract features from detected faces based on the NPU, obtaining the number of faces and the actual face feature vectors. Compared with the standard feature vectors of pre-stored authorized personnel The relationship between faces and the temporal information of face appearance;
[0125] Risk decision-making unit, used to determine the number of faces and actual facial feature vectors. Compared with the standard feature vectors of pre-stored authorized personnel Based on the relationship between the faces and the timing information of their appearance, we can determine whether there are pre-set risk scenarios and make risk decisions.
[0126] The early warning output module is used to trigger an early warning when the risk decision-making unit determines that there is a risk.
[0127] When the risk decision-making unit identifies a risk event, it triggers a local audible and visual alarm or sends a warning message via the network.
[0128] The tracking and feature extraction unit is the core processing unit of the system and is deployed on edge computing devices.
[0129] The risk scenarios identified by the risk decision-making unit include at least one of the following: risk of personnel replacement, risk of multiple people appearing in the same frame, and risk of leaving midway.
[0130] Using the RK3588 edge computing device as an example, the system framework of one embodiment of this application is described. Figure 2 As shown, the system can be divided into a hardware platform layer, a core algorithm layer, and a platform application layer. The hardware platform layer includes the RK3588, camera module, storage, and I / O. The core algorithm layer includes an image acquisition module, a face detection unit, a quality assessment unit, a tracking unit, a face recognition unit, and a risk decision-making unit; the tracking unit and face recognition unit are the face tracking and feature extraction units described in the previous embodiment. The platform application layer includes risk scene monitoring, early warning output, and system management.
[0131] In some embodiments, the overall processing flow of the system in this application is as follows:
[0132] S1. Startup and Registration: After the device is powered on, it prompts authorized personnel to face the camera. The system captures 3-5 high-quality facial images, sends them to the ArcFace algorithm model to extract the corresponding feature vectors, calculates the average value to obtain standard features, and stores them in local non-volatile storage.
[0133] S2. Enter monitoring mode: The system begins processing the real-time video stream.
[0134] S3. Looping Frame Processing: For each frame, the SCRFD algorithm model is first run on the NPU to output the bounding boxes of all faces.
[0135] S4. Quality Filtering: For example, calculate the area of the bounding box for each face. .like If the face is too small and of low quality, it will be ignored. Face size threshold, such as 48×48 pixels².
[0136] Then, calculate the overall quality score. ,like Higher than the preset quality threshold It is then identified as a high-quality face bounding box.
[0137] S5. Face Tracking and Feature Extraction: The centroid position of the high-quality face bounding box is passed to the face tracking and feature extraction unit, and the tracking ID is updated or assigned. For tracking IDs of... The authorized personnel's faces are scanned to confirm their presence, and the system is reset. For newly emerging tracking IDs, the corresponding feature vectors are extracted using the ArcFace algorithm model. .
[0138] S6. Risk assessment, i.e., risk scenario decision-making:
[0139] Multiple faces in the same frame: If the number of faces detected in the current video frame is... If the value is greater than 1, an alert for multiple people appearing in the same frame will be triggered immediately.
[0140] Risk detection for leaving midway: If It has not been detected for a long time. Start accumulating.
[0141] when (e.g., 10 seconds) and This triggers a warning for leaving midway.
[0142] Replacement risk detection: When the tracking ID is The authorized personnel were not present, but other tracking IDs were present on site. For unauthorized personnel, the calculation is... .like If the value is 0.6, it is determined to be a player substitution and an alert is triggered.
[0143] S6. Calculate the risk level, determine the threshold, and output multi-level early warnings.
[0144] Figure 3 This paper illustrates the interaction relationships between the customer, application, and detection module in each stage of a business scenario (personnel change midway) in one embodiment of this application, as well as the risk control triggering logic: customer initiates business → system starts monitoring → customer is replaced → system intercepts business → early warning notification → business ends / resumes.
[0145] Taking bank self-service terminal business as an example, the process in the above business scenario includes:
[0146] When a customer clicks the "Withdraw" button, the system activates its risk control module, taking and caching the customer's face (e.g., that of an authorized person). );
[0147] During the transaction, the customer was replaced by an unauthorized person (e.g.) );
[0148] The system detected The authorized personnel left the scene. Calculate the cosine similarity between authorized and unauthorized personnel entering the premises. It was determined that there was a risk of having to replace the player.
[0149] The system immediately intercepts the withdrawal transaction, triggers a local audible and visual alarm, and notifies the bank's management platform.
[0150] Once the management platform confirms the risk, it will arrange for customer service to intervene remotely to prevent financial losses.
[0151] Effect test:
[0152] Comparative tests were conducted on several mainstream algorithms:
[0153] (1) Face detection model:
[0154] A comparison was made with MTCNN, RetinaFace, YOLOv5s-Face, and the SCRFD series. Detailed performance comparison data are shown in Tables 1-1 and 1-2, and the actual test data on the RK3588 platform is shown in Table 2. As can be seen from Tables 1-1, 1-2, and 2, SCRFD-2.5G achieves the best balance between accuracy (WIDERFACEHard: 85.9%), model size (10.2MB), and inference performance on the RK3588 NPU (15.6ms). In particular, its accuracy retention rate after NPU quantization is as high as 99.2%, far exceeding other models, which is crucial for ensuring the actual effect of edge deployment.
[0155] ;
[0156] ;
[0157] Note: WIDERFACE is a commonly used benchmark dataset for face detection in the field of computer vision (it can be directly translated as WIDER face dataset). The corresponding Easy, Medium, and Hard datasets are three sub-datasets of this dataset divided according to the detection difficulty (easy, medium, and hard). FPS (Frames Per Second) is a core metric for measuring the speed of video processing / model inference, representing the number of video frames that the system can process per second or the number of inferences that the model can complete per second.
[0158] ;
[0159] (2) Facial feature extraction model:
[0160] A comparison was made between FaceNet, CosFace, SphereFace, and the ArcFace series. Detailed performance comparison data, training efficiency comparison, and RK3588 platform compatibility are shown in Tables 3, 4, and 5, respectively. As can be seen from Tables 3, 4, and 5, ArcFace-R50 achieved the highest accuracy on multiple authoritative benchmark sets, including LFW (99.82%) and CFP-FP (98.27%). On the RK3588 platform, although its inference time (18.3ms) was slightly higher than the lightweight model, its quantization accuracy retention rate (99.1%) and feature extraction stability were the best, ensuring robustness in identity recognition.
[0161] ;
[0162] Among them, the LFW dataset, CFP-FP dataset, and AgeDB-30 dataset are three different face datasets. The LFW dataset is a field-labeled face dataset, the CFP-FP dataset is a cross-pose face frontal-lateral dataset, and the AgeDB-30 dataset is an age-varying face dataset.
[0163] ; ;
[0164] The innovativeness and technological contribution of this application:
[0165] 1. Fully edge-driven and highly real-time: All computations are performed on local miniaturized devices without network transmission, resulting in extremely low system response latency and the ability to react instantly to risk events.
[0166] 2. Low cost and high reliability: No expensive cloud servers and high bandwidth are required, reducing deployment and operation costs. The system operates stably and reliably, unaffected by network fluctuations.
[0167] 3. High data security: Facial data does not leave the local device, effectively protecting user privacy and data security.
[0168] 4. High performance and high precision: By selecting the SCRFD-2.5G face detection model and ArcFace-R50 face feature extraction model, which perform excellently on the RK3588 NPU, and performing quantization conversion from ONNX to RKNN, fast inference at the edge is achieved (detection <16ms, recognition <19ms) while maintaining extremely high precision (quantization precision retention rate >99%), ensuring the accuracy of risk identification.
[0169] 5. Innovative risk quantification model: The risk level calculation formula proposed in this application, based on the number of faces, identity similarity, and time away from the post, transforms risk assessment from a simple binary judgment (yes / no) into a continuous quantitative indicator. This makes it easier to set different warning thresholds according to different scenarios, resulting in higher system flexibility and intelligence.
[0170] Through the above implementation methods, this application has successfully built a powerful, responsive and highly reliable risk management system on an edge computing device, effectively solving the various problems mentioned in the background art.
[0171] Example 2:
[0172] This embodiment provides an electronic device, including: a memory and a processor;
[0173] The memory is used to store computer programs;
[0174] The processor is configured to invoke the computer program to execute the method as described in Embodiment 1.
[0175] Example 3:
[0176] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is run on an electronic device, it causes the electronic device to perform the method described in Embodiment 1.
[0177] Example 4:
[0178] This embodiment provides a computer program product, including a computer program that, when run on an electronic device, causes the electronic device to perform the method described in Embodiment 1.
[0179] The specific implementation of the system, electronic device, computer-readable storage medium, and computer program product provided in this application can be referred to the specific embodiments of the above methods, and will not be repeated here.
[0180] The technical content of the above embodiments can be referred to each other. For the same or similar technical features, appropriate omissions have been made in some embodiments to avoid repeated descriptions.
[0181] Obviously, those skilled in the art should understand that the various units or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device, or fabricating them separately as individual integrated circuit modules, or fabricating multiple modules or steps into a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0182] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. An edge computing-based real-time risk management method, characterized in that, The application is applied to an edge computing device with a neural network processing unit (NPU) and comprises the following steps: real-time capture of a video frame sequence; face detection on the real-time captured video frame based on the NPU to obtain a face bounding box; performing quality assessment on the detected face bounding box, calculating a comprehensive quality score ; comprising: firstly, the size score, the definition score and the posture score of the face bounding box are calculated respectively; The calculation method of the face boundary box size score is as follows. ; In the formula, is the area of the face bounding box; and are the minimum effective area and ideal area, respectively. Sharpness score The calculation method is: ; wherein is a Laplacian response value computed for an image within the face bounding box; is a minimum sharpness threshold, is an ideal sharpness; Pose score The calculation method is: ; In the formula, is the head deflection angle; the head deflection angle is calculated by the face key points, including the horizontal deflection angle and the vertical deflection angle, and the maximum value of the absolute values of the two is taken as the head deflection angle ; is the maximum acceptable deflection angle; The comprehensive quality score is calculated by weighting the face boundary box size score, the definition score and the pose score : ; In the formula, , , and are respectively the weights of the face boundary box size score, the sharpness score and the pose score, and the weights are valued according to experience. retained comprehensive quality score a face bounding box greater than a preset quality threshold Based on the NPU, the face in the retained face bounding box is tracked and feature extracted to obtain the number of faces and actual face feature vectors The relationship with the standard feature vector of the authorized personnel stored in advance and the timing information of the face appearance; Based on the number of faces, actual face feature vectors The relationship with the standard feature vector of the pre-stored authorized personnel And the timing information of the appearance of the face, whether a preset risk scenario exists is judged when the judgment result is that there is a risk, a warning is triggered.
2. The method of claim 1, wherein, The face detection on the real-time captured video frame based on the NPU to obtain a face bounding box comprises the following steps: Based on the NPU, face detection is performed on real-time captured video frames, and the detected face bounding boxes and corresponding confidence scores are output. The confidence level This refers to the probability value that the face bounding box corresponds to a real face; after outputting the face bounding box, it is calculated based on the set confidence threshold. Filter the test results: retain only Face bounding box, discard The bounding box of the face.
3. The method of claim 1, wherein, detecting a face through a face detection model; the face detection model adopts an SCRFD algorithm model; extracting an actual face feature vector through a face feature extraction model ; the face feature extraction model adopts an ArcFace algorithm model.
4. The method of claim 3, wherein, the edge computing device is a RUIKUI micro device; the SCRFD algorithm model and the ArcFace algorithm model are converted from the ONNX format to the RKNN format and quantized to adapt to the acceleration inference on the NPU.
5. The method of claim 1, wherein, The judgment of whether there is a preset risk scenario comprises at least one of the following: determining whether there is a risk of substitution: based on the actual face feature vector and the standard feature vector of the authorized personnel pre-stored , calculating the cosine similarity between the actual face feature vector and the standard feature vector of the authorized personnel pre-stored , if the cosine similarity is less than a preset similarity threshold, it is determined that there is a risk of substitution judgment of whether there is a multi-person same frame risk: based on the number of faces obtained through tracking, if the number of detected faces in the current video frame is greater than 1, it is determined that there is a multi-person same frame risk; judgment of whether there is a midway leaving risk: based on the time sequence information of the face, a timer is maintained to record the duration of the authorized personnel not being detected, and if the duration exceeds a preset off-duty time threshold, it is determined that there is a midway leaving risk.
6. An edge computing-based real-time risk management and control system, characterized in that, The system is deployed on an edge computing device with a neural network processing unit (NPU) and is used to implement the method of any one of claims 1-5, comprising: an image acquisition module for real-time capture of a video frame sequence; a face detection unit for processing the real-time captured video frame, detecting a face through a face detection model; a face quality evaluation unit configured to evaluate a quality of the detected face bounding box, and calculate a comprehensive quality score , and retain the face bounding box with the comprehensive quality score greater than a preset quality threshold The face tracking and feature extraction unit is configured to track and extract features of the detected face based on the NPU to obtain the number of faces, actual face feature vectors and the relationship between the pre-stored standard feature vectors of authorized personnel and the timing information of the appearance of the face. a risk decision unit configured to determine whether a preset risk scenario exists based on a number of faces, actual face feature vectors and timing information of the faces, and a relationship between the faces and standard feature vectors of pre-stored authorized personnel a warning output module for triggering a warning when the judgment result of the risk decision unit is that there is a risk.
7. An electronic device, comprising: comprise: a memory and a processor; the memory is used to store a computer program; the processor is used to call the computer program to execute the method of any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program runs on an electronic device to make the electronic device implement the method of any one of claims 1-5.
9. A computer program product comprising a computer program, characterized in that, The computer program runs on an electronic device to make the electronic device implement the method of any one of claims 1-5.
Citation Information
Patent Citations
Artificial intelligence convolutional neural network face recognition system
CN110414305A
vision-based Fall risk assessment system
CN112784662A