Face recognition emergency method and equipment based on display terminal and medium
Through the multi-layer cascade face detection model and silent liveness detection technology, the problem of insufficient face recognition feature extraction in complex environments is solved, efficient liveness recognition and precise positioning in outdoor environments are achieved, and emergency response efficiency is improved.
Patent Information
- Application Number
- CN202510871801.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
Smart Images

Figure CN120808411A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of portrait recognition, and in particular to a face recognition emergency method based on a display terminal, a device and a medium. BACKGROUND
[0002] With the acceleration of urbanization and the increase in the frequency of large-scale activities, such as large-scale concerts, sports events and other scenarios, the emergency needs of personnel evacuation, lost search and rescue, dangerous personnel tracking and the like are increasingly urgent. Traditional emergency methods rely on manual investigation, broadcast search or limited monitoring systems, which have problems such as slow response, low efficiency, poor accuracy and the like, and are difficult to meet the emergency processing needs in complex scenarios.
[0003] Face recognition technology has become an important means of personnel positioning and rescue command due to its non-contact and efficient characteristics. However, the existing face recognition technology is difficult to accurately capture target information when dealing with complex environments such as changes in light, dense personnel and high fluidity of crowds, due to insufficient feature extraction capability in complex environments and reliance on cooperative liveness detection, thereby reducing emergency efficiency. SUMMARY
[0004] The embodiments of the present application provide a face recognition emergency method based on a display terminal, a device and a medium, which are used to solve the technical problem that the existing face recognition technology has insufficient feature extraction capability in complex environments and relies on cooperative liveness detection, thereby making it difficult to accurately capture target information and reducing emergency efficiency.
[0005] The embodiments of the present application adopt the following technical solutions:
[0006] The embodiments of the present application provide a face recognition emergency method based on a display terminal. It includes acquiring a face image appearing in a display terminal, detecting the face image based on a multi-level cascade face detection model group, and outputting face information; wherein the face information at least includes face position and face key point information; performing silent liveness detection on the face image through a liveness detection fusion model, and in the case of passing the detection, performing feature extraction on the face image through a preset convolutional network, and outputting an age prediction value based on the extracted features; based on the face information and the age prediction value, constructing a face information set associated with the display terminal; in response to a current emergency instruction, performing face matching in the face information set based on target portrait information uploaded by a user; in the case of successful matching, publishing the position information of the target personnel on each display terminal in a preset area based on the position of the matched display terminal.
[0007] The embodiment of the application solves the interference of outdoor complex illumination, shielding and the like, and completes the living body recognition without the cooperation of the user, is suitable for an emergency scene with a large number of people, and avoids the response delay of the traditional cooperative detection. Secondly, the embodiment of the application enhances the fine-grained feature extraction capability and reduces the false detection rate by using the center difference convolution and the like while ensuring the real-time performance of the layer-cascaded model. A multi-dimensional information set is constructed, the closed-loop emergency response of the geographic distribution of the display terminal is combined, the position is dynamically marked on the surrounding screen, the rescue personnel are guided to accurately position, and the emergency efficiency is improved.
[0008] In an implementation manner of the application, a face image is detected based on a multi-layer cascaded face detection model group, and face information is output, specifically including: performing multi-scale transformation on the face image at the input layer of the multi-layer cascaded face detection model group to generate an image pyramid containing different scale images; inputting the image pyramid to a first-level subnetwork to generate a plurality of candidate target region frames; inputting the candidate target region frames to a second-level subnetwork to perform primary screening and frame regression on the candidate target region frames based on a preset screening condition to screen out negative example regions with a confidence lower than a threshold; inputting the screened candidate target region frames to a third-level subnetwork to perform secondary discrimination and frame regression, and outputting a final face detection result and face information.
[0009] In an implementation manner of the application, a face image is detected based on a multi-layer cascaded face detection model group, and face information is output, specifically including: performing multi-scale transformation on the face image at the input layer of the multi-layer cascaded face detection model group to generate an image pyramid containing different scale images; inputting the image pyramid to a first-level subnetwork to generate a plurality of candidate target region frames; inputting the candidate target region frames to a second-level subnetwork to perform primary screening and frame regression on the candidate target region frames based on a preset screening condition to screen out negative example regions with a confidence lower than a threshold; inputting the screened candidate target region frames to a third-level subnetwork to perform secondary discrimination and frame regression, and outputting a final face detection result and face information.
[0010] In an implementation manner of the application, a face image is detected based on a multi-layer cascaded face detection model group, and face information is output, specifically including: performing multi-scale transformation on the face image at the input layer of the multi-layer cascaded face detection model group to generate an image pyramid containing different scale images; inputting the image pyramid to a first-level subnetwork to generate a plurality of candidate target region frames; inputting the candidate target region frames to a second-level subnetwork to perform primary screening and frame regression on the candidate target region frames based on a preset screening condition to screen out negative example regions with a confidence lower than a threshold; inputting the screened candidate target region frames to a third-level subnetwork to perform secondary discrimination and frame regression, and outputting a final face detection result and face information.
[0011] In an implementation manner of the present application, the face image is subjected to feature extraction by a preset convolutional network, and an age prediction value is output based on the extracted features, specifically comprising: based on the face information, the face image is subjected to face alignment processing; the face image after the alignment processing is input to a backbone convolutional neural network MobileFaceNets; wherein the backbone convolutional neural network comprises a network architecture corresponding to a fast down-sampling and dimension reduction strategy, and a 1x1 linear convolutional layer is connected after a global deep convolutional layer to output a face feature vector.
[0012] In an implementation manner of the present application, an age prediction value is output based on the extracted features, specifically comprising: the face feature vector is input to a prediction layer; wherein the prediction layer is subjected to model training based on a joint loss function of a classification loss and a regression loss, and the joint loss function is a weighted sum of the classification loss and the regression loss; by Softmax classification, a probability distribution of different age categories is output; by expectation regression, the probability distribution is multiplied by a corresponding age label to obtain an age prediction value.
[0013] In an implementation manner of the present application, in response to a current emergency instruction, based on target portrait information uploaded by a user, face matching is performed in a face information set, specifically comprising: receiving the current emergency instruction, and receiving the target portrait information uploaded by the user; based on a feature similarity algorithm, the target portrait information is matched with the face information set; in the case of successful matching, display terminal identification information associated with the matching result is output, and an emergency response operation is triggered; wherein the face information set at least contains face information, age prediction value, living body detection result and face feature vector.
[0014] In an implementation manner of the present application, based on the location of the matched display terminal, the location information of the target personnel is published on each display terminal in a preset area, specifically comprising: based on the display terminal identification information, the geographic location information of the display terminal is determined; based on the geographic location information of the display terminal, the emergency area range is determined, and a plurality of display terminals in the emergency area range are determined; according to the priority of the current emergency instruction, the publishing frequency and display time length of the emergency information on the display terminal are adjusted; based on the adjusted strategy, the emergency information is displayed on the plurality of display terminals in the emergency area range.
[0015] The embodiment of the application provides a face recognition emergency based on a display terminal, including: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: acquire a face image appearing in the display terminal, detect the face image based on a multi-layer cascade face detection model group, and output face information; wherein the face information at least includes face position and face key point information; perform silent live body detection on the face image through a live body detection fusion model, in the case of passing the detection, perform feature extraction on the face image through a preset convolution network, and output an age prediction value based on the extracted features; based on the face information and the age prediction value, construct a face information set associated with the display terminal; in response to a current emergency instruction, perform face matching in the face information set based on target portrait information uploaded by a user; in the case of successful matching, publish position information of the target personnel on each display terminal in a preset area based on the position of the matched display terminal.
[0016] The embodiment of the application provides a nonvolatile computer storage medium, which stores computer executable instructions, and the computer executable instructions are set to: acquire a face image appearing in the display terminal, detect the face image based on a multi-layer cascade face detection model group, and output face information; wherein the face information at least includes face position and face key point information; perform silent live body detection on the face image through a live body detection fusion model, in the case of passing the detection, perform feature extraction on the face image through a preset convolution network, and output an age prediction value based on the extracted features; based on the face information and the age prediction value, construct a face information set associated with the display terminal; in response to a current emergency instruction, perform face matching in the face information set based on target portrait information uploaded by a user; in the case of successful matching, publish position information of the target personnel on each display terminal in a preset area based on the position of the matched display terminal.
[0017] The above at least one technical scheme adopted by the embodiment of the application can achieve the following beneficial effects: the embodiment of the application solves the interference of outdoor complex light, shielding and the like through the multi-layer cascade face detection model and the silent live body detection technology, can complete live body recognition without user cooperation, is suitable for an emergency scene with dense personnel, and avoids the response delay of the traditional cooperative detection. Secondly, the embodiment of the application enhances the fine-grained feature extraction capability through center difference convolution and the like while ensuring real-time through the layer cascade model, and reduces the false detection rate. A multi-dimensional information set is constructed, a closed-loop emergency response is combined with the geographic distribution of the LED screen, and the position is dynamically marked on the surrounding screen, so that the rescue personnel can be guided to accurately position, and the emergency efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to make the technical solutions in the present application or prior art clearer, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor. In the drawings:
[0019] Figure 1 A flow chart of a face recognition emergency method based on a display terminal is provided for the embodiments of the present application.
[0020] Figure 2 A structural schematic diagram of a face recognition emergency device based on a display terminal is provided for the embodiments of the present application.
[0021] Reference signs:
[0022] 200: face recognition emergency device based on a display terminal, 201: processor, 202: memory. DETAILED DESCRIPTION
[0023] The embodiments of the present application provide a face recognition emergency method, device and medium based on a display terminal.
[0024] In order to make those skilled in the art better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0025] The technical solutions of the embodiments of the present application will be described in detail below with reference to the drawings.
[0026] Figure 1 A flow chart of a face recognition emergency method based on a display terminal is provided for the embodiments of the present application, as shown in Figure 1 The face recognition emergency method based on a display terminal includes the following steps:
[0027] S101, acquiring a face image appearing in a display terminal, detecting the face image based on a multi-layer cascade face detection model group, and outputting face information.
[0028] In one implementation manner of the present application, the face information at least includes face position and face key point information.
[0029] In an implementation form of the present application, a multi-scale transformation is performed on the face image at the input layer of the multi-level cascaded face detection model group to generate an image pyramid containing images of different scales. The image pyramid is input to a first level subnetwork to generate a plurality of candidate target region frames. The candidate target region frames are input to a second level subnetwork to perform preliminary screening and frame regression on the candidate target region frames based on a preset screening condition to screen out negative example regions with a confidence lower than a threshold. The screened candidate target region frames are input to a third level subnetwork to perform secondary discrimination and frame regression, and output the final face detection result and face information.
[0030] Specifically, an image of a current scene is captured by a shooting device arranged on the display terminal to obtain a face image appearing in the display terminal. A multi-scale transformation is performed on the input face image to generate a series of images of different sizes through scaling operation, and the images are arranged in the form of an image pyramid according to the scale size.
[0031] Further, the image pyramid is input to a first level subnetwork (P-Net, Proposal Network). The P-Net performs sliding window processing on the images of each scale, extracts features through convolution operation, and performs binary classification and frame regression on each sliding window position to generate a large number of candidate target region frames. These candidate frames preliminarily label the regions where the face may exist and have a preliminary confidence score.
[0032] Further, the candidate target region frames generated by the first level subnetwork are input to a second level subnetwork (R-Net, Refine Network). The R-Net performs more detailed feature extraction and classification on the images in each candidate frame, performs preliminary screening on the candidate frames based on a preset screening condition, and screens out negative example regions with a confidence lower than a threshold. At the same time, frame regression is performed on the retained candidate frames to adjust their position and size so as to be closer to the real face boundary.
[0033] Further, the candidate target region frames screened and optimized by the second level subnetwork are input to a third level subnetwork (O-Net, Output Network). The O-Net performs more strict secondary discrimination on the candidate frames to accurately distinguish the face and non-face regions, and performs more detailed frame regression to further adjust the position and size of the face frame. At the same time, the O-Net also predicts the key points of the face to output the final face detection result, including the accurate face frame coordinates, confidence and face key point information.
[0034] S102, performing silent liveness detection on the face image through the liveness detection fusion model, and in the case of passing the detection, performing feature extraction on the face image through the preset convolutional network, and outputting an age prediction value based on the extracted features.
[0035] In an implementation form of the present application, the intensity information and gradient information of the face image are aggregated by the living body detection fusion model to obtain the face fine-grained feature. The face fine-grained feature is input into the preset detector to perform the silent living body detection on the face image. The living body detection fusion model is set as a cascade network structure, and the living body detection fusion model is composed of a feature extraction layer based on center difference convolution, a network layer optimized by combining neural architecture search, and a multi-scale attention fusion module.
[0036] Specifically, the living body detection fusion model first performs multi-dimensional feature extraction and aggregation on the input face image. The feature extraction layer based on center difference convolution uses a specially designed convolution kernel to extract the intensity information and gradient information of the image respectively. The intensity information reflects the gray value distribution of the image and retains the texture details; the gradient information captures the rate of change of pixel values and highlights the edge features. The layer also introduces a residual connection to ensure that low-level features are not lost. Finally, the face fine-grained feature containing rich details is generated through channel splicing and nonlinear activation of the feature map.
[0037] Further, the extracted face fine-grained feature is input into the preset detector, which judges whether the face is a real living body based on the trained classification model. For example, if the detector uses a Softmax classifier, it outputs the probability distribution of belonging to a real face and a fake face, and combines the feature similarity comparison mechanism to enhance the robustness. Finally, the silent living body detection result is output by integrating the classification probability, similarity score and micro-dynamic feature.
[0038] In an implementation form of the present application, the intensity information and gradient information of the face image are extracted by the feature extraction layer based on center difference convolution. The intensity information and gradient information are fused based on adaptive weights to obtain a preliminary feature map. The network architecture parameters required for the living body detection task corresponding to the preliminary feature map are determined by the network layer optimized by neural architecture search, and the preliminary feature map is optimized by the network architecture parameters. The multi-scale face feature map corresponding to the optimized feature map is generated by the multi-scale attention fusion module, the multi-scale face feature map is processed in parallel, and the multi-scale feature map is adaptively weighted and fused by the attention mechanism to obtain the face fine-grained feature.
[0039] Specifically, the cascade structure of the living body detection fusion model in the embodiment of the present application optimizes feature expression through three-layer networks. The first layer of feature extraction based on central difference convolution completes the preliminary fusion of intensity and gradient information. The second layer of neural architecture search optimized network layer automatically finds the optimal network parameters through reinforcement learning and dynamically adjusts the channel number and convolution kernel size of feature extraction. The third layer of multi-scale attention fusion module divides the feature map into multiple scale branches for processing and assigns weights to different regions through spatial attention mechanism. The three-layer network refines the features layer by layer, and finally outputs discriminant features highly sensitive to fake attacks, improving the accuracy of silent living body detection.
[0040] Further, the multi-scale attention fusion module in the embodiment of the present application generates multiple groups of feature maps with different resolutions through convolution or pooling operations with different strides, respectively capturing local details and global structures. For each group of feature maps, spatial attention and channel attention mechanisms are applied in turn: spatial attention gives higher weights to key parts such as eyes and mouth by calculating the pixel correlation of each region; channel attention filters out the most relevant feature channels for living body detection through global average pooling and fully connected layers. Finally, the multi-scale feature maps are merged through adaptive weighted fusion to generate face fine-grained features containing cross-scale details and anti-fake attack.
[0041] Multi-scale feature fusion method based on attention mechanism:
[0042] F = a · F global + β · F local + γ · F texture;
[0043] Where F is the fused feature, F global is the global feature, F local is the local feature, and F texture is the texture feature; a is the weight corresponding to the global feature, β is the weight corresponding to the local feature, and γ is the weight corresponding to the texture feature.
[0044] In an implementation manner of the present application, based on face information, face alignment processing is performed on the face image. The face image after alignment processing is input to the backbone convolutional neural network MobileFaceNets. The backbone convolutional neural network includes a network architecture corresponding to fast down-sampling and dimension reduction strategy, and a 1x1 linear convolutional layer is connected after the global deep convolutional layer to output a face feature vector.
[0045] Specifically, the multi-task cascade network (such as MTCNN) is used to detect key feature points in the face image, which usually includes 5 key point coordinates, i.e., both pupils, nose tip, and left and right corners of the mouth. Based on these key points, a similarity transformation matrix is calculated to map faces with different poses to a standard front face view.
[0046] Further, when the aligned face image is input into the MobileFaceNets, a fast downsampling strategy is adopted in the front end of the network to reduce the amount of calculation. Through the large stride convolution operation in the early layers, the feature map resolution is quickly reduced while the key features are retained. Thus, the spatial dimension is reduced in the early stage of the network, avoiding the calculation redundancy of the deep network, and the feature expression ability is maintained through the small dilation factor inverted residual module, reducing the parameter amount compared with the traditional network.
[0047] Further, in the back end of the network, a global depth convolution (such as GDConv) is adopted instead of the traditional global average pooling, and the global features are adaptively extracted through the learnable convolution kernel. GDConv performs channel-by-channel convolution on the entire feature map, and each convolution kernel learns the weight distribution at different positions, which can capture more detailed feature differences compared with the average pooling. Then, a 1x1 linear convolution layer is connected as the feature output layer.
[0048] In an implementation of the present application, the face feature vector is input into the prediction layer; wherein the prediction layer is trained based on a joint loss function of classification loss and regression loss, and the joint loss function is the weighted sum of the classification loss and the regression loss. Through Softmax classification, the probability distribution of different age categories is output. Through expectation regression, the probability distribution is multiplied by the corresponding age label and summed to obtain the age prediction value.
[0049] Specifically, before the face feature vector is input into the prediction layer, the model needs to be trained through the joint loss function. The joint loss function is the weighted sum of the classification loss and the regression loss, that is:
[0050] L cls =L cls +λ·L reg ;
[0051] Wherein L cls is the cross-entropy classification loss, which is used to constrain the probability distribution of the age category, L reg is the mean square error regression loss, which is used to optimize the accuracy of age prediction, and λ is the coefficient. During training, first, the labeled age label is mapped to the preset multiple age categories, such as 0-1 years old, 2-4 years old, etc. After obtaining the probability distribution through Softmax classification, the classification loss is calculated, and at the same time, the probability distribution is multiplied by the median of each age category, such as 0.5 years old for 0-1 years old, and the sum is taken to obtain the expected age value. The regression loss is calculated with the true age. Through back propagation, the two sets of losses are simultaneously optimized, so that the model has both classification robustness and can correct the deviation of the probability distribution through regression, thereby avoiding the problem that the expected value deviates due to a small probability in the 90+ age category.
[0052] Further, when the trained model receives the face feature vector, it is first processed by the Softmax classification unit. This unit contains a fully connected layer with the number of neurons consistent with the number of age categories, which maps the feature vector to category scores through linear transformation, and then normalizes it to a probability distribution P = [p1, p2, …, p10] through the Softmax function, where pi represents the probability that the input face belongs to the i-th age category.
[0053] Further, after obtaining the age category probability distribution, the final age prediction value is calculated through the expectation regression formula. This process forces the expectation of the probability distribution to approximate the true age through the regression loss L reg optimization, forcing the expectation of the probability distribution to approximate the true age, especially for the small probability case of extreme age categories, and constraining its contribution to the expectation through the Mean Squared Error (MSE) loss to avoid prediction bias.
[0054] S103, based on the face information and the age prediction value, constructing a face information set associated with the display terminal.
[0055] In an implementation of the present application, the unique identifier of the display terminal and its geographic location information are obtained. Through the pre-configured device management system, a dedicated ID is allocated to each display terminal. At the same time, a mapping relationship between the display terminal ID and the coverage area is established to facilitate subsequent accurate positioning and information publishing. By detecting the source device identifier of the face image, each face record is associated with the display terminal ID that collected the image, thereby constructing a face information set associated with the display terminal.
[0056] S104, in response to the current emergency instruction, performing face matching in the face information set based on the target portrait information uploaded by the user.
[0057] In an implementation of the present application, the current emergency instruction is received, and the target portrait information uploaded by the user is received. Based on the feature similarity algorithm, the target portrait information is matched with the face information set. In the case of successful matching, the display terminal identifier information associated with the matching result is output, and the emergency response operation is triggered. The face information set at least contains face information, age prediction value, living body detection result and face feature vector.
[0058] Specifically, first, the emergency instruction and the target portrait information are received through a special interface, and the uploaded content is format-verified and preprocessed; then the feature vector of the target portrait is extracted, multi-dimensional similarity calculation is performed with the records in the face information set, a threshold is set to screen the candidate matching items and NMS is used for deduplication. After successful matching, based on the device association information in the face record, the corresponding display terminal identifier is quickly located. Finally, the hierarchical emergency response mechanism is triggered, which not only displays the target information on the associated LED screen, but also links the surrounding monitoring equipment and security system to realize multi-dimensional emergency resource collaborative scheduling. The embodiment of the application integrates face multi-modal features and Internet of Things device networks to improve the response speed and disposal efficiency of emergency events.
[0059] S105, in the case of successful matching, based on the location of the matched display terminal, the location information of the target personnel is published on each display terminal in the preset area.
[0060] In an implementation manner of the application, the geographic location information of the display terminal is determined based on the display terminal identifier information. Based on the geographic location information of the display terminal, the emergency area range is determined, and a plurality of display terminals in the emergency area range are determined. According to the priority of the current emergency instruction, the publishing frequency and display duration of the emergency information on the display terminal are adjusted. Based on the adjusted strategy, the emergency information is displayed on the plurality of display terminals in the emergency area range.
[0061] Specifically, first, the geographic data such as the coordinates and area number of the screen are accurately obtained through the mapping relationship between the device and the geographic location; then the emergency area range is dynamically defined according to the screen position combined with the preset rules, and all LED screens in the area are screened out. At the same time, according to the priority of the emergency instruction, the emergency information publishing strategy is intelligently adjusted: high-priority instructions increase the publishing frequency and prolong the display duration to ensure high-frequency exposure of information; low-priority instructions moderately reduce the frequency and shorten the duration to avoid resource waste. Finally, according to the adjusted strategy, the emergency information is batch-pushed to the display terminals in the emergency area through the Internet of Things communication protocol, realizing differentiated and accurate emergency information publishing, efficiently coordinating resources, and improving emergency response efficiency.
[0062] Figure 2 A structure schematic diagram of a face recognition emergency device based on a display terminal is provided for the embodiment of the application. As shown in FIG. 1, the device mainly includes a display terminal 100, a face recognition module 200, a communication module 300, a storage module 400 and a power supply module 500. Figure 2As shown, the face recognition emergency equipment based on the display terminal 200 comprises: at least one processor 201; and a memory 202 in communication connection with the at least one processor 201; wherein the memory 202 stores instructions executable by the at least one processor 201, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to: acquire a face image appearing in a display terminal, detect the face image based on a multi-layer cascade face detection model group, and output face information; wherein the face information at least comprises face position and face key point information; perform silent liveness detection on the face image through a liveness detection fusion model, in the case of passing the detection, perform feature extraction on the face image through a preset convolutional network, and output an age prediction value based on the extracted features; based on the face information and the age prediction value, construct a face information set associated with the display terminal; in response to a current emergency instruction, perform face matching in the face information set based on target portrait information uploaded by a user; in the case of successful matching, publish position information of a target person on each display terminal in a preset area based on a position of the matched display terminal.
[0063] The non-volatile computer storage medium provided by the embodiments of the present application stores computer executable instructions, and the computer executable instructions are configured to: acquire a face image appearing in a display terminal, detect the face image based on a multi-layer cascade face detection model group, and output face information; wherein the face information at least comprises face position and face key point information; perform silent liveness detection on the face image through a liveness detection fusion model, in the case of passing the detection, perform feature extraction on the face image through a preset convolutional network, and output an age prediction value based on the extracted features; based on the face information and the age prediction value, construct a face information set associated with the display terminal; in response to a current emergency instruction, perform face matching in the face information set based on target portrait information uploaded by a user; in the case of successful matching, publish position information of a target person on each display terminal in a preset area based on a position of the matched display terminal.
[0064] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts of each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. Especially, the device, equipment and non-volatile computer storage medium embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0065] The above only describes the embodiments of the present application and is not used to limit the present application. The embodiments of the present application can be variously changed and modified by those skilled in the art. The modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A face recognition emergency method based on a display terminal, characterized in that: The method comprises: Acquire a face image that appears in the display terminal, detect the face image based on a multi-layer cascade face detection model group, and output face information; wherein the face information includes at least face position and face key point information; Performing silent liveness detection on the facial image using a liveness detection fusion model; if the detection passes, extracting features from the facial image using a preset convolutional network, and outputting an age prediction value based on the extracted features; constructing a face information set associated with the display terminal based on the face information and the age prediction value; In response to the current emergency instruction, based on the target portrait information uploaded by the user, face matching is performed in the face information set; In the case of a successful match, the location information of the target person is published on each display terminal within the preset area based on the location of the matched display terminal.
2. The face recognition emergency method based on a display terminal according to claim 1, characterized in that: The multi-layer cascade-based face detection model group detects the face image and outputs face information, specifically including: performing a multi-scale transformation on the face image at an input layer of the multi-layer cascade face detection model group to generate an image pyramid containing images of different scales; Inputting the image pyramid into the first-level sub-network to generate multiple candidate target region boxes; Input the candidate target area frame into the second-level sub-network, perform initial screening and bounding box regression on the candidate target area frame based on preset screening conditions, and filter out negative example areas with confidence levels lower than a threshold; The filtered candidate target area frame is input into the third-level sub-network for secondary discrimination and frame regression, and the final face detection result and face information are output.
3. The face recognition emergency method based on a display terminal according to claim 1, characterized in that: The performing silent liveness detection on the face image using the liveness detection fusion model specifically includes: Aggregating intensity information and gradient information of the face image using the liveness detection fusion model to obtain fine-grained features of the face; Inputting the fine-grained facial features into a preset detector to perform silent living body detection on the facial image; Among them, the liveness detection fusion model is set as a cascade network structure, and the liveness detection fusion model consists of a feature extraction layer based on center differential convolution, a network layer combined with neural architecture search optimization, and a multi-scale attention fusion module.
4. The face recognition emergency method based on a display terminal according to claim 3, characterized in that: The method of aggregating intensity information and gradient information of the face image using the liveness detection fusion model to obtain fine-grained facial features specifically includes: Through the feature extraction layer based on center difference convolution, intensity information and gradient information are extracted from the face image. fusing the intensity information and the gradient information based on an adaptive weight to obtain a preliminary feature map; Determining network architecture parameters required for the liveness detection task corresponding to the preliminary feature map by searching the optimized network layer using the neural architecture, and optimizing the preliminary feature map using the network architecture parameters; Through the multi-scale attention fusion module, a multi-scale facial feature map corresponding to the optimized feature map is generated, the multi-scale facial feature map is processed in parallel, and the multi-scale feature map is adaptively weighted fused through the attention mechanism to obtain the fine-grained facial features.
5. The face recognition emergency method based on a display terminal according to claim 1, characterized in that: The step of extracting features from the face image using a preset convolutional network and outputting an age prediction value based on the extracted features specifically includes: Based on the facial information, performing face alignment processing on the facial image; Inputting the aligned and processed facial image into the backbone convolutional neural network MobileFaceNets; Among them, the backbone convolutional neural network includes a network architecture corresponding to the fast downsampling and dimensionality reduction strategy, and a 1×1 linear convolution layer is connected after the global depth convolution layer to output the facial feature vector.
6. The face recognition emergency method based on a display terminal according to claim 1, characterized in that: Outputting an age prediction value based on the extracted features specifically includes: Inputting the facial feature vector into a prediction layer; wherein the prediction layer performs model training based on a joint loss function of classification loss and regression loss, wherein the joint loss function is a weighted sum of the classification loss and the regression loss; Through Softmax classification, the probability distribution of different age categories is output; The age prediction value is obtained by multiplying and summing the probability distribution and the corresponding age label through expected regression.
7. The face recognition emergency method based on a display terminal according to claim 1, characterized in that: The step of responding to the current emergency instruction and performing face matching in the face information set based on the target portrait information uploaded by the user specifically includes: Receive the current emergency instruction and receive the target portrait information uploaded by the user; Matching the target portrait information with the face information set based on a feature similarity algorithm; In the event of a successful match, the display terminal identification information associated with the matching result is output and an emergency response operation is triggered; The face information set includes at least face information, age prediction value, liveness detection result and face feature vector.
8. The face recognition emergency method based on a display terminal according to claim 7, characterized in that: The method of publishing the target person's location information on each display terminal within a preset area based on the location of the matched display terminal specifically includes: determining the geographical location information of the display terminal based on the display terminal identification information; Determining an emergency area based on the geographical location information of the display terminal, and determining a plurality of display terminals within the emergency area; Adjusting the frequency and duration of display of the emergency information on the display terminal according to the priority of the current emergency instruction; Based on the adjusted strategy, the emergency information is displayed on multiple display terminals within the emergency area.
9. A face recognition emergency device based on a display terminal, characterized in that: The device comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method, device and system for intelligent person search
CN107295294A
Method and device for human face identification
CN108197542A
An intelligent railway station people flow monitoring system and method
CN109948550A
Scenic spot people searching information publishing method and system and storage medium
CN110633656A
Scenic spot people searching system based on video detection
CN112163568A