Face image acquisition system
By designing a face image acquisition system, using individual and group abnormality analysis modules, combined with camera acquisition and deep learning technology, the accuracy of face recognition and behavior analysis in high-density crowd scenarios is solved, and efficient identification and judgment of abnormal people and behaviors is achieved.
Patent Information
- Application Number
- CN202411840690.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-09
AI Technical Summary
In high-density crowd scenarios, existing facial recognition and behavioral analysis technologies are difficult to accurately identify targets, making it difficult to quickly and accurately extract effective evidence from large amounts of video data in high-security places, and lock in illegal personnel or abnormal behaviors.
A face image acquisition system is designed, including an image acquisition module, an individual abnormality analysis module, a population abnormality analysis module and a comprehensive abnormality analysis module. Video image frames are collected by the camera, and the individual anomaly analysis model is used for preprocessing, feature extraction, fusion and processing, to obtain the first anomaly coefficient; the cluster density and movement consistency are analyzed through the clustering algorithm and the YOLO algorithm to obtain the second anomaly coefficient; finally, the abnormality category judgment is performed by comparing the comprehensive anomaly coefficient with the preset threshold.
It realizes efficient identification and judgment of abnormal personnel in image acquisition data, can accurately identify abnormal situations of individuals and groups, and improves the efficiency of locking illegal personnel and judging abnormal behaviors in high-security places.
Smart Images

Figure CN119964213A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face image acquisition, and in particular to a face image acquisition system. Background Art
[0002] With the development of artificial intelligence technology and the increasing maturity of face recognition technology, face image acquisition systems have been widely used in many fields, such as security monitoring, identity authentication, attendance systems, and mobile payments. These systems collect face images through cameras and use image processing and analysis algorithms to perform identity recognition, behavior monitoring, and data storage. In today's society, more and more scenarios rely on face image acquisition systems to improve security and convenience.
[0003] In the current monitoring mode, the identification personnel need to quickly analyze the faces and movements of the people in the monitoring video to achieve the effect of trial. However, in the scene of high-density crowds, the existing face recognition and behavior analysis technology often cannot accurately identify the target. This makes it difficult for the relevant identification personnel in gold shops, banks and other high-security places to quickly and accurately extract effective evidence from a large amount of video data, making it difficult to accurately identify illegal persons or abnormal behaviors.
[0004] Therefore, a face image acquisition system is proposed. Summary of the invention
[0005] The purpose of the present invention is to provide a face image acquisition system. The present invention relates to the field of face image acquisition technology, specifically a face image acquisition system, including an image acquisition module, an individual anomaly analysis module, a group anomaly analysis module and a comprehensive anomaly analysis module; first, each video image frame of a gold shop is obtained; secondly, each video image frame is input into an individual anomaly analysis model, and is analyzed and processed in turn through the preprocessing layer, feature extraction layer, feature fusion layer and feature processing layer of the model, and a first anomaly coefficient is obtained at the output layer; then, each detection video image frame is analyzed by a clustering algorithm to obtain a comprehensive group density and a comprehensive group movement consistency, and a second anomaly coefficient is obtained by the comprehensive group density and the comprehensive group movement consistency; finally, the first anomaly coefficient and the second anomaly coefficient are analyzed to obtain a comprehensive anomaly coefficient, and compared with a preset anomaly threshold, to obtain an abnormal category of each detection video image frame. The present invention is applied to a gold shop, and can realize efficient identification and determination of abnormal personnel in image acquisition data.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A facial image acquisition system, characterized in that it comprises:
[0008] An image acquisition module is used to shoot the interior and surrounding of the gold shop through a camera to obtain a first video of the gold shop, and divide the first video into frames according to a preset time interval to obtain individual video image frames;
[0009] The individual anomaly analysis module is used to input each video image frame into the individual anomaly analysis model, and sequentially pass through the preprocessing layer, feature extraction layer, feature fusion layer, feature processing layer and output layer of the model; the preprocessing layer preprocesses each video image frame to obtain each detection video image frame; the feature extraction layer extracts features from each detection video image frame to obtain a first face feature and a second behavior feature; and obtains a first comprehensive feature through the feature fusion layer; the feature processing layer analyzes the first comprehensive feature through an LSTM network structure that incorporates an attention mechanism to obtain a first comprehensive time series feature; and analyzes the first comprehensive time series feature to obtain a second comprehensive feature; the output layer analyzes the second comprehensive feature to obtain a first anomaly coefficient;
[0010] A group anomaly analysis module, used for analyzing each of the detection video image frames by a clustering algorithm to obtain a comprehensive group density and a comprehensive group movement consistency, and obtaining a second anomaly coefficient by the comprehensive group density and the comprehensive group movement consistency;
[0011] The comprehensive abnormality analysis module analyzes the first abnormality coefficient and the second abnormality coefficient to obtain a comprehensive abnormality coefficient, and compares it with a preset abnormality threshold to obtain the abnormality category of each detected video image frame.
[0012] Preferably, the first facial features include expression features, facial posture and gaze direction; the second behavioral features include gait features, gesture features and body posture;
[0013] The first facial feature is:
[0014] F face (i,j)={f face,i ,f face,2 ,...f face,M};
[0015] Among them, F face (i,j) represents the first facial feature of the jth individual in the i-th detected video image frame; M is the number of first facial features;
[0016] The second behavior feature is:
[0017] F behavior (i,j)={f behavior,i ,f behavior,2 ,...f behavior,N};
[0018] Among them, F behavior (i, j) represents the second behavior feature of the jth individual in the i-th detected video image frame; N is the number of second behavior features.
[0019] Preferably, the first comprehensive feature is:
[0020] F combined (i,j)=concat(F face (i,j),F behavior (i,j));
[0021] Among them, F combined (i,j) represents the first comprehensive feature of the jth individual in the i-th detected video image frame; concat() represents a vector concatenation operation.
[0022] Preferably, the second comprehensive feature is:
[0023]
[0024] F optimized (i,j)=ReLu(W*LSTM(F combined (i-1,j),F combined (i,j))+b);
[0025] Among them, F attention (i,j) represents the second comprehensive feature of the jth individual in the i-th detected video image frame; γ k (i, j) represents the weight calculated by the attention mechanism, reflecting the importance of each feature; k represents the comprehensive feature index; F k optimized (i, j) represents the first comprehensive time series feature; RELU() represents the activation function; W represents the weight matrix; LSTM() represents the time series feature relationship function, which is used to extract the time dependency of the first comprehensive feature between the previous and next frames; b represents the bias term.
[0026] Preferably, the first abnormal coefficient is:
[0027]
[0028] Among them, R first (i,j) represents the first abnormal coefficient of the jth individual in the i-th detected video image frame; k represents the comprehensive feature index; M is the number of first facial features; N is the number of second behavioral features; ω k represents the abnormal weight of the kth component; F attention (i, j) represents the second comprehensive feature of the jth individual in the i-th detected video image frame.
[0029] Preferably, the comprehensive population density is:
[0030] D over_density =max(D 1 density (i),D 2 density (i),D 3 density (i),...D k density (i));
[0031]
[0032] Among them, D over_density (i) is the comprehensive population density of the i-th detection video image frame; D k density (i) represents the population density of the kth population in the i-th video image frame; R k represents the total number of individuals in the kth group of the i-th detection video image frame; (x j (i),y j (i)) represents the two-dimensional coordinates of the jth individual in the kth group in the i-th detection video image frame; (x u (i),y u (i)) represents the two-dimensional coordinates of the u-th individual in the k-th group in the i-th detection video image frame; Dist() represents the Euclidean distance function.
[0033] Preferably, the calculation formula for the comprehensive group movement consistency is:
[0034] V over_sync (i) = min(V 1 sync (i),V 2 sync (i),V 3 sync (i),...,V k sync (i));
[0035]
[0036] Among them, V over_sync (i) represents the comprehensive group movement consistency of the i-th detection video image frame; V k sync (i) represents the movement consistency of the kth group in the i-th detected video image frame; θ k j (i) represents the angle between the moving speed vector of the jth individual in the kth group and the average moving speed vector of the kth group; Vel kj (i) represents the moving speed vector of the jth individual in the kth group in the i-th detection video image frame; Vel k avg (i) represents the average moving speed vector of the kth group in the i-th detection video image frame; R k represents the total number of individuals in the kth group; x k j (i) represents the horizontal coordinate of the jth individual in the kth group in the i-th detection video image frame; y k j (i) represents the ordinate of the jth individual in the kth group in the ith detection video image frame; Δt represents the time interval between the ith detection video image frame and the i-1th detection video image frame.
[0037] Preferably, the second abnormal coefficient is:
[0038] R second (i) = ξ density *D over_density (i)+ψ sync *V over_sync (i);
[0039] Among them, R second (i) represents the second abnormal coefficient of the i-th detected video image frame; ξ density represents the comprehensive population density weight; ψ sync Represents the comprehensive movement consistency weight.
[0040] Preferably, the comprehensive abnormality coefficient is:
[0041]
[0042] Among them, R com (i) represents the comprehensive abnormality coefficient of the i-th detected video image frame; R first (i,j) represents the first abnormal coefficient of the jth individual in the i-th detected video image frame; R represents the number of individuals in the i-th detected video image frame; φ fir represents the first abnormal coefficient weight; δ se Represents the second abnormal coefficient weight; R second (i) represents the second abnormality coefficient of the i-th detected video image frame.
[0043] Preferably, the comprehensive abnormality coefficient is compared with a first preset abnormality coefficient threshold, and the detected video image frame is classified into abnormal categories, and the abnormal categories include abnormal video image frames and normal video image frames.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] 1. The present invention obtains a first video of a gold shop, divides the first video into frames according to a preset time interval, obtains each video image frame, and constructs an individual anomaly analysis model, inputs each video image frame into the model, and sequentially passes through the preprocessing layer, feature extraction layer, feature fusion layer, feature processing layer and output layer of the model; preprocesses the first video image frame through the preprocessing layer to obtain each detection video image frame; and extracts features from each detection video image frame through the feature extraction layer to obtain a first face feature and a second behavior feature; and uses the feature fusion layer to fuse the first face feature and the second behavior feature to obtain a first comprehensive feature; the feature processing layer analyzes the first comprehensive feature according to the LSTM network structure integrated with the attention mechanism to obtain a first comprehensive time series feature; and analyzes the first comprehensive time series feature to obtain a second comprehensive feature; finally, the second comprehensive feature is analyzed through the output layer to obtain a first anomaly coefficient; by constructing the individual anomaly analysis model, the individual anomalies of each video image frame in the image acquisition data can be effectively analyzed, which is conducive to improving the efficiency of identifying and judging illegal personnel.
[0046] 2. The present invention analyzes each detection video image frame through the YOLO algorithm and the clustering algorithm to obtain the comprehensive group density and comprehensive group movement consistency of each detection video image frame. The abnormal crowd gathering situation can be effectively identified through the comprehensive group density, and the abnormal group behavior can be effectively identified according to the comprehensive group movement consistency. Through the introduction of these two comprehensive factors, the second abnormal coefficient is obtained, the abnormal characteristics of the group can be reasonably captured, and the abnormal group situation of each video image frame in the image acquisition data can be effectively analyzed, so as to improve the efficiency of identifying and judging illegal personnel.
[0047] 3. The present invention obtains a comprehensive abnormality coefficient by comprehensively considering the first abnormality coefficient and the second abnormality coefficient. The comprehensive abnormality coefficient can comprehensively evaluate when illegal behavior occurs and identify the individual abnormalities of each video image frame in the image acquisition data, thereby realizing abnormal classification of the image acquisition data and efficiently identifying abnormal video image frames and non-abnormal video image frames, which is beneficial to improving the efficiency of identifying and judging illegal persons. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A structural diagram of a facial image acquisition system provided by an embodiment of the present invention;
[0049] Figure 2 An individual abnormality analysis model diagram provided by an embodiment of the present invention;
[0050] Figure 3 A second abnormal coefficient calculation flow chart provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] With the development of artificial intelligence technology and the increasing maturity of face recognition technology, face image acquisition systems have been widely used in many fields, such as security monitoring, identity authentication, attendance systems, and mobile payments. These systems collect face images through cameras and use image processing and analysis algorithms to perform identity recognition, behavior monitoring, and data storage. In today's society, more and more scenarios rely on face image acquisition systems to improve security and convenience.
[0053] In the case of densely populated scenes, existing face recognition and behavior analysis technologies often fail to accurately identify targets. This makes it difficult for relevant identification personnel to quickly and accurately extract effective evidence from a large amount of video data in gold shops, banks, and other high-security places, making it difficult to accurately identify illegal persons or abnormal behaviors. The present invention will be described in detail with the help of the following two embodiments to demonstrate that the system proposed in the present invention has certain utility in illegal anomaly identification.
[0054] Embodiment 1
[0055] Gold Shop A is a store that sells gold. It has recently been hit by illegal activities. In order to improve the efficient screening and judgment of illegal and abnormal persons, the relevant identification personnel applied a facial image acquisition system;
[0056] A face image acquisition system, the structure of which is as follows Figure 1 As shown, including:
[0057] An image acquisition module is used to shoot the interior and surroundings of the gold shop A through a camera to obtain a first video of the gold shop, and to divide the first video into frames according to a preset time interval to obtain each video image frame; in this embodiment, the preset time interval is 0.1s;
[0058] The individual anomaly analysis module is used to input each video image frame into the individual anomaly analysis model, and sequentially pass through the preprocessing layer, feature extraction layer, feature fusion layer, feature processing layer and output layer of the model; the individual anomaly analysis model is shown in FIG. Figure 2 As shown;
[0059] The preprocessing layer preprocesses each of the video image frames to obtain each detection video image frame; the preprocessing includes denoising and contrast enhancement;
[0060] The feature extraction layer first detects the face area of each detection video image frame through the MTCNN algorithm, and performs feature extraction to obtain the first face feature and the second behavior feature;
[0061] Furthermore, the first facial features include expression features, facial posture and gaze direction; the second behavioral features include gait features, gesture features and body posture;
[0062] The first facial feature is:
[0063] F face (i,j)={f face,i ,f face,2 ,...f face,M};
[0064] Among them, F face (i, j) represents the first facial feature of the jth individual in the i-th detected video image frame; M is the number of first facial features; in this embodiment, M is 3; respectively represent expression features, facial posture and gaze direction;
[0065] The expression features are acquired through a pre-trained expression recognition model to obtain a high-dimensional feature vector, and the high-dimensional feature vector is reduced in dimension through principal component analysis to obtain a 25*3 feature vector;
[0066] The vector representation of the facial expression feature is:
[0067]
[0068] Among them, f face,1 is the vector representation of facial expression features; n is 25;
[0069] Facial posture uses a 3D posture estimation model to obtain the rotation angle of the head in three-dimensional space; gaze direction indicates the gaze direction of the eyes in three-dimensional space;
[0070] The vector representation of the facial posture is:
[0071]
[0072] Among them, f face,2 is the vector representation of facial posture; θ pitch represents the pitch angle, indicating the up and down rotation of the head; θ yaw represents the yaw angle, indicating the left-right rotation of the head; θ pitch represents the roll angle, which indicates the left-right tilt of the head;
[0073] The vector of the gaze direction is expressed as:
[0074]
[0075] Among them, f face,3 G is the vector representation of the gaze direction; x ,G y ,G z Represent the components of the gaze direction on the x, y, and z axes respectively;
[0076] The second behavior feature is:
[0077] F behavior (i,j)={f behavior,i ,f behavior,2 ,...f behavior,N};
[0078] Among them, F behavior (i, j) represents the second behavior feature of the jth individual in the i-th detected video image frame; N is the number of second behavior features.
[0079] Wherein, the second behavior feature includes gait feature, gesture feature and body posture;
[0080] The gait feature is represented by a vector formed by the spatial position of the ankle joint of the individual in the video image frame; the gesture feature is represented by a vector formed by the spatial position of the palm center of the individual in the video image frame; the body posture is represented by a vector formed by the spatial position of the key points of the body (shoulders, elbows and knees);
[0081] The vector representation of the gait feature is:
[0082]
[0083] Among them, F gait A vector representation representing the gait feature; x left_ankle ,y left_ankle ,z left_ankle Respectively represent the three-dimensional space coordinates of the left ankle joint; x right_ankle ,y right_ankle ,z right_ankle Respectively represent the three-dimensional spatial coordinates of the right ankle joint;
[0084] The vector representation of the gesture feature is:
[0085]
[0086] Among them, F gesture A vector representation representing the gait feature; x left_palm ,y left_palm ,z left_palm They represent the three-dimensional coordinates of the center of the left palm; xleft_palm ,y left_palm ,z left_palm Respectively represent the three-dimensional space coordinates of the center of the right palm;
[0087] The vector representation of the body posture is:
[0088]
[0089] Among them, F posture A vector representation representing the gait feature; x left_shoulder ,y left_shoulder ,z left_shoulder Respectively represent the three-dimensional space coordinates of the left shoulder joint; right_shoulder ,y right_shoulder ,z right_shoulder Respectively represent the three-dimensional space coordinates of the right shoulder joint; x left_elbow ,y left_elbow ,z left_elbow Respectively represent the three-dimensional space coordinates of the left elbow joint; right_elbow ,y right_elbow ,z right_elbow Respectively represent the three-dimensional space coordinates of the right elbow joint; x left_knee ,y left_knee ,z left_knee Respectively represent the three-dimensional spatial coordinates of the left knee joint; right_knee ,y right_knee ,z right_knee Respectively represent the three-dimensional spatial coordinates of the right knee joint;
[0090] Furthermore, the feature fusion layer performs feature fusion on the first face feature and the second behavior feature to obtain a first comprehensive feature; the feature processing layer analyzes the first comprehensive feature through an LSTM network structure incorporating an attention mechanism to obtain a first comprehensive time series feature; and analyzes the first comprehensive time series feature to obtain a second comprehensive feature; the output layer analyzes the second comprehensive feature to obtain a first abnormality coefficient;
[0091] Furthermore, the first comprehensive feature is:
[0092] F combined (i,j)=concat(F face (i,j),F behavior (i,j));
[0093] Among them, F combined (i,j) represents the first comprehensive feature of the jth individual in the i-th detected video image frame; concat() represents a vector concatenation operation.
[0094] Furthermore, the second comprehensive feature is:
[0095]
[0096] F optimized (i,j)=ReLu(W*LSTM(F combined (i-1,j),F combined (i,j))+b);
[0097] Among them, F attention (i,j) represents the second comprehensive feature of the jth individual in the i-th detected video image frame; γ k (i, j) represents the weight calculated by the attention mechanism, reflecting the importance of each feature; k represents the comprehensive feature index; L1 represents the first face feature dimension; L2 represents the first face feature dimension; k represents the comprehensive feature dimension index; F k optimized (i,j) represents the k-th dimensional row vector of the first comprehensive time series feature; RELU() represents the activation function; W represents the weight matrix; LSTM() represents the time series feature relationship function, which is used to extract the time dependency of the first comprehensive feature between the previous and next frames; b represents the bias term.
[0098] Furthermore, the first abnormal coefficient is:
[0099]
[0100] Among them, R first (i,j) represents the first abnormal coefficient of the jth individual in the i-th detected video image frame; k represents the comprehensive feature index; ω k represents the abnormal weight of the kth component; F attention (i, j) represents the second comprehensive feature of the jth individual in the i-th detected video image frame.
[0101] In this embodiment, a first video of a gold shop is obtained, and the first video is divided into frames according to a preset time interval to obtain each video image frame, and an individual anomaly analysis model is constructed, and each video image frame is input into the model, and passes through the preprocessing layer, feature extraction layer, feature fusion layer, feature processing layer and output layer of the model in sequence; the first video image frame is preprocessed by the preprocessing layer to obtain each detection video image frame; and the feature extraction layer is used to extract features from each detection video image frame to obtain a first facial feature and a second behavioral feature; and the feature fusion layer is used to fuse the first facial feature and the second behavioral feature to obtain a first comprehensive feature; the feature processing layer analyzes the first comprehensive feature according to the LSTM network structure integrated with the attention mechanism to obtain a first comprehensive time series feature; and the first comprehensive time series feature is analyzed to obtain a second comprehensive feature; finally, the second comprehensive feature is analyzed by the output layer to obtain a first anomaly coefficient; by constructing an individual anomaly analysis model, the individual anomalies of each video image frame in the image acquisition data can be effectively analyzed, which is conducive to improving the efficiency of identifying and judging illegal persons.
[0102] A group anomaly analysis module, used to analyze each of the detection video image frames through the YOLO algorithm and the clustering algorithm to obtain a comprehensive group density and a comprehensive group movement consistency, and obtain a second anomaly coefficient through the comprehensive group density and the comprehensive group movement consistency;
[0103] The calculation process of the second abnormal coefficient is as follows: Figure 3 As shown;
[0104] The YOLO algorithm is used to detect the positions of all individuals in each detection video image frame; extract the spatial coordinates of each individual; cluster the detected individuals and classify the close individuals into a group; calculate the density of each group; the density reflects the degree of aggregation between individuals; and calculate the group movement consistency, which is used to measure the consistency of the movement direction and speed of individuals in each group, reflecting the coordination of individuals in the group;
[0105] Furthermore, the comprehensive population density is:
[0106] D over_density =max(D 1 density (i),D 2 density (i),D 3 density (i),...D k density (i));
[0107]
[0108] Among them, Dover_density (i) is the comprehensive population density of the i-th detection video image frame; D k density (i) represents the population density of the kth population in the i-th video image frame; R k represents the total number of individuals in the kth group of the i-th detection video image frame; (x j (i),y j (i)) represents the two-dimensional coordinates of the jth individual in the kth group in the i-th detection video image frame; (x u (i),y u (i)) represents the two-dimensional coordinates of the u-th individual in the k-th group in the i-th detection video image frame; Dist() represents the Euclidean distance function.
[0109] Furthermore, the calculation formula for the comprehensive group movement consistency is:
[0110] V over_sync (i) = min(V 1 sync (i),V 2 sync (i),V 3 sync (i),...,V k sync (i));
[0111]
[0112] Among them, V over_sync (i) represents the comprehensive group movement consistency of the i-th detection video image frame; V k sync (i) represents the movement consistency of the kth group in the i-th detected video image frame; θ k j (i) represents the angle between the moving speed vector of the jth individual in the kth group and the average moving speed vector of the kth group; Vel k j (i) represents the moving speed vector of the jth individual in the kth group in the i-th detection video image frame; Vel k avg (i) represents the average moving speed vector of the kth group in the i-th detection video image frame; R k represents the total number of individuals in the kth group; x k j (i) represents the horizontal coordinate of the jth individual in the kth group in the i-th detection video image frame; y k j(i) represents the ordinate of the jth individual in the kth group in the ith detection video image frame; Δt represents the time interval between the ith detection video image frame and the i-1th detection video image frame.
[0113] Furthermore, the second abnormal coefficient is:
[0114] R second (i) = ξ density *D over_density (i)+ψ sync *V over_sync (i);
[0115] Among them, R second (i) represents the second abnormal coefficient of the i-th detected video image frame; ξ density represents the comprehensive population density weight; ψ sync Represents the comprehensive movement consistency weight.
[0116] This embodiment uses the YOLO algorithm and clustering algorithm to analyze each detection video image frame to obtain the comprehensive group density and comprehensive group movement consistency of each detection video image frame. The comprehensive group density can effectively identify abnormal crowd gatherings, and the comprehensive group movement consistency can effectively identify abnormal group behavior. By introducing these two comprehensive factors, the second abnormal coefficient is obtained, which can reasonably capture the abnormal characteristics of the group and effectively analyze the group abnormality of each video image frame in the image acquisition data, thereby improving the efficiency of identifying and judging illegal personnel.
[0117] The comprehensive abnormality analysis module analyzes the first abnormality coefficient and the second abnormality coefficient to obtain a comprehensive abnormality coefficient, and compares it with a preset abnormality threshold to obtain the abnormality category of each detected video image frame.
[0118] Furthermore, the comprehensive abnormality coefficient is:
[0119]
[0120] Among them, R com (i) represents the comprehensive abnormality coefficient of the i-th detected video image frame; R first (i,j) represents the first abnormal coefficient of the jth individual in the i-th detected video image frame; R represents the number of individuals in the i-th detected video image frame; φ fir represents the first abnormal coefficient weight; δ se Represents the second abnormal coefficient weight; R second (i) represents the second abnormality coefficient of the i-th detected video image frame.
[0121] Furthermore, the comprehensive abnormality coefficient is compared with a first preset abnormality coefficient threshold, and the detected video image frame is classified into abnormal categories, where the abnormal categories include abnormal video image frames and normal video image frames.
[0122] This embodiment obtains a comprehensive abnormality coefficient by comprehensively considering the first abnormality coefficient and the second abnormality coefficient. The comprehensive abnormality coefficient can comprehensively evaluate when illegal behavior occurs and identify the individual abnormalities of each video image frame in the image acquisition data, thereby realizing abnormal classification of the image acquisition data and efficiently identifying abnormal video image frames and non-abnormal video image frames, which is beneficial to improving the efficiency of identifying and judging illegal persons.
[0123] Embodiment 2
[0124] Gold Shop B is another store that sells gold. It has also been the target of illegal activities recently. In order to improve the efficiency of identifying illegal and abnormal persons, the relevant identification personnel applied a facial image acquisition system.
[0125] A face image acquisition system, the structure of which is as follows Figure 1 As shown, including:
[0126] An image acquisition module is used to shoot the interior and surroundings of the gold shop B through a camera to obtain a first video of the gold shop, and to divide the first video into frames according to a preset time interval to obtain individual video image frames; in this embodiment, the preset time interval is 0.1s;
[0127] The individual anomaly analysis module is used to input each video image frame into the individual anomaly analysis model, and sequentially pass through the preprocessing layer, feature extraction layer, feature fusion layer, feature processing layer and output layer of the model; the individual anomaly analysis model is shown in FIG. Figure 2 As shown;
[0128] The preprocessing layer preprocesses each of the video image frames to obtain each detection video image frame; the preprocessing includes denoising and contrast enhancement;
[0129] The feature extraction layer first detects the face area of each detection video image frame through the MTCNN algorithm, and performs feature extraction to obtain the first face feature and the second behavior feature;
[0130] Furthermore, the first facial features include expression features, facial posture and gaze direction; the second behavioral features include gait features, gesture features and body posture;
[0131] The first facial feature is:
[0132] F face (i,j)={f face,i ,fface,2 ,...f face,M};
[0133] Among them, F face (i, j) represents the first facial feature of the jth individual in the i-th detected video image frame; M is the number of first facial features; in this embodiment, M is 3; respectively represent expression features, facial posture and gaze direction;
[0134] The expression features are acquired through a pre-trained expression recognition model to obtain a high-dimensional feature vector, and the high-dimensional feature vector is reduced in dimension through principal component analysis to obtain a 25*3 feature vector;
[0135] The vector representation of the facial expression feature is:
[0136]
[0137] Among them, f face,1 is the vector representation of facial expression features; n is 25;
[0138] Facial posture uses a 3D posture estimation model to obtain the rotation angle of the head in three-dimensional space; gaze direction indicates the gaze direction of the eyes in three-dimensional space;
[0139] The vector representation of the facial posture is:
[0140]
[0141] Among them, f face,2 is the vector representation of facial posture; θ pitch represents the pitch angle, indicating the up and down rotation of the head; θ yaw represents the yaw angle, indicating the left-right rotation of the head; θ pitch represents the roll angle, which indicates the left-right tilt of the head;
[0142] The vector of the gaze direction is expressed as:
[0143]
[0144] Among them, f face,3 G is the vector representation of the gaze direction; x ,G y ,G z Represent the components of the gaze direction on the x, y, and z axes respectively;
[0145] The second behavior feature is:
[0146] F behavior (i,j)={f behavior,i ,f behavior,2 ,...f behavior,N};
[0147] Among them, F behavior (i, j) represents the second behavior feature of the jth individual in the i-th detected video image frame; N is the number of second behavior features.
[0148] Wherein, the second behavior feature includes gait feature, gesture feature and body posture;
[0149] The gait feature is represented by a vector formed by the spatial position of the ankle joint of the individual in the video image frame; the gesture feature is represented by a vector formed by the spatial position of the palm center of the individual in the video image frame; the body posture is represented by a vector formed by the spatial position of the key points of the body (shoulders, elbows and knees);
[0150] The vector representation of the gait feature is:
[0151]
[0152] Among them, F gait A vector representation representing the gait feature; x left_ankle ,y left_ankle ,z left_ankle Respectively represent the three-dimensional space coordinates of the left ankle joint; x right_ankle ,y right_ankle ,z right_ankle Respectively represent the three-dimensional spatial coordinates of the right ankle joint;
[0153] The vector representation of the gesture feature is:
[0154]
[0155] Among them, F gesture A vector representation representing the gait feature; x left_palm ,y left_palm ,z left_palm They represent the three-dimensional coordinates of the center of the left palm; x left_palm ,y left_palm ,z left_palm Respectively represent the three-dimensional space coordinates of the center of the right palm;
[0156] The vector representation of the body posture is:
[0157]
[0158] Among them, F posture A vector representation representing the gait feature; x left_shoulder ,y left_shoulder ,z left_shoulder Respectively represent the three-dimensional space coordinates of the left shoulder joint; right_shoulder ,yright_shoulder ,z right_shoulder Respectively represent the three-dimensional space coordinates of the right shoulder joint; x left_elbow ,y left_elbow ,z left_elbow Respectively represent the three-dimensional space coordinates of the left elbow joint; right_elbow ,y right_elbow ,z right_elbow Respectively represent the three-dimensional space coordinates of the right elbow joint; x left_knee ,y left_knee ,z left_knee Respectively represent the three-dimensional spatial coordinates of the left knee joint; right_knee ,y right_knee ,z right_knee Respectively represent the three-dimensional spatial coordinates of the right knee joint;
[0159] Furthermore, the feature fusion layer performs feature fusion on the first face feature and the second behavior feature to obtain a first comprehensive feature; the feature processing layer analyzes the first comprehensive feature through an LSTM network structure incorporating an attention mechanism to obtain a first comprehensive time series feature; and analyzes the first comprehensive time series feature to obtain a second comprehensive feature; the output layer analyzes the second comprehensive feature to obtain a first abnormality coefficient;
[0160] Furthermore, the first comprehensive feature is:
[0161] F combined (i,j)=concat(F face (i,j),F behavior (i,j));
[0162] Among them, F combined (i,j) represents the first comprehensive feature of the jth individual in the i-th detected video image frame; concat() represents a vector concatenation operation.
[0163] Furthermore, the second comprehensive feature is:
[0164]
[0165] F optimized (i,j)=ReLu(W*LSTM(F combined (i-1,j),F combined (i,j))+b);
[0166] Among them, F attention (i,j) represents the second comprehensive feature of the jth individual in the i-th detected video image frame; γ k(i, j) represents the weight calculated by the attention mechanism, reflecting the importance of each feature; k represents the comprehensive feature index; L1 represents the first face feature dimension; L2 represents the first face feature dimension; k represents the comprehensive feature dimension index; F k optimized (i,j) represents the k-th dimensional row vector of the first comprehensive time series feature; RELU() represents the activation function; W represents the weight matrix; LSTM() represents the time series feature relationship function, which is used to extract the time dependency of the first comprehensive feature between the previous and next frames; b represents the bias term.
[0167] Furthermore, the first abnormal coefficient is:
[0168]
[0169] Among them, R first (i,j) represents the first abnormal coefficient of the jth individual in the i-th detected video image frame; k represents the comprehensive feature index; ω k represents the abnormal weight of the kth component; F attention (i, j) represents the second comprehensive feature of the jth individual in the i-th detected video image frame.
[0170] A group anomaly analysis module, used to analyze each of the detection video image frames through the YOLO algorithm and the clustering algorithm to obtain a comprehensive group density and a comprehensive group movement consistency, and obtain a second anomaly coefficient through the comprehensive group density and the comprehensive group movement consistency;
[0171] The calculation process of the second abnormal coefficient is as follows: Figure 3 As shown;
[0172] The YOLO algorithm is used to detect the positions of all individuals in each detection video image frame; extract the spatial coordinates of each individual; cluster the detected individuals and classify the close individuals into a group; calculate the density of each group; the density reflects the degree of aggregation between individuals; and calculate the group movement consistency, which is used to measure the consistency of the movement direction and speed of individuals in each group, reflecting the coordination of individuals in the group;
[0173] Furthermore, the comprehensive population density is:
[0174] D over_density =max(D 1 density (i),D 2 density (i),D 3 density (i),...D k density (i));
[0175]
[0176] Among them, D over_density (i) is the comprehensive population density of the i-th detection video image frame; D k density (i) represents the population density of the kth population in the i-th video image frame; R k represents the total number of individuals in the kth group of the i-th detection video image frame; (x j (i),y j (i)) represents the two-dimensional coordinates of the jth individual in the kth group in the i-th detection video image frame; (x u (i),y u (i)) represents the two-dimensional coordinates of the u-th individual in the k-th group in the i-th detection video image frame; Dist() represents the Euclidean distance function.
[0177] Furthermore, the calculation formula for the comprehensive group movement consistency is:
[0178] V over_sync (i) = min(V 1 sync (i),V 2 sync (i),V 3 sync (i),...,V k sync (i));
[0179]
[0180]
[0181] Among them, V over_sync (i) represents the comprehensive group movement consistency of the i-th detection video image frame; V k sync (i) represents the movement consistency of the kth group in the i-th detected video image frame; θ k j (i) represents the angle between the moving speed vector of the jth individual in the kth group and the average moving speed vector of the kth group; Vel k j (i) represents the moving speed vector of the jth individual in the kth group in the i-th detection video image frame; Vel k avg (i) represents the average moving speed vector of the kth group in the i-th detection video image frame; R k represents the total number of individuals in the kth group; x k j(i) represents the horizontal coordinate of the jth individual in the kth group in the i-th detection video image frame; y k j (i) represents the ordinate of the jth individual in the kth group in the ith detection video image frame; Δt represents the time interval between the ith detection video image frame and the i-1th detection video image frame.
[0182] Furthermore, the second abnormal coefficient is:
[0183] R second (i) = ξ density *D over_density (i)+ψ sync *V over_sync (i);
[0184] Among them, R second (i) represents the second abnormal coefficient of the i-th detected video image frame; ξ density represents the comprehensive population density weight; ψ sync Represents the comprehensive movement consistency weight.
[0185] The comprehensive abnormality analysis module analyzes the first abnormality coefficient and the second abnormality coefficient to obtain a comprehensive abnormality coefficient, and compares it with a preset abnormality threshold to obtain the abnormality category of each detected video image frame.
[0186] Furthermore, the comprehensive abnormality coefficient is:
[0187]
[0188] Among them, R com (i) represents the comprehensive abnormality coefficient of the i-th detected video image frame; R first (i,j) represents the first abnormal coefficient of the jth individual in the i-th detected video image frame; R represents the number of individuals in the i-th detected video image frame; φ fir represents the first abnormal coefficient weight; δ se Represents the second abnormal coefficient weight; R second (i) represents the second abnormality coefficient of the i-th detected video image frame.
[0189] Furthermore, the comprehensive abnormality coefficient is compared with a first preset abnormality coefficient threshold, and the detected video image frame is classified into abnormal categories, where the abnormal categories include abnormal video image frames and normal video image frames.
[0190] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A facial image acquisition system, characterized in that: include: An image acquisition module is used to shoot the interior and surrounding of the gold shop through a camera to obtain a first video of the gold shop, and divide the first video into frames according to a preset time interval to obtain individual video image frames; An individual anomaly analysis module, used to input each video image frame into an individual anomaly analysis model, and sequentially pass through the preprocessing layer, feature extraction layer, feature fusion layer, feature processing layer and output layer of the model; The preprocessing layer preprocesses each of the video image frames to obtain each detection video image frame; The feature extraction layer extracts features from each of the detection video image frames to obtain a first face feature and a second behavior feature; and obtains a first comprehensive feature through the feature fusion layer; The feature processing layer analyzes the first comprehensive feature through the LSTM network structure incorporating the attention mechanism to obtain a first comprehensive time series feature; and analyzes the first comprehensive time series feature to obtain a second comprehensive feature; the output layer analyzes the second comprehensive feature to obtain a first abnormality coefficient; A group anomaly analysis module, used to analyze each of the detection video image frames through the YOLO algorithm and the clustering algorithm to obtain a comprehensive group density and a comprehensive group movement consistency, and obtain a second anomaly coefficient through the comprehensive group density and the comprehensive group movement consistency; The comprehensive abnormality analysis module analyzes the first abnormality coefficient and the second abnormality coefficient to obtain a comprehensive abnormality coefficient, and compares it with a preset abnormality threshold to obtain the abnormality category of each detected video image frame.
2. A facial image acquisition system according to claim 1, characterized in that: The first facial features include expression features, facial posture and gaze direction; the second behavioral features include gait features, gesture features and body posture; the first facial features are: F face (i,j)={f face,i ,f face,2 ,...f face,M }; Among them, F face (i,j) represents the first facial feature of the jth individual in the i-th detected video image frame; M is the number of first facial features; The second behavior feature is: F behavior (i,j)={f behavior,i ,f behavior,2 ,...f behavior,N }; Among them, F behavior (i, j) represents the second behavior feature of the jth individual in the i-th detected video image frame; N is the number of second behavior features.
3. The facial image acquisition system according to claim 1, characterized in that: The first comprehensive feature is: F combined (i,j)=concat(F face (i,j),F behavior (i,j)); Among them, F combined (i,j) represents the first comprehensive feature of the jth individual in the i-th detected video image frame; concat() represents a vector concatenation operation.
4. The facial image acquisition system according to claim 1, characterized in that: The second comprehensive feature is: F optimized (i,j)=ReLu(W*LSTM(F combined (i-1,j),F combined (i,j))+b); Among them, F attention (i,j) represents the second comprehensive feature of the jth individual in the i-th detected video image frame; γ k (i, j) represents the weight calculated by the attention mechanism, reflecting the importance of each feature; k represents the comprehensive feature index; L1 represents the first face feature dimension; L2 represents the first face feature dimension; k represents the comprehensive feature dimension index; F k optimized (i,j) represents the k-th dimensional row vector of the first comprehensive time series feature; RELU() represents the activation function; W represents the weight matrix; LSTM() represents the time series feature relationship function, which is used to extract the time dependency of the first comprehensive feature between the previous and next frames; b represents the bias term.
5. The facial image acquisition system according to claim 1, characterized in that: The first abnormal coefficient is: Among them, R first (i,j) represents the first abnormal coefficient of the jth individual in the i-th detected video image frame; k represents the comprehensive feature index; ω k represents the abnormal weight of the kth component; F attention (i, j) represents the second comprehensive feature of the jth individual in the i-th detected video image frame.
6. The facial image acquisition system according to claim 1, characterized in that: The comprehensive population density is: D over_density =max(D 1 density (i),D 2 density (i),D 3 density (i),...D k density (i)); Among them, D over_density (i) is the comprehensive population density of the i-th detection video image frame; D k density (i) represents the population density of the kth population in the i-th video image frame; R k represents the total number of individuals in the kth group of the i-th detection video image frame; (x j (i),y j (i)) represents the two-dimensional coordinates of the jth individual in the kth group in the i-th detection video image frame; (x u (i),y u (i)) represents the two-dimensional coordinates of the u-th individual in the k-th group in the i-th detection video image frame; Dist() represents the Euclidean distance function.
7. The facial image acquisition system according to claim 1, characterized in that: The calculation formula for the comprehensive group movement consistency is: V over_sync (i)=min(V 1 sync (i),V 2 sync (i),V 3 sync (i),...,V k sync (i)); Among them, V over_sync (i) represents the comprehensive group movement consistency of the i-th detection video image frame; V k sync (i) represents the movement consistency of the kth group in the i-th detected video image frame; θ k j (i) represents the angle between the moving speed vector of the jth individual in the kth group and the average moving speed vector of the kth group; Vel k j (i) represents the moving speed vector of the jth individual in the kth group in the i-th detection video image frame; Vel k avg (i) represents the average moving speed vector of the kth group in the i-th detection video image frame; R k represents the total number of individuals in the kth group; x k j (i) represents the horizontal coordinate of the jth individual in the kth group in the i-th detection video image frame; y k j (i) represents the ordinate of the jth individual in the kth group in the ith detection video image frame; Δt represents the time interval between the ith detection video image frame and the i-1th detection video image frame.
8. The facial image acquisition system according to claim 1, characterized in that: The second abnormal coefficient is: R second (i)=ξ density *D over_density (i)+ψ sync *V over_sync (i); Among them, R second (i) represents the second abnormal coefficient of the i-th detected video image frame; ξ density represents the comprehensive population density weight; ψ sync Represents the comprehensive movement consistency weight.
9. The facial image acquisition system according to claim 1, characterized in that: The comprehensive anomaly coefficient is: Among them, R com (i) represents the comprehensive abnormality coefficient of the i-th detected video image frame; R first (i,j) represents the first abnormal coefficient of the jth individual in the i-th detected video image frame; R represents the number of individuals in the i-th detected video image frame; φ fir represents the first abnormal coefficient weight; δ se Represents the second abnormal coefficient weight; R second (i) represents the second abnormality coefficient of the i-th detected video image frame.
10. The facial image acquisition system according to claim 1, characterized in that: The comprehensive abnormality coefficient is compared with a first preset abnormality coefficient threshold, and the detected video image frame is classified into abnormal categories, where the abnormal categories include abnormal video image frames and normal video image frames.