A warehouse attendance management system based on AI vision
By generating local motion stability maps and adversarial declarative network models, the problems of low facial recognition accuracy and low attendance management efficiency during rapid passage in warehouse environments are solved, achieving efficient and accurate employee identification and attendance recording.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 安徽云易智能技术有限公司
- Filing Date
- 2026-03-24
- Publication Date
- 2026-07-28
AI Technical Summary
When existing technologies are used for rapid passage in warehouse environments, facial recognition is not accurate and reliable enough, resulting in low attendance management efficiency and excessive consumption of computing resources.
By generating local motion stability maps and combining motion blur features with an adversarial declarative network model, the missing areas of facial structure are reconstructed, enabling efficient and accurate employee identification.
It improves the accuracy of facial recognition and the integrity of attendance data during fast passage, reduces computing resource consumption, and ensures the real-time performance and reliability of attendance management.
Smart Images

Figure CN121920979B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of facial recognition technology, specifically to a warehouse attendance management system based on AI vision. Background Technology
[0002] With the rapid development of the warehousing and logistics industry, employee attendance management has gradually become a crucial aspect of ensuring efficient and standardized warehouse operations. Traditional warehouse attendance methods mainly include fingerprint recognition, card swiping, and static facial recognition. Fingerprint recognition and card swiping suffer from issues such as susceptibility to proxy attendance and difficulty in guaranteeing data accuracy. Traditional static facial recognition requires employees to pause briefly for image capture, which can cause congestion during peak hours or when personnel are densely packed, reducing efficiency and even impacting overall warehouse operations. Furthermore, during rapid passage, facial movements can cause blurred facial images, missing key features, and facial distortions, severely affecting the accuracy and reliability of AI-based visual facial recognition, further limiting the practical application of warehouse attendance technology.
[0003] At the same time, existing motion vision-based facial recognition technologies mainly focus on image processing in static or weak motion states. They lack targeted analysis methods and effective processing strategies for the problems of missing local facial structures and motion blur caused by employees passing through quickly in warehouse environments. It is difficult to achieve efficient, accurate and stable employee identification while ensuring reasonable consumption of computing resources.
[0004] Therefore, how to achieve accurate, real-time, and reliable automated management of employee attendance in a fast-moving state without excessively increasing the system's computational complexity and resource consumption has become a key technical problem that urgently needs to be solved in warehouse operation management. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a warehouse attendance management system based on AI vision.
[0006] To achieve the above objectives, the present invention provides a warehouse attendance management system based on AI vision, comprising:
[0007] The map generation module is used to generate a local motion stability map based on the cross-frame displacement relationship of facial key points in the initial multi-frame facial images.
[0008] The feature acquisition module is used to extract motion blur features from the initial multi-frame face images and combine them with the displacement features of key points in the local motion stability map to determine the type of motion blur.
[0009] The data processing module is used to select the corresponding adversarial deblurring network model based on the motion blur type to generate a sharpened face image, and to determine the missing areas of the face structure based on the sharpened face image;
[0010] The attendance recognition module is used to locate the temporal change trajectory of the missing facial structure region in the initial multi-frame facial images based on the local motion stability map, and reconstruct the regional motion path according to the temporal change trajectory to generate a local structure reconstruction map. Based on the cleared facial image and the local structure reconstruction map, the module identifies the employee's identity and generates attendance records.
[0011] Compared with the prior art, the beneficial effects of the present invention are:
[0012] This invention constructs a local motion stability map and utilizes the displacement relationship of key points in multiple consecutive frames of face images to effectively characterize the real motion state of each region of the face when employees pass through, avoiding the misidentification problem caused by single-frame image or overall motion blur analysis, and improving the accuracy of motion state analysis of personnel passing through quickly in a warehouse environment.
[0013] This invention determines the type of motion blur by combining local motion stability maps and gradient rhythm features. Based on the motion state of local regions, it reduces the risk of misjudgment in motion blur type identification and ensures the targeted and accurate processing of facial image sharpening under different blur modes.
[0014] This invention improves the level of facial structure integrity restoration by performing contour structure analysis on local structurally missing areas based on a cleared facial image and using a cross-frame trajectory reconstruction mechanism to compensate for the local structurally missing areas, thereby ensuring the reliability of personnel identification and the integrity of attendance data in warehouse attendance records. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a block diagram of the system of the present invention. Detailed Implementation
[0017] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figure 1 This embodiment provides a warehouse attendance management system based on AI vision, including:
[0019] The map generation module 101 is used to generate a local motion stability map based on the cross-frame displacement relationship of facial key points in the initial multi-frame facial images.
[0020] It should be noted that the initial multi-frame facial images are a collection of multiple images of the same employee captured within a continuous time slice. There is a clear temporal sequence between the frames, and the spatial position of the face changes continuously within each frame. This collection of multi-frame images is used to reflect the employee's actual movement during the passage process.
[0021] It should also be noted that a compliance authorization confirmation process is pre-set during the facial image acquisition phase of this system. The system demonstrates an informed consent mechanism through an interactive terminal, and the acquisition process only starts after obtaining explicit authorization confirmation from the person being collected. Furthermore, the system adheres to the principle of minimization in its technical logic, extracting only feature vector information for identity recognition and not storing the original high-resolution portrait image. All extracted feature data is stored and transmitted using encryption algorithms, ensuring the legality of facial information acquisition through these technical processes.
[0022] Specifically, generating the local motion stability map includes:
[0023] Based on the relative displacement relationship of facial key points in adjacent frames in the initial multi-frame facial images, displacement correlation features reflecting the cooperative motion state of local regions are extracted.
[0024] Specifically, the system performs key point detection on the facial regions of employees in each frame of the image to obtain stable and accurate facial feature information.
[0025] For example, the key points include typical feature points with relatively fixed positions, such as the corners of the eyes, the tip of the nose, the corners of the mouth, the brow peaks, and the angles of the jaw. The system can use a pre-trained deep learning facial key point detection model (e.g., HRNet or Hourglass network) to automatically detect the above key points, specifically obtaining the two-dimensional coordinate position of each key point in each frame of the image.
[0026] In one specific embodiment, for any two adjacent frames in the initial multi-frame image (e.g., the first... Frame and the (frame), the system calculates the displacement vector of each keypoint between two adjacent frames, and the specific calculation formula is as follows: ;in, Indicates the first The key point is from the first Frame to the The displacement vector of the frame. They represent the first The first frame Two-dimensional coordinates of key points.
[0027] Furthermore, in order to fully express the cooperative motion state of key points in a local area, the system combines several spatially adjacent key points into a local spatial neighborhood and calculates the covariance matrix of the displacement vectors of all key points in the neighborhood to quantify the cooperative motion characteristics of key points in the area.
[0028] For example, for containing The formula for calculating the covariance matrix of the local neighborhood of each key point is: ,in, This represents the average value of the displacement vectors of all key points within this local neighborhood.
[0029] In the displacement correlation features, the relative displacement maintenance relationship within the spatial neighborhood and the cross-frame displacement reversal relationship are identified in parallel, and a set of local motion constraints is formed accordingly.
[0030] The formation of the local motion constraint set includes:
[0031] Based on the mutual constraints of displacement correlation features in the spatial neighborhood, neighborhood cooperative displacement constraints are extracted.
[0032] Specifically, the system calculates the angle between the displacement vectors of any two key points within a spatial neighborhood to determine whether the motion of key points within that local neighborhood remains consistent. The formula for calculating the angle is as follows: ;
[0033] When the included angle Less than or equal to a preset threshold (e.g.) When determining key points With key points The neighborhood cooperative displacement constraint is satisfied between them.
[0034] Based on the directional change relationship of displacement correlation features on the time axis, time-reversal displacement constraints are extracted;
[0035] In practice, the system analyzes the direction change of the displacement vector of a single key point across consecutive frames, and the calculation formula is as follows: ;
[0036] When the calculation result When the value is less than 0, it indicates that the direction of motion of the key point has been reversed, and the system extracts the time-reversal displacement constraint accordingly.
[0037] By combining neighborhood cooperative displacement constraints with temporal reversal displacement constraints, a set of local motion constraints is formed.
[0038] Preferably, the system performs a logical union combination of the above two types of constraints. Specifically, when a key point satisfies the neighborhood cooperative displacement constraint, or when a motion mutation occurs that satisfies the time-reversal displacement constraint, the motion vector of that key point is included in the local motion constraint set. By including both the "cooperative constraint" representing smooth motion and the "reversal constraint" representing nonlinear disturbance in the set, the system comprehensively covers the complex motion states of employee traffic, ensuring the completeness of subsequent constraint screening.
[0039] Based on the set of local motion constraints, displacement-related features are constrained and filtered to obtain a feature set for characterizing the stable state of motion in a local region.
[0040] The local motion constraint set includes key point displacement vectors that satisfy the neighborhood cooperative displacement constraint or the time-reversal displacement constraint.
[0041] Specifically, the system scores the matching degree between the displacement correlation features of all key points and the set of local motion constraints. For example, cosine similarity can be used as the scoring index, and the calculation formula is as follows: ;in, For the first Scoring of key point displacement correlation features For the first in the set of local motion constraints Constraint vectors, The number of constraints in the constraint set.
[0042] Subsequently, the system sets a scoring threshold (e.g., 0.75) and retains only displacement-related features with scores greater than or equal to this threshold to form a highly stable feature set.
[0043] A local motion stability map is generated based on the feature set;
[0044] In practice, the system will visualize the filtered feature set in the form of a two-dimensional heatmap.
[0045] For example, regions with motion stability scores greater than or equal to a set threshold are marked in red, regions with stability scores less than the set threshold are marked in blue, and other regions are represented by intermediate colors (such as green or yellow). The system obtains a local motion stability map in this manner to visually represent the differences in motion stability of different local regions of an employee's face across multiple frames of images.
[0046] The feature acquisition module 102 is used to extract motion blur features based on the initial multi-frame face images and combine them with the displacement features of key points in the local motion stability map to determine the motion blur type.
[0047] It should be noted that in the initial multi-frame face images, the employees' rapid passage caused varying degrees of motion blur in the face region across consecutive frames. Therefore, it is necessary to first quantitatively analyze the grayscale change trend of the face region pixels in order to extract effective motion blur features.
[0048] Specifically, determining the motion blur type includes:
[0049] Based on the gray-level gradient evolution process of the face region in the initial multi-frame face images, gradient rhythm features describing the local gray-level change rhythm are extracted;
[0050] In practice, the system first converts each frame of the face image into a grayscale image to reduce the complexity of subsequent calculations. Then, it calculates the position of each pixel within the face region. Calculate the grayscale gradient change value between consecutive frames. Specifically, if the first frame... Frame and the The pixel grayscale values between frames are respectively and The corresponding grayscale gradient change is ;
[0051] Then, the system summarizes the grayscale gradient change values of all consecutive frames to form a gradient rhythm feature sequence that describes the grayscale change trend of each pixel.
[0052] Based on the local motion constraint state reflected in the local motion stability map, the gradient rhythm features are constrained and recombined to form a gradient rhythm representation modulated by motion constraints.
[0053] In practice, to improve the reliability of gradient rhythm features, the system needs to use the generated local motion stability map to perform constrained reorganization of the gradient rhythm features.
[0054] The constrained recombination of gradient rhythm features includes:
[0055] The gradient rhythm features are mapped to a rhythm sequence space with time as the main axis; and rhythm segments in the rhythm sequence space are selectively preserved and suppressed according to the local motion constraint state; wherein, the rhythm sequence space is a set of feature vectors arranged in timestamp order and reflecting the evolution trend of pixel gray-level gradient; the rhythm segment is the feature component in the feature vector set corresponding to a specific time frame;
[0056] For example, the system can arrange the gradient rhythm features of each frame in chronological order for each pixel group within a local region as follows: ;in, Indicates the first [unit] within this local region The pixel at the th point The grayscale gradient value corresponding to the frame. This indicates the number of pixels.
[0057] In practice, the system uses the stability scores of each region in the local motion stability map to selectively process the gradient rhythm features of the corresponding region.
[0058] For example, the system sets a stability score threshold of 0.75. When the stability score of a region is greater than or equal to 0.75, the corresponding rhythmic features are fully preserved. When the stability score is less than 0.75, the rhythmic features corresponding to that region are suppressed so that they do not participate in subsequent processing. In this way, the system can effectively screen gradient rhythmic features.
[0059] A gradient rhythm representation modulated by motion constraints is generated based on the rhythm segments after selective preservation and suppression.
[0060] In practice, the system recombines the retained gradient rhythm features using a weighted combination method to achieve a precise expression of motion constraints. The specific calculation formula is as follows: ;in, Indicates the first Gradient rhythmic feature representation of frames after motion constraint modulation; Indicates the first The local motion stability score of the region where each pixel is located is determined by the local motion stability map generated in the map generation module S101. Indicates the first The pixel in the first Gradient rhythm characteristics of frames; This represents the total number of pixels involved in the combination.
[0061] The motion fuzziness type is determined based on the gradient rhythm representation modulated by motion constraints;
[0062] Specifically, the system needs to perform amplitude integration on the modulated gradient rhythm features in both horizontal and vertical orthogonal directions to determine the specific motion blur type. The specific implementation is as follows:
[0063] First, calculate the horizontal motion blur amplitude. Vertical motion blur amplitude : ;in, Indicates pixel position The horizontal gradient component in frame t; Indicates pixel position In the Vertical gradient components of a frame.
[0064] Subsequently, the system determines the motion blur type based on the following judgment rules:
[0065] If the conditions are met If so, the system determines it to be lateral motion ambiguity;
[0066] If the conditions are met: If so, the system determines it to be longitudinal motion ambiguity;
[0067] If the difference between the horizontal and vertical motion blur amplitudes does not meet the above two conditions, but simultaneously meets the conditions If so, the system determines it to be rotational motion fuzziness. Wherein, the constant... and Determined based on experimental data.
[0068] The data processing module 103 is used to select the corresponding adversarial deblurring network model according to the motion blur type to generate a sharpened face image, and to determine the missing areas of the face structure based on the sharpened face image;
[0069] Understandably, since images with different motion blur types differ in texture, edge orientation, and information distribution, a specialized adversarial deblurring network model is pre-trained for each specific blur type to ensure deblurring effect.
[0070] In a specific embodiment, the system pre-constructs multiple specially trained adversarial deblurring network models (GAN models), corresponding to horizontal motion blur, vertical motion blur, and rotational motion blur, respectively. For example, when the feature acquisition module 102 determines that the current face image has horizontal motion blur, the system selects a GAN model specifically trained for horizontal motion blur to perform image deblurring processing.
[0071] It should be noted that the GAN model corresponding to each type of motion blur includes two parts: a generator network and a discriminator network.
[0072] The generative network is responsible for transforming the input motion-blurred image into a sharpened image;
[0073] The discriminant network is used to identify the difference between the sharpened image output by the generator network and the real sharp image, thereby prompting the generator network to produce more realistic sharpened images.
[0074] In some specific embodiments, the training objective function of the generative network of the GAN model can be defined as: ,in, To counteract the loss function, used to optimize the realism of the generated image; For reconstruction loss functions, such as L1 loss or perceptual loss, the generated image is used to ensure the similarity between the generated image and the real image at the pixel level or feature level; and The weighting coefficients are determined based on experimental data.
[0075] Specifically, determining the facial structure missing regions based on the sharpened facial image includes:
[0076] Based on the continuous directional relationship of the facial contour in the sharpened facial image, structural transition features describing the changes in the contour direction are extracted.
[0077] In practice, the system first uses an image edge detection algorithm (such as the Canny edge detection algorithm) to extract facial contour lines from a sharpened facial image, obtaining a series of contour point sequences. Next, the system analyzes the contour point sequence point by point, calculates the directional changes between adjacent contour points, and then extracts structural transition features.
[0078] In one specific embodiment, for any three consecutive contour points Their spatial coordinates are represented as follows:
[0079]
[0080] The system calculates the angle between two continuous contour segments. The calculation formula is as follows: ;in, These represent two consecutive direction vectors on the contour.
[0081] The system will calculate a series of included angles as described above. As a set of structural transition features, it is used to describe the changes in the contour direction.
[0082] In the structural transition features, the relationship between the interruption of the contour direction and the superposition of the contour transition are analyzed in parallel to form the basis for judging the structural integrity.
[0083] The criteria for determining the integrity of the formed structure include:
[0084] Based on the contour breakage pattern that appears during contour direction changes, the contour interruption judgment condition is extracted.
[0085] In one specific embodiment, the system first calculates the Euclidean distance between adjacent contour points along the contour point sequence. For example, for two adjacent contour points... The distance calculation formula is: ;
[0086] The system presets a distance threshold, for example, set to twice the average distance between contour points. This threshold is used when the distance between two adjacent points on the contour... When the value is greater than or equal to the threshold, it is determined that there is a break in the contour direction at that position, that is, the contour interruption determination condition is met.
[0087] Based on the spatial superposition distribution of contour transitions, contour superposition judgment conditions are extracted.
[0088] Specifically, the system uses a sliding window approach to move along the sequence of contour points, and calculates the contour turning features (angles) within each sliding window. The variance of () is calculated using the following formula: ;in, This represents the set of contour points within the current sliding window; This represents the number of contour points within the sliding window. This represents the average value of the turning features (angles) within the window.
[0089] The system presets a variance threshold, for example, a value of When the variance of the contour transition features within the window is greater than or equal to the variance threshold, the system determines that there is a dense superposition of contour transition features within the window area, which satisfies the contour superposition determination condition.
[0090] The criteria for determining contour interruption are combined with those for determining contour superposition to form the basis for determining structural integrity.
[0091] Preferably, the system combines the above two conditions using a logical "OR" combination. That is, when the contour region simultaneously meets at least one of the contour interruption judgment condition or the contour overlay judgment condition, the region is marked as having an abnormality in structural integrity, thereby forming a basis for structural integrity judgment.
[0092] Based on the aforementioned structural integrity determination criteria, the areas with missing facial structures are identified.
[0093] In practice, the system determines the closed region containing the above-mentioned abnormal points based on the contour point positions that meet the criteria for structural integrity judgment, using the minimum circumcircle or minimum bounding box algorithm. This region is then used as the missing facial structure area for further reconstruction of the facial structure across subsequent frames.
[0094] The attendance recognition module 104 is used to locate the temporal change trajectory of the missing facial structure region in the initial multi-frame facial images based on the local motion stability map, and reconstruct the region motion path according to the temporal change trajectory to generate a local structure reconstruction map. Based on the cleared facial image and the local structure reconstruction map, the module identifies the employee's identity and generates attendance records.
[0095] Specifically, the generation of the local structure reconstruction map includes:
[0096] The temporal change trajectory of the missing facial structure region in multiple frames was determined based on the local motion stability map.
[0097] In practice, the system first determines the initial center coordinates of the structurally missing regions in the sharpened face image. Then, the local motion stability map generated in the map generation module 101 is used to obtain the cross-frame displacement information of each key point, thereby determining the continuous position coordinates of the structural missing region in each frame image.
[0098] For example, if the system needs to determine the location coordinates of the missing region in frame t... It can be calculated using the following formula: ,in, For the first The coordinates of the missing structural regions in the frame; This indicates that the region determined using the local motion stability map is from the [missing information]. Frame to the The displacement vector of the frame.
[0099] Using the above method, the system can obtain the temporal change trajectory of the structurally missing region in all image frames.
[0100] In the temporally changing trajectory, the direction preservation relationship and the neighborhood cooperative displacement relationship of the trajectory are analyzed in parallel to form the structural reconstruction trajectory constraint;
[0101] The formation of structural reconstruction trajectory constraints includes:
[0102] Based on the directional continuity of the temporal change trajectory on the time axis, trajectory direction constraints are extracted;
[0103] Specifically, the system performs a cross-frame displacement vector analysis on two consecutive frames. and The angle between The calculation is performed using the following formula: ;
[0104] When the included angle When the angle is less than or equal to a set threshold (e.g., 15°), the system determines that the trajectory at that position satisfies the directional continuity relationship, and thus determines the trajectory direction constraint.
[0105] Based on the cooperative displacement relationship between the temporal change trajectory and the adjacent complete structural regions, neighborhood cooperative constraints are extracted.
[0106] In implementation, the system calculates the similarity of the displacement vectors of the missing structural region and its adjacent complete region in frame t to determine their cooperative motion relationship. For example, the similarity can be represented using cosine similarity: ;in, For the structurally missing region in the first The displacement vector of the frame; For adjacent complete structural regions in the first The displacement vector of the frame.
[0107] When similarity When the value is greater than or equal to a set threshold (e.g., 0.8), the neighborhood collaboration constraint is determined to be satisfied.
[0108] By combining trajectory direction constraints with neighborhood cooperative constraints, a structure reconstruction trajectory constraint is formed.
[0109] In practice, the system combines the two types of constraints using a logical AND operation. That is, only when a frame satisfies both the trajectory direction constraint and the neighborhood cooperation constraint is the trajectory of that frame determined to satisfy the structural reconstruction trajectory constraint, so as to ensure the accuracy of the structural reconstruction trajectory.
[0110] Based on the constraints of the structural reconstruction trajectory, the temporal variation trajectory is reconstructed to generate a local structural reconstruction map.
[0111] In practice, based on the above-mentioned structural reconstruction trajectory constraints, the system uses optimization methods to make overall corrections to the temporal change trajectory of the structurally missing region, so as to improve the stability and continuity of the trajectory and thus ensure the accurate reconstruction of the structurally missing region.
[0112] In one specific embodiment, the system first defines a trajectory constraint optimization objective function to quantify the trajectory reconstruction error. The specific optimization objective function can be expressed as: ;in, To optimize the total error of the trajectory; Total number of frames; Indicates the first Frame-corrected trajectory position Compared with the original trajectory position The Euclidean distance between them; For the first Frame neighborhood collaborative constraint similarity value;
[0113] , The weighting coefficient is determined based on experimental data and is used to adjust the degree of influence of positional deviation and synergy; for example, both are set to 0.5.
[0114] The system can use gradient descent, Levenberg-Marquardt algorithm or other optimization methods to solve the above optimization objective function, thereby obtaining the precise location coordinates of the corrected missing region in all frames, and then generating a local structure reconstruction map.
[0115] Finally, the system extracts global feature vectors of employee faces based on the sharpened face images. Simultaneously, for the correction regions marked by the local structure reconstruction map, the system inputs them into a pre-defined lightweight feature extraction network (such as MobileNet or SqueezeNet) to extract the geometric compensation feature vector of that region after temporal reconstruction. The two types of feature vectors are then concatenated to form a complete face feature vector F, which is represented as: ,in, This indicates a vector concatenation operation.
[0116] Next, the system compares the obtained feature vector F with the pre-stored facial feature database and uses a similarity measurement method (such as cosine similarity) to determine identity. When the feature similarity is greater than or equal to a predetermined threshold (e.g., 0.85), the system confirms the employee's identity; otherwise, the system indicates recognition failure and requires manual confirmation or re-collection.
[0117] Finally, the system will automatically record the employee's identity, passage time and location after successful identification into the warehouse attendance database, generating a valid attendance record.
[0118] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A warehouse attendance management system based on AI vision, characterized in that, include: The map generation module is used to generate a local motion stability map based on the cross-frame displacement relationship of facial key points in the initial multi-frame facial images. The feature acquisition module is used to extract motion blur features from initial multi-frame face images and, in conjunction with the displacement features of key points in the local motion stability map, determine the type of motion blur; including: Based on the gray-level gradient evolution process of the face region in the initial multi-frame face images, gradient rhythm features describing the local gray-level change rhythm are extracted; the gradient rhythm features are the summation of gray-level gradient change values of all consecutive frames to form a gradient rhythm feature sequence describing the gray-level change trend of each pixel. Based on the local motion constraint state reflected in the local motion stability map, the gradient rhythm features are constrained and recombined to form a gradient rhythm representation modulated by motion constraints. The motion fuzziness type is determined based on the gradient rhythm representation modulated by motion constraints; The data processing module is used to select the corresponding adversarial deblurring network model based on the motion blur type to generate a sharpened face image, and to determine the missing areas of the face structure based on the sharpened face image; The attendance recognition module is used to locate the temporal change trajectory of the missing facial structure region in the initial multi-frame facial images based on the local motion stability map, and reconstruct the regional motion path according to the temporal change trajectory to generate a local structure reconstruction map. Based on the cleared facial image and the local structure reconstruction map, the module identifies the employee's identity and generates attendance records.
2. The warehouse attendance management system based on AI vision according to claim 1, characterized in that, The generation of the local motion stability map includes: Based on the relative displacement relationship of facial key points in adjacent frames in the initial multi-frame facial images, displacement correlation features reflecting the cooperative motion state of local regions are extracted. In the displacement correlation features, the relative displacement maintenance relationship within the spatial neighborhood and the cross-frame displacement reversal relationship are identified in parallel, and a set of local motion constraints is formed accordingly. Based on the set of local motion constraints, displacement-related features are constrained and filtered to obtain a feature set for characterizing the stable state of motion in a local region. A local motion stability map is generated based on the feature set.
3. The warehouse attendance management system based on AI vision according to claim 2, characterized in that, The formation of the local motion constraint set includes: Based on the mutual constraints of displacement correlation features in the spatial neighborhood, neighborhood cooperative displacement constraints are extracted. Based on the directional change relationship of displacement correlation features on the time axis, time-reversal displacement constraints are extracted; By combining neighborhood cooperative displacement constraints with temporal reversal displacement constraints, a set of local motion constraints is formed.
4. The warehouse attendance management system based on AI vision according to claim 3, characterized in that, The constrained recombination of gradient rhythm features includes: The gradient rhythm features are mapped to a rhythm sequence space with time as the main axis; and rhythm segments in the rhythm sequence space are selectively preserved and suppressed according to the local motion constraint state; wherein, the rhythm sequence space is a set of feature vectors arranged in timestamp order and reflecting the evolution trend of pixel gray-level gradient; the rhythm segment is the feature component in the feature vector set corresponding to a specific time frame; A gradient rhythm representation modulated by motion constraints is generated based on rhythm segments that have undergone selective preservation and suppression.
5. A warehouse attendance management system based on AI vision according to claim 4, characterized in that, The step of determining the missing facial structure region based on the sharpened facial image includes: Based on the continuous directional relationship of the facial contour in the sharpened facial image, structural transition features describing the changes in the contour direction are extracted. In the structural transition features, the relationship between the interruption of the contour direction and the superposition of the contour transition are analyzed in parallel to form the basis for judging the structural integrity. Based on the aforementioned structural integrity criteria, the areas with missing facial structures are determined.
6. A warehouse attendance management system based on AI vision according to claim 5, characterized in that, The criteria for determining the integrity of the formed structure include: Based on the contour breakage pattern that appears during contour direction changes, the contour interruption judgment condition is extracted. Based on the spatial superposition distribution of contour transitions, contour superposition judgment conditions are extracted. The criteria for determining contour interruption are combined with those for determining contour overlay to form the basis for determining structural integrity.
7. A warehouse attendance management system based on AI vision according to claim 6, characterized in that, The generation of the local structure reconstruction map includes: The temporal change trajectory of the missing facial structure region in multiple frames was determined based on the local motion stability map. In the temporally changing trajectory, the direction preservation relationship and the neighborhood cooperative displacement relationship of the trajectory are analyzed in parallel to form the structural reconstruction trajectory constraint; Based on the constraints of the structural reconstruction trajectory, the temporal change trajectory is reconstructed to generate a local structural reconstruction map.
8. A warehouse attendance management system based on AI vision according to claim 7, characterized in that, The formation of the structural reconstruction trajectory constraint includes: Based on the directional continuity of the temporal change trajectory on the time axis, trajectory direction constraints are extracted; Based on the cooperative displacement relationship between the temporal change trajectory and the adjacent complete structural regions, neighborhood cooperative constraints are extracted. By combining trajectory direction constraints with neighborhood cooperative constraints, a structure reconstruction trajectory constraint is formed.