An enterprise attendance management system based on deep learning and multi-source data fusion
Patent Information
- Application Number
- CN202611212141.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-11
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明的目的在于提供一种基于深度学习和多源数据融合的企业考勤管理系统,以解决现有技术中多源考勤数据无法实现深度语义关联、考勤异常判定精度不足、系统鲁棒性差的技术问题
1、本发明通过构建语义关联图谱,将来自门禁、定位、考勤机和请假系统等多源异构考勤数据进行深度语义关联,实现不同来源考勤数据间的语义映射与协同分析,解决现有技术中多源数据仅能进行简单消解而无法深度融合的技术难题;语义关联图谱中的时间邻近规则、空间共现规则、行为序列规则和业务逻辑规则相互配合,能够从多维度捕捉考勤行为间的内在联系,为后续的融合决策提供丰富的语义支撑。
Smart Images

Figure CN122840807A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent attendance management systems, specifically relating to an enterprise attendance management system based on deep learning and multi-source data fusion. Background Technology
[0002] In corporate management practice, attendance management is a core component of human resource management, directly impacting employee behavior, performance evaluation, and organizational operational efficiency. With the deepening of enterprise informatization, traditional attendance methods relying on manual recording and paper sign-ins have been gradually replaced by digital systems. In recent years, the rapid development of artificial intelligence, big data, and the Internet of Things (IoT) technologies has brought opportunities for intelligent upgrades to enterprise attendance management, making data-driven, precise attendance analysis and decision support possible.
[0003] Among them, intelligent attendance management technology based on deep learning and multi-source data fusion is gradually becoming a research hotspot. This technology aims to integrate heterogeneous attendance data from different collection terminals and business systems, and realize semantic association and collaborative analysis between data through deep learning algorithms, thereby constructing a comprehensive and accurate employee attendance profile and providing more efficient and intelligent decision support for enterprise management.
[0004] While some existing technologies mention the concept of multi-source data fusion, their implementation is merely a simple resolution based on hierarchical adjudication rules. This fails to achieve true deep semantic association and collaborative analysis between attendance data from different sources, resulting in insufficient accuracy in attendance anomaly detection and difficulty in handling complex and ever-changing attendance patterns in real-world scenarios. Most existing technologies, such as GPS-based location-based attendance and fingerprint-based desk matching for identity verification, only process single or a few data sources. This leads to severe data fragmentation between systems, creating significant data silos and preventing the extraction of attendance features from multiple dimensions and levels to form a comprehensive and multi-dimensional attendance profile.
[0005] Furthermore, single-data-source solutions lack the ability to intelligently handle conflicts between multiple data sources. When faced with missing data, noise interference, or abnormal data collection, the system's robustness and reliability are difficult to guarantee. These technical limitations severely restrict the development of enterprise attendance management towards intelligence and refinement, urgently requiring an attendance management solution capable of achieving deep semantic association and multi-source collaborative analysis. Summary of the Invention
[0006] The purpose of this invention is to provide an enterprise attendance management system based on deep learning and multi-source data fusion, so as to solve the technical problems in the prior art that multi-source attendance data cannot achieve deep semantic association, the accuracy of attendance anomaly judgment is insufficient, and the system has poor robustness.
[0007] The technical solution of the present invention includes: The semantic association graph construction unit is used to construct a semantic association graph with employees as the main nodes and attendance behaviors corresponding to standardized attendance data from multiple heterogeneous data sources as the association edges. The semantic association graph construction unit has a preset semantic rule library, which includes time proximity rules, spatial co-occurrence rules, behavior sequence rules and business logic rules, which are used to determine the semantic association relationship between different attendance records and generate corresponding association edges. The deep learning fusion decision unit includes a fusion weight calculation module and a fusion result output module. The fusion weight calculation module uses a bidirectional long short-term memory network based on an attention mechanism to dynamically calculate the contribution weight of each data source in the fusion decision based on three dimensions: data quality, historical reliability, and real-time status. The fusion result output module is used to weight and fuse the contribution weights of each data source with the feature vectors of the corresponding attendance records to generate a comprehensive attendance feature vector.
[0008] Furthermore, when the semantic association graph construction unit performs the edge construction phase, it pairs the attendance records under each employee node and sequentially substitutes them into the time proximity rule, spatial co-occurrence rule, behavior sequence rule, and business logic rule for condition matching; when the condition of any rule is met, an independent association edge is created for the rule that triggers the association, and the rule type and association weight are written into the attribute field of the association edge.
[0009] Furthermore, the temporal proximity rule is used to associate attendance records from different data sources occurring within a preset 15-minute time threshold; the spatial co-occurrence rule is used to associate attendance records occurring simultaneously within a preset 50-meter spatial range, with the ground distance of the spatial location calculated using the Haversine formula; the behavioral sequence rule is used to associate time-series attendance records that conform to a regular pattern within a preset 7-day behavioral pattern window; and the business logic rule is used to associate business-related records that conform to the company's attendance system, including at least the association between leave approval and attendance records, the association between meeting room check-in and attendance anomalies, and the association between overtime applications and extended departure.
[0010] Furthermore, the data quality score calculated by the fusion weight calculation module is a weighted comprehensive score based on three dimensions: data integrity, data timeliness, and data consistency; the historical credibility score calculated by the fusion weight calculation module is obtained by statistically calculating the pass rate of the data source's historical attendance records over the past 30 days; and the real-time status score calculated by the fusion weight calculation module is assigned different scores based on the different states of the current data source: normal collection, delayed collection, or collection interruption.
[0011] Furthermore, it also includes a conflict detection and intelligent resolution unit for receiving the comprehensive attendance feature vector; The conflict detection and intelligent resolution unit includes a conflict detection module. The conflict detection module adopts an anomaly detection algorithm based on confidence propagation, compares the comprehensive attendance feature vector to be detected with the standard feature vector model extracted from historical normal attendance records, calculates the standardized deviation score of each dimension feature, and when the standardized deviation score is greater than the preset anomaly threshold, it determines that there is an anomaly in that dimension and triggers a conflict marker.
[0012] Furthermore, the conflict detection and intelligent resolution unit also includes an intelligent resolution module, which is equipped with an evidence priority strategy, a majority voting strategy, and a spatiotemporal consistency strategy. The intelligent resolution module attempts to resolve the marked conflicts in sequence according to the preset priority order of the evidence priority strategy, the majority voting strategy, and the spatiotemporal consistency strategy.
[0013] Furthermore, it also includes an attendance profile generation unit, which is used to generate a multi-dimensional attendance profile for each employee based on the comprehensive attendance feature vector, including the attendance stability dimension, spatiotemporal compliance dimension, behavioral regularity dimension, and abnormal risk dimension. The abnormal risk dimension is based on the comprehensive analysis results of the attendance stability dimension, spatiotemporal compliance dimension, and behavioral regularity dimension. It uses a pre-trained machine learning classification model to predict the probability of employees having attendance abnormalities within a preset time period in the future.
[0014] Furthermore, when constructing the behavioral regularity dimension, the attendance profile generation unit extracts the time-series features of the employee's historical attendance behavior, constructs a behavioral fingerprint model that includes daily behavior patterns, weekend behavior patterns, and holiday behavior patterns, and calculates the similarity between the employee's recent attendance behavior features and the behavioral fingerprint model to obtain a behavioral regularity score.
[0015] Furthermore, it also includes an intelligent early warning and decision support unit, which includes an intelligent early warning module and a decision support module; The intelligent early warning module sets multiple early warning thresholds. When the score of any dimension in the attendance profile is lower than the corresponding level of early warning threshold, an early warning signal containing the employee number, early warning level, and abnormal dimension identifier is automatically generated and sent synchronously through multiple channels. The decision support module is used to generate at least one of the following based on the comprehensive analysis results of the attendance profile: attendance rule optimization suggestions, abnormal employee interview suggestions, and shift adjustment suggestions.
[0016] Furthermore, the intelligent early warning module adopts a three-level early warning system. When the score of any dimension in the attendance profile is lower than 0.3, a level 3 early warning is triggered; when it is between 0.3 and 0.6, a level 2 early warning is triggered; and when it is between 0.6 and 0.85, a level 1 early warning is triggered.
[0017] In summary, this application includes at least one of the following beneficial technical effects: 1. This invention constructs a semantic association graph to deeply semantically associate heterogeneous attendance data from multiple sources such as access control, positioning, attendance machines, and leave systems. This enables semantic mapping and collaborative analysis between attendance data from different sources, solving the technical problem in existing technologies where multi-source data can only be simply decomposed and cannot be deeply integrated. The temporal proximity rules, spatial co-occurrence rules, behavioral sequence rules, and business logic rules in the semantic association graph work together to capture the inherent connections between attendance behaviors from multiple dimensions, providing rich semantic support for subsequent fusion decisions.
[0018] 2. This invention achieves adaptive dynamic fusion of multi-source attendance data through a deep learning fusion decision unit; the fusion weight calculation module comprehensively considers three dimensions: data quality, historical reliability, and real-time status, and can dynamically adjust the fusion weight according to the actual situation of each data source, avoiding the problem of fixed weight allocation in traditional methods that causes the fusion result to deviate from the actual situation; this data-driven adaptive fusion mechanism significantly improves the accuracy and reliability of attendance determination.
[0019] 3. This invention effectively improves the robustness of the system when facing data conflicts, noise interference, and acquisition anomalies through conflict detection and intelligent resolution units; the confidence propagation algorithm can accurately identify the abnormal dimensions in the fusion results, and the multi-strategy resolution mechanism can select the optimal resolution scheme according to the specific situation; compared with the prior art, the system of this invention no longer makes wrong judgments due to the anomaly of a single data source, but can make more reasonable decisions by comprehensively weighing multiple pieces of evidence.
[0020] 4. This invention uses an attendance profile generation unit to build a comprehensive and three-dimensional multi-dimensional attendance profile for each employee. The four dimensions of attendance stability, spatiotemporal compliance, behavioral regularity, and abnormal risk complement each other to form a complete employee attendance feature system. This multi-dimensional profile can not only accurately reflect the employee's attendance status, but also predict potential abnormal risks, providing more intelligent decision support for enterprise human resource management. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the overall technical solution for an enterprise attendance management system based on deep learning and multi-source data fusion. Figure 2 This is a schematic diagram illustrating the core principles of constructing a semantic association graph from multiple sources of data. Figure 3 This is a logical flowchart of a deep learning fusion decision unit; Figure 4 This is a flowchart of the conflict detection and intelligent resolution module. Figure 5 This is a schematic diagram of the output framework for generating multi-dimensional attendance profiles and intelligent early warning. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the following description is provided in conjunction with the appendix. Figure 1 To be continued Figure 5 Specific embodiments are provided to further illustrate the present invention in detail.
[0023] The overall architecture of the enterprise attendance management system provided in this embodiment includes a multi-source data acquisition unit, a semantic association graph construction unit, a feature extraction and vectorization module, a deep learning fusion decision-making unit, a conflict detection and intelligent resolution unit, an attendance profile generation unit, and an intelligent early warning and decision support unit. These units are connected sequentially according to the data flow direction, forming a complete attendance data processing chain. The specific implementation methods of each unit are described in detail below.
[0024] The multi-source data acquisition unit is used to collect raw attendance data from multiple heterogeneous data sources. Through standardized data interfaces, the multi-source data acquisition unit establishes data connections with the access control record module, mobile terminal positioning and tracking module, attendance machine clock-in record module, leave approval system module, and meeting room check-in system module, respectively. The acquisition frequency and data format of each data source are independent of each other.
[0025] The access control system continuously collects employee card-swiping records at a sampling frequency of one second, recording employee ID, access time, access control location, and direction of passage. The mobile terminal positioning and tracking module collects employee mobile phone location data every 30 seconds, recording employee ID, positioning time, latitude and longitude, and positioning accuracy. The attendance machine clock-in / clock-out recording module collects data triggered by each clock-in / clock-out event, recording employee ID, clock-in / clock-out time, clock-in / clock-out machine number, and clock-in / clock-out method.
[0026] The leave approval system module collects data according to the approval process, recording information including employee ID, leave start time, leave end time, leave type, and approval status. The meeting room check-in system module collects data based on each meeting check-in event, recording employee ID, meeting number, check-in time, and check-in location.
[0027] The five data source modules mentioned above cover attendance-related information in five dimensions: employee access, location trajectory, clock-in records, leave approval, and meeting check-in. Together, they form a multi-source heterogeneous attendance data collection system.
[0028] The multi-source data acquisition unit integrates a unified standardized interface and format conversion module. The standardized interface supports API calls based on the HTTP / RESTful architecture style, and the data exchange format adopts JSON, ensuring that data from various data sources can be seamlessly accessed in accordance with a unified format specification. The standardized interface also supports message queue access based on the AMQP protocol to adapt to the asynchronous data reception requirements in high-throughput scenarios.
[0029] The format conversion module maintains a complete set of data mapping rules, which define the correspondence between source data fields and target standardized fields. When heterogeneous data enters the format conversion module, the module first identifies the data source type, and then converts the source data format into a standardized format according to the corresponding data mapping rules.
[0030] The standardized data structure includes a unified timestamp field, a data source identifier field, a data quality marker field, and a data integrity marker field. The unified timestamp field uses a millisecond-level precision timestamp format to ensure accurate alignment and comparison of time information from different data sources. The data source identifier field identifies which data source module each record originates from. The data quality marker field records quality-related information generated during the data collection process. The data integrity marker field indicates whether a data field is complete. These standardized fields provide fundamental support for subsequent data quality assessment and fusion decisions in the modules.
[0031] As a preferred implementation, the multi-source data acquisition unit also collects employee performance data and project task data. Employee performance data is collected through the enterprise human resources system interface, including employee ID, performance evaluation cycle, performance score, performance level, and evaluation time. Project task data is collected through the project management system interface, including employee ID, project ID, task start time, task end time, task completion status, and task duration. The collected employee performance data and project task data are synchronously transmitted to the attendance profile generation unit via an internal message queue, supplementing the business feature dimensions of the attendance profile.
[0032] The semantic association graph construction unit is used to construct a semantic association graph between multiple data sources based on standardized attendance data collected by the multi-source data acquisition unit.
[0033] The semantic association graph construction unit uses employees as the main nodes and attendance behaviors from various data sources as the association edges. A pre-defined semantic rule base defines the semantic mapping relationships between different attendance behaviors. The graph uses a graph database as its underlying storage engine, and both the node and edge data structures are defined according to a unified graph data model. Each employee node stores basic attribute information such as employee ID, name, and department; each association edge stores attribute information such as source node identifier, target node identifier, association rule type, association weight, and creation timestamp.
[0034] The semantic rule base contains four core semantic rules: temporal proximity rules, spatial co-occurrence rules, behavioral sequence rules, and business logic rules. These four rules define the semantic association conditions between attendance records from different dimensions, and together they constitute the basis system for semantic association judgment.
[0035] The time proximity rule is used to associate attendance records from different data sources that occur within a preset time threshold. In this embodiment, the time threshold for the time proximity rule is set to 15 minutes. When the time difference between two attendance records from different data sources is less than or equal to 15 minutes, the semantic association graph construction unit determines that these two records have a time proximity relationship. For example, if an employee enters the office building through the access control at 9:00 AM and clocks in at the attendance machine at 9:05 AM, the time difference between these two records is 5 minutes, which is less than the 15-minute time threshold, and therefore they are determined to have a time proximity relationship.
[0036] Spatial co-occurrence rules are used to associate attendance records that occur simultaneously within a preset spatial range. In this embodiment, the spatial range threshold for the spatial co-occurrence rules is set to 50 meters. When the actual ground distance between the location information of two attendance records is less than or equal to 50 meters, the semantic association graph construction unit determines that these two records have a spatial co-occurrence relationship. The ground distance of spatial locations is calculated using the Haversine formula, which takes into account the curvature of the Earth and can accurately calculate the spherical distance between two points on the Earth's surface. The Haversine formula is expressed as follows:
[0037] in, This represents the spherical distance between two points, expressed in meters. This represents the Earth's average radius, with a value of 6,371,000 meters. and These represent the latitudes of the first and second points, respectively, in radians; This represents the difference in latitude between two points, i.e. The unit is radians; and These represent the longitudes of the first and second points, respectively, in radians. This represents the difference in longitude between two points, i.e. The unit is radians. The semantic association graph construction unit substitutes the collected location latitude and longitude data into the above formula to calculate the ground distance, and compares the calculation result with a spatial threshold of 50 meters to determine whether two attendance records meet the spatial co-occurrence association condition.
[0038] Behavioral sequence rules are used to associate time-series attendance records that conform to preset behavioral patterns. In this embodiment, the behavioral pattern window for the behavioral sequence rules is set to 7 days. The semantic association graph construction unit analyzes the employee's attendance behavioral sequences within the 7-day window period to identify regular patterns. For example, if an employee arrives at the company, enters the office building, and clocks in within similar time periods on several consecutive workdays, this behavioral sequence is identified as a regular behavioral pattern. When a new attendance record conforms to this behavioral pattern, the semantic association graph construction unit automatically establishes a behavioral sequence association relationship.
[0039] Business logic rules are used to associate business-related records that conform to the company's attendance system. These rules are based on the company's pre-defined attendance business logic and primarily include three categories: the association between leave approval and attendance records, the association between meeting room check-in and attendance exceptions, and the association between overtime applications and extended leave. For example, if an employee submits a leave application for the morning session and the application is approved, then the employee's access control records and attendance machine check-in records for that morning will be marked as associated with the leave application, so that the time period will be excluded from normal attendance assessment during subsequent attendance evaluation.
[0040] The construction process of the semantic association graph includes four stages: graph initialization, node insertion, edge construction, and graph update. Each stage is executed sequentially to complete the construction of the complete semantic association graph from the original attendance data.
[0041] During the graph initialization phase, the semantic association graph construction unit creates a blank graph storage structure upon startup. Specifically, the semantic association graph construction unit sends an initialization command to the graph database, creating data tables or datasets to store employee nodes, attendance record nodes, and associated edges, and establishing index relationships between the various storage structures. After initialization, the graph is in a writable state, awaiting the injection of new data.
[0042] During the node insertion phase, the semantic association graph construction unit inserts employee information as the main nodes into the graph and establishes association pointers between each employee node and attendance records from various data sources. Specifically, the semantic association graph construction unit obtains standardized attendance data from the multi-source data acquisition unit, extracts the employee ID field, and creates corresponding employee nodes for employees who do not yet have nodes in the graph; for employees with existing nodes, it directly obtains their node identifier. Each employee node is assigned a globally unique node identifier upon creation, and the graph maintains the pointing relationships between this node and its related attendance records.
[0043] During the edge construction phase, the semantic association graph construction unit traverses standardized attendance data based on temporal proximity rules, spatial co-occurrence rules, behavioral sequence rules, and business logic rules in the semantic rule base. It identifies attendance record pairs that meet the rule conditions and generates corresponding association edges. Specifically, the semantic association graph construction unit reads standardized attendance data according to a preset time window, pairs attendance records under each employee node, and sequentially substitutes each record pair into the four semantic rules for condition matching. When a rule's condition is met, the semantic association graph construction unit creates an association edge in the graph from the employee node to the attendance record node, and writes the rule type that triggered the association, the association weight, and the creation timestamp into the edge's attribute fields. The same record pair may simultaneously satisfy multiple semantic rules. In this case, the semantic association graph construction unit creates an independent association edge for each satisfying rule and records its respective rule type.
[0044] During the graph update phase, when new data arrives, the semantic association graph construction unit inserts new nodes and edges into the existing graph using an incremental update method, while simultaneously updating the graph's statistical information. The incremental update is triggered when the multi-source data acquisition unit completes a standardized data output and sends a data update signal to the semantic association graph construction unit, which then initiates the incremental update process. Specifically, the semantic association graph construction unit only performs node insertion and edge construction operations on newly received standardized attendance data, avoiding redundant processing of historical data already existing in the graph. This method controls the graph size and maintains update efficiency. After the incremental update is completed, the semantic association graph construction unit synchronously updates the affected statistical information in the graph, including the number of associated edges for each node and the frequency statistics of triggering various rule types.
[0045] Through the sequential execution of the above four stages, the semantic association graph construction unit organizes attendance records from multiple heterogeneous data sources, such as access control, location tracking, attendance tracking, leave approval, and meeting room check-in, into semantically related nodes, with employees as the main nodes. This forms structured graph data that can be used for subsequent feature extraction and fusion decision-making. The multi-dimensional semantic association information contained in this graph data, such as temporal proximity, spatial co-occurrence, behavioral sequences, and business logic, provides a rich source of input features for the feature extraction and vectorization modules.
[0046] The feature extraction and vectorization module is used to extract features from each main node and associated edge in the semantic association graph, generating a high-dimensional feature vector.
[0047] The feature extraction and vectorization module receives the graph data output by the semantic association graph construction unit and performs four operations on it: temporal feature extraction, spatial feature extraction, behavioral feature extraction, and business feature extraction. The four feature extractions represent the original attendance data from the time, space, behavior, and business dimensions, respectively. After generating corresponding feature vectors for each, they are concatenated to form a unified high-dimensional feature vector for use by the subsequent fusion decision unit.
[0048] Temporal feature extraction employs a Long Short-Term Memory (LSTM) network to extract periodic patterns and abnormal fluctuations from time-series attendance data. The LTM network addresses the vanishing and exploding gradient problems inherent in traditional recurrent neural networks when processing long-sequence data through three gating mechanisms: a forget gate, an input gate, and an output gate.
[0049] The network requires training before use. Training data is extracted from a semantic association graph. Each employee's attendance records are segmented into time-series samples according to a 30-day time window. Each sample is labeled with the employee's attendance status for the next day, including five categories: normal attendance, late, early departure, absent, and leave. The loss function uses a combination of cross-entropy loss and L2 regularization. The training objective is to minimize this loss function, enabling the network to accurately predict employees' future attendance status based on the input time series. Training uses the backpropagation algorithm, with Adam as the optimizer. Training terminates early when the validation set loss no longer decreases after 10 consecutive training iterations.
[0050] After training, the feature extraction and vectorization module segments the employees' historical attendance time series data into time windows. Each segment is input into the trained Long Short-Term Memory (LSTM) network, and the output vector of the last hidden layer of the network is extracted as the temporal feature vector. This vector captures the periodic patterns of employees' attendance times, including fixed attendance days per week and fixed attendance times per day, and can also identify sudden abnormal fluctuations such as lateness, early departure, or absence.
[0051] Spatial feature extraction employs graph neural networks to extract topological features from the relationships between spatial locations. Graph neural networks capture the topological relationships between nodes through the information transmission mechanism between neighboring nodes.
[0052] This network requires training before use. Training data extracts spatial location nodes and their associated edges from a semantic association graph, forming a spatial location relationship subgraph. Each node's label represents the compliance of the attendance behavior corresponding to that location. The loss function is the cross-entropy loss for node classification, and the training objective is to minimize this loss function, enabling the network to determine the attendance compliance of each location node based on its characteristics and its topological relationships with neighboring nodes. Training employs the backpropagation algorithm, and the optimizer is Adam.
[0053] After training, the feature extraction and vectorization module constructs a graph structure from the spatial location nodes and the edges connecting them in the semantic association graph. This graph structure is then input into the trained graph neural network to extract the embedding vectors of each spatial location node and perform average pooling to obtain the spatial feature vector. This vector captures employees' common attendance routes, including the path from home to the company and movement patterns between different areas within the company.
[0054] Behavioral feature extraction employs an attention mechanism to adaptively assign weights to different types of attendance behaviors.
[0055] The feature extraction and vectorization module encodes employee attendance behaviors in different scenarios into sequences of behavioral feature vectors. Attendance behavior types include five categories: access control, location-based check-in, clock-in / out records, leave applications, and meeting check-in. Each behavior type is encoded based on its recorded fields. After encoding, the attention mechanism calculates a similarity score between the learnable query vector and each behavioral feature vector. After normalization, a weight distribution is obtained. Finally, the behavioral feature vectors are weighted and summed to generate a weighted behavioral feature vector. The attention mechanism automatically adjusts the contribution weight of each behavior type according to the needs of the current analysis scenario. For example, when determining whether attendance time is compliant, clock-in / out and access control behaviors are given higher weights.
[0056] Business feature extraction is based on a business rule engine, extracting compliance features that conform to the company's attendance system.
[0057] The business rule engine stores attendance policy rules in the form of IF-THEN rules. The feature extraction and vectorization module matches employee attendance records one by one with the rules in the rule base to identify compliant and non-compliant behaviors and generate business feature vectors. These vectors specifically include quantitative indicators such as the number of normal attendance days, the number of times late, the number of times early departure, the number of leave days, the number of absences, and the amount of overtime.
[0058] The feature extraction and vectorization module concatenates the feature vectors from the four dimensions into a unified high-dimensional feature vector in a preset order. The total dimension is set according to the actual application scenario, typically ranging from several hundred to several thousand dimensions. After concatenation, the module outputs this vector to the deep learning fusion decision unit as input for subsequent adaptive fusion of multi-source data.
[0059] The deep learning fusion decision unit in this embodiment is used to make adaptive fusion decisions based on high-dimensional feature vectors generated by the feature extraction and vectorization module.
[0060] The deep learning fusion decision unit consists of two parts: a fusion weight calculation module and a fusion result output module. The fusion weight calculation module is responsible for determining the contribution of each data source in the fusion decision, while the fusion result output module completes the fusion of feature vectors according to the determined weights.
[0061] The fusion weight calculation module employs a bidirectional long short-term memory network based on an attention mechanism. It dynamically calculates the contribution weight of each data source in the fusion decision based on three dimensions: data quality, historical reliability, and real-time status. The bidirectional long short-term memory network is a neural network structure that adds a backpropagation path to the long short-term memory network, enabling it to simultaneously capture both forward and backward information from the sequence data, thus providing a more comprehensive understanding of the time-series dependencies between the feature vectors of each data source.
[0062] This bidirectional long short-term memory network requires training before use. The training data comes from historical attendance records. The specific construction method is as follows: historical feature vectors from various data sources are obtained from the feature extraction and vectorization module. Feature vectors from different data sources at the same time are combined into an input sequence for one time step. A complete sample contains feature sequences from 60 consecutive time steps. The label of each sample is the optimal weight allocation vector determined by manual verification for each data source at the corresponding time.
[0063] The specific process for manual verification is as follows: Attendance management experts, based on the final confirmation status in historical attendance data, combined with the data integrity of each data source at that moment, its consistency with the actual situation, and the reasonableness of the collection timestamp, comprehensively determine the confidence level of each data source and normalize the confidence level into a weight allocation vector. The weight allocation vector is an N-dimensional non-negative real-number vector, where N is the total number of data sources, and the sum of all dimensions is 1. The loss function is designed as the mean squared error loss of weight prediction, and the training objective is to minimize this loss function so that the network can accurately predict reasonable weight allocation based on the characteristics of each input data source and its real-time status.
[0064] The network structure employs a two-layer bidirectional long short-term memory (LSTM) network, with each hidden layer having a dimension of 256. A Dropout layer with a dropout rate of 0.5 is appended after each LSM layer. During training, backpropagation is used to update network parameters. The optimizer is Adam, with an initial learning rate of 0.001, a batch size of 64, and a maximum training epoch of 200. An early stopping strategy is employed, terminating training prematurely when the validation set loss no longer decreases within 10 consecutive epochs. A gradient pruning strategy is also used, clipping the gradient norm to below 1.0 to prevent gradient explosion.
[0065] After training, the weight calculation module receives feature vectors from each data source from the feature extraction and vectorization module, as well as data quality scores from the data quality assessment module, credibility scores from the historical credibility statistics module, and status scores from the real-time status monitoring module. These scores, along with the feature vectors, are then input into the trained bidirectional long short-term memory network. The network learns the optimal weight allocation for each data source under different contexts through an attention mechanism.
[0066] The weight distribution output by the attention mechanism reflects the relative importance of each data source in the current fusion decision. Data sources with high data quality, high historical reliability, and good real-time status will receive higher fusion weights, while those with lower data quality and better real-time status will receive lower fusion weights.
[0067] Data quality assessment is based on a comprehensive score across three dimensions: data integrity, data timeliness, and data consistency. In this embodiment, the weight for data integrity assessment is set to 0.3, the weight for data timeliness assessment is set to 0.4, and the weight for data consistency assessment is set to 0.3. The specific calculation method for data integrity assessment is as follows: check for missing data fields, calculate the ratio of the number of data fields actually collected to the total number of data fields that should have been collected, and this ratio is the data integrity score.
[0068] The specific calculation method for data timeliness assessment is as follows: Examine the difference between the data's timestamp and the current time. The smaller the difference, the fresher the data, and the higher the timeliness score. Specifically, the reciprocal of the normalized time difference is used as the timeliness score. The specific calculation method for data consistency assessment is as follows: Check whether the records of the same attendance event are consistent across different data sources. The higher the proportion of consistent source data, the higher the consistency score. The scores of the above three dimensions are weighted and summed according to their respective weights to obtain the comprehensive data quality score.
[0069] Historical credibility is calculated statistically based on the pass rate of historical attendance records from this data source. In this embodiment, the statistical calculation period for historical credibility is set to 30 days. The fusion weight calculation module counts the number of attendance records submitted by this data source in the past 30 days that have been confirmed as correct after manual verification or cross-validation, and calculates the ratio of the number of correct records to the total number of records, which serves as the historical credibility score for this data source.
[0070] Real-time status is determined based on the current data source's acquisition status. Acquisition status includes three types: normal acquisition, delayed acquisition, and acquisition interruption. Normal acquisition indicates that the data source is currently acquiring data at a preset frequency, with a corresponding real-time status score of 1.0; delayed acquisition indicates that the data source's acquisition frequency is lower than the preset frequency but is still operating, with a corresponding real-time status score of 0.5; acquisition interruption indicates that the data source is currently unable to provide data, with a corresponding real-time status score of 0.0.
[0071] The fusion result output module weights and corresponding feature vectors of each data source to generate a fused comprehensive attendance feature vector. The fusion process uses a weighted summation method: the feature vector of each data source is multiplied by its fusion weight, and then all weighted feature vectors are summed element-wise to obtain the comprehensive attendance feature vector. This comprehensive attendance feature vector fuses information from multiple data sources, including access control, location tracking, attendance tracking, leave approval, and meeting room check-in. Compared to feature vectors from a single data source, it has stronger representational capabilities, providing a more comprehensive data foundation for subsequent attendance determination.
[0072] After completing the above fusion operation, the deep learning fusion decision unit outputs the comprehensive attendance feature vector to the conflict detection and intelligent resolution unit, which then performs anomaly detection and conflict resolution processing on the fusion result.
[0073] The conflict detection and intelligent resolution unit performs conflict detection and intelligent resolution on the comprehensive attendance feature vector generated by the deep learning fusion decision unit. This unit comprises two parts: a conflict detection module and an intelligent resolution module. The conflict detection module identifies potential anomalies or conflict dimensions in the comprehensive attendance feature vector, while the intelligent resolution module employs corresponding resolution strategies for identified conflicts. These two modules work in tandem to complete the entire processing flow from anomaly identification to result correction.
[0074] The conflict detection module employs an anomaly detection algorithm based on confidence propagation to calculate the degree of deviation between each dimension of the comprehensive attendance feature vector and the normal pattern. This algorithm models and performs probabilistic inference on the dependencies between feature dimensions using a factor graph model.
[0075] The specific construction method of the factor graph is as follows: each dimension of the comprehensive attendance feature vector is defined as a variable node, and the state of the variable node is whether the dimension is abnormal, with a value set of {normal, abnormal}; the constraint relationship between dimensions is defined as a factor node, which includes unary factor nodes, paired factor nodes, and global factor nodes.
[0076] Unary factor nodes connect to single variable nodes, and their potential function is mapped to the prior probability of anomalies based on the standardized deviation score of that dimension using the sigmoid function. Paired factor nodes connect correlated feature dimension pairs, and their potential function is constructed based on the conditional probability of both dimensions occurring simultaneously in historical data. The correlation includes the relationship between the time and spatial dimensions and the temporal dependency between adjacent behavioral sequence dimensions. Global factor nodes connect all variable nodes, and their potential function is used to constrain the total number of anomalous dimensions at the same time to not exceed a preset upper limit.
[0077] The message passing mechanism is implemented using a sum-product algorithm to perform iterative message propagation on the factor graph. In each iteration, each variable node sends a message to its neighboring factor nodes, and the message content is the probability distribution vector of the variable node's state. Each factor node sends a message to its neighboring variable nodes, and the message content is the confidence information of the variable's state calculated by the factor node based on its potential function and the received variable messages.
[0078] The update formula for message passing is as follows: the message passed by a variable node to a factor node is equal to the product of the messages received by the variable node from other adjacent factor nodes; the message passed by a factor node to a variable node is equal to the factor node multiplying the joint potential function of all its adjacent variable nodes by the product of the messages passed by other adjacent variable nodes.
[0079] Before the message passing iteration begins, each variable node initializes its prior confidence based on its own standardized deviation score. During the iteration process, messages from each node are updated synchronously. In this embodiment, the convergence tolerance is set to 1×10⁻. 4 When the sum of the absolute values of the confidence changes of all nodes in two adjacent iterations is less than the tolerance, the algorithm is considered to have converged and the iteration is stopped. If the algorithm has not converged after the number of iterations reaches the preset maximum number of iterations of 10, the iteration is forcibly terminated and the current confidence of each node is used as the final inference result.
[0080] After the iteration converges, the conflict detection module calculates the posterior anomaly probability for each feature dimension. When the posterior anomaly probability is greater than the preset anomaly threshold of 0.75, it determines that there is an anomaly in that dimension and triggers a conflict marker.
[0081] During the conflict detection process, the conflict detection module first establishes a standard feature vector model for normal attendance patterns. This standard model is constructed as follows: It extracts comprehensive attendance feature vectors corresponding to a large number of attendance records confirmed as normal from historical attendance data. The mean and standard deviation of these feature vectors are calculated for each dimension. The mean of each dimension forms the standard feature vector, and the standard deviation of each dimension forms the confidence interval for that vector. This standard model represents the typical value distribution of each dimension's features for employees under normal attendance conditions.
[0082] After the standard model is constructed, the conflict detection module compares the comprehensive attendance feature vector to be detected with the standard model and calculates the degree of deviation of each dimension. The degree of deviation is calculated as follows: for each dimension, the difference between the value of the target vector in that dimension and the mean of the corresponding dimension of the standard vector is calculated, and then divided by the standard deviation of that dimension to obtain the standardized deviation score of that dimension.
[0083] In this embodiment, the anomaly threshold is set to 0.75. When the standardized deviation score of a feature in a certain dimension is greater than 0.75, the conflict detection module determines that the dimension is anomaly and triggers a conflict marker. The confidence propagation algorithm uses the message passing mechanism of the confidence propagation network to iteratively propagate the confidence information between nodes of each dimension on the factor graph, with a maximum number of iterations set to 10. After the iteration converges, the conflict detection module outputs the conflict marker information for each dimension. The conflict marker information includes an anomaly dimension identifier, an anomaly degree value, and a confidence score, where the confidence score indicates the credibility of the anomaly judgment.
[0084] The intelligent conflict resolution module resolves marked conflicts based on preset resolution strategies. These strategies include three types: evidence-first strategy, majority voting strategy, and spatiotemporal consistency strategy, each suitable for different conflict scenarios.
[0085] The evidence-first strategy prioritizes the data source records with the highest confidence level. When a conflict occurs, the intelligent resolution module compares the confidence scores of each relevant data source, selects the data source with the highest confidence level as the resolution basis, and uses the attendance records from that data source as the final judgment result.
[0086] The majority voting strategy arbitrates based on the consensus of the majority of data sources. When a conflict arises, the intelligent resolution module tallies the attendance determination results from each relevant data source and adopts the majority determination as the final resolution. For example, in determining whether an employee is late, if two of the three data sources—access control data, location data, and attendance machine data—determine whether an employee is late and one does not, the final resolution result is late.
[0087] The spatiotemporal consistency strategy combines contextual information from both the temporal and spatial dimensions for comprehensive judgment. When a conflict occurs, the intelligent resolution module analyzes the temporal and spatial information of the relevant attendance records and selects the record that best matches the overall spatiotemporal context as the resolution result.
[0088] For example, an employee's mobile phone location data shows that they are near the company, but access control records show that they failed to pass through the access control. In this case, the spatiotemporal consistency strategy prioritizes the location trajectory data because the location data is consistent with the spatial context in which the employee appeared near the company.
[0089] When performing conflict resolution, the intelligent resolution module attempts each strategy sequentially according to a preset priority order. The priority order is as follows: first, the evidence-first strategy is tried; if this strategy successfully resolves the conflict, the resolution result is output directly. If the evidence-first strategy fails, the majority voting strategy is tried. If the majority voting strategy also fails, the spatiotemporal consistency strategy is tried last. This multi-strategy progressive resolution mechanism ensures reasonable resolution results in various conflict scenarios.
[0090] After completing the conflict detection and resolution process, the conflict detection and intelligent resolution unit outputs the processed comprehensive attendance feature vector to the attendance profile generation unit, which then generates a multi-dimensional attendance profile for each employee.
[0091] The attendance profile generation unit is used to generate a multi-dimensional attendance profile for each employee based on the fused attendance feature vector processed by the conflict detection and intelligent resolution unit.
[0092] The attendance profile generated by the attendance profile generation unit includes four dimensions: attendance stability, spatiotemporal compliance, behavioral regularity, and abnormal risk. These four dimensions characterize employee attendance from the perspectives of attendance frequency and time regularity, spatial location compliance, behavioral pattern stability, and potential risk probability, respectively, and together constitute a complete employee attendance feature system.
[0093] The attendance stability dimension is quantitatively evaluated based on the statistical characteristics of attendance frequency and attendance time.
[0094] Attendance frequency is calculated by comparing the actual number of days an employee attended with the number of days they were expected to attend within the statistical period, then calculating the ratio to arrive at the attendance rate. Attendance time is calculated by recording the time of each employee's first attendance record each day, obtaining the time distribution of that employee's first attendance record over a period of time, and then calculating statistical indicators such as the average attendance time, the standard deviation of attendance time, and the skewness of attendance time. The average attendance time reflects the typical arrival time of employees, the standard deviation of attendance time reflects the degree of fluctuation in employee arrival time, and the skewness of attendance time reflects the distribution pattern of employee arrival times; a positive skewness indicates that employees arrive relatively late most of the time, while a negative skewness indicates that employees arrive relatively early most of the time.
[0095] The attendance stability dimension ultimately outputs a quantitative score, with the score ranged from 0 to 1. The higher the score, the more stable the attendance.
[0096] The spatiotemporal compliance dimension is based on the matching analysis of the consistency between employee location trajectories and preset attendance areas.
[0097] The preset attendance areas include different types of restricted attendance areas such as office areas, factory areas, and warehouse areas. The boundaries of each area are pre-stored in the system database as geographic coordinate polygons. The attendance profile generation unit acquires the employee's location trajectory data, performs spatial overlay analysis on it with the preset attendance areas, determines whether each location point falls within the corresponding attendance area boundary, and counts the duration and frequency of the employee's stay within the required attendance area to assess the employee's spatiotemporal compliance. If the employee's location trajectory mostly falls within the preset attendance area, the spatiotemporal compliance score is high; if the employee frequently appears outside the preset attendance area, the score is low.
[0098] The spatiotemporal compliance dimension outputs a quantitative score, ranging from 0 to 1. A higher score indicates better spatiotemporal compliance.
[0099] The behavioral regularity dimension calculates similarity based on the periodic patterns of employees' historical attendance behavior.
[0100] The attendance profile generation unit extracts the time-series features of employees' historical attendance behavior to construct an employee behavioral fingerprint model. The behavioral fingerprint model includes behavioral templates at three different time granularities: daily behavior patterns, weekend behavior patterns, and holiday behavior patterns, corresponding to typical attendance behavior characteristics of employees during weekdays, weekends, and statutory holidays, respectively. The behavioral fingerprint model is constructed by performing cluster analysis on employees' historical attendance records according to the three time granularities, and extracting the most frequent attendance behavior sequence at each time granularity as the behavioral fingerprint for that time granularity.
[0101] When generating attendance profiles, the attendance profile generation unit matches the employee's recent attendance behavior characteristics with the three behavioral fingerprint models mentioned above, calculates the matching score for each behavior pattern, and then obtains a comprehensive behavioral regularity score according to the preset weighting coefficients.
[0102] The behavioral regularity dimension outputs a quantitative score, ranging from 0 to 1. A higher score indicates better behavioral regularity.
[0103] The abnormal risk dimension is based on a comprehensive analysis of the attendance stability dimension, spatiotemporal compliance dimension, and behavioral regularity dimension to predict the probability of employees having attendance abnormalities.
[0104] The calculation of the anomaly risk dimension uses a machine learning classification model. This model needs to be trained before use.
[0105] The training data is constructed as follows: a large number of samples are extracted from historical attendance data. The input features of each sample consist of three parts: basic features, statistical features, and contextual features. Basic features include: employee attendance stability score, spatiotemporal compliance score, behavioral regularity score, attendance rate, average attendance time deviation, attendance time standard deviation, number of days with spatiotemporal compliance deviation, and number of behavioral regularity deviations within a 30-day time window. Statistical features include: the number and frequency of various attendance anomalies within a 30-day time window, including types such as lateness, early departure, absence, abnormal departure, and false sign-in; the number and frequency of conflict markers from various data sources within a 30-day time window; and the total number of attendance anomalies, average interval days, longest consecutive normal attendance days, and longest consecutive abnormal attendance days for the employee in the past 90 days. Contextual features include: the historical average anomaly rate of the employee's department, the attendance baseline pattern for the employee's position, the employee's length of service, whether it is a seasonal peak business period, the day of the week, and the week of the month.
[0106] The feature engineering process for the input features is as follows: For numerical continuous features, Z-score normalization is used, which calculates the mean and standard deviation for each feature dimension, subtracts the mean from the feature value, and divides by the standard deviation to ensure that each feature dimension has zero mean and unit variance, thus eliminating the influence of different units on model training; for categorical discrete features, one-hot encoding is used to map each category to a binary vector; for features with missing values, median imputation is used to complete the imputation. After feature engineering, all input features are concatenated into a 78-dimensional numerical vector.
[0107] The label for each sample indicates whether the employee actually committed an attendance violation within the subsequent 7-day time window. The label value is either 0 or 1, where 0 indicates no violation and 1 indicates a violation. In this embodiment, the time window length is set to 30 days, and the subsequent prediction window length is set to 7 days. The training data covers the historical records of at least 500 employees. For each employee, sampling is performed in 15-day increments along the timeline to ensure that multiple training samples are generated for each employee, with a total sample size of at least 3000. All samples are divided into training, validation, and test sets in an 8:1:1 ratio.
[0108] The loss function used is binary cross-entropy loss, and the training objective is to minimize this loss function so that the classification model can accurately predict the probability of employee attendance irregularities within a future period based on the four dimensions of the input features. Backpropagation is used to update the model parameters during training.
[0109] After training, the attendance profile generation unit inputs the current employee's attendance stability score, spatiotemporal compliance score, behavioral regularity score, and historical anomaly records into the classification model. The model outputs a predicted probability of the employee experiencing attendance anomalies within the next 7 days. The anomaly risk dimension outputs a quantitative score, ranging from 0 to 1, with higher scores indicating a greater risk of attendance anomalies.
[0110] The attendance profile generation unit outputs a structured data format containing fields such as employee ID, generation time, attendance stability score, spatiotemporal compliance score, behavioral regularity score, anomaly risk score, and overall score. The overall score is obtained by weighting and summing the scores from the four dimensions according to preset weights, and is used to reflect the overall attendance status of the employee. The specific calculation method for the overall score is as follows: the attendance stability dimension weight is set to 0.3, the spatiotemporal compliance dimension weight is set to 0.25, the behavioral regularity dimension weight is set to 0.25, and the anomaly risk dimension weight is set to 0.2. The scores of the four dimensions are multiplied by their respective weights and then summed to obtain the overall score.
[0111] Attendance profile data is transmitted to the intelligent early warning and decision support unit via an internal interface for subsequent early warning and decision support use. This profile data provides the basis for early warning judgments for the intelligent early warning module and the data foundation for generating management recommendations for the decision support module.
[0112] Finally, the intelligent early warning and decision support unit responds to the employee attendance profiles output by the attendance profile generation unit, and outputs early warning signals and decision suggestions based on the profile data. The intelligent early warning and decision support unit consists of two parts: an intelligent early warning module and a decision support module. The intelligent early warning module is responsible for triggering early warning signals based on the scoring data of the attendance profiles, while the decision support module is responsible for generating management suggestions based on the comprehensive analysis of the profiles. The two modules are independent yet complementary, together forming a complete closed loop from risk identification to management response.
[0113] The intelligent early warning module sets multiple early warning thresholds. When the evaluation result of any dimension in the attendance profile falls into the corresponding early warning range, an early warning signal of the corresponding level is generated.
[0114] In this embodiment, the intelligent early warning module adopts a three-level early warning system: Level 1 early warning corresponds to minor anomalies with a threshold of 0.3; Level 2 early warning corresponds to moderate anomalies with a threshold of 0.6; and Level 3 early warning corresponds to severe anomalies with a threshold of 0.85. A Level 3 early warning is triggered when the score for any dimension in the attendance profile is below 0.3, a Level 2 early warning is triggered when it is between 0.3 and 0.6, and a Level 1 early warning is triggered when it is between 0.6 and 0.85. An early warning signal is automatically triggered when the score for any dimension falls below the corresponding threshold.
[0115] The warning signal contains four fields: employee ID, warning level, anomaly dimension identifier, and warning generation time. The warning level indicates the severity of the anomaly, while the anomaly dimension identifier specifies whether the problem lies in attendance stability, temporal and spatial compliance, behavioral patterns, or anomaly risk, allowing managers to quickly pinpoint the source of the problem.
[0116] The warning signal is simultaneously sent to pre-defined administrators through three channels: system message push to send the warning information to the administrator's system inbox; SMS notification to send a warning summary to the administrator's pre-bound mobile phone number; and email notification to send the complete warning report to the administrator's pre-bound email address. These three channels serve as backups for each other, ensuring that warning information reaches administrators in a timely manner.
[0117] The decision support module generates targeted management recommendations based on the comprehensive analysis results of attendance profiles. These recommendations include three categories: suggestions for optimizing attendance rules, suggestions for interviewing employees with abnormal behavior, and suggestions for adjusting work schedules. These recommendations provide decision-making references for enterprise management from the perspectives of systems, personnel, and arrangements, respectively.
[0118] The attendance rule optimization suggestions are generated based on statistical analysis of the attendance profiles of all employees. The decision support module aggregates the attendance profile data of all employees, analyzes the distribution of scores across various dimensions, and identifies unreasonable rules in the attendance system that may lead to employee violations. For example, when statistics show that a large number of employees' attendance times are concentrated before the earliest attendance time stipulated in the attendance system, it indicates a discrepancy between the current earliest attendance time setting and employees' actual attendance habits. The decision support module then suggests to management that the earliest attendance time be appropriately advanced to reduce the waiting time for employees arriving early.
[0119] The abnormal employee interview suggestion is generated for employees with high scores in the abnormal risk dimension of the attendance profile. The decision support module sorts all employees in descending order according to the abnormal risk score, automatically filters out the list of employees with an abnormal risk score higher than 0.7, and proposes an interview suggestion to the management. The interview suggestion includes the employee's basic information, attendance abnormality history, abnormality type analysis, and key points for the interview, helping the management understand the specific abnormal situation of the employee before the interview, improving the targeting and efficiency of the interview.
[0120] The scheduling adjustment recommendations are generated based on the analysis results of employees' temporal and spatial compliance dimensions and behavioral patterns. The decision support module identifies employees who frequently have compliance issues in specific time periods or areas, analyzes the conflict between the actual needs of their work tasks and the requirements of the attendance system, and proposes scheduling adjustment recommendations to management. The scheduling adjustment recommendations specifically include the type of shift to be adjusted, the time range of the adjustment, and an explanation of the reasons for the adjustment.
[0121] The intelligent early warning and decision support unit is the final output stage of the attendance management system of this invention. This unit receives structured profile data from the attendance profile generation unit, promptly notifies managers of potential risks in the form of early warning signals through the intelligent early warning module, and transforms the data analysis results into actionable management suggestions through the decision support module. Thus, the seven functional units—multi-source data acquisition unit, semantic association graph construction unit, feature extraction and vectorization module, deep learning fusion decision unit, conflict detection and intelligent resolution unit, attendance profile generation unit, and intelligent early warning and decision support unit—are sequentially connected, forming a complete attendance data processing chain from data acquisition, semantic association, feature extraction, fusion decision, conflict resolution, profile generation to early warning output.
[0122] The enterprise attendance management system in this embodiment achieves deep semantic association, adaptive dynamic fusion, intelligent conflict resolution, and multi-dimensional attendance profile generation through the collaborative work of the seven functional units described above. The system can not only accurately determine employee attendance status but also predict potential attendance anomaly risks, providing intelligent decision support for enterprise human resource management.
[0123] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects.
[0124] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. An enterprise attendance management system based on deep learning and multi-source data fusion, characterized in that, include: The semantic association graph construction unit is used to construct a semantic association graph with employees as the main nodes and attendance behaviors corresponding to standardized attendance data from multiple heterogeneous data sources as the association edges. The semantic association graph construction unit has a preset semantic rule library, which includes time proximity rules, spatial co-occurrence rules, behavior sequence rules and business logic rules, which are used to determine the semantic association relationship between different attendance records and generate corresponding association edges. The deep learning fusion decision unit includes a fusion weight calculation module and a fusion result output module; The fusion weight calculation module uses a bidirectional long short-term memory network based on an attention mechanism to dynamically calculate the contribution weight of each data source in the fusion decision based on information from three dimensions: data quality, historical reliability, and real-time status. The fusion result output module is used to weight and fuse the contribution weights of each data source with the feature vectors of the corresponding attendance records to generate a comprehensive attendance feature vector.
2. The enterprise attendance management system according to claim 1, characterized in that, When the semantic association graph construction unit performs the edge construction phase, it pairs the attendance records under each employee node and sequentially substitutes them into the time proximity rule, spatial co-occurrence rule, behavior sequence rule and business logic rule for condition matching; when the condition of any rule is met, an independent association edge is created for the rule that triggers the association, and the rule type and association weight are written into the attribute field of the association edge.
3. The enterprise attendance management system according to claim 2, characterized in that, The temporal proximity rule is used to associate attendance records from different data sources that occur within a preset 15-minute time threshold; the spatial co-occurrence rule is used to associate attendance records that occur simultaneously within a preset 50-meter spatial range, with the ground distance of the spatial location calculated using the Haversine formula; the behavioral sequence rule is used to associate time-series attendance records that conform to a regular pattern within a preset 7-day behavioral pattern window; the business logic rule is used to associate business-related records that conform to the company's attendance system, including at least the association between leave approval and attendance records, the association between meeting room check-in and attendance exceptions, and the association between overtime applications and delayed departures.
4. The enterprise attendance management system according to claim 1, characterized in that, The data quality score calculated by the fusion weight calculation module is a weighted comprehensive score based on three dimensions: data integrity, data timeliness, and data consistency. The historical credibility score calculated by the fusion weight calculation module is obtained by statistically calculating the pass rate of the data source's historical attendance records over the past 30 days. The real-time status score calculated by the fusion weight calculation module is based on the different scores assigned to the data source in different states: normal collection, delayed collection, or collection interruption.
5. The enterprise attendance management system according to claim 1, characterized in that, It also includes a conflict detection and intelligent resolution unit, used to receive the comprehensive attendance feature vector; The conflict detection and intelligent resolution unit includes a conflict detection module. The conflict detection module adopts an anomaly detection algorithm based on confidence propagation, compares the comprehensive attendance feature vector to be detected with the standard feature vector model extracted from historical normal attendance records, calculates the standardized deviation score of each dimension feature, and when the standardized deviation score is greater than the preset anomaly threshold, it determines that there is an anomaly in that dimension and triggers a conflict marker.
6. The enterprise attendance management system according to claim 5, characterized in that, The conflict detection and intelligent resolution unit also includes an intelligent resolution module. The intelligent resolution module is equipped with an evidence priority strategy, a majority voting strategy, and a spatiotemporal consistency strategy. It attempts to resolve the marked conflicts in sequence according to the preset priority order of the evidence priority strategy, the majority voting strategy, and the spatiotemporal consistency strategy.
7. The enterprise attendance management system according to claim 1, characterized in that, It also includes an attendance profile generation unit, which is used to generate a multi-dimensional attendance profile for each employee based on the comprehensive attendance feature vector, including the attendance stability dimension, spatiotemporal compliance dimension, behavioral regularity dimension, and abnormal risk dimension. The abnormal risk dimension is based on the comprehensive analysis results of the attendance stability dimension, spatiotemporal compliance dimension, and behavioral regularity dimension. It uses a pre-trained machine learning classification model to predict the probability of employees having attendance abnormalities within a preset time period in the future.
8. The enterprise attendance management system according to claim 7, characterized in that, When constructing the behavioral regularity dimension, the attendance profile generation unit extracts the time-series features of employees' historical attendance behavior, constructs a behavioral fingerprint model that includes daily behavior patterns, weekend behavior patterns, and holiday behavior patterns, and calculates the similarity between the employee's recent attendance behavior features and the behavioral fingerprint model to obtain a behavioral regularity score.
9. The enterprise attendance management system according to claim 7, characterized in that, It also includes an intelligent early warning and decision support unit, which comprises an intelligent early warning module and a decision support module; The intelligent early warning module sets multiple early warning thresholds. When the score of any dimension in the attendance profile is lower than the corresponding level of early warning threshold, an early warning signal containing the employee number, early warning level, and abnormal dimension identifier is automatically generated and sent synchronously through multiple channels. The decision support module is used to generate at least one of the following based on the comprehensive analysis results of the attendance profile: attendance rule optimization suggestions, suggestions for interviewing abnormal employees, and suggestions for adjusting work schedules.
10. The enterprise attendance management system according to claim 9, characterized in that, The intelligent early warning module adopts a three-level early warning system. When the score of any dimension in the attendance profile is lower than 0.3, a level 3 early warning is triggered; when it is between 0.3 and 0.6, a level 2 early warning is triggered; and when it is between 0.6 and 0.85, a level 1 early warning is triggered.