A staff image matching and archiving method and system based on face recognition
By introducing multi-source auxiliary information and evidence chain analysis into the facial recognition system, the problem of identity fragmentation in complex environments of traditional systems has been solved, realizing automated and accurate image archiving and improving the efficiency and data quality of power industry archive management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
- Filing Date
- 2026-04-17
- Publication Date
- 2026-08-04
AI Technical Summary
Traditional facial recognition systems are prone to identity splitting in complex environments, especially in frontline operations in the power industry. Uneven lighting and facial occlusion can cause feature distortion, resulting in the same employee's image being incorrectly archived as multiple independent identities, affecting data accuracy and management efficiency.
By introducing multi-source auxiliary information on the basis of facial recognition, a chain of evidence is constructed for multi-dimensional analysis, the confidence level of file merging is calculated, and image files are automatically merged or given an early warning, reducing manual intervention.
It has enabled the automated and accurate archiving of employee image files, reduced manual intervention, improved the efficiency and data quality of file management, and ensured the uniformity and integrity of image files.
Smart Images

Figure CN122049967B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of records management technology, and more specifically, to a method and system for archiving employee images based on facial recognition. Background Technology
[0002] In modern enterprise management systems, intelligent management of employee image data has become a key aspect of improving operational efficiency. Traditional image archiving systems based on facial recognition typically employ a static feature comparison mechanism. Their core logic involves binary classification of acquired images based on a preset similarity threshold: if a newly acquired image's feature similarity to an employee in the archive exceeds the threshold, it is directly archived; otherwise, it is treated as a new, independent identity. This simplistic approach performs adequately in controlled environments, but it reveals serious flaws in real-world industrial scenarios.
[0003] Taking the power industry as an example, on-site workers are required to wear protective equipment for extended periods due to safety regulations, resulting in key facial features being obscured by masks, goggles, and other items. Existing systems often generate distorted feature vectors when extracting features from these images due to the lack of effective biometrics. More seriously, when the system compares these distorted features with standard ID photo features, even if the employee belongs to the same person, the similarity score will plummet below the threshold due to the occlusion effect. At this point, the system may incorrectly trigger an incremental learning mechanism, creating a redundant "new identity" file for that employee, resulting in a "one person, multiple files" identity split.
[0004] Existing technology suffers from three core flaws: First, it relies solely on the similarity of a single facial feature, failing to consider the feature degradation patterns under operational scenarios; second, the incremental learning mechanism lacks error correction capabilities, and incorrectly archived images contaminate the original database; third, it lacks effective automatic repair methods after identity fragmentation, requiring manual intervention. These problems lead to the generation of numerous fragmented image archives in practical application scenarios such as substation inspections and line repairs, resulting in the scattered storage of formal activity images and work record images of the same employee, severely disrupting data correlation. When enterprises conduct safety audits or performance evaluations, the system cannot fully present the employee's comprehensive work trajectory, creating management blind spots.
[0005] Current solutions primarily address the issue from a hardware perspective, such as requiring employees to remove protective gear when taking work photos or deploying costly multispectral acquisition equipment. These methods not only violate safety regulations but are also impractical in scenarios like nighttime repairs and high-risk operations. Another direction for software improvement is to lower the similarity threshold, but this would lead to a surge in mismatches between different employees, creating a more serious identity confusion problem.
[0006] There is currently no effective technical solution to the above problems. Summary of the Invention
[0007] The purpose of this invention is to provide a method and system for matching and archiving employee images based on facial recognition. It aims to solve the technical problem that traditional facial recognition systems are prone to identity splitting in complex environments, realize the automated and accurate archiving of employee image files, reduce the need for manual intervention, and significantly improve the efficiency and data quality of file management.
[0008] In a first aspect, the present invention provides a face recognition-based employee image matching and archiving method, which is applied to an employee image matching and archiving server. The employee image matching and archiving server has a pre-set real image archive and an incremental image archive. The real image archive is used to store real image archives of known employees. The incremental image archive is used to store image archives to be verified created through incremental learning. The employee image matching and archiving method based on facial recognition includes the following steps: S1. For each known employee, calculate the facial feature similarity between their real image file and each image file to be examined; S2. Identify all the image files to be examined whose facial feature similarity is lower than a preset similarity threshold and higher than a preset association threshold, and form a set of images to be examined for the corresponding known employees. S3. Determine whether there is a split identity for the corresponding known employee based on the image data of the image archives in the image set to be examined, and collect multi-source auxiliary information related to the image set to be examined for the known employee with a split identity. S4. For each image file in the image set to be examined, a chain of evidence for identity verification is constructed by performing multi-dimensional analysis on multi-source auxiliary information; the chain of evidence includes the analysis results of each dimension. S5. Based on the analysis results of each dimension in the chain of evidence, calculate the file merging confidence of each file in the set of images to be examined; S6. If the confidence level of file merging is greater than or equal to the preset merging threshold, then the corresponding image file to be examined in the image set to be examined will be merged into the real image file of the corresponding known employee. S7. If the confidence level for merging the files is less than the merging threshold, an early warning message will be generated and submitted for manual review.
[0009] The employee image matching and archiving method based on face recognition provided by this invention effectively avoids the problem of identity splitting caused by feature occlusion, improves the accuracy and automation of image archiving, and reduces the need for manual intervention.
[0010] Secondly, the present invention provides an employee image matching and archiving system based on face recognition, which is applied to an employee image matching and archiving server. The employee image matching and archiving server has a real image archive and an incremental image archive. The real image archive is used to store real image archives of known employees. The incremental image archive is used to store image archives to be verified created through incremental learning. The employee image matching and archiving system based on facial recognition includes: The first calculation module is used to calculate the facial feature similarity between the real image archives of each known employee and the image archives to be examined. The recognition module is used to identify all the image files to be examined whose facial feature similarity is lower than a preset similarity threshold and higher than a preset association threshold, and to form a set of images to be examined for the corresponding known employees. The judgment module is used to determine whether a known employee has a split identity based on the image data of the image archives in the image set to be examined, and to collect multi-source auxiliary information related to the image set to be examined for known employees with split identities. The analysis module is used to construct a chain of evidence for identity verification by performing multi-dimensional analysis on each image file in the image set to be examined through multi-source auxiliary information; the chain of evidence includes the analysis results of each dimension. The second calculation module is used to calculate the file merging confidence of each file in the image set to be examined based on the analysis results of each dimension in the chain of evidence. The merging module is used to merge the corresponding image files in the image set to be examined into the real image files of the corresponding known employees if the confidence level of file merging is greater than or equal to the preset merging threshold. The early warning module is used to generate early warning information and submit it for manual review if the confidence level of file merging is less than the merging threshold.
[0011] As can be seen from the above, the employee image matching and archiving method based on face recognition provided by the present invention achieves automatic archiving or early warning through steps such as calculating facial feature similarity, constructing a set to be verified, constructing an evidence chain using multi-source auxiliary information, and calculating confidence. It effectively solves the problem of identity splitting caused by feature occlusion, improves the accuracy and automation of image archiving, and reduces the need for manual intervention.
[0012] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description
[0013] Figure 1This is a flowchart of an employee image matching and archiving method based on face recognition provided in an embodiment of the present invention.
[0014] Figure 2 This is a schematic diagram of a structure of an employee image matching and archiving system based on face recognition provided in an embodiment of the present invention.
[0015] Label Explanation: 100. First calculation module; 200. Identification module; 300. Judgment module; 400. Analysis module; 500. Second calculation module; 600. Merging module; 700. Early warning module. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0017] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0018] In traditional facial recognition systems applied to employee image matching and archiving, when input image data is collected in complex working environments, such as unstable lighting conditions or partial facial obscuration by protective equipment, the extracted facial feature information is often incomplete and contains interference. This results in facial feature similarity calculations falling below a preset similarity threshold, triggering an incremental learning mechanism to create new image archives for verification. The essence of this problem lies in the system's reliance on a single facial feature similarity for identity determination, failing to effectively handle feature distortion caused by environmental factors. This leads to identity splitting—the same employee's image is incorrectly archived into multiple independent identity archives. This identity splitting reduces the accuracy and completeness of the image archive database, impacting retrieval efficiency and management reliability.
[0019] For example, in the front-line production departments of a power group, field engineers are required to wear safety helmets and reflective goggles to meet safety regulations when conducting equipment inspections at substations. When an engineer uses their personal mobile phone to take work-related photos outdoors in strong sunlight, the key areas of their face are obscured due to the shadow of the helmet brim and the reflection from the goggles, resulting in incomplete feature information extracted by the system. When this photo is uploaded to the image matching and archiving system and compared with the engineer's standard ID photo in the archive, the similarity score is below the similarity threshold. The system identifies the person as unknown and creates a new image archive for verification. As the engineer subsequently uploads more work photos under similar conditions, these photos are continuously archived into the newly created archive, causing their identity file to split into two: one associated with formal activity photos and the other associated with on-site work photos. When the safety management department retrieves the engineer's work records for auditing, the system can only retrieve a portion of the archive; a large number of on-site photos cannot be associated due to archiving errors, resulting in missing information.
[0020] If the above problems are not addressed, the image archive will gradually accumulate a large amount of incorrectly archived data, forming information silos. This will prevent searches based on name or employee number from returning complete image records, leading to the omission of critical information in the safety production traceability process. In key business processes such as personnel performance evaluation and safety audits, data incompleteness and inconsistency will introduce management risks, affecting the accuracy and timeliness of decision-making. In the long run, the contamination of the archive will weaken the system's credibility and increase the workload of manual review.
[0021] For reference, see the appendix. Figure 1 This invention provides a face recognition-based employee image matching and archiving method, applied to an employee image matching and archiving server. The employee image matching and archiving server has a pre-set real image archive and an incremental image archive; the real image archive is used to store real image archives of known employees; the incremental image archive is used to store image archives to be verified created through incremental learning. The employee image matching and archiving method based on facial recognition includes the following steps: S1. For each known employee, calculate the facial feature similarity between their real image file and each image file to be examined; S2. Identify all the image files to be examined whose facial feature similarity is lower than a preset similarity threshold and higher than a preset association threshold, and form a set of images to be examined for the corresponding known employees. S3. Determine whether there is a split identity among the known employees based on the image data of the image archives in the image archives to be examined, and collect multi-source auxiliary information related to the image archives to be examined for the known employees with split identities; the multi-source auxiliary information includes the shooting timestamp and geographical location contained in the image data of the image archives to be examined, as well as the team task data and behavior record data; S4. For each image file in the image set to be verified, a chain of evidence for identity verification is constructed by performing multi-dimensional analysis on multi-source auxiliary information; the chain of evidence includes the analysis results of each dimension; the dimensions used for analysis include time dimension, geographical location dimension, work group task dimension, and behavioral pattern dimension; S5. Based on the analysis results of each dimension in the chain of evidence, calculate the file merging confidence of each file in the set of images to be examined; S6. If the confidence level for merging the archives is greater than or equal to the preset merging threshold, then the corresponding archives of the images to be examined in the set of images to be examined are merged into the real archives of the corresponding known employees, and the comprehensive feature representation of the real archives of the corresponding employees is updated; the steps for updating the comprehensive feature representation of the real archives of the corresponding employees include: S61. Select multiple face images with a quality rating higher than a preset quality threshold from the merged image data and extract the depth feature vectors of the multiple face images; S62. Perform weighted fusion processing on the deep feature vectors to generate a global feature template for representing the real identity profile and use it as a comprehensive feature representation; S7. If the confidence level for merging the files is less than the merging threshold, an early warning message will be generated and submitted for manual review.
[0022] For ease of understanding, the following explains some key terms in this embodiment: Employee Image Matching and Archiving Server: This server is the core execution unit of this method, and it has a pre-set real image archive and an incremental image archive. Its main functions are to receive, process, store, and manage various types of employee image data, and to perform intelligent matching and archiving based on facial recognition results and multi-source auxiliary information.
[0023] Real Image Archive: This archive stores real image archives of known employees. These archives typically include standard ID photos, high-quality work photos, etc., and serve as the foundational data for the system's identity verification and comparison.
[0024] Incremental Image Archive: This archive stores image files to be verified, created through incremental learning. When the system determines, through facial feature recognition, that a person in an image does not belong to any known employee in the existing image archive, it automatically creates an image file to be verified through incremental learning, associating it with the identified "new person." These "new people's" images may resemble the facial features of a known employee, but because they do not meet the preset similarity threshold, they are classified as "associated new people," awaiting further identity verification.
[0025] Image archives awaiting verification: These refer to image records created by the system through incremental learning in the incremental image archive database that require further verification of their true identity. These archives are usually images where facial recognition similarity is insufficient due to complex environments (such as uneven lighting or facial occlusion), but which still have a certain correlation.
[0026] Multi-source auxiliary information refers to various data sources used to assist in identity verification, in addition to facial features. This information includes the shooting timestamps and geographical locations contained in the image data of the image file to be verified, as well as team task data and behavioral record data. This auxiliary information provides the system with a more comprehensive context, compensating for the limitations of single facial feature recognition.
[0027] Chain of evidence: refers to a logical sequence of connections constructed for identity verification after multi-dimensional analysis of multiple sources of auxiliary information. The chain of evidence includes the analysis results of various dimensions, such as time, geographical location, work group tasks, and behavioral patterns; these results collectively support the determination of identity.
[0028] File merging confidence level: This refers to the quantitative indicator calculated by the system based on the analysis results of each dimension of the evidence chain, indicating the probability that the image file to be verified and the actual image file belong to the same person. This confidence level is used to guide the system to automatically merge files or submit them for manual review.
[0029] Comprehensive feature representation: This refers to the process of processing the merged image data after file merging, selecting high-quality facial images, extracting depth feature vectors, and then performing weighted fusion processing to generate a global feature template that represents the authentic identity file. This representation can more comprehensively and accurately reflect the facial features of employees in different scenarios.
[0030] This application proposes a face recognition-based employee image matching and archiving method, applied to an employee image matching and archiving server. This method aims to address the problem of identity splitting caused by misidentification in complex environments by face recognition systems. The employee image matching and archiving server has a pre-set real image archive and an incremental image archive. The real image archive stores real image archives of known employees, while the incremental image archive stores pending image archives created through incremental learning. When the system determines, based solely on facial feature recognition, that a person in an image does not belong to any known employee in the real image archive, it automatically creates a pending image archive to associate with the identified "new person" through incremental learning. This "new person" has similar facial features to a known employee, but is classified as an "associated new person" because it does not meet a preset similarity threshold. In traditional systems, the image files of "newcomers" awaiting verification require manual review and archiving by staff. This application, however, uses facial feature recognition to filter out image files of "newcomers" with similar characteristics for each known employee. It further combines multi-source auxiliary information to determine if the "newcomer" is a known employee. If identified as such, the image file is automatically archived into the corresponding known employee's real image file. Compared to traditional methods, this application not only considers the similarity of facial features but also incorporates more auxiliary information to calculate confidence levels, using these confidence levels to accurately archive the image files.
[0031] Specifically, the method includes the following steps: In step S1, for each known employee, the facial feature similarity between their real image file and each image file to be examined is calculated. The purpose of this step is to initially screen out image files to be examined that have a certain correlation with the facial features of the known employees. For example, various algorithms such as cosine similarity and Euclidean distance can be used to calculate the similarity between facial feature vectors. In practical applications, the system extracts the comprehensive facial feature template of the known employees from the real image file database and compares it with the facial features of each image file to be examined in the incremental image file database.
[0032] In step S2, all candidate image files with facial feature similarity below a preset similarity threshold but above a preset association threshold are identified, forming a candidate image set for the corresponding known employee. This step aims to identify images where facial feature similarity is insufficient to directly confirm the same identity, but a certain degree of correlation exists. For example, the similarity threshold can be set to 0.85, and the association threshold can be set to 0.5. When the similarity between a candidate image file and a known employee's actual image file is between 0.5 and 0.85, the candidate image file is included in the candidate image set for that known employee. This screening method effectively focuses on potential identity splits, avoiding complex, indiscriminate analysis of all images.
[0033] In step S3, the system determines whether a known employee has a split identity based on the image data of the image files in the image set to be examined. For known employees with split identities, multi-source auxiliary information related to their corresponding image sets to be examined is collected. This multi-source auxiliary information includes the shooting timestamps and geographical locations contained in the image data of the image files to be examined, as well as team task data and behavioral record data. This step is one of the core elements of this method, aiming to assist in identity determination using non-facial feature information. For example, the system continuously monitors the updates to employee image files and periodically scans all employee files. For each employee, the system checks whether there are multiple independent image sets under their name. Although these sets are currently identified as different identities by the system, the facial feature similarity between them remains within a low but non-zero range. Simultaneously, the system also checks whether the images in these independent image sets have significant overlap in shooting timestamps and geographical location information. For example, if two image sets labeled with different identities are found, and more than 60% of the images within these sets were captured within the same week, and the average distance between their geographical coordinates is less than 500 meters, the system will mark these two image sets as "potential identity splits" and trigger the subsequent file merging and review process. This monitoring mechanism can proactively identify "shadow files" created by the system misjudging low-quality images caused by uneven lighting, facial occlusion, and other factors in frontline operations in the power industry as new identities, thus providing the triggering conditions for subsequent merging. To reduce resource consumption, this step can determine whether identity splits exist based on only some dimensions, such as time and geographical location.
[0034] In step S4, for each image file in the image set to be verified, a chain of evidence for identity verification is constructed by performing multi-dimensional analysis on multi-source auxiliary information. The chain of evidence includes the analysis results of each dimension, including time, geographic location, work group / task, and behavioral pattern. This step aims to integrate multi-source auxiliary information to form a comprehensive basis for identity verification. For example, the system collects auxiliary data related to these potentially fragmented files from multiple information sources within the enterprise and integrates them to form a "chain of evidence." These information sources include, but are not limited to, image metadata (such as shooting time and GPS coordinates in EXIF information), work group and task allocation records in the Enterprise Resource Planning (ERP) system, and personnel entry / exit and clock-in records in access control and attendance systems. This integration of multi-source information aims to compensate for the reduced accuracy of single facial features in complex environments such as strong light, weak light, and obstructions from safety helmets, goggles, and masks during frontline power operations, by introducing more stable behavioral and environmental information to enhance the reliability of identity verification. To ensure accurate results when performing this step, a more comprehensive analysis is needed. The rational allocation of resources can improve system efficiency and facilitate the processing of large amounts of employee image data.
[0035] In step S5, based on the analysis results of each dimension in the evidence chain, the combined confidence level of each image file in the image set to be examined is calculated. This step aims to quantify the reliability of identity verification. For example, the system assigns a weight to each of the above evidence chains (time series analysis, geographical location association, comparison of team and task information, historical behavior pattern analysis) and calculates a score based on its analysis results (e.g., overlap, distance, matching rate). The combined confidence level C can be calculated by weighted summation: C = w_time * S_time + w_geo * S_geo + w_team * S_team + w_behavior * S_behavior. Where w_time, w_geo, w_team, and w_behavior are preset weights, for example, w_time = 0.25, w_geo = 0.25, w_team = 0.3, w_behavior = 0.2. S_time is the score for the time dimension, S_geo is the score for the geographical location dimension, S_team is the score for the team / task dimension, and S_behavior is the score for the behavior pattern dimension. For example, if the time overlap is greater than 0.7, S_time = 1, otherwise it is 0.
[0036] In step S6, if the file merging confidence level is greater than or equal to a preset merging threshold, the corresponding image file to be examined in the image set to be examined is merged into the real image file of the corresponding known employee, and the comprehensive feature representation of the real image file of the corresponding employee is updated. This step aims to achieve automated file archiving and feature optimization. When the calculated merging confidence level C exceeds a preset threshold (e.g., 0.8), the system will trigger an automatic merging operation. The specific steps of automatic merging include: selecting a master file, usually a file with more images, higher quality, or earlier creation time as the master file; data migration, migrating all image data, metadata, and associated behavioral records in all "shadow files" to the master file; feature update, the system will recalculate the comprehensive facial feature representation of the master file. The steps of updating the comprehensive feature representation of the real image file of the corresponding employee include: S61. Selecting multiple facial images with a quality evaluation higher than a preset quality threshold from the merged image data and extracting the depth feature vectors of multiple facial images; S62. Performing weighted fusion processing on the depth feature vectors to generate a global feature template for representing the real identity file and using it as the comprehensive feature representation. This can be achieved by averaging, weighted averaging, or using more sophisticated feature fusion techniques (such as multi-instance learning) on the facial features of all merged images to ensure that the new composite features better represent the employee's appearance in various complex scenarios. For example, for occlusion situations common in frontline power operations, the new composite features will include more stable feature information from unoccluded areas. Finally, "shadow profiles" are marked as merged or deleted to ensure data uniqueness.
[0037] In step S7, if the file merging confidence level is less than the merging threshold, an early warning message is generated and submitted for manual review. This step aims to ensure that the system introduces human intervention when there is uncertainty, avoiding erroneous archiving. If the calculated merging confidence level C is lower than the automatic merging threshold (e.g., 0.8) but higher than an early warning threshold (e.g., 0.5), the system generates an "awaiting manual review" early warning. This early warning is sent to the designated file manager via the company's internal messaging system (e.g., DingTalk, WeChat Work, or email). The early warning message includes: the names of the employees involved (if known) and temporary identifiers for all relevant files; the merging confidence level calculated by the system; detailed analysis results and scores for each chain of evidence, such as a time overlap of 0.65, a geographical distance of 300 meters, and a team matching rate of 0.8; thumbnails and key metadata of all relevant images.
[0038] The core technical concept of this method lies in its departure from the limitations of relying solely on facial features for identity verification in complex scenarios. Instead, it employs a strategy of "multi-source information backtracking and evidence chain analysis" to proactively correct identity discrepancies in the personnel files. When the system detects a potential misidentification of the same employee with multiple identities, it initiates a comprehensive review process. This process systematically collects and integrates information such as the image capture time, geographical location, employee's work group and task assignments, and attendance and access control records, weaving this auxiliary information into a robust "chain of evidence." By quantitatively assessing the strength of this chain of evidence, the system can intelligently determine whether to automatically merge these fragmented files or submit them for manual review, thereby ensuring the uniformity and integrity of employee image files in challenging environments such as frontline operations in the power industry, characterized by variable lighting and severe facial occlusion.
[0039] The following example will provide a more detailed explanation of the above technical solution: Employee A's existing image archive contains their standard ID photo and some clear meeting photos. Because Employee A frequently performs frontline power line inspections, photos of them wearing a safety helmet and goggles outdoors are uploaded to the system. Traditional facial recognition systems, due to uneven lighting and facial occlusion, identify these work photos as belonging to "newcomers" who do not belong to the known Employee A. Through incremental learning, a new image archive is created for this new employee and added to the incremental image archive database. As Employee A uploads more similar work photos, the system continuously archives these photos under this newly created "newcomer" archive, leading to a "split identity" for Employee A within the system.
[0040] First, in step S1, the system calculates the facial feature similarity between the known employee A's real image file and all the image files to be verified in the incremental image file database. The system finds that although the similarity between these image files to be verified and employee A's real image file is lower than the conventional identity verification threshold (e.g., 0.85), they are not completely dissimilar.
[0041] Next, in step S2, the system identifies all image files to be examined that have a facial feature similarity lower than a preset similarity threshold (e.g., 0.85) and higher than a preset association threshold (e.g., 0.5). These image files to be examined constitute the image set corresponding to the known employee A. This indicates that the system initially determines that these "newcomer" files may be associated with employee A.
[0042] Subsequently, in step S3, the system determines whether employee A has a split identity based on the image data of the image files in the image set to be examined. By analyzing the shooting timestamps and geographical locations of the images in the image set, the system finds that these images are highly concentrated in time (e.g., over 60% of the images were taken within the same week), and the average distance between their geographical coordinates is very close (e.g., less than 500 meters), which closely matches employee A's daily inspection routes and working hours. Therefore, the system determines that employee A has a split identity and begins collecting multi-source auxiliary information related to the image set to be examined, including the image shooting timestamps, geographical locations, and employee A's team task data obtained from the Enterprise Resource Planning (ERP) system and behavioral record data obtained from the attendance system.
[0043] In step S4, the system constructs a chain of evidence for identity verification by performing multi-dimensional analysis on the collected multi-source auxiliary information for each image file in the image set to be verified. Specifically: In the time dimension analysis, the system extracts the shooting timestamps of employee A's actual image archive and the image archive to be examined, and calculates the overlap between them. For example, the system finds that the number of overlapping images between the two sets of timestamps within a preset time window (e.g., 1 hour) is very high, with a time overlap exceeding 0.7.
[0044] In the geographic location dimension analysis, the system extracts the geographic location coordinates of the images and calculates the distance between the geographic center points of the real image archive and the image archive to be examined. For example, the system finds that the distance between two center points is less than 500 meters, and most of the images fall within the geofence of a substation managed by employee A.
[0045] In the team task dimension analysis, the system compares the team task data of employee A's real image file and the image file to be verified. For example, the system finds that the image uploader or the person identified in the image file to be verified belongs to "Power Inspection Team A" and is assigned to "Substation Maintenance Task" in the same time period as employee A during the image shooting period.
[0046] In the behavioral pattern analysis, the system compares the behavioral record data of employee A's real image file and the image file to be verified. For example, the system found that the identities associated with the two files had more than 80% of the same attendance clock-in times on the same day, or both participated in the same mandatory security training.
[0047] These analytical results together constitute a chain of evidence for identity verification.
[0048] Next, in step S5, the system calculates the combined confidence level of each image file in the image set to be examined based on the analysis results of each dimension in the evidence chain. The system assigns preset weights to the time dimension, geographical location dimension, team task dimension, and behavioral pattern dimension (e.g., w_time=0.25, w_geo=0.25, w_team=0.3, w_behavior=0.2), and performs a weighted sum based on the scores of each dimension. For example, the calculated combined confidence level C is 0.9.
[0049] Finally, in step S6, since the file merging confidence score of 0.9 is greater than or equal to the preset merging threshold of 0.8, the system automatically merges the corresponding image files to be examined from the image set to the known real image files of employee A. The system selects employee A's original real image files as the master file and migrates all image data, metadata, and associated behavioral records from all "new employee" files to the master file. At the same time, the system updates the comprehensive feature representation of employee A's real image files. Specifically, the system selects multiple face images with quality evaluations higher than the preset quality threshold from the merged image data and extracts the depth feature vectors of these images. Subsequently, these depth feature vectors are weighted and fused to generate a global feature template representing employee A's true identity, which serves as the new comprehensive feature representation. This new comprehensive feature representation can better represent employee A's facial features in complex scenarios such as wearing a safety helmet and goggles. The original "new employee" files are marked as merged.
[0050] Through the above process, the method of this application successfully corrects the "identity split" problem caused by facial recognition misjudgment due to complex environment, ensuring the uniformity and integrity of employee image archives.
[0051] The technical concept of this application, by introducing multi-source auxiliary information and multi-dimensional analysis, constructs a chain of evidence and calculates the confidence level of file merging, thereby automatically handling identity splitting problems in complex environments and avoiding misjudgments caused by relying solely on facial feature similarity, achieving efficient and accurate archiving. Compared with traditional methods, the advantages of this application are: First, traditional systems are prone to misjudgment when facing complex environments (such as uneven lighting and facial occlusion in frontline power operations), relying solely on facial feature recognition. This can lead to the same employee being incorrectly identified as having multiple identities, resulting in "identity splitting." For example, a work photo of employee A wearing a safety helmet and goggles will show a significantly lower similarity between their facial features and a standard ID photo. Traditional systems would identify them as a "newcomer" and create a new profile for them. This application, through steps S1 and S2, identifies image files with similarity scores below a conventional threshold but above a correlation threshold, based on a preliminary calculation of facial feature similarity. This allows the system to focus on potential identity splitting situations, rather than simply treating all low-similarity images as irrelevant.
[0052] Secondly, traditional systems typically require manual review and archiving of these misidentified "newcomer" files, which is not only costly in terms of manpower but also inefficient, easily leading to information silos and missing records. In step S3 of this application, based on the shooting timestamps and geographical locations of the image data in the image set to be examined, it determines whether there is identity discrepancy and collects multi-source auxiliary information, such as team task data and behavioral record data. This introduction of multi-source information compensates for the decreased accuracy of single facial features in complex environments, providing richer context for subsequent identity verification. For example, even if the face is obscured, if the shooting time and location of the image highly match employee A's team tasks and attendance records, it strongly suggests that these images belong to employee A.
[0053] Furthermore, in step S4, this application constructs an evidence chain for identity verification by performing multi-dimensional analysis of multi-source auxiliary information across time, geographic location, work group tasks, and behavioral patterns. For example, by analyzing the overlap between the image capture time and employee A's work time series, the correlation between the image's geographic location and employee A's work area, the matching degree between the personnel in the image and employee A's work group tasks, and the consistency between the personnel's behavioral patterns in the image and employee A's historical behavioral records, a comprehensive judgment basis is formed. This enables the system to cross-verify identity from multiple perspectives, greatly improving the accuracy and robustness of identity verification.
[0054] Finally, in steps S5 and S6, this application calculates the file merging confidence level based on the analysis results of each dimension in the evidence chain, and automatically performs file merging and updates the comprehensive feature representation based on the confidence level. For example, when the calculated merging confidence level reaches a preset threshold, the system automatically merges the image file to be examined into the real image file, and generates a more representative global feature template by screening high-quality images and performing weighted fusion. This not only solves the identity split problem, but also dynamically optimizes the comprehensive feature representation of employees, improving the accuracy of subsequent identification. For cases with insufficient confidence, the system generates an early warning message in step S7 and submits it for manual review, ensuring accuracy under uncertain conditions.
[0055] In summary, this application effectively solves the technical problem of identity splitting caused by traditional facial recognition systems in complex environments through the strategy of "multi-source information backtracking and evidence chain analysis," realizing the automated and accurate archiving of employee image files and significantly improving the efficiency and data quality of file management.
[0056] In some embodiments, step S3, which involves determining whether a known employee has a split identity based on the image data of the image files in the image set to be examined, includes: S31. Based on the image data of the image archives to be examined in the image collection, obtain the shooting timestamp and geographical coordinates of each image archive to be examined; S32. A spatiotemporal density clustering algorithm is used to identify image clusters that meet the preset spatiotemporal density conditions by analyzing the shooting timestamps and geographic coordinates; S33. Based on the identified image clusters, determine whether there is a split identity among the corresponding known employees.
[0057] Obtaining the capture timestamps and geographic coordinates of each image archive to be examined aims to extract key metadata from the image data. A capture timestamp refers to the specific point in time the image was recorded, usually accurate to the second, reflecting the time sequence of image generation. Geographic coordinates refer to the latitude and longitude information of the image at the time it was captured, reflecting the spatial location of the image. This information can be obtained in several ways. One common method is to parse the EXIF (Exchange Image File Format) data embedded in the image file, which typically contains the model of the capturing device, the capture time, GPS location information, etc. Another method is to query the database associated with the image archive, which may have already extracted and stored this metadata during the initial image upload or processing phase.
[0058] A spatiotemporal density clustering algorithm is employed to identify image clusters that meet preset spatiotemporal density conditions by analyzing capture timestamps and geographic coordinates. Spatiotemporal density clustering is an analytical method that can effectively process large-scale spatiotemporal data and identify densely distributed areas of data points in time and space. Its core idea is to discover groups with similar spatiotemporal characteristics based on the density of data points within a preset spatiotemporal neighborhood. The algorithm can be implemented in ways including, but not limited to: an extension based on DBSCAN (Density-Based Spatial Clustering of Applications with Noise), which defines a metric that comprehensively considers temporal and spatial distance to identify core points, boundary points, and noise points, thereby forming image clusters; or a grid-based clustering method, which discretizes a continuous spatiotemporal region into a series of grid cells, then determines the density based on the number of images within each grid cell, and connects adjacent grid cells that meet the density conditions to form image clusters.
[0059] Based on the identified image clusters, this step determines whether there is identity splitting among known employees. This step aims to determine whether known employees have identity splitting based on the output of a spatiotemporal density clustering algorithm. Image clusters represent highly concentrated activity patterns in image data within a specific time period and geographical area. The determination can be based on the following logic: if a known employee's actual image archive typically exhibits a stable spatiotemporal activity pattern, while its corresponding set of images to be examined (i.e., images with low facial feature similarity but not completely excluded) forms another or more image clusters that are highly concentrated spatiotemporally and do not completely overlap with the actual image archive pattern, then it can be inferred that the known employee has identity splitting. This indicates that the employee's images in different scenarios (e.g., wearing protective equipment while working on-site) are misidentified by the system as having different identities.
[0060] This application's solution, after identifying image files whose facial feature similarity is below a preset similarity threshold but above a preset association threshold, proposes a scheme to determine identity splitting based on a spatiotemporal density clustering algorithm to address the problem of inaccurate or inefficient judgments caused by the lack of efficient spatiotemporal data analysis mechanisms in traditional methods. Specifically, the scheme first obtains the shooting timestamp and geographic coordinates of each image file in the image set. This metadata objectively reflects the spatiotemporal background of image generation. Subsequently, the system employs a spatiotemporal density clustering algorithm to conduct in-depth analysis of these shooting timestamps and geographic coordinates. This algorithm can intelligently identify image groups that are highly concentrated in time and space, i.e., image clusters that meet the preset spatiotemporal density conditions. These image clusters represent the activity trajectories and patterns of employees within a specific spatiotemporal range. Finally, based on the identified image clusters, the system determines whether there is identity splitting for the corresponding known employees. For example, if a known employee's authentic image archive is mainly concentrated in a certain office area, but their image set to be verified forms a spatiotemporal image cluster at a specific work site, this strongly indicates that the employee's identity has been split in the system. In this way, the proposed solution can proactively and efficiently identify identity splits caused by inaccurate facial recognition due to complex environments, without relying on manual judgment. This spatiotemporal density clustering method enables the system to automatically discover potential identity split clues from massive amounts of image data, providing accurate triggering conditions and a data foundation for subsequent collection of multi-source auxiliary information, construction of evidence chains, and calculation of archive merging confidence. This not only improves the accuracy and efficiency of identity split judgment but also lays a solid foundation for the eventual automatic archiving of image archives to be verified and the updating of authentic image archives, thereby effectively maintaining the uniformity and integrity of employee image archives.
[0061] As a specific implementation, when the system continuously monitors the update status of a known employee's image file and discovers that multiple independent image sets exist under that employee's name, and the facial feature similarity between these sets, although lower than the conventional identity verification threshold (e.g., 0.85), remains within a low but non-zero range (e.g., between 0.5 and 0.7), and the images in these independent image sets significantly overlap in terms of shooting timestamps and geographical location information (e.g., more than 60% of the images were taken within the same week, and the average distance between geographical coordinates is less than 500 meters), the system will trigger the identity split judgment process in this step. In this process, firstly, the system automatically extracts the shooting timestamps (e.g., date and time information accurate to the second) and geographical location coordinates (e.g., GPS latitude and longitude data) from the image data of these image files marked as potentially having identity splits. For example, an image might be recorded as "November 15, 2023, 10:35:22, 30.5 degrees North latitude, 104.1 degrees East longitude". Subsequently, the system employs a grid-based spatiotemporal density clustering algorithm. Specifically, the system discretizes the timeline into hourly time segments and the geographic space into square grid cells with sides of 100 meters. Then, the system counts the number of image files to be examined within each spatiotemporal grid cell. Next, the system identifies "core spatiotemporal grid cells" whose number of image files exceeds a preset threshold (e.g., at least 3 images per grid cell). Finally, starting from these core spatiotemporal grid cells, the system connects and expands other adjacent spatiotemporal grid cells or those within a certain spatiotemporal distance (e.g., a time difference of 2 hours or a spatial distance of 500 meters) that also meet the density conditions, thereby forming one or more image clusters. For example, if multiple image files to be examined were taken between 10:00 AM and 12:00 PM on the same day, and all within a specific area of a substation, they will be clustered into one image cluster. Ultimately, the system uses these identified image clusters to determine the extent of image fragmentation. For example, if the system discovers that a known employee's authentic video profile primarily records their activities in the company headquarters office, but then a new video cluster consisting of numerous unverified video profiles appears, with its spatiotemporal characteristics highly concentrated at a specific outdoor substation, and the capture time closely matching the employee's daily working hours, then the system will determine that the known employee has a split identity. This indicates that, due to limitations in facial recognition, the employee's image was incorrectly categorized as a "newcomer" while working outdoors, leading to a split in their identity profile.
[0062] Through the aforementioned technical solution, this application introduces a spatiotemporal density clustering algorithm to determine identity splitting, significantly improving the accuracy and efficiency of employee image matching and archiving methods. Traditional methods, when facing complex environments, often struggle to accurately identify "shadow files" of the same employee due to limitations in facial recognition, resulting in heavy manual review burdens and low archiving efficiency due to a lack of effective analysis of image spatiotemporal information. This solution, by obtaining the shooting timestamps and geographic coordinates of the images to be examined, provides an objective and quantitative data foundation for subsequent spatiotemporal analysis, avoiding biases in subjective judgment. Furthermore, by using a spatiotemporal density clustering algorithm to analyze this spatiotemporal data, image clusters that meet preset spatiotemporal density conditions can be efficiently identified. This means the system can automatically discover employee activity patterns within specific time periods and geographical areas; even if facial feature recognition is obstructed, the correlation of images can be inferred through the concentration of their spatiotemporal trajectories. This intelligent clustering mechanism greatly improves the efficiency and accuracy of detecting potential identity splitting, solving the problem of inaccurate judgments caused by insufficient spatiotemporal data analysis in traditional methods. Ultimately, by identifying image clusters, the system determines the extent of identity splits, enabling it to accurately identify image files belonging to the same known employee but whose facial recognition failed due to complex working environments (such as uneven lighting or facial occlusion). This mechanism ensures that only employees with genuine identity splits are flagged, triggering subsequent multi-source auxiliary information collection and evidence chain construction processes. This avoids unnecessary resource consumption and improves the overall efficiency of the archiving system. When facial feature similarity is insufficient to confirm identity, this solution provides a powerful supplementary judgment mechanism based on spatiotemporal behavioral patterns. It allows the system to proactively identify and correct identity splits caused by the limitations of facial recognition, ensuring the integrity and consistency of employee image files and significantly improving the automation and reliability of employee image matching and archiving in complex scenarios such as frontline operations in the power industry.
[0063] In some embodiments, the specific steps in step S32 include: S321. Discretize the shooting timestamp and geographic coordinates into spatiotemporal grid cells; S322. Count the number of image files in each spatiotemporal grid unit; S323. Identify core spatiotemporal grid units with a density exceeding a preset threshold; S324. Starting from the core spatiotemporal grid cell, connect and expand adjacent spatiotemporal grid cells or those within a certain spatiotemporal distance that meet the density conditions to obtain an image cluster.
[0064] Step S321 aims to transform continuous, complex spatiotemporal data into a discrete, easily managed, and computationally efficient structure. Specifically, the capture timestamp can be discretized according to a preset time granularity, such as dividing a day into 24 one-hour segments or a week into 7 day segments. Simultaneously, geographic coordinates can be discretized according to a preset spatial granularity, such as dividing the Earth's surface into a fixed-size latitude and longitude grid, with each grid representing a specific geographic region, or pre-defining geofences as spatial units based on the actual work area (such as a substation or power line segment). Combining the discretized time units and geospatial units forms a spatiotemporal grid unit, such as a specific geographic grid area within a specific date. This discretization process significantly improves the efficiency of subsequent data processing and reduces the consumption of computational resources.
[0065] Step S322 is used to quantify the data density distribution of each spatiotemporal grid cell, providing basic data support for subsequent identification of high-density areas. Specifically, a counting method can be used: for each spatiotemporal grid cell, all image files to be examined are traversed; if the image's capture timestamp and geographic coordinates fall within that cell, the counter for that cell is incremented. Alternatively, a hash mapping method can be used, using the identifier of the spatiotemporal grid cell as the key and the number of image files within that cell as the value, thereby achieving rapid statistics and queries. By counting the number of image files, the system can intuitively understand the density of image data in different spatiotemporal regions.
[0066] Step S323 aims to select reliable clustering starting points, ensuring that the clustering process is based on high-probability regions, thereby effectively avoiding interference from low-density noise. Specifically, a fixed threshold for the number of image files can be preset; any spatiotemporal grid cell containing more than this threshold is identified as a core spatiotemporal grid cell. Alternatively, this threshold can be dynamically adjusted based on the statistical characteristics of the overall image file distribution (such as average density and standard deviation), for example, by setting it to a multiple of the average density. Identifying core cells is a crucial step in the density clustering algorithm, ensuring the effectiveness and accuracy of subsequent expansion.
[0067] Step S324 aims to form complete and meaningful image clusters through dynamic expansion, while considering spatiotemporal proximity and density continuity to prevent fragmented clustering. Specifically, it can start with an identified core spatiotemporal grid cell and search for its direct neighbors in both the temporal and geographic dimensions. If these neighbors also meet density conditions (e.g., their image archive count exceeds a certain secondary threshold, or they are themselves core cells), they are added to the current image cluster, and expansion continues from newly added cells until no further expansion is possible. Another approach is to define a certain spatiotemporal distance. For a core spatiotemporal grid cell, all non-core cells within that distance that meet the density conditions are searched and added to the current image cluster. This process can be iterative until all cells meeting the conditions are included in a cluster. In this way, the system can aggregate continuous image data taken by the same employee within a specific spatiotemporal range to form a complete image cluster.
[0068] Traditional methods for spatiotemporal density clustering of large amounts of image data often suffer from inefficiency and are prone to misjudgments, especially in the complex and ever-changing environment of frontline operations in the power industry. This application addresses this issue by discretizing the capture timestamps and geographic coordinates into spatiotemporal grid cells, transforming continuous spatiotemporal data into structured discrete units. This significantly simplifies the complexity of data processing and lays a high-efficiency foundation for subsequent density calculations. Furthermore, the system counts the number of image files within each spatiotemporal grid cell, quantifying the image density of each spatiotemporal region and enabling the clear identification of high-density areas. Subsequently, by identifying core spatiotemporal grid cells with densities exceeding a preset threshold, the system can accurately locate the regions where image data is most concentrated. These regions serve as reliable starting points for clustering, effectively avoiding misjudgments that may occur when clustering begins in sparse or noisy regions. Finally, starting from these core spatiotemporal grid cells, the system intelligently connects and expands adjacent spatiotemporal grid cells or those within a certain spatiotemporal distance that meet density conditions, thereby dynamically constructing a complete image cluster. This expansion mechanism considers not only the temporal and spatial proximity of images but also the continuity of density, ensuring that the formed image clusters are complete and meaningful, effectively preventing fragmented clustering. The above steps in this scheme are closely integrated with step S32 of the basic method, which determines whether a known employee has a split identity. By efficiently and accurately identifying image clusters, this scheme provides a solid data foundation for the subsequent step S33, which determines whether a corresponding known employee has a split identity based on the identified image clusters. For example, if the image of a known employee is clustered into multiple image clusters that are significantly separated in time and space, it strongly indicates the possibility of the employee having a split identity. This refined spatiotemporal density clustering method enables the system to accurately identify potential split identity problems by analyzing the spatiotemporal distribution patterns of images, even when facing inaccurate face recognition due to uneven lighting, facial occlusion, and other factors in frontline operations in the power industry. This significantly improves the accuracy and robustness of the entire employee image matching and archiving method.
[0069] The following is a concrete example. Suppose that in a power line inspection scenario, the system receives a large number of image files to be inspected from a certain employee. The timestamps and geographic coordinates of these images are complexly distributed. As a specific implementation method, the system first discretizes the timestamps of these images into time periods in hours, for example, dividing a day into 24 time periods. Simultaneously, the geographic coordinates are discretized into geographic grid units of 100 meters × 100 meters. By combining time and spatial dimensions, a spatiotemporal grid unit is formed, for example, "within a 100 meter × 100 meter grid at substation A from 14:00 to 15:00 on October 26, 2023". Then, the system counts the number of image files contained in each spatiotemporal grid unit. For example, the system might count 50 images in a specific spatiotemporal grid unit, while in another unit it might only count 2. Next, the system identifies core spatiotemporal grid units with a density higher than a preset threshold. For example, if the preset threshold is 30 images, then a cell containing 50 images will be identified as a core spatiotemporal grid cell, while a cell containing 2 images will not. Finally, starting from these core spatiotemporal grid cells, the system connects and expands adjacent spatiotemporal grid cells or those within a certain spatiotemporal distance that meet density conditions to obtain image clusters. For example, starting from a core cell (e.g., "grid X in substation A, October 26, 2023, 14:00-15:00"), the system checks its temporally adjacent cells (e.g., "grid X, October 26, 2023, 15:00-16:00") and spatially adjacent cells (e.g., "grid Y, October 26, 2023, 14:00-15:00"). If these adjacent cells also meet density conditions (e.g., more than 10 images), they are added to the current image cluster, and the system continues to expand until it can no longer expand. In this way, the system can aggregate all the images taken by the employee within a specific time and space range (for example, an entire afternoon of continuous inspection at a substation) into a complete image cluster, thus providing a clear basis for determining whether there is a split identity.
[0070] Through the above technical solution, this application can efficiently and accurately achieve spatiotemporal density clustering, thus effectively solving the problems of low efficiency and misjudgment that may occur in the spatiotemporal density clustering process of traditional methods when processing large amounts of employee image data. Specifically, discretizing the shooting timestamp and geographic coordinates into spatiotemporal grid units significantly reduces the complexity of data processing and the consumption of computing resources, improving the running efficiency of the clustering algorithm. Statistical analysis of the number of image files within each spatiotemporal grid unit provides a quantitative basis for subsequent density analysis. Identifying core spatiotemporal grid units with densities higher than a preset threshold ensures that the clustering process starts from high-confidence regions, effectively avoiding interference from noisy data on the clustering results. Starting from the core spatiotemporal grid unit, connecting and expanding adjacent spatiotemporal grid units or those within a certain spatiotemporal distance that meet the density conditions makes the formed image clusters more complete and accurate, avoiding fragmented clustering, and thus more reliably reflecting the actual activity trajectory of employees within a specific spatiotemporal range. These improvements enable the acquisition of more accurate and reliable image cluster information when determining whether known employees have split identities. For example, in frontline operations in the power industry, even if facial recognition is limited by factors such as lighting and occlusion, this solution can still accurately identify the image set of the same employee in different times and spaces through refined analysis of the spatiotemporal distribution of images. This effectively avoids the "identity split" phenomenon caused by image quality issues, ensures the integrity and accuracy of employee image files, and greatly improves the efficiency of file management and the reliability of data traceability.
[0071] In some embodiments, the specific steps in step S4 include: S41. Perform time-dimensional analysis on multi-source auxiliary information according to the following steps: S411. Obtain the shooting timestamp contained in the image data in the real image archive; S412. Determine the number of overlapping images whose absolute value is less than or equal to a preset time window based on the absolute value of the difference between the shooting timestamp contained in the image data in the real image archive and the shooting timestamp contained in the image data in the image archive to be examined. S413. Based on the number of overlapping images, determine the temporal overlap between the image archive to be examined and the real image archive, and determine the time dimension score based on the temporal overlap. S42. Perform geographic location dimension analysis on multi-source auxiliary information according to the following steps: S421. Obtain the geographic coordinates contained in the image data in the real image archive; S422. Based on the relative distance between the geographic coordinates contained in the image data in the real image archive and the geographic coordinates contained in the image data in the image archive to be examined; S423. Determine the geographic location dimension score based on relative distance; S43. Perform team task dimension analysis on multi-source auxiliary information according to the following steps: S431. Obtain team task data from real image archives; S432. By comparing the team task data of real image archives and image archives to be examined, determine the team task matching degree between the two, and determine the team task dimension score based on the team task matching degree. S44. Perform behavioral pattern dimension analysis on multi-source auxiliary information according to the following steps: S441. Obtain behavioral record data from real image archives; S442. By comparing the behavioral record data of real image archives and image archives to be examined, determine the degree of matching of behavioral patterns between the two, and determine the behavioral pattern dimension score based on the degree of matching of behavioral patterns; S45. Construct a chain of evidence based on scores in the time dimension, geographical location dimension, team task dimension, and behavioral pattern dimension.
[0072] In the above steps, S411 aims to extract the shooting time information from known real image archives. Shooting timestamps can be read directly from the image file's metadata (e.g., EXIF information) or retrieved from database records associated with the image. For example, when an image is uploaded to the system, the system can automatically parse the shooting date and time from its EXIF data and store it as an attribute of the image archive. S412 is used to quantify the temporal correlation between the real image archive and the image archive to be examined. The preset time window is a configurable parameter, such as 1 hour, 30 minutes, or 15 minutes, used to define the time range within which two image shooting times are considered "overlapping." The system iterates through each shooting timestamp in the real image archive and compares it with all shooting timestamps in the image archive to be examined, counting how many pairs of images have an absolute difference in shooting time that falls within this preset time window. S413 converts the quantified number of overlapping images into a standardized temporal overlap index and further maps it to a time dimension score. Temporal overlap can be calculated by dividing the number of overlapping images by the minimum number of images in the two files, or by using more complex statistical methods (such as the Jaccard similarity coefficient). The temporal dimension score can be determined by setting a mapping function based on the temporal overlap; for example, a higher score is given when the temporal overlap reaches a certain threshold, and vice versa.
[0073] S421 aims to extract location information from real image archives. Geographic coordinates are typically expressed in latitude and longitude and can be obtained from the EXIF metadata of the image file (such as GPS tags), or from location service records of image uploading devices (such as smartphones), or retrieved from databases linked to known geographic areas (such as substations or power line sections) through manual annotation. S422 calculates the geographic proximity between the real image archive and the image archive to be examined. Various geographic distance algorithms can be used to calculate the relative distance; for example, Euclidean distance can be used for shorter distances, while the Haversine or Vincenty formulas can be used to calculate the great circle distance for two points on the Earth's surface to obtain more accurate results. S423 converts the calculated relative distance into a standardized geographic location dimension score. The score can be determined based on the inverse relationship of distance, i.e., the closer the distance, the higher the score; or multiple distance thresholds can be set to map distance ranges to different score intervals. For example, a distance less than 500 meters receives a high score, 500 meters to 1 kilometer receives a medium score, and greater than 1 kilometer receives a low score.
[0074] S431 aims to obtain team and task information associated with the actual image archives. This data is typically stored in the enterprise resource planning (ERP) system, task management system, or scheduling system. The system can query via API or database to retrieve the team ID, task ID, and task description based on information such as the image's capture time and the photographer (if known). S432 is used to evaluate the association between the actual image archive and the image archive to be examined in terms of team and task. Team-task matching degree can be determined by directly comparing whether the team ID and task ID are consistent, or by comparing the similarity of task descriptions through semantic analysis. For example, if two archives are associated with the same team and the same task within the image capture time period, the matching degree is high. The team-task dimension score can be set according to the matching degree; the higher the matching degree, the higher the score.
[0075] S441 aims to obtain employee behavior records associated with real video archives. This data can come from attendance systems (e.g., clock-in times), access control systems (e.g., entry / exit records), training management systems (e.g., training participation status), etc. The system can extract relevant behavioral events and time sequences from these systems based on information such as the video's capture time and the person who captured it. S442 is used to evaluate the similarity of behavioral patterns between real video archives and the video archive to be examined. The behavioral pattern matching degree can be determined by comparing the behavioral event sequences, event types, and event frequencies of the two archives within a specific time period. For example, if the two archives have similar clock-in times on the same day, or similar access control entry / exit records within the same week, the behavioral pattern matching degree is high. The behavioral pattern dimension score can be set according to the matching degree; the higher the matching degree, the higher the score.
[0076] S45 aims to integrate the analysis results from the four dimensions mentioned above to form a structured chain of evidence. This chain of evidence can be a vector containing scores from all dimensions, or a more complex weighted scoring model. Its purpose is to provide a comprehensive, multi-faceted assessment basis for subsequent identity verification, rather than relying solely on single facial feature similarity.
[0077] This application's solution addresses the accuracy degradation of facial recognition in complex environments by systematically analyzing time, geolocation, work group tasks, and behavioral patterns. Specifically, in the time dimension analysis, by acquiring the shooting timestamps of real and candidate image files and calculating the number of overlapping images and temporal overlap, the temporal correlation between the two files is quantified, enabling the identification of images taken within similar time periods. In the geolocation dimension analysis, by acquiring the geographic coordinates of the images and calculating their relative distance, the spatial proximity of the two files is assessed, which is crucial for determining whether employees are active in the same work area. In the work group task dimension analysis, by comparing the work group task data of real and candidate image files, it is confirmed whether employees participated in the same work group and tasks within a specific time period, providing strong work context information. In the behavioral pattern dimension analysis, by comparing behavioral record data, such as attendance and access control records, it is further verified whether the employees associated with the two files have consistent behavioral trajectories. The scores obtained from the above dimension analyses collectively constitute a chain of evidence for identity verification. This multi-dimensional and systematic analysis method allows the system to move beyond simply relying on facial feature similarity and comprehensively consider the multiple relationships between employees across time, space, tasks, and behaviors. Compared to the basic approach that only determines image attribution through facial feature recognition and merely proposes building an evidence chain in step S4, this approach provides a specific and actionable analysis path, greatly enhancing the reliability and comprehensiveness of the evidence chain. In this way, even when facial feature recognition is limited, the system can establish a robust identity association through other auxiliary information, effectively solving the problem of inaccurate facial recognition caused by complex environments, avoiding identity fragmentation, and ensuring the accuracy and completeness of employee image archives.
[0078] As a specific implementation method, suppose that a power group's employee image matching and archiving server contains a real image file of a known employee, Zhang San. This file includes images of Zhang San conducting routine inspections at substation A. The timestamps of these images are concentrated between 9:00 AM and 11:00 AM, and their geographical coordinates are all within the geofence of substation A. Furthermore, the shift task data shows that Zhang San was assigned the task of "inspecting equipment at substation A" that day. Simultaneously, the incremental image archive contains an image file to be verified. This file contains images showing a person wearing a safety helmet and goggles working at substation A. The facial recognition system initially determines that the facial feature similarity between this person and Zhang San is below a preset similarity threshold but above a preset association threshold, thus classifying them as a "new person with association." To confirm whether this "new person" is indeed Zhang San, the system will initiate multi-dimensional analysis to construct a chain of evidence. First, in the time dimension analysis, the system will obtain the timestamps of all images in Zhang San's real image file, for example, {9:05, 9:30, 9:55, 10:20, 10:45}. Simultaneously, the system obtains the capture timestamps of the images in the image archive to be examined, for example, {10:00, 10:30, 11:00}. The system can set a preset time window of 15 minutes. Through comparison, it is found that the 10:00 image in the image archive to be examined overlaps with the 9:55 image in the actual archive (difference of 5 minutes), and the 10:30 image overlaps with the 10:20 image in the actual archive (difference of 10 minutes). Therefore, the number of overlapping images is 2. Based on this number, the system can calculate the degree of temporal overlap and determine a higher time dimension score, such as 0.8. Secondly, in the geographic location dimension analysis, the system obtains the geographic location coordinates of the images in Zhang San's actual image archive. These coordinates are all located near the central area of substation A. The images in the image archive to be examined also contain geographic location coordinates, for example, GPS positioning shows that they are located in the same area of substation A. The system can calculate the Haversine distance between the average latitude and longitude of all coordinates in the actual image archive and the average latitude and longitude of all coordinates in the image archive to be examined. If the calculated relative distance is less than a preset threshold (e.g., 500 meters), a higher geographical location dimension score is determined, such as 0.9. Next, in the team task dimension analysis, the system retrieves Zhang San's team task data for that day (when the image was captured) from the enterprise ERP system, showing that he belongs to "Maintenance Team 1" and his task is "Inspection of Substation A Equipment". Simultaneously, the system attempts to retrieve his team task data from the metadata or related information of the image file to be inspected. If it finds that the image file to be inspected is also associated with "Maintenance Team 1" and the "Inspection of Substation A Equipment" task, the team task matching degree is high, thus determining a higher team task dimension score, such as 0.95.Finally, in the behavioral pattern dimension analysis, the system retrieves Zhang San's behavioral record data for the day from the attendance system and access control system. For example, entering substation A at 8:30 AM and leaving substation A at 5:30 PM, with no abnormal absences during this period. If the related information in the image file to be examined (e.g., if the "newcomer" has independent access control records) also shows similar entry and exit records within the same time period, the behavioral pattern matching degree is high, thus determining a high behavioral pattern dimension score, such as 0.85. Ultimately, the system integrates these time dimension scores, geographical location dimension scores, team task dimension scores, and behavioral pattern dimension scores to construct a complete chain of evidence, providing a comprehensive basis for subsequent calculation of the file merging confidence level.
[0079] Through the above technical solution, this application effectively solves the problem that in the process of multi-source auxiliary information analysis, the lack of specific analysis steps leads to unsystematic and incomplete evidence chain construction, thus affecting the accuracy and efficiency of identity verification. This solution ensures a comprehensive and robust evidence chain construction by conducting detailed and quantitative analysis of time, geographical location, work group / task, and behavioral pattern dimensions. Specifically, time dimension analysis accurately captures the temporal correlation of images, avoiding misjudgments caused by discontinuous time information; geographical location dimension analysis enhances the evidence strength of the spatial dimension, effectively preventing isolated geographical location information from affecting judgment; work group / task dimension analysis ensures the effective integration of task information, reducing identity fragmentation caused by task mismatch; and behavioral pattern dimension analysis adds behavioral-level verification, compensating for the shortcomings of single-dimensional analysis. These specific analysis steps enable the system to cross-verify employee identities from multiple perspectives, forming a multi-layered, highly reliable evidence system. Compared to the basic solution that only proposes building an evidence chain without providing specific analysis methods, this solution provides a clear and operable implementation path, greatly improving the systematicness and effectiveness of the evidence chain. This multi-dimensional analysis method, combined with facial recognition technology in the basic solution, can provide strong supplementary evidence through auxiliary information when the accuracy of facial feature recognition decreases due to the influence of complex environments. This significantly improves the accuracy and robustness of employee image matching and archiving, effectively avoids the occurrence of "identity splitting," and ensures the integrity and consistency of employee image archives.
[0080] In some embodiments, the specific steps in step S5 include: S51. For each image file in the image set to be examined, obtain the job task type and job environment characteristics associated with the image file to be examined; S52. Based on the type of work task and the characteristics of the work environment, determine the weights of the time dimension, geographical location dimension, team task dimension, and historical behavior pattern dimension in the multi-source auxiliary information; S53. Based on the analysis results of each dimension in the chain of evidence, and combined with the determined weights, calculate the archive merging confidence of each archive in the set of images to be examined.
[0081] The acquisition of job task types and work environment characteristics associated with the image archives to be examined refers to identifying the specific work category performed by employees and the physical environmental conditions they are in when performing tasks. Job task types may include, but are not limited to, "equipment inspection," "line erection," "fault repair," and "substation maintenance." These can be acquired through enterprise resource planning (ERP) systems, task assignment systems, or field work management systems, or by inferring them through image recognition analysis of the image content to identify tools, equipment, or scenes in the image. Work environment characteristics may include "outdoor high-altitude," "indoor machine room," "nighttime operation," "rainy or snowy weather," and "high-temperature environment." These can be acquired through image metadata (such as GPS information, timestamps combined with weather data), sensor data (such as ambient light sensors and temperature sensors), or by analyzing the visual characteristics of the image (such as brightness, contrast, color saturation, and background elements) using machine learning models for identification.
[0082] Determining the weights of the time, geographic location, work group task, and historical behavior pattern dimensions in multi-source auxiliary information involves assigning different importance coefficients to these four dimensions. These weights are not fixed but dynamically adjusted based on the acquired task type and environmental characteristics. For example, in a "nighttime emergency repair" task, insufficient lighting may hinder facial recognition; therefore, the weights of the time dimension (e.g., whether it's outside of working hours) and the behavior pattern dimension (e.g., emergency response records) can be increased to reflect their higher reliability in this specific scenario. Conversely, in an "indoor equipment inspection" task, the weights of the geographic location dimension (e.g., fixed machine room location) and the work group task dimension (e.g., fixed work group personnel) may be higher. Weight determination can be achieved through matching with a pre-defined rule base or through a machine learning model. Inputting the task type and environmental characteristics, the model outputs the weights for each dimension. This model can be trained using historical data to optimize the accuracy of weight allocation.
[0083] Calculating the file merging confidence score for each image file in the image set to be examined refers to obtaining a quantitative value by comprehensively considering the analysis results (scores) of each dimension in the evidence chain and dynamically determined weights. This value is used to assess the likelihood of merging the image file to be examined with known employee image files. The calculation method can employ a weighted summation model; for example, the merging confidence score C can be calculated using a weighted summation method. Alternatively, more complex fusion algorithms, such as Bayesian networks, fuzzy logic, or neural networks, can be used, taking the scores and weights of each dimension as input and outputting the merging confidence score to adapt to more complex decision-making scenarios.
[0084] This application's solution first obtains the job task type and work environment characteristics associated with the image file to be examined before calculating the confidence score for file merging. This provides crucial scene information for subsequent weight adjustment. Based on this contextual information, the system dynamically adjusts the weights of the time dimension, geographical location dimension, team task dimension, and historical behavior pattern dimension in the multi-source auxiliary information. This dynamic weight adjustment mechanism is crucial because in complex and variable environments such as front-line operations in the power industry, the relevance and reliability of each dimension can vary significantly depending on the task and environment. For example, in scenarios with insufficient lighting or severe facial occlusion, the accuracy of facial feature recognition may decrease, and the importance of auxiliary information such as behavior patterns or task assignments should be increased. Finally, the system uses these dynamically adjusted weights, combined with the analysis results of each dimension in the evidence chain constructed in step S4, to calculate the confidence score for file merging. This weighted combination ensures that the final confidence score can more accurately reflect the matching probability between the image file to be examined and the known real image files of employees, thereby overcoming the limitations of fixed weights in traditional methods under variable scenarios. In this way, the solution proposed in this application can more intelligently and accurately assess the confidence level of identity matching, effectively reducing the problems of identity splitting and erroneous archiving caused by environmental complexity.
[0085] The following is a concrete example to illustrate this. Suppose an electrical engineer is performing emergency repairs at night, and their image is identified by the system as a "newcomer," creating a pending image file. The system first obtains the job task type associated with this pending image file as "nighttime emergency repair" and the job environment characteristic as "outdoor high-altitude." Based on this information, the system dynamically adjusts the weights of each dimension. For example, due to poor lighting at night, facial recognition is difficult, so the system may increase the weights of the time dimension (such as whether the emergency work was performed outside of working hours) and the behavioral pattern dimension (such as emergency response records and records of specific repair tool usage), for example, setting them to 0.35 respectively; while the weights of the geographical location dimension (possibly in remote areas with unstable GPS signals) and the team task dimension (possibly a temporarily formed repair team) are relatively decreased, for example, setting them to 0.15 respectively. Subsequently, the system combines the analysis results of each dimension obtained in step S4 (e.g., time overlap score 0.9, geographical location score 0.7, team task score 0.8, behavioral pattern score 0.9) with these dynamically determined weights to calculate the file merging confidence score. For example, C = 0.35*0.9 + 0.15*0.7 + 0.15*0.8 + 0.35*0.9. This dynamic weighting method makes the confidence calculation more consistent with actual operating scenarios and avoids misjudgments caused by fixed weights in complex environments.
[0086] Through the above technical solution, this application can dynamically adjust the weights of each dimension in multi-source auxiliary information according to the type of work task and the characteristics of the work environment, thereby overcoming the limitations of inaccurate reliability assessment by traditional fixed-weight resetting in complex and ever-changing environments such as front-line operations in the power industry. This makes the calculation of the confidence level of file merging more accurate and robust, effectively reducing the risk of identity splitting and erroneous archiving, and ensuring the integrity and accuracy of employee image files. In particular, it significantly improves the adaptability and reliability of the system when facial feature recognition is affected by environmental factors or protective equipment.
[0087] Reference Appendix Figure 2 This invention provides a face recognition-based employee image matching and archiving system (this face recognition-based employee image matching and archiving system adopts the face recognition-based employee image matching and archiving method of the above embodiment, and the specific process is referred to the corresponding steps above), applied to an employee image matching and archiving server. The employee image matching and archiving server has a preset real image archive and an incremental image archive; the real image archive is used to store real image archives of known employees; the incremental image archive is used to store image archives to be verified created through incremental learning; The employee image matching and archiving system based on facial recognition includes: The first calculation module 100 is used to calculate the facial feature similarity between the real image archives of each known employee and the image archives to be examined. The recognition module 200 is used to identify all the image files to be examined whose facial feature similarity is lower than a preset similarity threshold and higher than a preset association threshold, and to form a set of images to be examined corresponding to the known employees. The judgment module 300 is used to determine whether a known employee has a split identity based on the image data of the image archives in the image set to be examined, and to collect multi-source auxiliary information related to the image set to be examined for the known employee with a split identity. Analysis module 400 is used to construct a chain of evidence for identity verification by performing multi-dimensional analysis on each image file in the image set to be verified through multi-source auxiliary information; the chain of evidence includes the analysis results of each dimension. The second calculation module 500 is used to calculate the file merging confidence of each file in the image set to be examined based on the analysis results of each dimension in the chain of evidence. The merging module 600 is used to merge the corresponding image files in the image set to be examined into the real image files of the corresponding known employees if the confidence level of file merging is greater than or equal to the preset merging threshold. The early warning module 700 is used to generate early warning information and submit it for manual review if the confidence level of file merging is less than the merging threshold.
[0088] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.
[0089] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for matching and archiving employee images based on face recognition, applied to an employee image matching and archiving server, characterized in that, The employee image matching and archiving server has a pre-set real image archive and an incremental image archive; the real image archive is used to store real image archives of known employees; the incremental image archive is used to store image archives to be verified created through incremental learning. The employee image matching and archiving method based on facial recognition includes the following steps: S1. For each known employee, calculate the facial feature similarity between their real image file and each image file to be examined; S2. Identify all the image files to be examined whose facial feature similarity is lower than a preset similarity threshold and higher than a preset association threshold, and form a set of images to be examined for the corresponding known employees. S3. Determine whether there is a split identity for the corresponding known employee based on the image data of the image archives in the image set to be examined, and collect multi-source auxiliary information related to the image set to be examined for the known employee with a split identity. S4. For each image file in the image set to be examined, a chain of evidence for identity verification is constructed by performing multi-dimensional analysis on multi-source auxiliary information; the chain of evidence includes the analysis results of each dimension. S5. Based on the analysis results of each dimension in the chain of evidence, calculate the file merging confidence of each file in the set of images to be examined; S6. If the confidence level of file merging is greater than or equal to the preset merging threshold, then the corresponding image file to be examined in the image set to be examined will be merged into the real image file of the corresponding known employee. S7. If the confidence level for merging the files is less than the merging threshold, an early warning message will be generated and submitted for manual review. Step S3, which involves determining whether a known employee has a split identity based on the image data of the image files in the image set to be examined, includes: S31. Based on the image data of the image archives to be examined in the image collection, obtain the shooting timestamp and geographical coordinates of each image archive to be examined; S32. A spatiotemporal density clustering algorithm is used to identify image clusters that meet the preset spatiotemporal density conditions by analyzing the shooting timestamps and geographic coordinates; S33. Based on the identified image clusters, determine whether there is a split identity among the corresponding known employees; The specific steps in step S32 include: S321. Discretize the shooting timestamp and geographic coordinates into spatiotemporal grid cells; S322. Count the number of image files in each spatiotemporal grid unit; S323. Identify core spatiotemporal grid units with a density exceeding a preset threshold; S324. Starting from the core spatiotemporal grid cell, connect and expand adjacent spatiotemporal grid cells or those within a certain spatiotemporal distance that meet the density condition to obtain an image cluster; In step S4, the dimensions used for analysis include time dimension, geographical location dimension, team task dimension, and behavioral pattern dimension; The specific steps in step S4 include: S41. Perform time-dimensional analysis on multi-source auxiliary information according to the following steps: S411. Obtain the shooting timestamp contained in the image data in the real image archive; S412. Determine the number of overlapping images whose absolute value is less than or equal to a preset time window based on the absolute value of the difference between the shooting timestamp contained in the image data in the real image archive and the shooting timestamp contained in the image data in the image archive to be examined. S413. Based on the number of overlapping images, determine the temporal overlap between the image archive to be examined and the real image archive, and determine the time dimension score based on the temporal overlap. S42. Perform geographic location dimension analysis on multi-source auxiliary information according to the following steps: S421. Obtain the geographic coordinates contained in the image data in the real image archive; S422. Based on the relative distance between the geographic coordinates contained in the image data in the real image archive and the geographic coordinates contained in the image data in the image archive to be examined; S423. Determine the geographic location dimension score based on relative distance; S43. Perform team task dimension analysis on multi-source auxiliary information according to the following steps: S431. Obtain team task data from real image archives; S432. By comparing the team task data of real image archives and image archives to be examined, determine the team task matching degree between the two, and determine the team task dimension score based on the team task matching degree. S44. Perform behavioral pattern dimension analysis on multi-source auxiliary information according to the following steps: S441. Obtain behavioral record data from real image archives; S442. By comparing the behavioral record data of real image archives and image archives to be examined, determine the degree of matching of behavioral patterns between the two, and determine the behavioral pattern dimension score based on the degree of matching of behavioral patterns; S45. Construct a chain of evidence based on scores in the time dimension, geographical location dimension, team task dimension, and behavioral pattern dimension. The specific steps in step S5 include: S51. For each image file in the image set to be examined, obtain the job task type and job environment characteristics associated with the image file to be examined; S52. Based on the type of work task and the characteristics of the work environment, determine the weights of the time dimension, geographical location dimension, team task dimension, and historical behavior pattern dimension in the multi-source auxiliary information; S53. Based on the analysis results of each dimension in the chain of evidence, and combined with the determined weights, calculate the archive merging confidence of each archive in the set of images to be examined.
2. The employee image matching and archiving method based on face recognition according to claim 1, characterized in that, Multi-source auxiliary information includes the shooting timestamps and geographical locations contained in the image data in the image archive to be examined, as well as team task data and behavior record data.
3. The employee image matching and archiving method based on face recognition according to claim 1, characterized in that, The specific steps in step S6 include: If the confidence level of file merging is greater than or equal to the preset merging threshold, the corresponding image file to be examined in the image set to be examined will be merged into the real image file of the corresponding known employee and the comprehensive feature representation of the real image file of the corresponding employee will be updated.
4. The employee image matching and archiving method based on face recognition according to claim 3, characterized in that, The steps for updating the comprehensive feature representation of the corresponding employee's real image file include: S61. Select multiple face images with a quality rating higher than a preset quality threshold from the merged image data and extract the depth feature vectors of the multiple face images; S62. Perform weighted fusion processing on the deep feature vectors to generate a global feature template for representing the real identity profile and use it as a comprehensive feature representation.
5. A face recognition-based employee image matching and archiving system employing the face recognition-based employee image matching and archiving method as described in any one of claims 1-4, applied to an employee image matching and archiving server, characterized in that, The employee image matching and archiving server has a pre-set real image archive and an incremental image archive; the real image archive is used to store real image archives of known employees; the incremental image archive is used to store image archives to be verified created through incremental learning. The employee image matching and archiving system based on facial recognition includes: The first calculation module is used to calculate the facial feature similarity between the real image archives of each known employee and the image archives to be examined. The recognition module is used to identify all the image files to be examined whose facial feature similarity is lower than a preset similarity threshold and higher than a preset association threshold, and to form a set of images to be examined for the corresponding known employees. The judgment module is used to determine whether a known employee has a split identity based on the image data of the image archives in the image set to be examined, and to collect multi-source auxiliary information related to the image set to be examined for known employees with split identities. The analysis module is used to construct a chain of evidence for identity verification by performing multi-dimensional analysis on each image file in the image set to be examined through multi-source auxiliary information; the chain of evidence includes the analysis results of each dimension. The second calculation module is used to calculate the file merging confidence of each file in the image set to be examined based on the analysis results of each dimension in the chain of evidence. The merging module is used to merge the corresponding image files in the image set to be examined into the real image files of the corresponding known employees if the confidence level of file merging is greater than or equal to the preset merging threshold. The early warning module is used to generate early warning information and submit it for manual review if the confidence level of file merging is less than the merging threshold.