Digital portrait construction method based on multi-source heterogeneous data

By constructing a multi-source heterogeneous dataset and adopting an adaptive weight allocation model and dynamic update mechanism, the problems of incomplete integration of multi-source heterogeneous data and lagging profile updates in existing technologies are solved, realizing the accuracy and dynamic adaptability of employee digital profiles, and supporting real-time decision-making and visual analysis in employee management.

CN121998506APending Publication Date: 2026-05-08WEST CHINA HOSPITAL SICHUAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WEST CHINA HOSPITAL SICHUAN UNIV
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for building digital profiles suffer from incomplete integration of multi-source heterogeneous data, one-sided design of quantitative indicators, lack of dynamic adaptability in weight allocation, and lag in profile updates. As a result, the accuracy of employee digital profiles is insufficient, failing to fully reflect employees' true abilities and development trajectories, and making it difficult to support employee management decisions.

Method used

By acquiring structured and unstructured data from employees, a multi-source heterogeneous dataset is established. After standardization, quantitative indicators are constructed. An adaptive weight allocation model combined with attention mechanisms and hierarchical analysis is used to dynamically update the digital profiles of employees and display core capabilities and development potential in multi-dimensional charts.

Benefits of technology

It achieves deep integration of multi-source heterogeneous data, improves the accuracy and dynamic adaptability of quantitative assessment of profiles, supports real-time management decisions, and provides comprehensive and visualized analysis of employee capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998506A_ABST
    Figure CN121998506A_ABST
Patent Text Reader

Abstract

The invention discloses a digital portrait construction method based on multi-source heterogeneous data, and the method comprises the steps: obtaining employee structured and unstructured data, and building a multi-source heterogeneous data set; after standardization processing, four types of quantitative indexes are constructed; designing an adaptive weight distribution model based on an attention mechanism, fusing subjective and objective weights, and obtaining a comprehensive ability score through feature fusion; dynamically generating employee digital portraits containing core capabilities, advantage shortages and development potential by combining data real-time updating, and displaying the employee digital portraits in a multi-dimensional chart; according to the method, accurate and dynamic presentation of employee portraits is realized, and data support is provided for employee management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data intelligent processing, specifically a method for constructing digital profiles based on multi-source heterogeneous data. Background Technology

[0002] As employee management enters a phase of digital transformation, the demand for accuracy in assessing employees' comprehensive capabilities and evaluating their development potential is increasing. Employee-related data is characterized by its multi-source heterogeneity, encompassing structured data such as annual performance reviews and research projects, as well as unstructured data such as academic papers and awards. Effectively integrating this type of data to construct a comprehensive and objective digital profile of employees has become a key requirement for improving the efficiency and scientific rigor of employee management.

[0003] Existing digital profile building methods are mostly designed for processing single types of data or simply splicing together data from multiple sources. The design of quantitative indicators is not systematic enough, and the weight allocation mostly adopts fixed rules or a single subjective / objective assignment method. There is a lack of dynamic mining of the correlation between data and employees' core competencies, and the profile update mechanism is lagging behind, with a relatively simple display format.

[0004] Such solutions suffer from problems such as incomplete data integration, one-sided indicator quantification, lack of adaptability in weight allocation, and insufficient dynamics in profiling. As a result, the constructed digital profiles fail to fully reflect the true capabilities and development trajectories of employees, and cannot provide accurate and real-time data support for employee management decisions. Summary of the Invention

[0005] To address the shortcomings of the existing technologies, the present invention aims to provide a digital profile construction method based on multi-source heterogeneous data. This method solves the problems in the existing technologies, such as insufficient integration of multi-source heterogeneous data, one-sided design of quantitative indicators, lack of dynamic adaptability in weight allocation, and lagging profile updates and single display formats. These problems result in insufficient accuracy of employee digital profiles and an inability to fully reflect the true abilities and development trajectories of employees, making it difficult to support scientific employee management decisions.

[0006] To achieve the above objectives, this invention discloses a method for constructing digital profiles based on multi-source heterogeneous data, the method comprising the following steps: S1: Obtain structured and unstructured data of employees; the structured data includes annual performance evaluation results, scientific research project information, work experience and training records, and the unstructured data includes academic papers, academic certificates of appointment and commendation documents; establish a multi-source heterogeneous employee dataset based on the structured and unstructured data; S2: Standardize the multi-source heterogeneous employee dataset to obtain a standardized dataset; S3: Construct quantitative indicators based on the standardized dataset; the quantitative indicators include assessment dimension indicators, scientific research dimension indicators, resume dimension indicators, and honor dimension indicators; S4: An adaptive weight allocation model is designed based on the attention mechanism, and the comprehensive weights of various indicators are obtained by combining subjective weights and objective weights; the comprehensive weights of various indicators after quantification of the standardized dataset are fused through feature fusion to obtain the comprehensive ability score; S5: Based on the comprehensive ability score and combined with the real-time data updates of various indicators, dynamically generate employee digital profiles. The employee digital profiles include core ability dimension scores, strengths and weaknesses analysis, and development potential assessment. The employee digital profiles are displayed in multi-dimensional charts, including comprehensive ability radar charts, indicator weight bar charts, dynamic trend charts, and attendance status quantitative charts.

[0007] Furthermore, the establishment of the multi-source heterogeneous employee dataset is specifically implemented as follows: Each employee is assigned a unique identifier, and the acquired structured and unstructured data are bound to these unique identifiers. The structured data is organized by standardizing the field formats of annual performance evaluations, research project information, work experience, and training records, and storing them as structured data tables categorized by unique identifiers. Natural language processing (NLP) technology is used to extract core metadata from the unstructured data, forming a metadata list. Using the unique identifier as an index, a mapping relationship is established between the structured data tables and the metadata list, clarifying the correspondence between fields and metadata for each employee. Simultaneously, the unstructured data is named according to the employee's unique identifier, data type, and time, forming a mapping index between the original data storage path and the metadata. The structured data tables, metadata list, mapping relationships, and original data storage index are integrated to construct a multi-source heterogeneous employee dataset containing data content, relationships, and storage paths.

[0008] Furthermore, the standardization process involves cleaning, aligning, and normalizing the multi-source heterogeneous employee dataset. The data cleaning includes missing value imputation and outlier removal. Missing value imputation uses the median of similar employee data, outlier removal uses the 3σ criterion, and data alignment is based on the unique employee identifier, achieving alignment of time and attribute dimensions between different data sources.

[0009] Furthermore, the quantitative formula for the aforementioned assessment dimension indicators is as follows: ; in, As assessment dimensions and indicators; Let be the grading coefficient for the k-th assessment. The time decay coefficient for the k-th assessment is denoted by m; m is the number of assessments. The quantitative formula for the scientific research dimension indicators is as follows: ; in, Indicators for scientific research; Let p be the level coefficient of the p-th research project; Let p be the participation coefficient of the p-th research project; Let q be the completion quality coefficient of the p-th research project; q is the number of research projects. The quantitative formula for the resume dimension indicators is as follows: ; in, d represents the resume dimension indicator; y represents the weighting coefficient for years of service; e represents the cumulative years of service; g represents the coefficient for job importance; The quantitative formula for the honor dimension indicators is as follows: ; in, For the honor dimension indicators; Let r be the level coefficient of the r-th honor; Let be the time decay coefficient for the acquisition of the r-th honor; n is the number of honors.

[0010] Furthermore, the adaptive weight allocation model consists of an attention mechanism calculation unit, a subjective weight generation unit, an objective weight calculation unit, and a weight fusion unit; by inputting the quantification results of the quantification indicators, it outputs the comprehensive weights of each quantification indicator. The attention mechanism calculation unit is based on the Transformer multi-head attention mechanism to explore the correlation between various quantitative indicators and employees' core competencies; The subjective weight generation unit is based on the analytic hierarchy process and is used to process information related to evaluation requirements; The objective weight calculation unit is based on the entropy weight method and is used to analyze the indicator distribution characteristics of the standardized dataset. The weight fusion unit is used to integrate subjective weights, objective weights, and attention-related results.

[0011] Furthermore, the expression for the comprehensive weight is: ; in, is the comprehensive weight of the j-th type of indicator; This is the subjective weighting adjustment coefficient; Subjective weights based on the analytic hierarchy process (AHP); For objective weights based on the entropy weight method; The subjective weights are obtained by constructing a judgment matrix using the analytic hierarchy process and calculating the eigenvector corresponding to the largest eigenvalue. The objective weights are calculated by determining the information entropy of various indicators, according to the entropy weight formula: ; in, Let be the information entropy of the j-th type of index.

[0012] Furthermore, the formula for feature fusion is expressed as: ; in, Employees are scored based on their overall capabilities; is the comprehensive weight of the j-th type of indicator; is the quantified value of the j-th type of indicator.

[0013] Furthermore, the dynamic update of the employee digital profile follows the data change-driven principle, and the update triggering conditions include the addition of structured data from multi-source heterogeneous employee datasets, the supplementation of unstructured data, and the adjustment of quantitative indicator calculation parameters; During the update process, the changed data is first checked for consistency to confirm the accuracy of the binding between the data and the employee's unique identifier and the reasonableness of the matching between the data attributes and the corresponding quantitative indicators. Then, based on the changed data that has passed the check, the quantitative indicator values ​​of the corresponding dimensions are recalculated. The comprehensive weight of various indicators is updated synchronously through an adaptive weight allocation model. Finally, the employee's comprehensive ability score and the core ability dimension scores, strengths and weaknesses analysis, and development potential assessment results contained in the employee's digital profile are refreshed according to the feature fusion formula. Each update to the employee digital profile is based on the data change trajectory, recording the original information of the changed data, the quantitative indicator values ​​before and after the change, the details of changes in comprehensive weight and comprehensive ability score, forming a traceable evolution trajectory of the employee digital profile, and supporting backtracking and querying the status of the employee digital profile at different stages by time dimension.

[0014] Compared with the prior art, the beneficial effects of the present invention are: (1) This invention achieves deep fusion of multi-source heterogeneous data, associates structured and unstructured data through unique identifiers, extracts core metadata of unstructured data by combining natural language processing, and constructs a complete dataset containing data content, relationships and storage paths. This solves the problems of fragmented data integration and low correlation in traditional methods, and provides comprehensive data support for profile construction.

[0015] (2) This invention proposes an adaptive weight allocation model based on attention mechanism, which integrates the subjective weight of the analytic hierarchy process and the objective weight of the entropy weight method. At the same time, it explores the correlation between indicators and employees' core competencies, breaks through the limitations of traditional fixed weights or single values, and makes the weight allocation of indicators more targeted and dynamically adaptable, thereby improving the accuracy of profile quantitative assessment.

[0016] (3) This invention constructs a dynamic profile update mechanism driven by data changes, synchronously records the profile evolution trajectory to support historical backtracking, and presents key information such as core capabilities and weight distribution in a multi-dimensional chart. It not only solves the defects of static profiles and delayed updates in traditional profiles, but also realizes the visualization and traceability of employee capabilities, providing real-time and intuitive data references for employee management decisions. Attached Figure Description

[0017] Figure 1 This is a flowchart of the method of the present invention.

[0018] Figure 2 This is a radar chart showing the overall capabilities of the method of the present invention.

[0019] Figure 3 This is a bar chart showing the index weights of the method of this invention.

[0020] Figure 4 This is a dynamic trend diagram of the method of the present invention.

[0021] Figure 5 This is a quantitative graph of attendance status in the method of the present invention. Detailed Implementation

[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Please see Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 This application provides a detailed description of the technical solutions provided in each embodiment.

[0024] This application provides a method for constructing digital profiles based on multi-source heterogeneous data. Specific steps are as follows: Figure 1 As shown.

[0025] S1: Obtain structured and unstructured data of employees; the structured data includes annual performance evaluation results, scientific research project information, work experience and training records, and the unstructured data includes academic papers, academic certificates of appointment and commendation documents; establish a multi-source heterogeneous employee dataset based on the structured and unstructured data; Assign unique employee IDs to 50 employees, in the format of G + 6 digits, such as G000001.

[0026] The acquired structured data is categorized by employee ID and data type, organized into standardized field formats, and then stored in MySQL structured data tables. The table names are named with "table" appended to the data type, for example, assessment_table.

[0027] Natural language processing technology is used to process the acquired unstructured data; the metadata extracted from the paper results includes the paper title, journal, year of publication, author ranking and keywords; the metadata extracted from the academic society appointment certificate includes the academic society name, position, start time and end time of appointment; the metadata extracted from the commendation document includes the commendation name, awarding unit, awarding time and commendation level.

[0028] Organize the extracted metadata into a metadata list in JSON format; an example is as follows: { Employee ID: “G000001” Data type: "Academic paper" Metadata list: { Paper title: "Research on Employee Performance Appraisal Optimization Based on Big Data" Published in the journal "Chinese Public Administration". Year of publication: 2025 Author ranking: "1" Keywords: ["Big Data", "Employee Performance Appraisal", "Quantitative Evaluation"] } } Using employee ID as the index, a mapping table is constructed to associate structured data tables with metadata lists, clarifying the correspondence between structured fields and unstructured metadata for the same employee. The naming rule for unstructured raw text is: employee ID_data type_timestamp, for example, G000001_academic paper_20251211.pdf. It is stored on the MinIO file server, generating the raw text storage path and establishing an association index with the metadata list to form a metadata-storage path mapping relationship.

[0029] It integrates structured data tables, JSON-formatted metadata lists, association mapping tables, and original text storage indexes to form a complete multi-source heterogeneous employee dataset, supporting multi-dimensional retrieval by employee ID, data type, and time dimension.

[0030] S2: Standardize the multi-source heterogeneous employee dataset to obtain a standardized dataset; for missing values, imputation is performed using the median of similar employee data. For example, if employee G000005's 2025 training duration is missing, the 2025 training duration data of 10 employees in the same position and at the same level are selected: 40h, 45h, 38h, 42h, 40h, 39h, 41h, 43h, 40h, and 37h respectively. The median is calculated to be 40h and used to imput the missing field; for outlier removal, the 3σ criterion is used. Taking the research project participation coefficient as an example, the mean μ=0.65 and standard deviation σ=0.12 of the participation coefficients of 50 employees are calculated, and the outlier threshold is μ±3σ. Employee G000012's participation coefficient for a certain project is 1.15, exceeding the threshold, and is therefore removed.

[0031] The data is aligned by using the year as a unified time unit in the time dimension, and by standardizing data from different periods to the year dimension in the attribute dimension. In the attribute dimension, the field names and value standards of different data sources are unified, and the attribute association and alignment between structured data and unstructured data is achieved by using employee ID as the benchmark.

[0032] Finally, data normalization is achieved using the min-max normalization method, mapping the original values ​​of various quantitative indicators to the [0,1] interval. The normalization formula is as follows: ; Where x is the original index value, This is the minimum value of the indicator. This is the maximum value of the indicator. For example, if the original value range for a research dimension indicator is 0-15, and an employee's original value is 12, then the normalized value is... .

[0033] S3: Construct quantitative indicators based on the standardized dataset; the quantitative indicators include assessment dimension indicators, scientific research dimension indicators, resume dimension indicators, and honor dimension indicators; The quantitative formulas for the assessment dimensions and indicators are as follows: ; in, As assessment dimensions and indicators; Let be the grading coefficient for the k-th assessment. The time decay coefficient for the k-th assessment is denoted by m; m is the number of assessments. 3 indicates excellent, 2 indicates good, and 1 indicates unsatisfactory. 0.8 indicates the past year, 0.5 indicates the past 2-3 years, 0.5 indicates the past 4-5 years, and 0.3 indicates the past 5 years or more. Calculation Example: Employee G000001's five performance evaluation results are: Good in 2021, Excellent in 2022, Excellent in 2023, Good in 2024, and Excellent in 2025. Substituting these results into the formula yields: ; After normalization .

[0034] The quantitative formula for the scientific research dimension indicators is as follows: ; in, Indicators for scientific research; Let p be the level coefficient of the p-th research project; Let p be the participation coefficient of the p-th research project; Let q be the completion quality coefficient of the p-th research project; q is the number of research projects. 1 represents the national level, 2 represents the provincial level, 1 represents the municipal level, and 0.5 represents the school level; 0.7 represents the host, 0.3 represents the top three core members, and 0.3 represents a regular member. 0.8 indicates excellent, 0.6 indicates good, and 0 indicates unsatisfactory. Calculation example: Employee G000001 has three projects: Provincial, Project Leader, and Excellent; National, Core Member, and Good; Municipal, General Member, and Qualified; Substituting into the formula: ; After normalization .

[0035] The quantitative formula for the resume dimension indicators is as follows: ; in, Here are the indicators for the resume dimension; d is the weighting coefficient for years of service; y is the cumulative years of service; e is the coefficient for job importance; g is the number of times a key position has been held; in this example, d is 0.4; e=1.0 represents a key position department head, 0.8 represents an important position business backbone head, and 0.5 represents an ordinary position; Calculation example: Employee G000001 has served a total of 13 years, y=12, and has served as a department head twice, g=2; Substituting into the formula: ; After normalization .

[0036] The quantitative formula for the honor dimension indicators is as follows: ; in, For the honor dimension indicators; Let r be the level coefficient of the r-th honor; Let be the time decay coefficient for the acquisition of the r-th honor; n is the number of honors. 1 represents the national level, 2 represents the provincial level, 1 represents the municipal level, and 0.5 represents the school level; 0.7 indicates the past year, 0.4 indicates the past 2-3 years, 0.4 indicates the past 4-5 years, and 0.2 indicates the past 5 years or more. Example calculation: Employee G000001 received two honors, the first being a provincial-level honor in 2024, and the second a municipal-level honor in 2025; substituting into the formula, we get: ; After normalization .

[0037] S4: An adaptive weight allocation model is designed based on the attention mechanism, and the comprehensive weights of various indicators are obtained by combining subjective weights and objective weights; the comprehensive weights of various indicators after quantification of the standardized dataset are fused through feature fusion to obtain the comprehensive ability score; The model consists of an attention mechanism calculation unit, a subjective weight generation unit, an objective weight calculation unit, and a weight fusion unit, and is implemented as follows: The attention mechanism calculation unit employs a Transformer multi-head attention mechanism. The input consists of normalized values ​​of four types of quantitative indicators, and the output is the correlation score between each indicator and the employee's core competencies. The model parameters are fine-tuned using training samples to ensure that the correlation scores reflect the contribution of each indicator to core competencies; for example, the research dimension has a higher attention weight for employees in research positions.

[0038] The subjective weight generation unit uses the Analytic Hierarchy Process (AHP) to construct a judgment matrix: Five employee management experts were invited to conduct pairwise comparisons of the importance of the four types of indicators, and the resulting judgment matrix is ​​shown in Table 1. Table 1 Judgment Matrix

[0039] By calculating the largest eigenvalue of the judgment matrix Consistency indicators Random Consistency Index Consistency ratio Verification passed; subjective weights were calculated, and the eigenvector corresponding to the largest eigenvalue was normalized to obtain the subjective weights. .

[0040] The objective weight calculation unit analyzes the indicator distribution characteristics of the standardized dataset based on the entropy weight method; based on the normalized data of four types of indicators from 50 employees, it calculates the information entropy of the j-th type of indicator. : ; in, Let i be the normalized weight of the i-th employee under the j-th type of metric. Let be the normalized value of the j-th category indicator for the i-th employee; The total number of employees; calculated , , and .

[0041] Calculate the objective weights using the entropy weight formula: ; in, For the first Information entropy of class indicators; The objective weights are based on the entropy weight method; substituting the data yields... , , and , Objective weight .

[0042] The weight fusion unit is dynamically adjusted according to the comprehensive weight formula, which is expressed as follows: ; in, is the comprehensive weight of the j-th type of indicator; This is the subjective weighting adjustment coefficient; Subjective weights based on the analytic hierarchy process (AHP); This is an objective weight based on the entropy weight method; in this example... Take 0.4; Substitute the subjective and objective weights mentioned above into the formula for calculation: ; ; ; ; Final composite weight .

[0043] Then, feature fusion is used to fuse the comprehensive weights of various indicators after quantification of the standardized dataset to obtain a comprehensive capability score; the formula for feature fusion is expressed as: ; in, Employees are scored based on their overall capabilities; is the comprehensive weight of the j-th type of indicator; is the quantified value of the j-th type of indicator.

[0044] Substitute the normalized value and overall weight of employee G000001: ; The overall ability score of employee G000001 is 0.4478.

[0045] S5: Based on the comprehensive ability score and combined with the real-time data updates of various indicators, dynamically generate employee digital profiles. The employee digital profiles include core ability dimension scores, strengths and weaknesses analysis, and development potential assessment. The employee digital profiles are displayed in multi-dimensional charts, including comprehensive ability radar charts, indicator weight bar charts, dynamic trend charts, and attendance status quantitative charts.

[0046] The dynamic update mechanism will be implemented, and the profile will be updated when the following conditions are met: (1) When adding structured data, the 2025 annual assessment results are entered and the information of newly established scientific research projects in 2025 is supplemented.

[0047] (2) When supplementing unstructured data, academic papers published in 2025 and provincial commendation documents obtained in 2025.

[0048] (3) When adjusting the calculation parameters of quantitative indicators, such as updating the time decay coefficient rules and adjusting the value of the job importance coefficient.

[0049] The update process is as follows: First, a consistency check is performed to confirm that the employee IDs of the newly added data are correctly bound and that the data attributes match the quantitative indicators. Then, a recalculation is performed to update the quantitative indicator values ​​of the corresponding dimensions based on the newly added data. The comprehensive weight is updated synchronously through an adaptive weight model to refresh the comprehensive capability score. Finally, a trajectory record is performed to store the indicator values, weights, and scores before the change, forming a profile evolution trajectory that supports backtracking queries by time dimension.

[0050] The profile content is presented in four types of charts through a web-based visualization interface: like Figure 2 The image shown is a radar chart of comprehensive capabilities. Taking employee G000001 as an example, the chart uses four axes: performance evaluation, scientific research, resume, and honors, to mark the scores of each dimension and intuitively display the distribution of core capabilities. like Figure 3The chart shown is a bar chart of indicator weights, which presents the comprehensive weights of the four types of indicators to clarify the evaluation focus. Here, the indicator weights for employee G000001 are displayed. like Figure 4 The chart shown is a dynamic trend graph, which uses a line graph to show the changing trend of employee G000001's comprehensive ability score and scores in each dimension from 2020 to 2025, presenting the employee's growth trajectory. like Figure 5 The chart shown is a quantitative representation of attendance status. It combines attendance records in structured data, such as the number of times late, early departure, and leave requests, and quantifies them into attendance scores. The attendance compliance rate of employee G000001 is displayed in a pie chart.

[0051] This embodiment constructs digital profiles of 50 employees, achieving 100% correlation between structured and unstructured data, eliminating data fragmentation issues, and increasing data coverage from 75% in traditional methods to 98%. Through adaptive weight allocation, the profile score matches the actual performance of employees by 89%, a 23% improvement over fixed-weight methods. The profile update response time after data changes is less than or equal to 10 minutes, supports historical trajectory backtracking, and solves the problem of static profiles in traditional methods. Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for constructing digital profiles based on multi-source heterogeneous data, characterized in that, The method steps include: S1: Obtain structured and unstructured data of employees; the structured data includes annual performance evaluation results, scientific research project information, work experience and training records, and the unstructured data includes academic papers, academic certificates of appointment and commendation documents; establish a multi-source heterogeneous employee dataset based on the structured and unstructured data; S2: Standardize the multi-source heterogeneous employee dataset to obtain a standardized dataset; S3: Construct quantitative indicators based on the standardized dataset; the quantitative indicators include assessment dimension indicators, scientific research dimension indicators, resume dimension indicators, and honor dimension indicators; S4: An adaptive weight allocation model is designed based on the attention mechanism, and the comprehensive weights of various indicators are obtained by combining subjective weights and objective weights; the comprehensive weights of various indicators after quantification of the standardized dataset are fused through feature fusion to obtain the comprehensive ability score; S5: Based on the comprehensive ability score and combined with the real-time data updates of various indicators, dynamically generate employee digital profiles. The employee digital profiles include core ability dimension scores, strengths and weaknesses analysis, and development potential assessment. The employee digital profiles are displayed in multi-dimensional charts, including comprehensive ability radar charts, indicator weight bar charts, dynamic trend charts, and attendance status quantitative charts.

2. The method for constructing a digital profile based on multi-source heterogeneous data according to claim 1, characterized in that, The specific implementation of establishing the multi-source heterogeneous employee dataset is as follows: Assign a unique identifier to each employee and bind the acquired structured and unstructured data to the unique identifiers respectively; organize the structured data, and organize the annual assessment results, scientific research project information, job resumes and training records into standardized field formats, and store them as structured data tables according to the unique identifiers; Natural language processing techniques are used to extract core metadata from unstructured data and form a metadata list; Using a unique identifier as an index, establish a mapping relationship between structured data tables and metadata lists, and clarify the correspondence between fields corresponding to each employee and metadata. Simultaneously, unstructured data originals are named according to the order of employee unique identifier, data type, and time, forming an index linking the original storage path and metadata; structured data tables, metadata lists, association mapping relationships, and original storage indexes are integrated to construct a multi-source heterogeneous employee dataset containing data content, association relationships, and storage paths.

3. The method for constructing a digital profile based on multi-source heterogeneous data according to claim 1, characterized in that, The standardization process involves cleaning, aligning, and normalizing the multi-source heterogeneous employee dataset. The data cleaning includes missing value imputation and outlier removal. Missing value imputation uses the median of similar employee data, and outlier removal uses the 3σ criterion. Data alignment is based on the unique employee identifier, achieving alignment of time and attribute dimensions between different data sources.

4. The method for constructing a digital profile based on multi-source heterogeneous data according to claim 1, characterized in that, The quantitative formulas for the assessment dimensions and indicators are as follows: ; in, As assessment dimensions and indicators; Let be the grading coefficient for the k-th assessment. The time decay coefficient for the k-th assessment is denoted by m; m is the number of assessments. The quantitative formula for the scientific research dimension indicators is as follows: ; in, Indicators for scientific research; Let p be the level coefficient of the p-th research project; Let p be the participation coefficient of the p-th research project; Let q be the completion quality coefficient of the p-th research project; q is the number of research projects. The quantitative formula for the resume dimension indicators is as follows: ; in, d represents the resume dimension indicator; y represents the weighting coefficient for years of service; e represents the cumulative years of service; g represents the coefficient for job importance; The quantitative formula for the honor dimension indicators is as follows: ; in, For the honor dimension indicators; Let r be the level coefficient of the r-th honor; Let be the time decay coefficient for the acquisition of the r-th honor; n is the number of honors.

5. The method for constructing a digital profile based on multi-source heterogeneous data according to claim 1, characterized in that, The adaptive weight allocation model consists of an attention mechanism calculation unit, a subjective weight generation unit, an objective weight calculation unit, and a weight fusion unit; by inputting the quantification results of the quantification indicators, it outputs the comprehensive weights of each quantification indicator. The attention mechanism calculation unit is based on the Transformer multi-head attention mechanism to explore the correlation between various quantitative indicators and employees' core competencies; The subjective weight generation unit is based on the analytic hierarchy process and is used to process information related to evaluation requirements; The objective weight calculation unit is based on the entropy weight method and is used to analyze the indicator distribution characteristics of the standardized dataset. The weight fusion unit is used to integrate subjective weights, objective weights, and attention-related results.

6. The method for constructing a digital profile based on multi-source heterogeneous data according to claim 4, characterized in that, The expression for the comprehensive weight is: ; in, is the comprehensive weight of the j-th type of indicator; This is the subjective weighting adjustment coefficient; Subjective weights based on the analytic hierarchy process (AHP); For objective weights based on the entropy weight method; The subjective weights are obtained by constructing a judgment matrix using the analytic hierarchy process and calculating the eigenvector corresponding to the largest eigenvalue. The objective weights are calculated by determining the information entropy of various indicators, according to the entropy weight formula: ; in, Let be the information entropy of the j-th type of index.

7. The method for constructing a digital profile based on multi-source heterogeneous data according to claim 1, characterized in that, The formula for feature fusion is expressed as follows: ; in, Employees are scored based on their overall capabilities; is the comprehensive weight of the j-th type of indicator; is the quantified value of the j-th type of indicator.

8. The method for constructing a digital profile based on multi-source heterogeneous data according to claim 1, characterized in that, The dynamic updating of the employee digital profile follows the data change-driven principle. The update triggering conditions include the addition of structured data from multi-source heterogeneous employee datasets, the supplementation of unstructured data, and the adjustment of quantitative indicator calculation parameters. During the update process, the consistency of the changed data is first checked to confirm the accuracy of the binding between the data and the employee's unique identifier and the reasonableness of the matching between the data attributes and the corresponding quantitative indicators. Based on the verified change data, the quantitative indicator values ​​of the corresponding dimensions are recalculated, and the comprehensive weight of various indicators is updated synchronously through the adaptive weight allocation model. Finally, the comprehensive ability score of employees and the core ability dimension scores, strengths and weaknesses analysis and development potential assessment results contained in the employee digital profile are refreshed according to the feature fusion formula. Each update to the employee digital profile is based on the data change trajectory, recording the original information of the changed data, the quantitative indicator values ​​before and after the change, the details of changes in comprehensive weight and comprehensive ability score, forming a traceable evolution trajectory of the employee digital profile, and supporting backtracking and querying the status of the employee digital profile at different stages by time dimension.