Method and system for generating and sharing talent enabling data based on digitization

Through digital talent empowerment data generation and sharing methods and systems, problems such as data non-interoperability and inconsistent definitions have been solved, standard integration and personalized empowerment of multi-source talent data have been achieved, a safe and controllable data sharing and optimization mechanism has been provided, and data value mining and privacy protection have been improved.

CN120672177APending Publication Date: 2025-09-19QINGDAO XIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510520478.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing talent empowerment data generation and sharing system has problems such as data non-interoperability, inconsistent definitions, asynchronous data updates, sharing difficulties, security and privacy risks, lack of intelligent analysis methods and insufficient data value mining.

Method used

A digital-based talent empowerment data generation and sharing method is adopted, including data collection and standardization, talent data modeling and analysis, data sharing and permission management, feedback and optimization. It uses ETL engines, stream processing engines, natural language processing and knowledge graphs and other technologies, combined with core components such as Apache NiFi, Spark ML, and Hyperledger Fabric to achieve standardization, modeling and secure sharing of multi-source data.

Benefits of technology

Break down data silos, achieve standardized integration of multi-source talent data, fully tap talent potential, provide personalized empowerment suggestions, achieve controllable data circulation and privacy protection, and optimize data models and empowerment strategies.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to a digitization-based talent enabling data generation sharing system, which comprises a data acquisition and standardization framework, a talent data modeling and analysis framework, a data sharing and authority management framework and a feedback and optimization framework. Standard fusion of multi-source talent data is realized, talent potentials are fully mined, personalized enabling suggestions are provided, controllable data circulation is realized based on authority management and privacy protection technologies, and a data model and an enabling strategy are continuously optimized through a feedback mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention particularly relates to a method and system for generating and sharing talent empowerment data based on digitization. Background Art

[0002] The existing talent empowerment data generation and sharing systems are generally designed to be relatively simple. Data between systems such as HRM, OA and CRM are not interoperable, and the same fields are inconsistently defined in different systems. For example, the "professional skill level" is a 5-level system in system A and a 10-level system in system B. The data protectionism of business departments leads to sharing difficulties, data updates are not synchronized, the data sharing mechanism is imperfect, there are security and privacy risks, there is a lack of intelligent analysis methods, and data value mining is insufficient. Summary of the Invention

[0003] The technical problem to be solved by the present invention is: in order to overcome the shortcomings of the existing technology, a method and system for generating and sharing talent empowerment data based on digitalization is provided.

[0004] The technical solution adopted by the present invention to solve the technical problem is: a method for generating and sharing talent empowerment data based on digitalization, including the following processes:

[0005] a. Data collection and standardization;

[0006] b. Talent data modeling and analysis;

[0007] c. Data sharing and authority management;

[0008] d. Feedback and optimization;

[0009] The data collection and standardization includes input data sources and data processing methods. The data sources include structured data, unstructured data and real-time data. The data processing methods include ETL engines, stream processing engines, natural language processing and knowledge graphs.

[0010] Structured data includes data from HR systems, performance management systems, and learning platforms;

[0011] Unstructured data includes project documents, social media presence, and skills certifications;

[0012] Real-time data includes equipment-monitored training participation and online assessment results;

[0013] The ETL engine is used for data cleaning, data deduplication, and data standardization;

[0014] Stream processing engines are used to process continuous data streams in real time;

[0015] Natural language processing is used to parse unstructured text such as resumes and project reports;

[0016] Knowledge graphs are used to build a network of connections among talent skills, experiences, and interests;

[0017] The basic attributes of talent data modeling include hard indicators, soft indicators, education, years of work experience, technical certifications, personality traits, communication skills, and stress tolerance;

[0018] The talent data modeling process includes data preprocessing, model training, and model validation;

[0019] Data preprocessing includes missing value processing, outlier processing and feature engineering;

[0020] Model validation uses Time-based Split validation. Key indicators for model validation include accuracy, AUC-ROC curve, and business indicator matching.

[0021] The talent data analysis includes four dimensions: capability assessment, career development, team matching, and demand forecasting. The capability assessment dimension analyzes technical level and potential prediction, using machine learning as the technical method. The career development dimension analyzes growth path recommendations, using reinforcement learning and knowledge graph reasoning as the technical methods. The team matching dimension analyzes optimal team composition, using graph computing as the technical method. The demand forecasting dimension analyzes future technology needs, using time series analysis as the technical method.

[0022] Data sharing modes include internal sharing, cross-organizational sharing, and personal data sovereignty;

[0023] Internal sharing includes data permission management and data desensitization based on role-based access control;

[0024] Cross-organizational sharing uses consortium chain technology and smart contracts;

[0025] Feedback and optimization include user feedback, A / B testing, and model iteration.

[0026] Furthermore, the missing value processing method is the multiple imputation method, the outlier detection method is the isolation forest method, and the feature engineering method is PCA dimensionality reduction.

[0027] Furthermore, the data cleaning processing system includes an outlier detection technology stack and a missing value processing strategy matrix; data deduplication includes a multi-dimensional similarity matching algorithm and distributed deduplication implementation; and the data standardization architecture includes a value domain standardization engine, natural language processing standardization, and spatiotemporal data unification.

[0028] Furthermore, the model verification method includes BERT consistency detection, code base reverse verification and network consistency analysis. BERT consistency detection is used to verify resume text, and adopts an error correction strategy that automatically aligns with JD requirements. Code base reverse verification is used to verify project experience, and adopts an error correction strategy that triggers superior confirmation. Network consistency analysis is used to verify social data, and adopts a weighted fusion error correction strategy.

[0029] Furthermore, the implementation steps of the multiple imputation method include missing pattern diagnosis, imputation process optimization and result merging rules.

[0030] A system using the above-described digital-based talent empowerment data generation and sharing method, including a data collection and standardization architecture, a talent data modeling and analysis architecture, a data sharing and rights management architecture, and a feedback and optimization architecture;

[0031] The core components of the data collection and standardization architecture are Apache NiFi and Debezium, and plug-in connectors are used for scalability.

[0032] The core components of the talent data modeling and analysis architecture are Spark ML and Flink ML, and a custom algorithm registration center is used for scalability design.

[0033] The core components of the data sharing and rights management architecture are Hyperledger Fabric and IPFS, and a multi-chain Internet gateway is used for scalability design.

[0034] Among them, the feedback and optimization architecture uses MLflow and Evidently, and adopts an automatic rollback mechanism for scalability design.

[0035] Furthermore, the data collection and standardization architecture includes multi-source data collection, intelligent access layer, ETL pipeline, NLP processing engine, data lake, stream processing engine and standardization service. Key components include CDC log capture, outlier detection, resume parser, speech-to-text, metadata management and data passport generation.

[0036] Furthermore, the talent data modeling and analysis architecture includes standard data collection, analysis engine, batch modeling, streaming analysis, model warehouse and service packaging. Model types include clustering model, prediction model, CEP rule engine, real-time feature calculation, REST API and GraphQL.

[0037] Furthermore, the data sharing and rights management architecture includes data service collection, policy execution point, ABAC engine, blockchain network, data export and consumer end, and rights management includes encryption and desensitization, dynamic watermark, smart contract and evidence hash.

[0038] Furthermore, the feedback and optimization architecture includes user behavior collection, feedback collector, effect evaluation, model iteration, data correction and version release. It has a closed-loop system, which includes an A / B testing framework, a multi-core correction platform and grayscale release.

[0039] The beneficial effect of the present invention is that the digital-based talent empowerment data generation and sharing method and system breaks the data silos, realizes the standard integration of multi-source talent data, fully taps the potential of talents, provides personalized empowerment suggestions, realizes controllable data circulation based on authority management and privacy protection technology, and continuously optimizes data models and empowerment strategies through feedback mechanisms. DETAILED DESCRIPTION

[0040] A digital-based talent empowerment data generation and sharing method includes the following processes:

[0041] a. Data collection and standardization;

[0042] b. Talent data modeling and analysis;

[0043] c. Data sharing and authority management;

[0044] d. Feedback and optimization;

[0045] The data collection and standardization includes input data sources and data processing methods. The data sources include structured data, unstructured data and real-time data. The data processing methods include ETL engines, stream processing engines, natural language processing and knowledge graphs.

[0046] Structured data includes data from HR systems, performance management systems, and learning platforms;

[0047] Unstructured data includes project documents, social media presence, and skills certifications;

[0048] Real-time data includes equipment-monitored training participation and online assessment results;

[0049] The ETL engine is used for data cleaning, data deduplication, and data standardization;

[0050] Stream processing engines are used to process continuous data streams in real time;

[0051] Natural language processing is used to parse unstructured text such as resumes and project reports;

[0052] Knowledge graphs are used to build a network of connections among talent skills, experiences, and interests;

[0053] The basic attributes of talent data modeling include hard indicators, soft indicators, education, years of work experience, technical certifications, personality traits, communication skills, and stress tolerance;

[0054] The talent data modeling process includes data preprocessing, model training, and model validation;

[0055] Data preprocessing includes missing value processing, outlier processing and feature engineering;

[0056] Model validation uses Time-based Split validation. Key indicators for model validation include accuracy, AUC-ROC curve, and business indicator matching.

[0057] The talent data analysis includes four dimensions: capability assessment, career development, team matching, and demand forecasting. The capability assessment dimension analyzes technical level and potential prediction, using machine learning as the technical method. The career development dimension analyzes growth path recommendations, using reinforcement learning and knowledge graph reasoning as the technical methods. The team matching dimension analyzes optimal team composition, using graph computing as the technical method. The demand forecasting dimension analyzes future technology needs, using time series analysis as the technical method.

[0058] Data sharing modes include internal sharing, cross-organizational sharing, and personal data sovereignty;

[0059] Internal sharing includes data permission management and data desensitization based on role-based access control;

[0060] Cross-organizational sharing uses consortium chain technology and smart contracts;

[0061] Feedback and optimization include user feedback, A / B testing, and model iteration.

[0062] Furthermore, the missing value processing method is the multiple imputation method, the outlier detection method is the isolation forest method, and the feature engineering method is PCA dimensionality reduction.

[0063] Furthermore, the data cleaning processing system includes an outlier detection technology stack and a missing value processing strategy matrix; data deduplication includes a multi-dimensional similarity matching algorithm and distributed deduplication implementation; and the data standardization architecture includes a value domain standardization engine, natural language processing standardization, and spatiotemporal data unification.

[0064] Furthermore, the model verification method includes BERT consistency detection, code base reverse verification and network consistency analysis. BERT consistency detection is used to verify resume text, and adopts an error correction strategy that automatically aligns with JD requirements. Code base reverse verification is used to verify project experience, and adopts an error correction strategy that triggers superior confirmation. Network consistency analysis is used to verify social data, and adopts a weighted fusion error correction strategy.

[0065] Furthermore, the implementation steps of the multiple imputation method include missing pattern diagnosis, imputation process optimization and result merging rules.

[0066] A system using the above-described digital-based talent empowerment data generation and sharing method, including a data collection and standardization architecture, a talent data modeling and analysis architecture, a data sharing and rights management architecture, and a feedback and optimization architecture;

[0067] The core components of the data collection and standardization architecture are Apache NiFi and Debezium, and plug-in connectors are used for scalability.

[0068] The core components of the talent data modeling and analysis architecture are Spark ML and Flink ML, and a custom algorithm registration center is used for scalability design.

[0069] The core components of the data sharing and rights management architecture are Hyperledger Fabric and IPFS, and a multi-chain Internet gateway is used for scalability design.

[0070] Among them, the feedback and optimization architecture uses MLflow and Evidently, and adopts an automatic rollback mechanism for scalability design.

[0071] Furthermore, the data collection and standardization architecture includes multi-source data collection, intelligent access layer, ETL pipeline, NLP processing engine, data lake, stream processing engine and standardization service. Key components include CDC log capture, outlier detection, resume parser, speech-to-text, metadata management and data passport generation.

[0072] Furthermore, the talent data modeling and analysis architecture includes standard data collection, analysis engine, batch modeling, streaming analysis, model warehouse and service packaging. Model types include clustering model, prediction model, CEP rule engine, real-time feature calculation, REST API and GraphQL.

[0073] Furthermore, the data sharing and rights management architecture includes data service collection, policy execution point, ABAC engine, blockchain network, data export and consumer end, and rights management includes encryption and desensitization, dynamic watermark, smart contract and evidence hash.

[0074] Furthermore, the feedback and optimization architecture includes user behavior collection, feedback collector, effect evaluation, model iteration, data correction and version release. It has a closed-loop system, which includes an A / B testing framework, a multi-core correction platform and grayscale release.

[0075] In the cross-organizational sharing model, the main role of consortium chain technology is to ensure data traceability and tamper-proofing. Smart contracts are mainly used to set automated data exchange rules, such as limiting salary data to authorized access only.

[0076] User feedback is used to provide feedback on corresponding results, such as satisfaction with the recommendation results;

[0077] A / B testing is mainly used to compare the effects of different empowerment strategies;

[0078] Model iteration is mainly for updating AI models through online learning.

[0079] This method maximizes the value of talent data through intelligent and secure data sharing.

[0080] The basic core functions of the ETL engine:

[0081] 1. Data integration hub role:

[0082] Multi-source heterogeneous connectivity: supports integration with multiple systems, including HRIS (such as Workday), ATS (such as Greenhouse), and LMS (such as Cornerstone).

[0083] Real-time / batch dual-mode acquisition: daily batch processing throughput can reach TB level, and real-time stream processing latency is less than 1 second;

[0084] Intelligent data recognition: Automatically detect 200+ resume file formats, including unstructured PDF / image resumes;

[0085] 2. Core carriers of data governance:

[0086] Quality rule engine: Built-in 300+ talent data verification rules, such as "salary-level-years of experience" triangulation verification;

[0087] Data lineage tracking: Field-level change traceability to meet GDPR / CCPA compliance requirements;

[0088] Version control: supports snapshot backtracking of talent data, which can be traced back to any historical version;

[0089] 3. Value transformation bridge

[0090] Indicator processing: Converting raw data into 58 standard talent indicators, such as turnover risk index and skill popularity coefficient;

[0091] Feature Engineering: Generate 400+ machine learning features, such as project complexity score and collaboration network density;

[0092] Service-ready: Outputs multiple consumption forms such as API, dataset, and real-time stream.

[0093] The modeling of talent data mainly constructs a digital expression of talent capabilities in a structured and quantitative manner.

[0094] It mainly realizes four functions:

[0095] 1. Accurately assess and objectively measure the capabilities of individuals or teams;

[0096] 2. Development forecast, identification of growth potential and future development direction;

[0097] 3. Intelligent matching to optimize job matching and team composition;

[0098] 4. Decision support, providing data basis for talent management.

[0099] Modeling quantitative evaluation models include weighted scoring models, clustering models, prediction models and recommendation models.

[0100] The weighted scoring model uses the AHP hierarchical analysis algorithm and is mainly used for basic ability assessment;

[0101] The clustering model uses K-Means and DBSCAN algorithms, which are mainly used for talent classification;

[0102] The prediction model uses XGBoost and LSTM algorithms, mainly used to predict turnover risk and promotion potential;

[0103] The recommendation model uses collaborative filtering and GNN algorithms, and is mainly used for training courses and job recommendations.

[0104] This system uses transfer learning and cross-job knowledge transfer methods to solve the problem of data sparsity;

[0105] The multi-source data cross-validation method is used to solve the objectivity problem of the evaluation;

[0106] Use SHAP value analysis to solve the problem of model interpretability;

[0107] The concept drift detection mechanism is used to solve the problem of data drift.

[0108] RPA process robots can also be used to automatically capture data from various systems;

[0109] Use computer vision detection to provide training participation analysis data;

[0110] Use voice emotion analysis software to analyze interviews and job descriptions;

[0111] IDE code commit pattern analysis through digital footprint tracking.

[0112] When a data format error occurs in the system, it is transferred to the dead letter queue for manual repair, which is mainly used when non-standard resume parsing fails;

[0113] When a system connection timeout occurs, the exponential backoff retry mechanism is used, which is mainly used for performance system API throttling.

[0114] When a business rule conflict occurs in the system, the approval workflow is triggered, which is mainly used when problems arise in cross-border work experience verification.

[0115] Missing value handling strategy:

[0116] When random missingness occurs, multiple imputation is used, such as supplementing missing items in employee satisfaction surveys;

[0117] When non-random missing data occurs, a missing flag field should be established, such as to handle unfilled reasons for leaving a job;

[0118] When structured missing information occurs, it is implemented through rule-based inference logic, such as estimating the years of work experience of recent graduates.

[0119] In the closed-loop quality monitoring system, the KPI dashboard for data quality is:

[0120] The cleaning efficiency is calculated using the formula: effective processing records / total number of problem records, with the compliance threshold being ≥95%;

[0121] Calculate the deduplication compression ratio using the formula: number of records after deduplication / number of original records, with the compliance threshold ≤ 85%;

[0122] The standardized consistency rate was calculated using the formula: number of consistent standard fields / total number of fields, with the compliance threshold being ≥98%.

[0123] Example 1

[0124] Talent pool construction project for a retail enterprise:

[0125] 1. Data Challenges

[0126] Merged HR systems of three acquired subsidiaries;

[0127] There is 23% duplicate candidate data;

[0128] There are 148 different descriptions of skills; 2. ETL solution:

[0129] Adopt NLP-driven skill standardization management; implement graph-based resume deduplication;

[0130] Establish a rule base for converting cross-border work transfer experiences; 3. Implementation effect:

[0131] The talent data quality index increased from 58 to 92; the recruitment process efficiency increased by 40%;

[0132] The accuracy of talent inventory increased by 35%.

[0133] Example 2

[0134] Multi-drop effect of a financial group

[0135] 1. Collection stage:

[0136] Resume parsing accuracy increased from 68% to 92%; data freshness was less than 5 minutes of delay;

[0137] 2. Modeling stage:

[0138] The AUC for turnover prediction increased from 0.89 to 0.93; feature engineering efficiency increased by 8 times;

[0139] 3. Sharing stage:

[0140] The time spent on cross-departmental data collaboration is reduced by 90%;

[0141] Achieve 100% compliance audit pass rate;

[0142] 4. Optimization stage:

[0143] The model iteration cycle is shortened from quarterly to weekly;

[0144] False positive rates dropped by 42%.

[0145] Example 3

[0146] Talent analysis platform project for a financial group:

[0147] 1. Data integration scope: Integrate 9 systems and 37 data sources;

[0148] 2. Improved processing efficiency: The resume parsing speed has been reduced from 15 resumes per second to 8 resumes per second, and the report generation cycle has been reduced from three weeks to real-time generation;

[0149] 3. Business value: The cycle for filling key positions was shortened by 40%, the accuracy of identifying high-potential talents increased by 35%, and annual labor costs were saved by US$2.3 million.

[0150] Compared with existing technologies, this digital talent empowerment data generation and sharing method and system breaks down data silos, realizes the standard integration of multi-source talent data, fully taps talent potential, provides personalized empowerment suggestions, realizes controllable data circulation based on permission management and privacy protection technology, and continuously optimizes data models and empowerment strategies through feedback mechanisms.

[0151] With the above-described preferred embodiments of the present invention as a guide, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical scope of this invention. The technical scope of this invention is not limited to the contents of the specification and must be determined according to the scope of the claims.

Claims

1. A method for generating and sharing talent empowerment data based on digitalization, characterized in that: The following processes are included: a. Data collection and standardization; b. Talent data modeling and analysis; c. Data sharing and authority management; d. Feedback and optimization; The data collection and standardization includes input data sources and data processing methods. The data sources include structured data, unstructured data and real-time data. The data processing methods include ETL engines, stream processing engines, natural language processing and knowledge graphs. Structured data includes data from HR systems, performance management systems, and learning platforms; Unstructured data includes project documents, social media presence, and skills certifications; Real-time data includes equipment-monitored training participation and online assessment results; The ETL engine is used for data cleaning, data deduplication, and data standardization; Stream processing engines are used to process continuous data streams in real time; Natural language processing is used to parse unstructured text such as resumes and project reports; Knowledge graphs are used to build a network of connections among talent skills, experiences, and interests; The basic attributes of talent data modeling include hard indicators, soft indicators, education, years of work experience, technical certifications, personality traits, communication skills, and stress tolerance; The talent data modeling process includes data preprocessing, model training, and model validation; Data preprocessing includes missing value processing, outlier processing and feature engineering; Model validation uses Time-based Split validation. Key indicators for model validation include accuracy, AUC-ROC curve, and business indicator matching. The talent data analysis includes four dimensions: capability assessment, career development, team matching, and demand forecasting. The capability assessment dimension analyzes technical level and potential prediction, using machine learning as the technical method. The career development dimension analyzes growth path recommendations, using reinforcement learning and knowledge graph reasoning as the technical methods. The team matching dimension analyzes optimal team composition, using graph computing as the technical method. The demand forecasting dimension analyzes future technology needs, using time series analysis as the technical method. Data sharing modes include internal sharing, cross-organizational sharing, and personal data sovereignty; Internal sharing includes data permission management and data desensitization based on role-based access control; Cross-organizational sharing uses consortium chain technology and smart contracts; Feedback and optimization include user feedback, A / B testing, and model iteration.

2. The method for generating and sharing digital talent empowerment data according to claim 1, wherein: The missing value processing method is the multiple imputation method, the outlier detection method is the isolation forest method, and the feature engineering method is PCA dimensionality reduction.

3. The method for generating and sharing digital talent empowerment data according to claim 1, wherein: The data cleaning processing system includes the outlier detection technology stack and the missing value processing strategy matrix. Data deduplication includes a multi-dimensional similarity matching algorithm and distributed deduplication implementation. The data standardization architecture includes a value domain standardization engine, natural language processing standardization, and spatiotemporal data unification.

4. The method for generating and sharing digital talent empowerment data according to claim 1, wherein: Model verification methods include BERT consistency detection, code base reverse verification and network consistency analysis. BERT consistency detection is used to verify resume text, and adopts an error correction strategy that automatically aligns with JD requirements. Code base reverse verification is used to verify project experience, and adopts an error correction strategy that triggers superior confirmation. Network consistency analysis is used to verify social data, and adopts a weighted fusion error correction strategy.

5. The method for generating and sharing digital talent empowerment data according to claim 2, wherein: The implementation steps of the multiple imputation method include missing pattern diagnosis, imputation process optimization and result merging rules.

6. A system using the method for generating and sharing digital talent empowerment data according to any one of claims 1 to 5, characterized in that: Including data collection and standardization architecture, talent data modeling and analysis architecture, data sharing and authority management architecture, and feedback and optimization architecture; The core components of the data collection and standardization architecture are Apache NiFi and Debezium, and plug-in connectors are used for scalability. The core components of the talent data modeling and analysis architecture are Spark ML and Flink ML, and a custom algorithm registration center is used for scalability design. The core components of the data sharing and rights management architecture are Hyperledger Fabric and IPFS, and a multi-chain Internet gateway is used for scalability design. Among them, the feedback and optimization architecture uses MLflow and Evidently, and adopts an automatic rollback mechanism for scalability design.

7. The system for generating and sharing digital talent empowerment data according to claim 6, characterized in that: The data collection and standardization architecture includes multi-source data collection, intelligent access layer, ETL pipeline, NLP processing engine, data lake, stream processing engine and standardization service. Key components include CDC log capture, outlier detection, resume parser, speech-to-text, metadata management and data passport generation.

8. The system for generating and sharing digital talent empowerment data according to claim 6, characterized in that: The talent data modeling and analysis architecture includes standard data collection, analysis engine, batch modeling, streaming analysis, model warehouse and service packaging. Model types include clustering models, predictive models, CEP rule engine, real-time feature calculation, REST API and GraphQL.

9. The system for generating and sharing digital talent empowerment data according to claim 6, characterized in that: The data sharing and rights management architecture includes data service collection, policy execution point, ABAC engine, blockchain network, data export and consumer end. Rights management includes encryption and desensitization, dynamic watermark, smart contract and evidence hash.

10. The system for generating and sharing digital talent empowerment data according to claim 6, characterized in that: The feedback and optimization architecture includes user behavior collection, feedback collector, effect evaluation, model iteration, data correction and version release. It has a closed-loop system that includes an A / B testing framework, a multi-core correction platform and grayscale release.