Talent cultivation recommendation method based on federal learning and natural language processing
Through the talent training recommendation method based on federated learning and natural language processing, problems such as insufficient accuracy, insufficient personalization, data silos and privacy in the existing job matching methods are solved, efficient, accurate and personalized job matching and course optimization are achieved, and the coordinated development of education and employment is promoted.
Patent Information
- Application Number
- CN202510100547.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing job matching methods have problems such as insufficient matching accuracy, insufficient personalization, data silos and privacy, insufficient dynamics and lack of interpretation, which is difficult to meet the needs of the modern education and employment market.
The talent training recommendation method based on federated learning and natural language processing is adopted. By collecting multi-source heterogeneous data, building talent knowledge graphs, anonymization and desensitization processing data, using federated learning framework for joint modeling, generating dynamic updated talent and job portraits, and matching them with multi-objective optimization models, and finally using interpretable AI technology for visual interpretation.
It improves matching accuracy and efficiency, breaks down data silos, protects privacy and security, dynamically adapts to changes in market demand, meets the personalized needs of multiple parties, and promotes the precise connection between education and employment.
Smart Images

Figure CN119988735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of employment matching technology, and in particular to a talent training recommendation method based on federated learning and natural language processing. Background Art
[0002] In the modern education and employment market, the precise matching of course content with job requirements is a key issue that needs to be urgently addressed in the vocational education and higher education systems. With the rapid iteration of industrial technology and the diversification of market demand, the traditional education model is unable to quickly respond to the requirements of enterprises for talent skills, which leads to a waste of educational resources and a decline in the employment competitiveness of graduates.
[0003] At present, the existing job matching methods include but are not limited to traditional methods based on keyword matching, matching methods based on scoring and weighting, matching methods based on recommendation systems, intelligent matching methods based on artificial intelligence, etc. However, the existing job matching methods have the following problems: 1) Insufficient matching accuracy. Traditional methods such as keyword matching lack a deep understanding of job requirements and candidate capabilities, and are prone to missing potential matches or generating mismatches. 2) Insufficient personalization. Most existing systems use general matching rules, ignoring the personalized needs of enterprises and candidates, such as corporate culture and candidates' career development intentions. 3) Data silos and privacy issues. Data between educational institutions, enterprises and talents are not effectively shared, resulting in incomplete matching information, and the security of personal and corporate sensitive information is difficult to guarantee when using data. 4) Insufficient dynamics. At present, most matching systems are statically designed, that is, fixed recommendation rules are generated after model training, and it is impossible to dynamically adjust course design or job recommendations according to rapidly changing market needs. 5) Lack of explainability: Although intelligent matching methods (such as deep learning) have good results, the matching results are difficult to explain, and it is difficult for enterprises and candidates to be convinced. Summary of the invention
[0004] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide a talent training recommendation method based on federated learning and natural language processing, which can not only improve matching accuracy and efficiency, break data silos, and protect privacy and security, but also dynamically adapt to changes in market demand, meet the personalized needs of multiple parties, and promote the precise connection between education and employment.
[0005] To achieve the above object, the present invention provides the following solution: a talent training recommendation method based on federated learning and natural language processing, comprising the following steps:
[0006] Collect relevant data from educational institutions, enterprises and third-party platforms to obtain multi-source heterogeneous data, store the multi-source heterogeneous data in a data lake, clean and integrate the multi-source heterogeneous data to obtain standard data, and store the standard data in a data warehouse;
[0007] Construct a talent knowledge graph related to talent development and job skills, and perform real-time incremental updates to the talent knowledge graph;
[0008] Anonymize and desensitize the sensitive data in the standard data, and use the federated learning framework to perform joint modeling to obtain a local model without sharing the original data between educational institutions and enterprises;
[0009] Based on the standard data, students and positions are deeply semantically characterized to generate dynamically updated talent portraits and position portraits, and a matching algorithm is preset in combination with multiple target factors and dynamic weights, and then the talent portraits, the position portraits and the matching algorithm are deployed to the local model to obtain a talent matching model;
[0010] Continuously monitor the similarity gap between the talent profile and the job profile, and dynamically generate a course optimization plan and a job optimization plan based on the similarity gap and the output of the talent matching model;
[0011] Explainable AI technology is used to provide visual explanations or keyword descriptions of the output of the talent matching model, the course optimization plan, and the job optimization plan.
[0012] Optionally, collecting relevant data from educational institutions, enterprises and third-party platforms to obtain multi-source heterogeneous data, storing the multi-source heterogeneous data in a data lake, and then performing data cleaning and integration on the multi-source heterogeneous data to obtain standard data, and storing the standard data in a data warehouse, including:
[0013] According to the unified data format and field definition, use API interface or crawler technology to obtain relevant data from educational institutions, enterprises and third-party platforms to obtain multi-source heterogeneous data;
[0014] A data lake is constructed by using a distributed storage system, and the multi-source heterogeneous data is partitioned and stored in the data lake according to the data source and type;
[0015] Processing missing values, outliers, duplicate data and data integration on the multi-source heterogeneous data to obtain standard data;
[0016] A data warehouse is constructed using a relational database, the standard data partitions are stored in the data warehouse according to data subjects, and a query index is created in the data warehouse.
[0017] Optionally, a talent knowledge graph related to talent training and job skills is constructed, and the talent knowledge graph is incrementally updated in real time, including:
[0018] Collecting position-related data, industry-related data and course-related data to obtain an initial text, and preprocessing the initial text to obtain a knowledge text;
[0019] Combine the BERT model to understand the contextual semantics of the knowledge text, use NLP technology to extract entities from the knowledge text, and then map the entities to a standardized vocabulary in combination with an external knowledge base to complete the entity standardization operation;
[0020] Use dependency parsing to extract relationships between entities and classify the relationships;
[0021] Select a graph database, import entities and relationships into the selected graph database, obtain a graph structure, use graph embedding technology to map nodes and relationships in the graph structure into a vector space, generate node embedding vectors, complete semantic matching operations, and obtain a talent knowledge graph;
[0022] The newly added data of the initial text is automatically parsed, and the talent knowledge graph is updated according to the parsing result.
[0023] Optionally, the sensitive data in the standard data is anonymized and desensitized, and on the premise that the educational institution and the enterprise do not share the original data, a federated learning framework is used for joint modeling to obtain a local model, including:
[0024] Classify the standard data into personal privacy information, enterprise sensitive data and general data;
[0025] The personal privacy information is encrypted using an irreversible hash algorithm, and a random salt value is added before the hash to prevent rainbow table attacks and complete the anonymization of the personal privacy information;
[0026] For the enterprise sensitive data, use regularized expressions to identify sensitive fields, perform hashing, shielding or fuzzification operations on the sensitive fields, and complete the desensitization of the enterprise sensitive data;
[0027] Educational institutions and enterprises are set as participants. Each participant uses its own data to train the model locally, obtains the local model and local parameters, and completes the deployment of the local node.
[0028] Encrypt the local parameters using homomorphic encryption to obtain encrypted parameters, receive the encrypted parameters through a central server, and perform aggregation operations on the encrypted parameters to obtain global parameters, thereby completing the deployment of cloud nodes;
[0029] The global parameters are sent to each of the participants using a central server to perform a global update of the local model.
[0030] Optionally, based on the standard data, students are deeply semantically profiled to generate dynamically updated talent portraits, including:
[0031] Collecting multi-source student data and storing the multi-source student data in the data warehouse; the multi-source student data includes student course grades, skill mastery, internship experience, career preferences and other student data;
[0032] Extracting features from the student multi-source data to obtain student numerical features, student classification features, and student text features, then standardizing and normalizing the student numerical features, encoding the student classification features, and vectorizing the student text features to complete feature conversion;
[0033] Using a fully connected layer, the student numerical features and the student classification features are converted into a structured data vector of fixed dimension, a pre-trained language model is used to obtain a text data vector of the student text features, and then the structured data vector and the text data vector are fused to obtain a multi-dimensional student vector representation;
[0034] Match the students' skills with the skill nodes in the talent knowledge graph to obtain the students' skill semantic vector, use the graph neural network to generate a high-dimensional vector of the students' skill semantic vector, and then fuse the high-dimensional vector of the students' skill semantic vector with the multi-dimensional student vector representation to complete the deep semantic characterization of the students and generate a talent portrait.
[0035] Optionally, based on the standard data, a deep semantic description of the position is performed to generate a dynamically updated position portrait, including:
[0036] Collect multi-source data of positions, and store the multi-source data of positions in the data warehouse; the multi-source data of positions include position requirements, skill requirements, industry attributes, corporate culture and other data of positions;
[0037] Extracting features from the multi-source data of the positions to obtain position numerical features, position classification features and position text features, then standardizing and normalizing the position numerical features, encoding the position classification features, and vectorizing the position text features to complete feature conversion;
[0038] Using a fully connected layer, the position numerical features and the position classification features are converted into a structured data vector of fixed dimension, a pre-trained language model is used to obtain a text data vector of the position text features, and then the structured data vector and the text data vector are fused to obtain a multi-dimensional position vector representation;
[0039] The skills required for the position are matched with the skill nodes in the talent knowledge graph to obtain the position skill semantic vector, and a high-dimensional vector of the position skill semantic vector is generated using a graph neural network. The high-dimensional vector of the position skill semantic vector is then fused with the multi-dimensional position vector representation to complete a deep semantic characterization of the position and generate a position portrait.
[0040] Optionally, a matching algorithm is preset in combination with multiple target factors and dynamic weights, and then the talent portrait, the position portrait and the matching algorithm are deployed to the local model to obtain a talent matching model, including:
[0041] Mapping the talent portrait and the job portrait into a unified vector space to generate a talent portrait vector and a job portrait vector;
[0042] Using cosine similarity, calculate the cosine value of the angle between the talent portrait vector and the job portrait vector, and using Euclidean distance, calculate the straight-line distance between the talent portrait vector and the job portrait vector to obtain the semantic similarity between the talent portrait vector and the job portrait vector;
[0043] Select multiple objective factors, assign dynamic weights to different objective factors, and then combine optimization algorithms and reinforcement learning to build a multi-objective optimization model;
[0044] The semantic similarity and the multi-objective optimization model are combined to obtain a matching algorithm for calculating a comprehensive matching score, and the matching algorithm is deployed in the local model to obtain a talent matching model.
[0045] Optionally, the expression of the matching algorithm is:
[0046]
[0047] Among them, T is the comprehensive matching score, w i is the dynamic weight, f i is the score of each target factor, and m is the total number of target factors.
[0048] Optionally, continuously monitoring the similarity gap between the talent profile and the job profile, and dynamically generating a course optimization plan and a job optimization plan based on the similarity gap and the output of the talent matching model, including:
[0049] Utilize the talent matching model to obtain target positions and target position skill requirements, and compare and analyze and update the target position skill requirements with relevant data of corresponding positions in the talent knowledge graph;
[0050] Evaluate and quantify students' current skills, compare them with the updated skill requirements of the target position, obtain skill gaps, and then use machine learning algorithms to analyze the skill gaps to recommend relevant courses, obtain course optimization plans, and predict and plan courses based on industry trends and student needs;
[0051] Monitor changes in job requirements and courses in real time, optimize the target jobs based on the monitoring results and multiple target factors, and obtain job optimization plans.
[0052] Optionally, interpretable AI technology is used to provide a visual explanation or keyword description of the output of the talent matching model, the course optimization plan, and the job optimization plan, including:
[0053] Use the feature contribution graph to show the contribution of each feature to the matching results and recommendation results, use the decision path graph to show the decision path of the talent matching model, and use the keyword cloud to show the driving factors of job recommendations;
[0054] Provide a dashboard for viewing data usage and recommendation process, and design privacy terms to achieve transparent management.
[0055] The present invention discloses the following technical effects by providing a talent training recommendation method based on federated learning and natural language processing:
[0056] 1. High matching accuracy: 1) The present invention can construct the four core entities of "skills-courses-positions-industries" and their relationships through the talent knowledge graph, supporting more accurate semantic matching and recommendations. At the same time, the semantic structure and graph embedding technology of the knowledge graph significantly improve the matching efficiency and accuracy of positions and talents, giving the method stronger semantic understanding and reasoning capabilities. 2) By combining multi-source data and knowledge graphs to construct multi-dimensional talent portraits and position portraits, the semantic expression of high-dimensional vectors can be achieved, the depth and accuracy of the portraits are enhanced, and the similarity calculation between talent portraits and position portraits is made more accurate. Combined with the multi-objective optimization process, the present invention achieves accurate position matching and course optimization, and improves the efficiency of recruitment and training.
[0057] 2. High security: The present invention can ensure data security and enhance data protection by anonymizing and desensitizing personal privacy information (such as name, contact information, ID number) and sensitive enterprise data (such as financial information, core technology keywords). By utilizing the federated learning framework, the model can be jointly trained without sharing data between educational institutions and enterprises, solving the data island problem while protecting data privacy.
[0058] 3. Personalized recommendations and dynamic adjustment plans: The present invention can compare students’ current skills with the changing skill requirements of target positions, generate dynamic course optimization plans and position optimization plans, provide each student with the most suitable position and course, realize personalized dynamic recommendations, promote the coordinated development of education and industry, and enhance the practicality and competitiveness of talent training.
[0059] 4. High trust: The present invention makes the data transparent and traceable by visually displaying and explaining the recommendation process and data usage process, which enhances the user's understanding and trust and promotes collaboration between educational institutions and enterprises.
[0060] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0062] Figure 1 A schematic diagram of a method flow chart provided by an embodiment of the present invention;
[0063] Figure 2 A schematic diagram of a talent and position portrait generation process provided by an embodiment of the present invention;
[0064] Figure 3 A schematic diagram of a job recommendation process provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0067] like Figure 1 As shown, the present invention provides a talent training recommendation method based on federated learning and natural language processing, comprising the following steps:
[0068] 1. Collect relevant data from educational institutions, enterprises and third-party platforms to obtain multi-source heterogeneous data, store the multi-source heterogeneous data in the data lake, clean and integrate the multi-source heterogeneous data to obtain standard data, and store the standard data in the data warehouse. Including:
[0069] 1.1 Based on the unified data format (such as JSON format) and field definition, use API interface or crawler technology to obtain relevant data from educational institutions, enterprises and third-party platforms to obtain multi-source heterogeneous data.
[0070] Educational institutions: student information (name, student ID, course grades, internship experience), course information (course name, course content, skill points).
[0071] Enterprises: job information (job title, job description, skill requirements, salary range), recruitment data (interview records, hiring status).
[0072] Third-party platforms: industry trend data (popular positions, salary levels), recruitment website data (job requirements, company evaluations).
[0073] API interface: Get data in real time through RESTful API or GraphQL interface.
[0074] Batch import: supports batch data upload in CSV, Excel, JSON and other formats.
[0075] Crawler technology: For third-party platforms that cannot directly obtain APIs, use crawler technology to collect public data.
[0076] Field definition and mapping are shown in Table 1 Field Mapping Table below:
[0077] Table 1 Field mapping table
[0078] Data Source Original field name Standard field names Educational institutions Student ID student_id Educational institutions Name name enterprise Job Title job_title Third-party platforms Skills required skill_requirements
[0079] 1.2 Use a distributed storage system (such as Hadoop HDFS, Amazon S3) to build a data lake, and partition and store the multi-source heterogeneous data in the data lake according to the data source and type.
[0080] 1.3 Perform missing value processing, outlier processing, duplicate data processing and data integration on the multi-source heterogeneous data to obtain standard data.
[0081] 1.31 Missing value processing includes:
[0082] Deletion method: For non-key fields (such as remark information), directly delete the records with missing values.
[0083] Filling method: For key fields (such as student grades and job salaries), use the following method to fill:
[0084] Mean filling: fill with the mean of similar data (such as average grade).
[0085] Interpolation method: interpolation based on time series data.
[0086] Machine Learning Prediction: Predict missing values using regression models.
[0087] 1.32 Outlier processing includes:
[0088] Rule detection: Set a reasonable range (such as the score range is 0-100, the salary range is 3000-50000), and mark the data outside the range as abnormal.
[0089] Statistical tests: Outliers were detected using box plots (IQR) or Z-scores.
[0090] Processing method: Outliers can be deleted or replaced with the median.
[0091] 1.33 Duplicate data processing includes:
[0092] Detect duplicate data using hashing algorithms or unique identifiers (e.g., student ID, position ID). Merge duplicate records to keep the latest or most complete data.
[0093] 1.34 Data integration includes:
[0094] Entity unified ID mapping: Generate a unique identifier (UUID) for each entity (such as student, position).
[0095] Data association and merging: Use primary and foreign keys to associate data from different data sources. Example: Merge a student's course grades with their internship experience.
[0096] 1.4 Use a relational database (such as PostgreSQL, Amazon Redshift) to build a data warehouse, store the standard data partitions in the data warehouse according to data topics (to support fast retrieval and analysis), and create query indexes (such as student_id, job_id) in the data warehouse to improve query efficiency.
[0097] 2. Construct a talent knowledge graph related to talent training and job skills, and perform real-time incremental updates on the talent knowledge graph. Including:
[0098] 2.1 Collect job-related data (job requirement documents provided by the enterprise or job information on recruitment websites), industry-related data (industry trend analysis reports provided by third-party platforms) and course-related data (course outlines, teaching objectives, skill points, etc. provided by educational institutions) to obtain initial text, and pre-process the initial text to obtain knowledge text.
[0099] Text preprocessing: segment text, remove stop words, and perform part-of-speech tagging. Example:
[0100] Input text: `"Requires proficiency in Python, machine learning, and data visualization."`
[0101] Preprocessing results: `["Requires","proficiency","in","Python",",","machine learning",",","data visualization"]`.
[0102] 2.2 Combine the BERT model to perform contextual semantic understanding on the knowledge text, and use NLP technology to extract entities (such as skills, positions, industries) from the knowledge text, and then combine with external knowledge bases (such as LinkedIn Skills, ONET) to map the entities to standardized lexicons (such as "Python" is mapped to "PythonProgramming" in the skills library) to complete the entity standardization operation.
[0103] 2.3 Use dependency parsing to extract the relationships between entities (Dependency Parsing) and classify the relationships (such as "need", "include", "depend"). Example:
[0104] Input text: `"Python is required for this position."`
[0105] Dependencies: `[("Python","required","Skill-Job")]`.
[0106] 2.4 Select a graph database (such as Neo4j), import entities and relationships into the selected graph database to obtain a graph structure, use graph embedding technology (such as Node2Vec, GraphSAGE) to map the nodes and relationships in the graph structure into a vector space, generate node embedding vectors, complete semantic matching operations, and obtain a talent knowledge graph.
[0107] Graph structure:
[0108] Node types: Skill, Job, Course, Industry.
[0109] Edge types: requires, contains, depends_on, belongs_to.
[0110] 2.5 Automatically parse the newly added data of the initial text, that is, when a new position or course is launched, the system automatically parses the text and extracts entities and relationships. According to the parsing results, the talent knowledge graph is updated.
[0111] 3. Anonymize and desensitize the sensitive data in the standard data, and use the federated learning framework (open source federated learning framework, such as TensorFlowFederated, PySyft) to perform joint modeling and obtain a local model, under the premise that educational institutions and enterprises do not share the original data. Including:
[0112] 3.1 Classify the standard data into personal privacy information (such as name, contact information, ID number, etc.), enterprise sensitive data (such as financial information, core technology keywords, etc.) and general data (such as job descriptions, course content, industry trends, etc.). Different protection strategies are adopted for different types of data to ensure privacy protection without affecting data availability.
[0113] 3.2 The personal privacy information is encrypted using an irreversible hash algorithm (such as SHA-256). Even if the data is leaked, the original information cannot be restored. A random salt value is added before the hash to prevent rainbow table attacks and complete the anonymization of the personal privacy information. Example:
[0114] Input: `Name: Zhang San`
[0115] Output: `Hash value: 3a7bd3e2360a3d5a6d8b6f4f1a1e5f6c`
[0116] Salting example:
[0117] Input: `Name: Zhang San + Salt value: random123`
[0118] Output: `Hash value: 9f8b3e2a6d7c4f1b2a3e5f6d8b6f4c1a`
[0119] 3.3 For the enterprise sensitive data, use regular expressions to identify sensitive fields, perform hashing, shielding or fuzzification operations on the sensitive fields, and complete the desensitization processing of the enterprise sensitive data.
[0120] Example of shielding processing:
[0121] Input: `Financial data: Revenue in 2024 is 50 million yuan`
[0122] Output: `Financial data: Revenue in 2024 is 10,000 yuan`
[0123] Fuzzification example:
[0124] Input: `Core technology: deep learning model optimization`
[0125] Output: `Core technology: deep learning`.
[0126] 3.4 Set educational institutions and enterprises as participants. Each participant uses its own data to train the model locally, obtains the local model and local parameters (such as weight matrix), and completes the deployment of local nodes.
[0127] 3.5 Encrypt the local parameters using homomorphic encryption to obtain encrypted parameters, for example:
[0128] Input: `Model parameters: [0.1, 0.2, 0.3]`
[0129] Output: `Encryption parameters: [Enc(0.1), Enc(0.2), Enc(0.3)]`.
[0130] The encryption parameters are received through the central server, and the encryption parameters are aggregated to obtain global parameters, thereby completing the deployment of cloud nodes; Example:
[0131] Input: `Enc(0.1)+Enc(0.2)+Enc(0.3)`
[0132] Output: `Enc(0.6)`.
[0133] 3.6 Using a central server (deployed in the cloud) to send the global parameters to each of the participants, the local model is globally updated. Example:
[0134] Input: `Global parameters: [Enc(0.6)]`
[0135] Output: `Update local model after decryption`.
[0136] 4. If Figure 2As shown, based on the standard data, students and positions are deeply semantically characterized to generate dynamically updated talent portraits and position portraits, and a matching algorithm is preset in combination with multiple target factors and dynamic weights. The talent portraits, position portraits and matching algorithm are then deployed to the local model to obtain a talent matching model.
[0137] 4.1 Generation of Talent Portraits:
[0138] 4.11 Collect student multi-source data and store the student multi-source data in the data warehouse; the student multi-source data includes:
[0139] Student course grades, students' grades and rankings in each course;
[0140] Skill mastery, based on data such as course grades, assessment exam results, and skill certifications;
[0141] Internship experience, internship company, position, internship time, internship performance, etc.;
[0142] Career preferences, students’ career interests, target positions, desired industries, etc.;
[0143] Other student data include interests and hobbies, soft skills evaluation, project experience, etc.
[0144] 4.12 Extract features from the student multi-source data to obtain:
[0145] Student numerical characteristics, such as course grades, internship duration, skill mastery scores, etc.
[0146] Student classification characteristics, including course categories, skill categories, internship company types, career preference categories, etc.;
[0147] Characteristics of student texts, such as course descriptions, internship descriptions, and career goal descriptions.
[0148] The student numerical features are then standardized and normalized, the student classification features are encoded (one-hot encoding or embedded encoding), and the student text features are vectorized to complete the feature conversion.
[0149] 4.13 Using a fully connected layer, the student numerical features and the student classification features are converted into a structured data vector of a fixed dimension, and a pre-trained language model is used to obtain a text data vector of the student text features. The structured data vector and the text data vector are then fused to obtain a multi-dimensional student vector representation.
[0150] 4.14 Match the student's skills with the skill nodes in the talent knowledge graph to obtain the student's skill semantic vector, use the graph neural network to generate a high-dimensional vector of the student's skill semantic vector, and then fuse the high-dimensional vector of the student's skill semantic vector with the multi-dimensional student vector representation to complete the deep semantic characterization of the student and generate a talent portrait.
[0151] 4.2 Generation of job profiles:
[0152] 4.21 Collect multi-source data of positions and store the multi-source data in the data warehouse; the multi-source data of positions includes:
[0153] Job requirements, job description, responsibilities, and required skills;
[0154] Skill requirements, required technical skills, soft skills, certification requirements;
[0155] Industry attributes, industry, industry development trends, and industry standards;
[0156] Corporate culture, corporate values, work environment, team structure, etc.;
[0157] Other job data, including salary range, work location, promotion path, etc.
[0158] 4.22 Extract features from the multi-source data of the positions and obtain:
[0159] Numerical characteristics of the position, salary range, years of experience required, etc.;
[0160] Job classification characteristics, job categories, skill categories, industry categories, etc.;
[0161] Job text characteristics, job description, job description, corporate culture description, etc.
[0162] Then the position numerical features are standardized and normalized, the position classification features are encoded, and the position text features are vectorized to complete the feature conversion.
[0163] 4.23 Using a fully connected layer, convert the position numerical features and the position classification features into a structured data vector of a fixed dimension, use a pre-trained language model to obtain a text data vector of the position text features, and then fuse the structured data vector and the text data vector to obtain a multi-dimensional position vector representation;
[0164] 4.24 Match the skills required for the position with the skill nodes in the talent knowledge graph to obtain the position skill semantic vector, use the graph neural network to generate a high-dimensional vector of the position skill semantic vector, and then fuse the high-dimensional vector of the position skill semantic vector with the multi-dimensional position vector representation to complete the deep semantic characterization of the position and generate a position portrait.
[0165] 4.3 As Figure 3 As shown, the generation of the talent matching model includes:
[0166] 4.31 Map the talent portrait and the job portrait into a unified vector space to generate a talent portrait vector and a job portrait vector.
[0167] 4.32 Using cosine similarity, calculate the cosine value of the angle between the talent portrait vector and the job portrait vector, and using Euclidean distance, calculate the straight-line distance between the talent portrait vector and the job portrait vector to obtain the semantic similarity between the talent portrait vector and the job portrait vector.
[0168] 4.33 Select multiple objective factors and assign dynamic weights to different objective factors, and then combine optimization algorithms and reinforcement learning to build a multi-objective optimization model.
[0169] Multiple target factors, such as: 1) Student personal wishes: including career preferences, expected salary, work location preferences, etc. Quantify student wishes through data sources such as questionnaires and course selection records. 2) Corporate recruitment preferences: including skill requirements, experience requirements, cultural fit, etc. Extract preference parameters based on job descriptions and historical recruitment data. 3) Potential development space: Evaluate the role of the position in promoting students' future career development, such as promotion opportunities, skill improvement, etc. Scores are obtained through analysis of industry development trends and job growth paths. 4) Other factors: such as market demand, job scarcity, etc.
[0170] Dynamic weighting: Identify based on market changes, enterprise urgency, seasonal factors, and other dynamic factors. Use stream processing frameworks (such as Apache Kafka and Apache Flink) to collect and process market and enterprise data in real time to trigger weighting adjustments. Example: The current market demand for "data science" is high, and the weight of related skills is increased to improve the matching score.
[0171] Optimization algorithms, such as linear weighted models (suitable for simple weighted scoring). Multi-objective optimization algorithms, such as Pareto optimization (ensuring balance between multiple objectives).
[0172] Reinforcement learning: Train the agent to dynamically adjust weights based on feedback to optimize matching results.
[0173] Example: Use deep reinforcement learning (such as DQN) to adjust weights based on historical matching results to improve the satisfaction of future matching.
[0174] 4.34 Combining the semantic similarity and the multi-objective optimization model, a matching algorithm for calculating a comprehensive matching score is obtained, and the matching algorithm is deployed in the local model to obtain a talent matching model.
[0175] The expression of the matching algorithm is:
[0176]
[0177] Among them, T is the comprehensive matching score, w i is the dynamic weight, f i is the score of each target factor, and m is the total number of target factors.
[0178] 5. If Figure 3 As shown, the similarity gap between the talent profile and the job profile is continuously monitored, and a course optimization plan and a job optimization plan are dynamically generated according to the similarity gap and the output of the talent matching model. Including:
[0179] 5.1 Use the talent matching model to obtain the target position and the skill requirements of the target position, and compare and analyze and update the skill requirements of the target position with the relevant data of the corresponding position in the talent knowledge graph.
[0180] 5.2 Evaluate and quantify students' current skills, compare the students' current skills with the updated skill requirements of the target positions to obtain skill gaps, and then use machine learning algorithms (such as cluster analysis, classification models) to analyze the skill gaps to recommend relevant courses, obtain course optimization plans, and predict and plan courses based on industry trends and student needs.
[0181] Evaluate and quantify students’ current skills: Evaluate students’ current skills based on their course grades, skill certifications, internship experiences, etc. Use skill nodes and graph embedding technology in the knowledge graph to quantify students’ skill levels.
[0182] 5.3 Monitor changes in job requirements and courses in real time, optimize the target jobs based on the monitoring results and multiple target factors, and obtain job optimization plans.
[0183] 6. Use explainable AI technology to provide visual explanations or keyword descriptions of the output of the talent matching model, the course optimization plan, and the job optimization plan. Including:
[0184] 6.1 Use feature contribution graphs (bar charts, pie charts, etc.) to show the contribution of each feature to the matching results and recommendation results. For example, a bar chart shows "skill matching: 40%, internship experience: 30%, cultural fit: 20%".
[0185] 6.2 Use a decision path diagram, a decision tree or a flowchart to show the decision path of the talent matching model. For example, the decision path diagram shows "skill matching degree > 80% → recommended position A".
[0186] 6.3 Use keyword cloud to display the driving factors for job recommendations. For example, the keyword cloud shows "skill match, job requirements, and cultural fit" as the basis for recommendation.
[0187] 6.4 provides a dashboard for viewing data usage and recommendation process. For example, after the user clicks "Skill Matching", they can view the matching status of specific skills.
[0188] 6.5 And design privacy terms to achieve transparent management. For example:
[0189] Data source description: clearly state the data source used (such as resumes submitted by students, job requirements posted by companies). For example, "The system uses your course grades, internship experience, and career preference information to generate job recommendations."
[0190] Data usage description: Detailed description of the data usage (such as for portrait generation, matching model training). For example, "Your data will be used to generate personalized job recommendations and will not be used for other commercial purposes."
[0191] Privacy clause design: Data access rights control, clearly specifying which users or institutions can access which data. For example, "Only authorized educational institutions and companies can access your public resume information."
[0192] Make the approval process transparent: Set up the approval process on the management side and record the approval records of each data access and recommended operation.
[0193] Therefore, the present invention provides a talent training recommendation method based on federated learning and natural language processing, which can not only improve matching accuracy and efficiency, break data silos, and protect privacy and security, but also dynamically adapt to changes in market demand, meet the personalized needs of multiple parties, and promote the precise connection between education and employment.
[0194] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0195] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A talent training recommendation method based on federated learning and natural language processing, characterized in that: The following steps are involved: Collect relevant data from educational institutions, enterprises and third-party platforms to obtain multi-source heterogeneous data, store the multi-source heterogeneous data in a data lake, clean and integrate the multi-source heterogeneous data to obtain standard data, and store the standard data in a data warehouse; Construct a talent knowledge graph related to talent development and job skills, and perform real-time incremental updates to the talent knowledge graph; Anonymize and desensitize the sensitive data in the standard data, and use the federated learning framework to perform joint modeling to obtain a local model without sharing the original data between educational institutions and enterprises; Based on the standard data, students and positions are deeply semantically characterized to generate dynamically updated talent portraits and position portraits, and a matching algorithm is preset in combination with multiple target factors and dynamic weights, and then the talent portraits, the position portraits and the matching algorithm are deployed to the local model to obtain a talent matching model; Continuously monitor the similarity gap between the talent profile and the job profile, and dynamically generate a course optimization plan and a job optimization plan based on the similarity gap and the output of the talent matching model; Explainable AI technology is used to provide visual explanations or keyword descriptions of the output of the talent matching model, the course optimization plan, and the job optimization plan.
2. The talent training recommendation method based on federated learning and natural language processing according to claim 1 is characterized in that: Collect relevant data from educational institutions, enterprises and third-party platforms to obtain multi-source heterogeneous data, store the multi-source heterogeneous data in a data lake, clean and integrate the multi-source heterogeneous data to obtain standard data, and store the standard data in a data warehouse, including: According to the unified data format and field definition, use API interface or crawler technology to obtain relevant data from educational institutions, enterprises and third-party platforms to obtain multi-source heterogeneous data; A data lake is constructed by using a distributed storage system, and the multi-source heterogeneous data is partitioned and stored in the data lake according to the data source and type; Processing missing values, outliers, duplicate data and data integration on the multi-source heterogeneous data to obtain standard data; A data warehouse is constructed using a relational database, the standard data partitions are stored in the data warehouse according to data subjects, and a query index is created in the data warehouse.
3. The talent training recommendation method based on federated learning and natural language processing according to claim 2 is characterized in that: Construct a talent knowledge graph related to talent development and job skills, and perform real-time incremental updates to the talent knowledge graph, including: Collecting position-related data, industry-related data and course-related data to obtain an initial text, and preprocessing the initial text to obtain a knowledge text; Combine the BERT model to understand the contextual semantics of the knowledge text, use NLP technology to extract entities from the knowledge text, and then map the entities to a standardized vocabulary in combination with an external knowledge base to complete the entity standardization operation; Use dependency parsing to extract relationships between entities and classify the relationships; Select a graph database, import entities and relationships into the selected graph database, obtain a graph structure, use graph embedding technology to map nodes and relationships in the graph structure into a vector space, generate node embedding vectors, complete semantic matching operations, and obtain a talent knowledge graph; The newly added data of the initial text is automatically parsed, and the talent knowledge graph is updated according to the parsing result.
4. The talent training recommendation method based on federated learning and natural language processing according to claim 3 is characterized in that: The sensitive data in the standard data is anonymized and desensitized, and on the premise that educational institutions and enterprises do not share the original data, a federated learning framework is used for joint modeling to obtain a local model, including: Classify the standard data into personal privacy information, enterprise sensitive data and general data; The personal privacy information is encrypted using an irreversible hash algorithm, and a random salt value is added before the hash to prevent rainbow table attacks and complete the anonymization of the personal privacy information; For the enterprise sensitive data, use regularized expressions to identify sensitive fields, perform hashing, shielding or fuzzification operations on the sensitive fields, and complete the desensitization of the enterprise sensitive data; Educational institutions and enterprises are set as participants. Each participant uses its own data to train the model locally, obtains the local model and local parameters, and completes the deployment of the local node. Encrypt the local parameters using homomorphic encryption to obtain encrypted parameters, receive the encrypted parameters through a central server, and perform aggregation operations on the encrypted parameters to obtain global parameters, thereby completing the deployment of cloud nodes; The global parameters are sent to each of the participants using a central server to perform a global update of the local model.
5. The talent training recommendation method based on federated learning and natural language processing according to claim 4 is characterized in that: Based on the standard data, students are deeply semantically profiled to generate dynamically updated talent portraits, including: Collecting multi-source student data and storing the multi-source student data in the data warehouse; the multi-source student data includes student course grades, skill mastery, internship experience, career preferences and other student data; Extracting features from the student multi-source data to obtain student numerical features, student classification features, and student text features, then standardizing and normalizing the student numerical features, encoding the student classification features, and vectorizing the student text features to complete feature conversion; Using a fully connected layer, the student numerical features and the student classification features are converted into a structured data vector of fixed dimension, a pre-trained language model is used to obtain a text data vector of the student text features, and then the structured data vector and the text data vector are fused to obtain a multi-dimensional student vector representation; Match the students' skills with the skill nodes in the talent knowledge graph to obtain the students' skill semantic vector, use the graph neural network to generate a high-dimensional vector of the students' skill semantic vector, and then fuse the high-dimensional vector of the students' skill semantic vector with the multi-dimensional student vector representation to complete the deep semantic characterization of the students and generate a talent portrait.
6. The talent training recommendation method based on federated learning and natural language processing according to claim 5 is characterized in that: Based on the standard data, the positions are deeply semantically characterized to generate dynamically updated position portraits, including: Collect multi-source data of positions, and store the multi-source data of positions in the data warehouse; the multi-source data of positions include position requirements, skill requirements, industry attributes, corporate culture and other data of positions; Extracting features from the multi-source data of the positions to obtain position numerical features, position classification features and position text features, then standardizing and normalizing the position numerical features, encoding the position classification features, and vectorizing the position text features to complete feature conversion; Using a fully connected layer, the position numerical features and the position classification features are converted into a structured data vector of fixed dimension, a pre-trained language model is used to obtain a text data vector of the position text features, and then the structured data vector and the text data vector are fused to obtain a multi-dimensional position vector representation; The skills required for the position are matched with the skill nodes in the talent knowledge graph to obtain the position skill semantic vector, and a high-dimensional vector of the position skill semantic vector is generated using a graph neural network. The high-dimensional vector of the position skill semantic vector is then fused with the multi-dimensional position vector representation to complete a deep semantic characterization of the position and generate a position portrait.
7. The talent training recommendation method based on federated learning and natural language processing according to claim 6 is characterized in that: Combining multiple target factors and dynamic weight preset matching algorithms, and then deploying the talent portrait, the job portrait and the matching algorithm to the local model, a talent matching model is obtained, including: Mapping the talent portrait and the job portrait into a unified vector space to generate a talent portrait vector and a job portrait vector; Using cosine similarity, calculate the cosine value of the angle between the talent portrait vector and the job portrait vector, and using Euclidean distance, calculate the straight-line distance between the talent portrait vector and the job portrait vector to obtain the semantic similarity between the talent portrait vector and the job portrait vector; Select multiple objective factors, assign dynamic weights to different objective factors, and then combine optimization algorithms and reinforcement learning to build a multi-objective optimization model; The semantic similarity and the multi-objective optimization model are combined to obtain a matching algorithm for calculating a comprehensive matching score, and the matching algorithm is deployed in the local model to obtain a talent matching model.
8. The talent training recommendation method based on federated learning and natural language processing according to claim 7 is characterized in that: The expression of the matching algorithm is: Among them, T is the comprehensive matching score, w i is the dynamic weight, f i is the score of each target factor, and m is the total number of target factors.
9. The talent training recommendation method based on federated learning and natural language processing according to claim 8 is characterized in that: Continuously monitor the similarity gap between the talent profile and the job profile, and dynamically generate a course optimization plan and a job optimization plan based on the similarity gap and the output of the talent matching model, including: Utilize the talent matching model to obtain target positions and target position skill requirements, and compare and analyze and update the target position skill requirements with relevant data of corresponding positions in the talent knowledge graph; Evaluate and quantify students' current skills, compare them with the updated skill requirements of the target position, obtain skill gaps, and then use machine learning algorithms to analyze the skill gaps to recommend relevant courses, obtain course optimization plans, and predict and plan courses based on industry trends and student needs; Monitor changes in job requirements and courses in real time, optimize the target jobs based on the monitoring results and multiple target factors, and obtain job optimization plans.
10. The talent training recommendation method based on federated learning and natural language processing according to claim 9 is characterized in that: Use explainable AI technology to provide visual explanations or keyword descriptions of the output of the talent matching model, the course optimization plan, and the job optimization plan, including: Use the feature contribution graph to show the contribution of each feature to the matching results and recommendation results, use the decision path graph to show the decision path of the talent matching model, and use the keyword cloud to show the driving factors of job recommendations; Provide a dashboard for viewing data usage and recommendation process, and design privacy terms to achieve transparent management.
Citation Information
Patent Citations
Occupational recommendation method and device, electronic equipment and storage medium
CN116501959A
Atlas federal learning privacy enhancement method for industrial terminal network flow detection
CN116701618A
Personalized learning recommendation system based on federal learning
CN118036774A
System and method for artificial intelligence based data integration of entities post market consolidation
US20200387529A1
Method for intelligently matching supply and demand in innovation and entrepreneurship services
WO2022252014A1
Cited By
Talent evaluation management method and system based on AI intelligence
CN120338741A
Business candidate person and member information management and storage method and system
CN120372669A
Teaching auxiliary system for improving job hunting ability
CN120374330A
A teaching aid system for job-seeking ability improvement
CN120374330B
Proprietary education fusion management method based on big data analysis
CN120563286A