Automated data management system and method for job skills
An NLP-based automated system analyzes job postings to address skills mismatch, offering precise talent matching and market insights, improving job search and recruitment efficiency.
Patent Information
- Application Number
- PCT/CN2025/072991
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
The challenge of skills supply and demand mismatch in the job market leads to inefficiencies in recruitment processes, with job seekers lacking clarity on required skills and companies facing difficulties in identifying suitable candidates, necessitating a streamlined solution for precise talent matching.
An automated data management system using natural language processing (NLP) to analyze job postings, extracting key skill information through web crawling and semantic analysis, and creating a skill classification scheme for real-time market insights.
Enhances job search efficiency for seekers and recruitment efficiency for companies by providing accurate skill demand data, reducing human resource costs and subjective biases, and facilitating strategic talent matching.
Smart Images

Figure CN2025072991_24072025_PF_FP_ABST
Abstract
Description
Automated data management system and method for job skillsTechnical Field
[0001] The present invention relates to the technical area of data processing and analysis, and particularly relates to an automated data management system and method for job skills.Background
[0002] In today's dynamic job market, the challenge of skills supply and demand mismatch is increasingly evident. Companies are encountering difficulties as they discover that recent graduates and job seekers often lack the necessary skills crucial for their operations, posing a significant obstacle in the recruitment process. Simultaneously, job seekers frequently struggle with a lack of clarity regarding the specific skill sets required for the positions they are pursuing, leading to inefficiencies in their job search endeavours. With a plethora of job listings available on various recruitment platforms, job seekers are required to dedicate substantial time and effort to gather and analyse pertinent information to comprehend the prevailing market demands.
[0003] The enduring issue of talent supply and demand mismatch has long been a prominent concern in the job market. While the ideal of optimizing skill utilization remains a goal, the practical implementation often falls short. Against the backdrop of profound structural shifts in the global economy, millions of job seekers are grappling with the challenge of securing positions that align with their skill sets, further exacerbating the existing mismatch.
[0004] Moreover, companies are increasingly realizing that only a small fraction of graduates or job applicants possess the requisite skills essential for driving business growth and development. Many of the graduates they onboard necessitate additional training to effectively contribute value. Conversely, a significant portion of students lack insight into the skills sought by companies and the skills they need to cultivate for future success. Employers observe that only a minority of graduates or applicants possess the fundamental skills crucial for business operations and expansion, underscoring the importance of supplementary training to enhance their efficacy.
[0005] Therefore, there is a compelling need for a streamlined solution that can apprise job seekers of the specific skills sought by companies in current job openings that align with their individual criteria. Such a tool would significantly streamline the job search process, eliminating the need for individuals to sift through numerous listings and independently synthesize skill requirements. This resource would not only increase the likelihood of finding a job that aligns with their preferences but also empower students by offering a comprehensive understanding of the skills requisite for their desired roles, facilitating strategic planning for skill development.Summary of the Invention
[0006] This patent application proposes a comprehensive solution to the inefficiencies prevalent in the job market by proposing a natural language processing (NLP) -based professional skill analysis system. It automates the analysis of job postings to identify critical skill requirements in demand in the current market. Employing various technologies including web crawling technology, the system gathers data from various recruitment websites and utilizes NLP techniques such as tokenization, word frequency analysis, and semantic analysis to extract key skill information from job descriptions. This approach not only saves job seekers time and increases their chances of securing ideal positions but also provides precise talent matching data for companies, enhancing recruitment efficiency and reducing human resource costs. Furthermore, the system offers data-driven decision support for the recruitment process, aiding companies in formulating targeted recruitment strategies, while also enabling educational institutions to stay abreast of market demands, optimize their course offerings, and cultivate talent tailored to current market needs.
[0007] According to the first aspect of the present application, there is provided an automated data management system for job skills, comprising: a data collection module for gathering raw job position data from public resources; a data preprocessing module for preprocessing the collected raw job position data to remove irrelevant information and / or symbols; a data identification module for identifying text related to job skill requirements from the pre-processed data and identifying keywords related to the job skills; a semantic analysis module for semantically analysing the identified keywords related to the job skills and forming a skill classification scheme; a data storage module for establishing a skill database based on the skill classification scheme and storing the skill classification scheme.
[0008] Preferably, the data preprocessing module is configured to perform at least one of the following operations: cleaning html format tags from the collected raw job position data, tokenizing, lemmatizing, identifying and removing stop words.
[0009] Preferably, the data identification module stores lists of "skill headings" and "skill words" and is configured to perform the following operations: utilizing a first moving window and a second moving window to traverse the pre-processed data from the beginning; evaluating if any word in the first moving window matches any skill heading in the list of skill headings and calculating the density of skill words in the first and second moving windows, where the density is the ratio of skill words to total words; identifying the skill requirement section from the pre-processed data; identifying keywords related to job skills from the identified skill requirement section.
[0010] Preferably, identifying the skill requirement section from the pre-processed data further includes: determining the beginning of the skill requirement section if a skill heading is found in the first moving window, or if the density of skill words in either the first or second moving window exceeds a first threshold; determining the end of the skill requirement section if the density of skill words falls below a second threshold in two consecutive windows.
[0011] Preferably, the first threshold and the second threshold are between 50%and 90%, and the lengths of the first moving window and the second moving window are fixed or adjustable.
[0012] Preferably, the semantic analysis module is configured to: convert keywords into numerical vectors; cluster the keywords using statistical clustering methods; align the clusters with relevant skill sets to create the skill classification scheme.
[0013] Preferably, the automated data management system further comprises one or more of the following modules: an update module for real-time or periodic updates of job data and skill requirement information to update the skill classification scheme; a user interaction module providing a user interface for communicating with users, displaying job skill-related query information to users, showing information related to job skill analysis, and / or collecting feedback on skill requirements; an intelligent matching module for searching the job skills database using string exact matching or fuzzy algorithms based on user query information to find matching job skills; a user customization module for optimizing search and matching results for job skills based on user feedback; a statistical analysis module for calculating the demand percentage for specific job skills; a job skill analysis report module for generating customized job skill analysis reports based on the skill classification scheme as per user request; a data classification module for categorizing collected raw job position data to differentiate different types of recruitment data and feeding the classified information into the data preprocessing module.
[0014] Preferably, the types of recruitment data include at least one of company introductions, job descriptions, benefits, skill requirements, and educational degree requirements.
[0015] According to the second aspect of the present application, there is provided an automated data management method for job skills, comprising the following steps: Step one, collecting raw job position data from public resources; Step two, preprocessing the collected raw job position data to remove irrelevant information and / or symbols; Step three, identifying text related to job skill requirements from the standardized data and identifying keywords related to the job skills; Step four, semantically analysing the identified keywords related to the job skills and forming a skill classification scheme; Step five, establishing a skill database based on the skill classification scheme.
[0016] Preferably, the preprocessing includes at least one of the following operations: cleaning html format tags from the collected raw job position data, tokenizing, lemmatizing, identifying and removing stop words.
[0017] Preferably, Step three further includes: utilizing a first moving window and a second moving window to traverse the pre-processed data from the beginning to find phrases matching the pre-stored lists of "skill headings" and "skill words" ; evaluating if any word in the first moving window matches any skill heading in the list of skill headings and calculating the density of skill words in the first and second moving windows, where the density is the ratio of skill words to total words; identifying the skill requirement section from the pre-processed data; identifying keywords related to job skills from the identified skill requirement section.
[0018] Preferably, identifying the skill requirement section from the pre-processed data further includes: determining the beginning of the skill requirement section if a skill heading is found in the first moving window, or if the density of skill words in either the first or second moving window exceeds a first threshold; determining the end of the skill requirement section if the density of skill words falls below a second threshold in two consecutive windows.
[0019] Preferably, the first threshold and the second threshold are between 50%and 90%, and the lengths of the first moving window and the second moving window are fixed or adjustable.
[0020] Preferably, Step four further includes: converting keywords into numerical vectors;clustering the keywords using statistical clustering methods; aligning the clusters with relevant skill sets to create the skill classification scheme.
[0021] Preferably, the automated data management method further comprises one or more of the following operations: real-time or periodic updates of job data and skill requirement information to update the skill classification scheme; retrieving job skill-related query information from users, displaying job skill analysis-related information to users, and / or collecting user feedback on skill requirements; combining user query information and searching the job skills database using string exact matching or fuzzy algorithms to find job skills matching user queries; obtaining user feedback to optimize search and matching results for job skills; classifying collected recruitment job data to differentiate different types of recruitment data; calculating the demand percentage for specific job skills;generating customized job skill analysis reports based on the skill classification scheme as per user request; categorizing collected raw job position data to differentiate different types of recruitment data and feeding the classified information into the data preprocessing module.
[0022] Preferably, the types of recruitment data include at least one of company introductions, job descriptions, benefits, skill requirements, and educational degree requirements.
[0023] According to the third aspect of the present application, there is provided An automated data management system comprising: a data storage resource configured to store data and instructions; and a processing circuitry configured to implement the method as defined above.
[0024] According to the fourth aspect of the present application, there is provided A computer-implemented automated data management method for job skills, the method being executed by one or more processors and comprising: collecting, at a data collection unit, raw job position data from public resources; preprocessing, at a data preprocessing unit, the collected raw job position data to remove irrelevant information and / or symbols; identifying, at a data identification unit, text related to job skill requirements from the standardized data and identifying keywords related to the job skills; semantically analysing, at a semantic analysis unit, the identified keywords related to the job skills and forming a skill classification scheme; establishing, at data storage unit, a skill database based on the skill classification scheme.
[0025] According to the fifth aspect of the present application, there is provided An automated data management system, comprising: one or more processors; and a computer-readable storage device coupled to the one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for automating data management, the operations comprising the steps as defined above.
[0026] According to the sixth aspect of the present application, there is provided A non-transitory computer readable medium storing a plurality of instructions that when executed control a computer including one or more processors to implement the methods as defined above.
[0027] This invention offers several key advantages. Firstly, it emphasizes automation and efficiency by leveraging NLP technology for automated job data analysis, enabling the system to exhibit high processing efficiency and effectively manage substantial amounts of job-related information. Additionally, the system provides real-time updates on market skill demands, ensuring that users have access to the most up-to-date data. Furthermore, the implementation of data-driven analytical methods enables objective analysis, reducing subjective biases and enhancing the accuracy of skill demand analysis. Lastly, the system's capability to simultaneously consider multiple job postings allows it to identify industry skill trends and demand shifts, ultimately facilitating improved matching between job seekers and companies.
[0028] Brief Description of The Figures
[0029] The accompanying figures, which are incorporated in and constitute a part of this specification, illustrate one or more embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0030] Figure 1 shows an automated data management system according to the first embodiment of the present application.
[0031] Figure 2 shows an automated data management method implemented by the automated data management system of Figure 1.
[0032] Figure 3 further illustrates an automated data management system for job skills based on another embodiment of the present application.
[0033] Figures 4A-4D display the results generated by an embodiment of the present application in relation to the percentage for generic skills.
[0034] Figures 5A-5C display other results generated by an embodiment of the present application the percentage for technical skills.
[0035] Figure 6 demonstrate an example a user interface of a real time job skill analyser employing the automated system and method of the present application.
[0036] Detailed Description of The Embodiments
[0037] Detailed reference is now made to the embodiments of the present application, with the figures illustrating one or more embodiments. The repeated use of figure labels in this specification and the figures is to indicate similar features or elements of the present application. It is important to note that the following content is provided to assist further understanding of the present application for those skilled in the art but does not limit the application in any form. It should be noted that various modifications and changes can be made by those skilled in the art without departing from the concept of the present application. For example, features shown or described as part of one embodiment may be used in conjunction with another embodiment to generate further implementations. Consequently, the present application is intended to encompass such variations and changes within the scope of the appended claims and their equivalents.
[0038] Figure 1 shows an automated data management system according to the first embodiment of the present application. As shown, the system comprises job position data collection module 110, job position data preprocessing module 120, job text data identification module 130 and text data semantic analysis module 140 as well as a data storage module 150. The job position data collection module 110 is responsible for collecting relevant data on publicly advertised job positions using internet and database technologies from public resources 160, encompassing both text and numerical data. The job position data preprocessing module 120 works on standardizing the various types of job position data for enhanced analysis by subsequent modules. The job text data identification module 130 employs natural language processing methods to tokenize and categorize relevant text data, as well as identify key terms. The text data semantic analysis module 140 conducts semantic analysis on the tokenized data to extract specific skills required for job positions, thereby generating a skill classification scheme. The data storage module 150 establishes a skill database based on the skill classification scheme and stores the skill classification scheme. The skill classification scheme is accessible to users 170.
[0039] Figure 2 shows an automated data management method implemented by the automated data management system of Figure 1. As shown in Figure 2, the method comprising the following steps: Step 210, collecting raw job position data from public resources; Step 220, preprocessing the collected raw job position data to remove irrelevant information and / or symbols; Step 230, identifying text related to job skill requirements from the standardized data and identifying keywords related to the job skills; Step 240, semantically analysing the identified keywords related to the job skills and forming a skill classification scheme; Step 250, establishing a skill database based on the skill classification scheme.
[0040] The automated data management system and method as shown in Figures 1 and 2 are detailed below.
[0041] Job Position Data Collection Module 110
[0042] Step 210 is implemented by job position data collection module 110. This module is designed to gather pertinent data from publicly advertised job listings, encompassing both textual and numerical information, by leveraging internet and database technologies. It predominantly functions through a data collection module, which employs a blend of manual collection techniques and web crawler methods to amass data.
[0043] The system's data collection functionality facilitates the transformation of information into the requisite data format via mechanisms like manual input, web crawling, and interfaces for human-machine interaction. Users can configure data collection frequencies and schedules to ensure prompt and convenient data retrieval.
[0044] The web crawler feature utilizes internet crawling languages and computer tools to access openly available data, significantly boosting data update speed and timeliness. Moreover, self-check and confirmation procedures enhance data accuracy while decreasing data acquisition expenses.
[0045] In unique scenarios, such as with handwritten job postings, manual collection approaches are adopted. Manual data collection is employed for sources that are challenging to pinpoint accurately and necessitate on-site investigation, phone inquiries, or other forms of manual intervention. Throughout the collection process, data is acquired and updated using scanning and image recognition technologies, manual data entry, or storage tools like Excel. Excel is utilized for support and direct data entry, enabling background manual data updates.
[0046] Job Position Data Preprocessing Module 120
[0047] Step 220 is implemented by job position data preprocessing module 120. In preparation for skill identification within a job advertisement, a requisite preliminary stage involves the meticulous cleaning and pre-processing of the advertisement text. This preparatory phase conventionally entails the removal of HTML format tags, tokenization, lemmatization, and the identification and subsequent elimination of stopwords. The identification of common English stopwords is facilitated through reference to the NLTK (Natural Language ToolKit) 's established list.
[0048] Data preprocessing procedures encompass a range of activities, including preprocessing tasks, the rectification of missing data, the correction of formatting errors, the resolution of logical inconsistencies, the deduplication of information, the elimination of extraneous data, and the verification of relevance. Notably, during the processing of job postings, the extraction of special characters followed by their archival in an array is conducted to prepare the data for subsequent analytical phases. These rigorous procedures ensure that the compiled job data is refined to yield the desired informational output.
[0049] The process of data cleaning entails a thorough review and validation of data to expunge duplicates, rectify errors, and enforce data consistency. It embodies the act of purging data impurities, symbolizing the conclusive phase of identifying and rectifying perceptible errors within the dataset. Given that databases amalgamate data from diverse origins, including historical records, encountering erroneous or conflicting data, colloquially termed "dirty data, " is an inevitable occurrence. The overarching objective is to purify this "dirty data" in accordance with predefined criteria, retaining solely the essential information. Data deemed unsuitable, characterized by incomplete, erroneous, or replicated entries, is systematically sieved, permitting refined data to advance to subsequent processing stages.
[0050] In a specific implementation scenario, the data preprocessing process is orchestrated through a series of grouped steps aimed at acquiring the target data. This methodology encompasses the delineation of cleaning conditions for each data module group, the application of these conditions to purify data within the respective groups, culminating in the extraction of the target data module group, thereby furnishing the final dataset based on refined data.
[0051] In the realm of talent acquisition for job vacancies, prospective employers frequently delineate detailed skill prerequisites within their online job postings. Presented herein are three illustrative instances of online job advertisements, typically partitioned into distinct sections.
[0052] Example 1
[0053] Index Equity Portfolio Manager
[0054] Full Time
[0055] Middle Management
[0056] 7 years exp
[0057] Banking and Finance
[0058] $15,000to$19,000 Monthly
[0059] 24 applications Posted 24 Nov 2022 Closing on 24 Dec 2022
[0060] Roles &Responsibilities
[0061] ETFs and Index Investments -Portfolio Management
[0062] Your Team
[0063] Are you ready to join the world's largest asset manager and lead the way in index investing?
[0064] BlackRock's ETFs and Index Investments team (EII) safeguard over $4 trillion in global index equity assets across Global Developed Markets, Emerging Markets, Commodities and Real Estate Investment Trusts. We offer investors one of the industry's biggest choices of index investments.
[0065] Our investment teams include professionals who have deep experience within their respective markets, looking after the assets of institutional and individual investors alike. Our clients benefit from our portfolio management expertise and our culture of performance and risk management.
[0066] Our purpose is to help more and more people achieve greater financial wellbeing and it's something we feel passionately about!
[0067] Your Role and Impact
[0068] We have a superb opportunity for a Portfolio Manager to join our ETF and Index Investments team. You'll be responsible for all aspects of index equity portfolio management, including the day to day portfolio management of our index portfolios, risk and performance analysis and business process reengineering. We're looking for individuals who are passionate about making a difference and have the know-how to make it happen -both as Students of the Market as well as Students of Technology!
[0069] Your Responsibilities
[0070] ·Perform daily portfolio management tasks; daily liquidity management, portfolio re-balancing, corporate action analysis, client activity; risk and performance monitoring
[0071] ·Performance and risk management
[0072] ·Work with our Global Trading teams to navigate rebalances and other large-scale investment events
[0073] ·Help establish portfolio management best practices that can be shared globally
[0074] ·Build our next generation investment platform
[0075] ·Identify and drive operational improvements that lower risk and increase efficiency across the global teams
[0076] ·Engage with business partners to provide thought leadership and advice to promote high quality client solutions
[0077] ·Contributing positively to our culture and supporting diversity, equity and inclusion You Have
[0078] We're looking for individuals, from a variety of backgrounds, who have:
[0079] ·A passion for financial markets and technology
[0080] ·Knowledge of both equities and the indexing ecosystem
[0081] ·7+ years of investment industry / portfolio management experience
[0082] ·Strong technical aptitude with an interest in technology solutions related to portfolio management, trading and data analytics. Python, SQL, or any coding experience preferred
[0083] ·Strong analytical and problem solving skills
[0084] ·High levels of self-motivation and a strong work-ethic
[0085] ·High attention to detail and accountability, organizational and project management skills
[0086] ·The ability to work quickly and accurately in a fast-paced environment
[0087] ·A desire to work collaboratively and great relationship-building skills
[0088] ·Excellent written and verbal communication skills
[0089] In the first example, information is organized under distinct headings such as "Roles &Responsibilities, " "Your Team, " "Your Role and Impact, " "Your Responsibilities, " and "You Have, " with the pertinent data for skill analysis typically nestled within the "You Have" section.
[0090] Results of the preprocessing is as below:
[0091] “index equity portfolio manager full time middle management years exp banking finance monthly applicationsposted nov dec roles responsibilities etfs index investments portfolio management team ready join world largest asset manager lead way index investing blackrock etfs index investments team eii safeguard global index equity assets global developed markets emerging markets commodities real estate investment trusts offer investors industry biggest choices index investments investment teams include professionals deep experience respective markets looking assets institutional individual investors alike clients benefit portfolio management expertise culture performance risk management purpose help people achieve greater financial wellbeing feel passionately role impact superb opportunity portfolio manager join etf index investments team responsible aspects index equity portfolio management including day day portfolio management index portfolios risk performance analysis business process reengineering looking individuals passionate making difference know make happen students market well students technology responsibilities perform daily portfolio management tasks daily liquidity management portfolio balancing corporate action analysis client activity risk performance monitoring performance risk management work global trading teams navigate rebalances large scale investment events help establish portfolio management best practices shared globally build next generation investment platform identify drive operational improvements lower risk increase efficiency global teams engage business partners provide thought leadership advice promote high quality client solutions contributing positively culture supporting diversity equity inclusion looking individuals variety backgrounds passion financial markets technology knowledge equities indexing ecosystem years investment industry portfolio management experience strong technical aptitude interest technology solutions related portfolio management trading data analytics python sql coding experience preferred strong analytical problem solving skills high levels self motivation strong work ethic high attention detail accountability organizational project management skills ability work quickly accurately fast paced environment desire work collaboratively great relationship building skills excellent written verbal communication skills”
[0092] Example 2
[0093] This position will assist the controllers in carrying out all finance and accounting functions.
[0094] DUTIES AND RESPONSIBILITIES
[0095] Internal / external reporting
[0096] Cash management
[0097] Regulatory and Tax matters
[0098] Assist with system implementation
[0099] QUALIFICATIONS
[0100] Job Specifications:
[0101] Ability to multi-task and work under pressure
[0102] Strong PC skills
[0103] Good command of written and spoken English
[0104] Education:
[0105] Qualified accountant with Bachelor degree in finance or related discipline
[0106] Experience:
[0107] A minimum of 5 years audit or accounting experience
[0108] General insurance accounting and audit background preferred
[0109] Interested parties please email full resume (with availability, current / last &expected salary stated) by clicking the below button "Apply Now" .
[0110] All personal data collected will be treated in strict confidence and used for recruitment purpose only.
[0111] The second job advertisement organizes information in two sections under the headings “DUTIES AND RESPONSIBLITLIES” and “QUALIFICATIONS” . The required skills are specified in the section “QUALIFICATIONS” .
[0112] Results of the preprocessing is as below:
[0113] “position assist controllers carrying finance accounting functions duties responsibilities internal external reporting cash management regulatory tax matters assist system implementation qualifications job specifications ability multi task work pressure strong pc skills good command written spoken english education qualified accountant bachelor degree finance related discipline experience minimum years audit accounting experience general insurance accounting audit background preferred interested parties please email full resume availability current last expected salary stated clicking button apply personal data collected treated strict confidence used recruitment purpose”
[0114] Example 3.
[0115] Skills & Requirements:
[0116]
[0117] Excellent command of written and spoken English and Chinese with native / near-native proficiency;
[0118] <span style=" "color: rgb (51, 51, 51) ; background-color: rgb (255, 255, 255) ; font-size:
[0119] 16px; font-family: Helvetica Neue" ", Helvetica, Calibri, Arial, sans-serif; " ">Ability to translate complicated technical information into an easily understood format
[0120] <span style=" "color: rgb (51, 51, 51) ; background-color: rgb (255, 255, 255) ; font-size:
[0121] 16px; font-family: Helvetica Neue" ", Helvetica, Calibri, Arial, sans-serif; " ">Sound planning and organisational skills
[0122] <span style=" "color: rgb (51, 51, 51) ; background-color: rgb (255, 255, 255) ; font-size:
[0123] 16px; font-family: Helvetica Neue" ", Helvetica, Calibri, Arial, sans-serif; " ">Presentation skills
[0124] Entry-level advertising, marketing, copywriting
[0125] Market trends research and reporting
[0126] <span style=" "color: rgb (0, 0, 0) ; background-color: rgb (255, 255, 255) ; font-size:
[0127] medium; font-family: Roboto, Helvetica, Arial, sans-serif; " ">HKID Holder
[0128]
[0129]
[0130] Nice to have:
[0131]
[0132] Knowledge and experience in cryptocurrency trading, blockchain / DLT;
[0133] <span style=" "color: rgb (0, 0, 0) ; background-color: rgb (255, 255, 255) ; font-size:
[0134] medium; font-family: Roboto, Helvetica, Arial, sans-serif; " ">Knowledge in the areas of risk management in financial services industries or fintech environment;
[0135] <span style=" "color: rgb (0, 0, 0) ; background-color: rgb (255, 255, 255) ; font-size:
[0136] medium; font-family: Roboto, Helvetica, Arial, sans-serif; " ">Broad knowledge of the financial services industries, familiarity with KYC and / or AML / CFT regulatory requirements
[0137]
[0138]
[0139] The third job advertisement contains the text for job skill requirements in the HTML format. The HTML tags must be cleaned before we identify any required skills.
[0140] Results of the preprocessing is as below:
[0141] “skills requirements excellent command written spoken english chinese native native proficiency ability translate complicated technical information easily understood format sound planning organisational skills presentation skills entry level advertising marketing copywriting market trends research reporting hkid holder nice knowledge experience cryptocurrency trading blockchain dlt knowledge areas risk management financial services industries fintech environment broad knowledge financial services industries familiarity kyc aml cft regulatory requirements”
[0142] Job Text Data Identification Module 130
[0143] Step 230 is implemented by job text data identification module 130. The provided examples of online job advertisements underscore the common practice among advertising entities to utilize textual content for delineating company background, work environment, job scope, roles, and responsibilities. Such textual information, though extensive, is deemed inconsequential for job skill analysis. The incorporation of this irrelevant text in analytical processes employing statistical models, like frequency analysis of keywords or semantic embeddings language models, can potentially introduce biases and inaccuracies into the evaluation of job skills stipulated in the job postings.
[0144] A pivotal advancement offered by the present application lies in the recognition of the criticality of segregating text specifying job skill requirements within online job advertisements from extraneous textual content. To achieve this, the present application have devised an algorithm and operationalized it leveraging natural language processing (NLP) methodologies.
[0145] The proposed algorithm hinges on two essential text features utilized in delineating skill prerequisites within online job postings. Firstly, such text typically forms a discrete section within the advertisement body, often denoted by section headings like "Skill requirements, " "Qualifications, " "You have, " and the like. Secondly, these skill requirements are commonly expressed not through complete sentences but rather by keywords or phrases. These characteristics aid advertising entities in lucidly articulating job skill prerequisites and facilitate the screening of potential candidates for the advertised positions. The proposed algorithm capitalizes on these characteristics to effectively segregate skill requirements text from extraneous content.
[0146] The algorithm's execution unfolds in three sequential steps. Initially, job text data identification module 130 compile lists of " skill headings" and "skill words (including phrases) " commonly employed to specify skill requirements in online job advertisements. These lists are constructed through an exhaustive review of numerous online job postings to collate the prevalent headings and terminologies denoting requisite skills.
[0147] Subsequently, a "moving window" traverses the advertisement body from its inception. job text data identification module 130 assess whether any word in the moving window aligns with the skill headings list and tabulate the count of skill words contained therein. The density of skill words within the window, computed as the ratio of skill words to total words, aids in identifying the commencement of the skill requirement section. The length of the moving window, either fixed or adaptable based on the advertisement context, is pivotal in this determination. The algorithm employs a threshold, typically ranging between 50%and 90%, to ascertain the presence of the skill requirement section based on skill word density.
[0148] Continuing, the moving window progresses through the entire advertisement. If the density of skill words falls below a stipulated threshold across two consecutive windows, the algorithm identifies the conclusion of the skill requirement section. This threshold, also adjustable within the 50%to 90%range, may differ from the threshold applied in the preceding step.
[0149] The proposed algorithm has demonstrated remarkable efficacy in segregating skill requirements text from extraneous content. In a dataset comprising 33,613 online job advertisements, the algorithm autonomously identified job skill requirements in 28,128 cases. Of the remaining 5,485 advertisements, manual review uncovered skill requirements in 2,567 instances, confirming the absence of such requirements in the remaining 2,918 ads. This attests to the algorithm's success rate of 91.6% (=28,128 / (33,613-2,918) ) . In an independent sample of 8,661 online job ads, the algorithm successfully discerned job skill requirements in 8,344 cases, yielding a success rate of 96.4% (=8,344 / 8,661) .
[0150] Text Data Semantic Analysis Module 140
[0151] Step 240 is implemented by text data semantic analysis module 140. Within this module, a semantic embeddings language model is utilized to convert keywords into numerical vectors, enabling their representation in a mathematical format. Subsequently, statistical clustering methods are applied to categorize these keywords into clusters, followed by the alignment of these clusters with pertinent skill sets within the relevant knowledge domain. The outcome is the creation of a skill classification scheme that is user-friendly and conducive to clear and effective communication.
[0152] Initially, the semantic embeddings language model is employed to transmute keywords into numeric vectors, encapsulating the semantic essence of these keywords in numerical form. These vectors serve to encapsulate the semantic relationships between the keywords, with proximate vectors signifying akin meanings. This process relies on pre-trained language models such as BERT, Word2Vec, or GloVe, which have gleaned semantic associations from vast text corpora.
[0153] Subsequently, statistical clustering techniques are employed to group these numeric representations of keywords into clusters based on their similarities. This method aggregates related keywords within these clusters, aiding in the identification of semantic groupings.
[0154] Following this, the clusters of keywords are mapped to relevant skill sets within the specific domain under analysis. For instance, a cluster featuring keywords like "Python, " "Data Analysis, " and "Machine Learning" could be linked to the skill set "Data Science. "
[0155] The culmination of this process is the establishment of a structured skill classification system. This system is designed for easy comprehension, fostering effective communication of identified skills. By organizing keywords into clusters and aligning them with pertinent skill sets, this classification scheme furnishes a coherent overview of requisite skills within the domain under scrutiny.
[0156] In essence, this automated approach entails the conversion of keywords into skill sets. It encapsulates keyword meaning through semantic embeddings, employs clustering methodologies to group semantically akin keywords, and subsequently associates these groups with specific skill sets, thereby facilitating the automatic classification of keywords into skills. This method's advantage lies in its capability to handle a substantial number of keywords, automatically identify potential skill relationships, enhance efficiency, and reduce manual annotation burdens.
[0157] One illustrative embodiment of the present application employs terms "generic / soft skills" and "technical / hard skills" . The term "generic / soft skills" encompasses attributes that aid individuals in assimilating into a work environment, including aspects related to personality, characteristics, adaptability, motivation, goals, and preferences (Heckman &Kautz, 2012) .
[0158] Building upon the work of de Freitas and Almendra (2021) , the present application further refines the classification of generic skills into social and cognitive skills. As emphasized by de Freitas and Almendra (2021) , gateway skills play a pivotal role in fostering the development of high-order skills. In contrast, "technical / hard skills" refer to proficiencies that equip individuals to excel in specific tasks, such as scientific knowledge, professional expertise, and technical proficiency (Laker &Powell, 2011) .
[0159] To illustrate, an exemplar of the skill classification scheme tailored for the accounting and finance profession is presented across the subsequent two tables. Table 1 catalogues 26 generic skills commonly referenced, while Table 2 delineates 23 technical skills typically sought by companies for roles within the accounting and finance domain. Each skill in these tables is linked with a corresponding set of keywords. In practice, for each job advertisement, the text detailing job skill prerequisites as outlined in Step 4 is extracted, and subsequently, the identified keywords are leveraged to ascertain the skills mandated by the advertising entity.
[0160] Table 1. Most frequent five keywords for each generic skill
[0161] Table 2. Most frequent five keywords for each technical skill
[0162] The embodiment depicted in Figure 3 further illustrates an automated data management system 300 for job skills based on another embodiment of the present application. This embodiment includes additional modules beyond those shown in Figure 1. The same modules as in Figure 1 are identified using the corresponding figure numbers, with their functions and operations mirroring those described in Figure 1.
[0163] As shown in Figure 3, the automated data management system 300 includes: update module 310, user interaction module 320, intelligent matching module 330, user customization module 340, statistical analysis module 350, and job skill analysis report module 360 as well as data classification module 370. Note, it is possible that the system 300 may include any additional modules to meet users’ requirements. It is also possible to remove any of these modules to reduce costs.
[0164] Update module 310 is responsible for facilitating real-time or periodic updates of job data and skill requirement information to ensure the accuracy and relevance of the skill classification scheme as per the evolving job market trends.
[0165] User interaction module 320 serves as a user interface for engaging with users. It displays job skill-related query information, offers insights into job skill analysis, and gathers feedback on skill requirements to enhance user experience and intelligent matching module system efficiency.
[0166] Intelligent matching module 330 utilizes string exact matching or fuzzy algorithms. This module searches the job skills database based on user query information to identify matching job skills, streamlining the job matching process for users.
[0167] User customization module 340 enables users to optimize search and matching results for job skills by incorporating user feedback, ensuring personalized and tailored job skill recommendations based on individual preferences and requirements.
[0168] Statistical analysis module 350 is dedicated to calculating the demand percentage for specific job skills. This module provides valuable insights into the job market trends and the popularity of various skill sets, aiding in strategic decision-making.
[0169] Job skill analysis report module 360 generates customized job skill analysis reports based on the skill classification scheme, catering to user requests and providing detailed insights into job skill trends and requirements for informed decision-making.
[0170] Data classification module 370 is responsible for categorizing collected raw job position data to differentiate various types of recruitment data, ensuring organized and structured data for further processing and analysis within the system.
[0171] Figures 4A-4D and 5A-5C display the results generated by the system of Figure 3 using at least one of the listed modules like statistical analysis module 350 and / or job skill analysis report module 360.
[0172] The job postings were collected from a public recruiter resource website under four job functions -Accounting, Banking &Finance, E-commerce, and Insurance, throughout four months -February, May, August, and November 2023. The number of job postings is presented in Table 3.
[0173] Table 3
[0174] The skills required are extracted using the automated method, and the keyword-skill mapping illustrated in Tables 1 and 2 is utilized to identify the generic skills and technical skills expected by each advertising company in job candidates. For each generic / technical skill, the number of job postings requiring the skill is counted, and the percentage of such job postings among all job postings in the same month, aggregated over all four job functions, is calculated.
[0175] Step 250 is implemented by a data storage module 150. The module 150 uses the skill classification scheme to build a skill database. The skill classification scheme is stored therein and can be accessed when required.
[0176] Figures 4A-4D demonstrate the percentage for all generic skills. According to Table 3, the generic skills are categorized as "Social0" , "Social" , "Cognitive0" , and "Cognitive" . The "Social0" category encompasses gateway social skills, the "Social" category includes high-order social skills, the "Cognitive0" category comprises gateway cognitive skills, and the "Cognitive" category involves high-order cognitive skills. The horizon index in each panel of Figures 4A-4D denotes the months, with "_8" for August, "_5" for May, "_2" for February, and "11" for November. Figures 5A-5C present the percentage for all technical skills, grouped into "Business" , "Technique" , and "Other" . The horizon index in each panel of Figures 5A-5C indicate the respective months.
[0177] The analysis was successfully applied to analyse nearly 50, 000 job postings collected from the public recruiter resource website. The generic and technical skills demanded in individual job postings were identified. The demand for each skill was assessed by calculating the percentage of job postings requiring the skill among all job postings. The figures indicate significant variations in skill demand across the Hong Kong job market, with relatively stable demand observed throughout the four months in 2023. The analysis and findings documented in the report provide supportive evidence that the natural language processing (NLP) based job skill analysis system yields valuable job skill information.
[0178] Figure 6 further demonstrate an example a user interface of a real time job skill analyser employing the automated system and method as described above. As shown, the user could easily find the required skills upon entering the desired job title or description,
[0179] In summary, this application proposes a professional skill analysis and management system and a method based on natural language processing. The system automates the analysis of job postings to identify the most in-demand skills in the current market. Through big data analysis, this invention accurately assesses the demand for various skills, providing valuable data support for job seekers and companies. Job seekers can use this system to understand the skills in demand in the market, thereby enhancing their competitiveness and increasing job opportunities. Companies can leverage the system to quickly identify capable candidates, thereby improving recruitment efficiency and reducing human resource costs. The system offers data support for the recruitment process, assisting companies in formulating more precise recruitment strategies. Additionally, the system enables educational institutions to gain timely insights into market demands, facilitating the optimization of their curriculum to nurture talent that meets market needs.
[0180] Generally, systems 100 and 300 both comprise a processor and memory. The processor may have one or more processing cores, such as a 4-core processor or an 8-core processor. It can be implemented using hardware like a digital signal processor (DSP) , a field-programmable gate array (FPGA) , or a programmable logic array (PLA) . The processor may also include a main processor and a coprocessor. The main processor, also known as the central processing unit (CPU) , handles data in a wake-up state, while the coprocessor is a low-power processor for standby state data. In some cases, the processor may be integrated with a graphics processing unit (GPU) to render and draw content for display screens. Additionally, the processor may have an artificial intelligence (AI) processor for machine learning tasks.
[0181] The memory may consist of one or more non-transitory computer-readable storage media. It can include high-speed random access memory and non-volatile memory like disk storage devices and flash storage devices. In some cases, the memory stores instructions that the processor executes to implement the method described in this application.
[0182] Therefore, the embodiments of this application provide a system with a processor and memory for storing instructions executed by the processor to perform the method shown in Figure 2. The application also includes a computer-readable storage medium that stores a program which, when executed by a processor, implements the method shown in Figure 2. Additionally, a computer program product containing an instruction enables a computer to perform the method shown in Figure 2 when run on a computer.
[0183] A person skilled in the art would understand that all or part of the blocks for implementing the described embodiments can be completed by hardware or by a program instructing related hardware. This program can be stored in a computer-readable storage medium, such as read-only memory, a magnetic disk, or an optical disk. The described embodiments are optional and not meant to limit this application. Any modifications, equivalent replacements, improvements, etc., made within the spirit and principle of this application are included within its scope of protection.
Claims
1.An automated data management system for job skills, comprising:a data collection module for gathering raw job position data from public resources;a data preprocessing module for preprocessing the collected raw job position data to remove irrelevant information and / or symbols;a data identification module for identifying text related to job skill requirements from the pre-processed data and identifying keywords related to the job skills;a semantic analysis module for semantically analysing the identified keywords related to the job skills and forming a skill classification scheme;a data storage module for establishing a skill database based on the skill classification scheme and storing the skill classification scheme.2.The automated data management system of claim 1, wherein the data preprocessing module is configured to perform at least one of the following operations: cleaning html format tags from the collected raw job position data, tokenizing, lemmatizing, identifying and removing stop words.3.The automated data management system of claim 1, wherein the data identification module stores lists of skill headings and skill words and is configured to perform the following operations:utilizing a first moving window and a second moving window to traverse the pre-processed data from the beginning;evaluating if any word in the first moving window matches any skill heading in the list of skill headings and calculating the density of skill words in the first and second moving windows, where the density is the ratio of skill words to total words;identifying the skill requirement section from the pre-processed data;identifying keywords related to job skills from the identified skill requirement section.4.The automated data management system of claim 3, wherein identifying the skill requirement section from the pre-processed data further includes:determining the beginning of the skill requirement section if a skill heading is found in the first moving window, or if the density of skill words in either the first or second moving window exceeds a first threshold;determining the end of the skill requirement section if the density of skill words falls below a second threshold in two consecutive windows.5.The automated data management system of claim 4, wherein the first threshold and the second threshold are between 50%and 90%, and the lengths of the first moving window and the second moving window are fixed or adjustable.6.The automated data management system of claim 1, wherein the semantic analysis module is configured to:convert keywords into numerical vectors;cluster the keywords using statistical clustering methods;align the clusters with relevant skill sets to create the skill classification scheme.7.The automated data management system of claim 1, further comprising one or more of the following modules:an update module for real-time or periodic updates of job data and skill requirement information to update the skill classification scheme;a user interaction module providing a user interface for communicating with users, displaying job skill-related query information to users, showing information related to job skill analysis, and / or collecting feedback on skill requirements;an intelligent matching module for searching the job skills database using string exact matching or fuzzy algorithms based on user query information to find matching job skills;a user customization module for optimizing search and matching results for job skills based on user feedback;a statistical analysis module for calculating the demand percentage for specific job skills;a job skill analysis report module for generating customized job skill analysis reports based on the skill classification scheme as per user request;a data classification module for categorizing collected raw job position data to differentiate different types of recruitment data and feeding the classified information into the data preprocessing module.8.The automated data management system of claim 7, wherein the types of recruitment data include at least one of company introductions, job descriptions, benefits, skill requirements, and educational degree requirements.9.An automated data management method for job skills, comprising the following steps:Step one, collecting raw job position data from public resources;Step two, preprocessing the collected raw job position data to remove irrelevant information and / or symbols;Step three, identifying text related to job skill requirements from the standardized data and identifying keywords related to the job skills;Step four, semantically analysing the identified keywords related to the job skills and forming a skill classification scheme;Step five, establishing a skill database based on the skill classification scheme.10.The automated data management method of claim 9, wherein the preprocessing includes at least one of the following operations: cleaning html format tags from the collected raw job position data, tokenizing, lemmatizing, identifying and removing stop words.11.The automated data management method of claim 9, wherein Step three further includes:utilizing a first moving window and a second moving window to traverse the pre-processed data from the beginning to find phrases matching the pre-stored lists of "skill headings" and "skill words" ;evaluating if any word in the first moving window matches any skill heading in the list of skill headings and calculating the density of skill words in the first and second moving windows, where the density is the ratio of skill words to total words;identifying the skill requirement section from the pre-processed data;identifying keywords related to job skills from the identified skill requirement section.12.The automated data management method of claim 11, wherein identifying the skill requirement section from the pre-processed data further includes:determining the beginning of the skill requirement section if a skill heading is found in the first moving window, or if the density of skill words in either the first or second moving window exceeds a first threshold;determining the end of the skill requirement section if the density of skill words falls below a second threshold in two consecutive windows.13.The automated data management method of claim 12, wherein the first threshold and the second threshold are between 50%and 90%, and the lengths of the first moving window and the second moving window are fixed or adjustable.14.The automated data management method of claim 11, wherein Step four further includes:converting keywords into numerical vectors;clustering the keywords using statistical clustering methods;aligning the clusters with relevant skill sets to create the skill classification scheme.15.The automated data management method of claim 9, further comprising one or more of the following operations:real-time or periodic updates of job data and skill requirement information to update the skill classification scheme;retrieving job skill-related query information from users, displaying job skill analysis-related information to users, and / or collecting user feedback on skill requirements;combining user query information and searching the job skills database using string exact matching or fuzzy algorithms to find job skills matching user queries;obtaining user feedback to optimize search and matching results for job skills;classifying collected recruitment job data to differentiate different types of recruitment data;calculating the demand percentage for specific job skills;generating customized job skill analysis reports based on the skill classification scheme as per user request;categorizing collected raw job position data to differentiate different types of recruitment data and feeding the classified information into the data preprocessing module.16.The automated data management method of claim 15, wherein the types of recruitment data include at least one of company introductions, job descriptions, benefits, skill requirements, and educational degree requirements.17.An automated data management system comprising:a data storage resource configured to store data and instructions; anda processing circuitry configured to implement the method as defined in any one of claims 10-15.18.A computer-implemented automated data management method for job skills, the method being executed by one or more processors and comprising:collecting, at a data collection unit, raw job position data from public resources;preprocessing, at a data preprocessing unit, the collected raw job position data to remove irrelevant information and / or symbols;identifying, at a data identification unit, text related to job skill requirements from the standardized data and identifying keywords related to the job skills;semantically analysing, at a semantic analysis unit, the identified keywords related to the job skills and forming a skill classification scheme;establishing, at data storage unit, a skill database based on the skill classification scheme.19.An automated data management system, comprising:one or more processors; anda computer-readable storage device coupled to the one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for automating data management, the operations comprising the steps as defined in any one of claims 10-15.20.A non-transitory computer readable medium storing a plurality of instructions that when executed control a computer including one or more processors to implement the methods as defined in any one of claims 10-15.
Citation Information
Patent Citations
A keyword classification processing system and method based on Internet mass information
CN109635180A
Business information classification and induction system based on big data integration
CN117332316A
System and method for improved job seeking
WO2006099307A2
Cited By
Regional industry talent skill gap identification equipment, method and device and medium
CN120832491A
Apparatus, method, device and medium for identifying regional industry talent skill gap
CN120832491B