A Job Title Clustering Method and Device Based on Big Data
Through the big data-based job conversion information and combination iterative clustering method, the problem of difficulty in capturing job titles is solved, more accurate job title clustering is achieved, and efficient career planning suggestions are provided.
Patent Information
- Application Number
- CN202410871695.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-06-28
AI Technical Summary
In the prior art, job name clustering methods fail to effectively capture the similarity between job names, resulting in planning bias and inaccuracy in career planning, especially in the context of big data, the problems of job names diversity and frequency imbalance have not been effectively solved.
Through big data-based job conversion information, manifold learning algorithms and density-based clustering methods, a job high-dimensional matrix is constructed, job features are extracted, and clustered through combination iterations to identify the intimacy and similarity between job names, forming a sparse job high-dimensional matrix, and finally obtaining job clusters of different density levels.
It realizes more accurate job title clustering, reduces workload, improves the accuracy of clustering, provides valuable career planning suggestions for individuals from the perspective of career mobility, and reduces dependence on job content description.
Smart Images

Figure CN118797385B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly to a method and device for clustering job titles based on big data. Background Art
[0002] In today's era of rapid technological development, individuals increasingly need to actively manage their career paths and conduct career planning, especially in rapidly developing industries, such as IT professionals. Therefore, systematic career planning and the design of a scientific development path will be very necessary.
[0003] Currently, the implementation methods of career planning services are mainly divided into three categories: The first category is mainly based on the offline traditional career development planning method based on expert knowledge. The evaluation and planning are carried out through the manual connection of experts offline. There are deficiencies such as high prices and low efficiency, which make it difficult to be quickly promoted on a large scale. At the same time, such career planning services are affected by the scope of artificial knowledge and their own professional levels, and even due to subjective factors, it is easy to cause certain planning deviations. When reflected on individuals, the impact will be amplified and cause a very bad experience. Therefore, it is of great practical significance to deeply explore the characteristics of individuals' personalities, educations, skills, etc. to avoid the emergence of such extreme samples. The second category is to recommend jobs online by combining keyword groups extracted from the job database with keyword groups input by users. The technology here mainly considers the matching degree between the user and a certain job at that time, and does not consider the user's career growth, personality characteristics, and soft characteristics not reflected in the resume, and cannot well support the relatively long-term planning problem of career development. The third category is to use big data or deep learning for career planning. This technology is mainly limited to only inputting the information of isolated jobs, without considering the associated transitions and changes between jobs. Therefore, through this method, only the suitability of a single node for employment or the skills and abilities that need to be supplemented can be guided, but the dynamic changes in career development and transitions are not considered, and it is not sufficient to form a comprehensive career development plan.
[0004] Based on this, various career rule methods and platforms have been proposed in the prior art.
[0005] For example, the Chinese invention patent with patent number ZL202211152160.9 discloses a method for career development planning based on a time-series knowledge graph. It extracts massive information by parsing to construct nodes and relationships of the knowledge graph, and uses this to build a graph neural network and career development planning model. Career planning analysis based on a time-series knowledge graph of massive data can provide comprehensive and rapid career development suggestions for relevant candidates, and also provide candidates with a career benchmark with a similar background for reference. It takes into account dynamic job transitions, can dynamically depict job development trends in combination with the time-series graph, and deduce career paths based on comprehensive career development and growth processes.
[0006] Another example is the Chinese invention patent with the patent number ZL202311667740.6, which discloses a career development prediction system and method based on big data analysis. It builds a knowledge graph and combines deep learning with multi-source data fusion and feature extraction to analyze a large amount of complex data, providing users with unprecedented insights. Through the prediction model optimization module of enhanced learning, it can maintain its leading position in prediction accuracy and provide users with the most optimized career development strategy.
[0007] It can be seen that in the process of career planning, it is necessary to first analyze the career and explore the career development path, and then combine the user's personal career-related data to make career rules. However, accurate analysis and mining of careers will involve unified classification of career names. But there is currently no unified classification standard to classify them. Even for the same position, usually each big brand company or brand platform has different job titles. Therefore, in order to better carry out career planning, it is imperative to uniformly cluster job titles.
[0008] However, most of the existing occupational cluster analysis methods are based on small sample data, which has the following limitations:
[0009] 1) Most studies focus on single career mobility events without considering entire career trajectories. An important area of interest is the correlation between specific types of career mobility and career success, as explored by Sheridan et al. (1990), Ishida et al. (2002), and Hamori and Kakarika (2009). For example, Hamori and Kakarika (2009) examined the correlation between the frequency of inter-organizational mobility and the acceleration of promotion to senior positions in large European and US companies. A key shortcoming of these studies is that they ignore the sequence and time information inherent in career development paths.
[0010] 2) When studying the career patterns of different groups, such as senior management positions (Koch et al., 2017), members of the top management team (TMT) (Biemann and Wolf, 2009), executives in large companies in Europe and the United States (Hamori and Kakarika, 2009), and IT professionals (Joseph et al., 2012), the analysis of career patterns is usually coarse-grained. These studies typically use a single job change indicator to determine career patterns, while ignoring the two-dimensional nature of job mobility (i.e., vertical and horizontal) (Vinkenburg and Weber, 2012). Only one of these two dimensions is regarded as a measure of job change, such as the measure of cross-level mobility in terms of vertical mobility (e.g., Blair-Loy (1999), Stovel et al. (1996)) and the measure of cross-functional area mobility in terms of horizontal mobility (e.g., Joseph et al. (2012)).
[0011] In other words, when dealing with limited data, it is not feasible to comprehensively analyze and extract career mobility patterns covering job functions and hierarchical aspects. In today's digital age, professional platforms have enabled the accumulation of a large number of real-world career trajectories. The massive resume data naturally contains information on diverse career development paths. If it is possible to obtain the common characteristics and attributes of the relevant career development of different user groups based on the career development evolution information provided by the massive data, and then refine targeted career planning suggestions and summarize the relevant educational backgrounds, skills, abilities, etc. information, this will be both efficient and accurate.
[0012] However, precisely because in Internet big data, even the same position (or the same job) can have different job titles, some previous work on job title clustering has been limited to comparing the lexical similarity between job titles (Iezzi et al., 2013), which belongs to the boundary category of comparing and clustering strings. Others consider the semantic meaning of job titles based on job description texts (Aken et al., 2009; Colihan and Burger, 1995; Quesenberry and Trauth, 2008; Lappas, 2020).
[0013] However, the information of both methods is insufficient to capture the correlation between job titles. In many cases, the lexical overlap between two very similar job titles may be small (e.g., "programmer" and "developer"). Incorporating the semantic similarity between job titles can effectively address this problem, because semantic meanings explain the tasks, responsibilities, functions, and duties of the positions. However, it is difficult to distinguish the job status between two positions in similar functional areas from semantic descriptions alone (e.g., "software engineer" and "senior software engineer"). In addition, another part of the relevant literature uses supervised methods to handle this task (Javed et al., 2015, 2016): Given a training data set with job category labels attached, a model is trained to assign corresponding labels to unknown job title data. However, appropriate training data is difficult and costly.
[0014] In view of this, there is an urgent need for a clustering method that can more precisely capture the similarity of job titles to cluster job titles, so as to provide reliable data support for career pattern mining, career planning, etc. Summary of the Invention
[0015] The purpose of the present invention is to provide a big data-based job title clustering method and device, which can partially solve or alleviate the above deficiencies in the prior art and can more precisely capture the similarity of job titles.
[0016] To solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions:
[0017] In the first aspect of the present invention, there is provided a big data-based job title clustering method, which includes the steps of:
[0018] Extracting job transition information based on multiple career transition sequences within a professional field; the job transition information includes: job title, job co-occurrence frequency, job co-occurrence closeness, job transition duration, and job transition direction;
[0019] Calculating the closeness representing the similarity between each job title in the career transition sequence based on the job transition information, and constructing a high-dimensional job matrix based on the closeness;
[0020] Associating each job title in the high-dimensional job matrix only with the job title having the greatest closeness to it to obtain a sparsified high-dimensional job matrix;
[0021] Using a manifold learning algorithm to extract job features from the high-dimensional job matrix to obtain job features of each job title;
[0022] Cluster the job titles in the job high-dimensional matrix according to the job characteristics and using a density-based clustering algorithm to obtain similar job clusters with different density levels;
[0023] Repeat the above steps of job characteristic extraction and clustering until no further clustering can be performed;
[0024] Calculate the intimacy representing the similarity between each similar job cluster, and use a density-based clustering algorithm to cluster multiple similar job groups to obtain different job clusters.
[0025] In some embodiments, use the following formula to calculate the intimacy representing the similarity between the positions in the career transition sequence: ; where represents the intimacy between the job title and the job title when transitioning from the job title to the job title in the career transition sequence, = , represents the intimacy between the job title and the job title when transitioning from the job title to the job title in the career transition sequence, is the total number of career transition sequences, is the total number of job titles in the th career transition sequence, and are the th and th job titles and the th job titles in the , , and , is the number of years of work required to jump from the th position to the th position, and are non-increasing functions with and as parametric parameters.
[0026] In some embodiments, before extracting job transition information, preprocessing is performed on each job name, specifically including: constructing a dataset of standard job names; extracting the core information of all job names in all career transition sequences; the core information includes: job level and core functions; using a fuzzy matching package to align the extracted core information with the dataset of standard job names.
[0027] In some embodiments, during the fuzzy matching process, job names with a similarity greater than a preset threshold are mapped to the corresponding standard job names, and job names with a similarity less than the preset threshold are classified as "unknown".
[0028] In some embodiments, the job name clustering method further includes the steps of: merging adjacent different job names in the same similar job cluster into one job name; encoding the job names in the merged career transition sequence to obtain the sequence length of the career transition sequence; screening out career transition sequences with a sequence length less than or equal to a preset length threshold from them to obtain a number of candidate career transition sequences; using a frequent sequence mining algorithm to perform frequent pattern mining to obtain frequent sequence patterns.
[0029] A second aspect of the present invention lies in providing a job name clustering device based on big data, which includes:
[0030] A database for pre-storing multiple career transition sequences in each different professional field;
[0031] A data acquisition module for extracting job transition information from multiple career transition sequences in any professional field; the job transition information includes: job name, job co-occurrence frequency, job co-occurrence closeness, job transition duration, and job transition direction;
[0032] An intimacy recognition module for calculating the intimacy representing the similarity between job names in the career transition sequence based on the job transition information obtained by the data acquisition module; and constructing a job high-dimensional matrix based on the intimacy;
[0033] A sparsification module for associating each job name in the job high-dimensional matrix only with the job name having the greatest intimacy with it to obtain a sparsified job high-dimensional matrix;
[0034] A first clustering module for using a manifold learning algorithm to extract job features of the job high-dimensional matrix to obtain job features of each job name; then clustering the job names in the job high-dimensional matrix according to the job features and using a density-based clustering algorithm to obtain similar job clusters with different density levels; and repeatedly performing job feature extraction and clustering in sequence multiple times until no further clustering can be performed;
[0035] The second clustering module calculates the intimacy representing the similarity between each similar position cluster, and uses a density-based clustering algorithm to cluster the multiple similar position groups to obtain different position clusters.
[0036] In some embodiments, the intimacy recognition module calculates the intimacy representing the similarity between each position in the career transition sequence by using the following formula: ; where represents the intimacy between the th position and the th position when transitioning from the th position to the th position in the career transition sequence, = , represents the intimacy between the th position and the th position when transitioning from the th position to the th position in the career transition sequence, N is the total number of career transition sequences, is the total number of position names in the nth career transition sequence, and are the th position name and the th position name in the nth career transition sequence, , , and , is the number of years of work required to jump from the ( )th position to the th position, is the number of years of work required to jump from the (( -1)th position name to the th position name in the nth career transition sequence, and are non-increasing functions with and as parameters respectively.
[0037] In some embodiments, the database is further configured to store a pre-constructed standard position name dataset;
[0038] The position clustering device further includes: a data preprocessing module, configured to extract the core information of all position names in all career transition sequences before the data acquisition module extracts the position transition information, and align the extracted core information with the standard position name dataset by using a fuzzy matching package; where the core information includes: position level and core function.
[0039] In some embodiments, during the fuzzy matching process, the data preprocessing module maps job names with a similarity greater than a preset threshold to corresponding standard job names, and classifies job names with a similarity less than the preset threshold as "unknown".
[0040] In some embodiments, the job name clustering device further includes:
[0041] An encoding module, configured to merge adjacent different job names in the same job name cluster into one job name; and encode the job names in the merged career transition sequence to obtain the sequence length of the career transition sequence;
[0042] A frequent pattern mining module, configured to filter out career transition sequences with a sequence length less than or equal to a preset length threshold to obtain a number of candidate career transition sequences; and perform frequent pattern mining using the frequent sequences to obtain frequent sequences.
[0043] Advantageous effects: The present invention obtains the similarity between different job names and clusters them through a clustering method based on job transition information and combinatorial iteration; it not only considers the similarity information between the job names included in the career transition sequence; at the same time, the combinatorial iteration method in job name clustering effectively balances the problem of uneven occurrence frequencies between different job names, thereby providing valuable insights for an individual's career planning from the perspective of career mobility.
[0044] The job name clustering method of the present invention is unsupervised and does not rely on job content description data of job names (e.g., intelligent descriptions); instead, it discerns the close relationship between job names from the career transition sequence. For example, usually an individual needs to experience several lateral transitions, indicating transitions between similar levels (i.e., job status) and functional roles. Therefore, such sequential information can be used to capture the close relationship information between job names. Compared with the prior art, the clustering method that only clusters based on the similarity of job names or only based on the similarity of intelligent descriptions not only greatly reduces the workload, but also has better clustering accuracy.
[0045] The prior art with the publication number CN109829500A discloses a job composition and automatic clustering method, which predefines a set of job feature templates. Then, semi-structured job sample data is collected from recruitment websites, feature information is extracted to fill the job templates, and company type information is extracted. At the same time, a job network is constructed using web link information. The job network is sampled by random walk to obtain sample paths, and then the distributed representations of nodes are trained using a language model. Finally, the distributed representations of job nodes and structured feature information are fused, and the K-means algorithm is used for clustering. However, this job network information is essentially the link information between different jobs, etc. The job sampling paths are obtained using this associated information between jobs and then the representations are obtained using a language model. Moreover, the "job sequence" is also used in the process of obtaining the representations, but it is a sequence constructed by random walk using the job link information network. In contrast, in the present invention, the most original job transition sequences of each individual in the big data are directly used, that is, the real job transition sequences of each individual. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. In all the drawings, similar elements or parts are generally denoted by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to actual scale. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of a job title clustering method based on big data according to an embodiment of the present invention;
[0048] Figure 2 It is a functional module diagram of a job title clustering device based on big data according to an embodiment of the present invention;
[0049] Figure 3A They are the original job transition sequences of different individuals in the same professional field obtained from big data;
[0050] Figure 3B Based on Figure 3A The job graph constructed from the multiple job transition sequences shown;
[0051] Figure 3C Using the job title clustering method of the present invention for Figure 3B The various job title clusters obtained by clustering different job titles in;
[0052] Figure 3D Based on Figure 3CThe cluster pairs of job titles shown Figure 3A The new job transition sequence after processing the original job transition sequence shown;
[0053] Figure 4A A schematic diagram for reflecting the sparsification of the distance high-dimensional matrix;
[0054] Figure 4B A schematic diagram for reflecting the clustering of the sparsified distance high-dimensional matrix to obtain cluster pairs of job titles;
[0055] Figures 5A - 5E The results of job title clustering in the IT industry using baseline methods 1-5 respectively;
[0056] Figure 5F The results of job title clustering in the IT industry based on the clustering method of the present invention. Detailed implementation manners
[0057] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] In this article, suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of the description of the present invention and have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.
[0059] In this article, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0060] In this article, unless otherwise clearly defined and limited, terms such as "installed", "provided with", "connected", etc. shall be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0061] In this article, "and / or" includes any and all combinations of one or more of the listed related items. In this article, "a plurality" means two or more, that is, it includes two, three, four, five, etc.
[0062] "Position Name" in this article: The same position or the same job may have different names. Therefore, there are a large number of different position names. For example, some people call those who complete the same job content computer programmers, while others call them software engineers.
[0063] "Function" in this article: It refers to a combination of a set of knowledge, skills, behaviors, and attitudes that can help improve an individual's work effectiveness, and then drive the enterprise's influence and competitiveness in the economy. It is also called "job function". Usually, when an enterprise recruits employees for a corresponding position, it will describe the function of this position. Therefore, it is also called function description or job responsibilities.
[0064] "Occupation Transition Sequence" in this article: It refers to a development path or direction formed by job changes, such as promotions, in each occupation. Therefore, it is also called career path. For example, a career path consists of a series of job titles , ,... to represent an individual's actual career movement history.
[0065] "Job Movement History" in this article: It refers to the historical data of an individual's transfer from one position to another, and can also be called "job transition history". Usually, a career transition sequence includes all the career movement history data of the corresponding individual.
[0066] "Job Co-occurrence Frequency" in this article: It refers to the frequency of two or more job names co-occurring in the same career transition sequence.
[0067] "Job Co-occurrence Closeness" in this article: It refers to the closeness between two job names that co-occur in the same career transition sequence. Specifically, the closeness between two job names is characterized by the absolute value of the distance between the two job names in the career transition sequence (that is, the number of job names between the two job names). For example, in the nth career transition sequence , the th job name and the th job name The closeness between them (that is, the job co-occurrence closeness between them) is expressed as (that is, the distance between the two job names, that is, the number of job names between the two job names). That is, the smaller, the closer the job name and the job name appear in the same career transition sequence.
[0068] "Job transition duration" or "job transition time" in this article: refers to the time required to transition from the position corresponding to one job title to the position corresponding to another job title in the same career transition sequence.
[0069] "Job transition direction" in this article: refers to the transition direction between different positions or different job titles. For example, from software engineer to senior software engineer is a direction. Based on all job jumps in this direction, the job title and job title intimacy can be calculated. Conversely, from job title to job title that is, from senior software engineer to software engineer, an intimacy can also be calculated. These two intimacies are the intimacies under two different career transition directions.
[0070] Although large-scale real career path data has the potential to mine fine-grained career patterns, there are still major technical challenges: 1) The same position can have different job titles, that is, the high dimensionality of job titles will greatly increase the workload in job title clustering, resulting in the inability to perform job title clustering well, that is, unable to perform job clustering well, and thus affecting the fine-grained mining of career patterns. 2) The occurrence frequencies of various positions are different, that is, the occurrence frequencies of job titles of different positions are different. For example, the occurrence frequency of lower-level positions is higher than that of higher-level positions. This imbalance in frequency hinders the effective clustering of job titles, thus affecting the fine-grained mining of career patterns.
[0071] To solve the above problems, in the context of big data, the present invention proposes a new job title clustering method, which clusters high-dimensional career names based on the conversion data between different positions in the job sequence and the combined iterative algorithm, and can cluster job titles more accurately and efficiently, thus providing reliable data support for career pattern mining or career planning, etc.
[0072] Example 1:
[0073] Referring to Figure 1 , the present invention provides a job title clustering method based on big data, which includes the steps:
[0074] S101, extract job transition information based on multiple career transition sequences in a professional field, and execute step S102.
[0075] In some embodiments, the job transition information includes: job title, job co-occurrence frequency, job co-occurrence closeness, job transition duration, and job transition direction.
[0076] Exemplarily, let represent a set of career paths (i.e., career transition sequences) of multiple professionals (i.e., multiple individuals) within any professional field, and let represent the set of job titles corresponding to the positions held by a certain professional. Among them, each career path (i.e., career transition sequence) is composed of a series of job titles , ,..., , that is represents the actual job movement history (or job transition history) of an individual. Therefore, based on the actual job movement history data represented by each career path, the co-occurrence frequency of different jobs, the job transition duration between different jobs, and the job transition direction can be statistically obtained.
[0077] S102. Calculate the intimacy representing the similarity between job titles in the career transition sequence based on job transition information, and construct a high-dimensional job matrix based on the intimacy, and then execute step S103.
[0078] If two job titles are at a similar job level and are responsible for similar job functions, then they are similar. If job titles frequently and closely appear together in the career transition sequence (i.e., both the job co-occurrence frequency and the job co-occurrence closeness are relatively high), they are very likely to be similar. In addition, there are two features that affect the similarity of job titles: 1) job transition duration, the shorter the transition time, the greater the similarity; 2) job transition balance, based on which two-way transitions between similar jobs should occur at a comparable ratio in the career transition sequence dataset.
[0079] Exemplarily, let represent the set of job titles assumed by an individual, represent the th job title and the th job title ; represent the job title intimacy matrix; then the intimacy between the th job title and the th job title is calculated by the following formula:
[0080] = (1).
[0081] Among them, represents the transition from the position corresponding to the job title to the position corresponding to the job title in a career transition sequence, and the job title The intimacy with the job title is the total number of career transition sequences, is the total number of job titles in the th career transition sequence, and and are the th and the th job titles and the th job titles in the th career transition sequence, and is the number of years of work required to jump from the position corresponding to the th job title in the ( )th career transition sequence to the position corresponding to the th job title in the th career transition sequence; represents the number of years required for an individual to jump from the position corresponding to the th job title in the th career transition sequence to the position corresponding to the th job title in the th career transition sequence (e.g., the number of years required to be promoted from the job title to the job title ); is a non-increasing function with as the parameter.
[0082] However, since the job-hopping pattern from the position corresponding to the job title to the position corresponding to the job title is different from the job-hopping pattern from the position corresponding to the job title to the position corresponding to the job title , therefore, is not equal to .
[0083] If the positions corresponding to the job titles and are similar (e.g., between two job titles of the same level, the number of two-way conversions between them is similar), then the intimacy value tends to be close to .
[0084] If the positions corresponding to the job titles and There is a potential promotion relationship between corresponding positions (for example, two positions at different levels, as the conversion usually occurs only in one direction). Therefore, will be greater than .
[0085] For example, if a person's work experience (i.e., career conversion sequence): Software Engineer → Embedded Software Engineer → Software Systems Engineer → Senior Software Engineer. Among them, the intimacy between Software Engineer and Embedded Software Engineer is higher than that with Senior Software Engineer, because it only takes one job hop from Software Engineer to Embedded Software Engineer, while it takes three job hops from Software Engineer to Senior Software Engineer. That is, the total number of transition records from Software Engineer to Senior Software Engineer is greater than the total number of transition records from Senior Software Engineer to Software Engineer, which indicates a potential promotion relationship between these two positions. Therefore, they cannot be classified into the same category.
[0086] Therefore, in order to avoid classifying two positions at different levels (i.e., different position states) into the same category, a penalty is imposed on the intimacy value between position names to obtain the final intimacy calculation formula:
[0087] (2).
[0088] In some embodiments, after calculating the intimacy between each pair of position names, a high-dimensional matrix of positions in this career direction can be constructed based on this intimacy. The high-dimensional matrix of positions includes multiple different position names, and the similarity between different position names is represented by the intimacy.
[0089] Preferably, for visualization and to facilitate subsequent calculations, a position graph is used to represent the high-dimensional matrix of positions. Each vertex in the graph represents a position name, and the connection (or edge) between position names is the intimacy between the two position names.
[0090] Exemplarily, referring to Figure 3A , let the career conversion sequence of professional A be , and the set of position names held by this professional A is: ; let the career conversion sequence of professional B be , and the set of position names held by this professional B is: ; let the career conversion sequence of professional C be , and the set of position names held by this professional C is: .
[0091] Referring to Figure 3B, visually represent the high-dimensional matrix of positions constructed based on intimacy through a position map, where the position names are respectively connected to the position names and the position names , that is, the position names are respectively connected to the position names and the position names have a certain degree of intimacy; similarly, the position names are also respectively connected to the position names and the position names have a certain degree of intimacy.
[0092] S103, associate each position name in the high-dimensional matrix of positions only with the position name having the greatest intimacy with it to obtain a sparsified high-dimensional matrix of positions, and perform step S104.
[0093] Generally, some common position names, such as software engineer, appear more frequently than some less common position names, such as embedded software engineer. This imbalance in the frequency of occurrence of position names hinders effective measurement of job title similarity. As mentioned above, the key factors determining the similarity between position names (i.e., the similarity between positions) include: position co-occurrence frequency, position co-occurrence closeness, position transition duration, and position transition direction.
[0094] However, the imbalance in the frequency of occurrence of different position names will cause the similarity measurement to be dominated by the position co-occurrence frequency. For example, frequently occurring job titles, such as software engineer and senior software engineer, will have relatively high similarity values even if they should not be classified into the same category. For example, if software engineer and senior software engineer often appear in the career transition sequence, while embedded software engineer and lead software engineer are two less common positions. Due to the high frequency of software engineer and the low frequency of lead software engineer, the similarity between software engineer and senior software engineer and the similarity between senior software engineer and lead software engineer show low differences. Therefore, referring to Figure 4A , these four position names may be misclassified into the same cluster.
[0095] To avoid this problem, a sparsified high-dimensional matrix of positions is obtained by associating each position name in the high-dimensional matrix of positions only with the position name having the greatest intimacy with it.
[0096] Furthermore, to facilitate subsequent feature extraction of this high-dimensional matrix of positions by the manifold learning algorithm, therefore, convert the above-mentioned intimacy into a distance. Specifically, convert the intimacy into a distance through the following method:
[0097] (3).
[0098] Among them, a radial basis function is used to calculate the weighted value.
[0099] Exemplarily, as described above, based on the intimacy, software engineers, senior software engineers, embedded software engineers, and supervisor software engineers are visualized to generate a position map, and then it is converted into a position distance map (or a distance high-dimensional matrix, abbreviated as a distance matrix) as Figure 4A shown, where each point represents a position name, and the connection line (or edge) between points represents the distance between two position names.
[0100] From Figure 4A it can be seen that the distance between a software engineer and a senior software engineer is 4, the distance between a software engineer and an embedded software engineer is 7, and the distance between a senior software engineer and a supervisor software engineer is 5. To sparsify this position distance map, by only retaining the edges between position names with the maximum intimacy (that is, the edge with the largest distance), therefore, there is no connected edge between the software engineer and the senior software engineer (because the distance of 4 between the software engineer and the senior software engineer is less than the distance of 7 between the software engineer and the embedded software engineer, so only the edge between the software engineer and the embedded software engineer is retained), so that these four position names are divided into two clusters, see Figure 4A the two clusters shown by the orange dashed box and the green dashed box in
[0101] S104. Use a manifold learning algorithm to extract position features from the position high-dimensional matrix to obtain the position features of each position name, and execute step S105.
[0102] In some embodiments, the manifold learning algorithm Isomap is used to obtain the embedded topology of the sparsified position high-dimensional matrix, which enables the analysis of the proximity information of each occupation name in the metric space. That is, through this manifold learning algorithm, the position features of each position name in this position high-dimensional matrix (for example, the sparsified position map) are vectorized, so as to obtain the multi-dimensional vectorized representation of each position name in this position high-dimensional matrix.
[0103] Preferably, the position features include: position name, distance (or intimacy) between positions, position co-occurrence frequency, position co-occurrence closeness, position transition duration, and position transition direction, etc.
[0104] S105. According to the position features and using a density-based clustering algorithm, cluster the position names in the above position high-dimensional matrix to obtain similar position clusters with different density levels, and execute step S106.
[0105] S106. Repeat the above steps of job feature extraction and clustering until no further clustering can be performed, and then execute step S107.
[0106] In this embodiment, by taking the sparsified high-dimensional job matrix (i.e., the distance matrix) as the input, then applying the Isomap algorithm for feature extraction, and vectorizing the extracted job features to obtain the low-dimensional embedding (i.e., the job feature vector) of each job name. Then, use the DBSCAN algorithm to identify the clusters of job names. Then, use the Isomap algorithm for feature extraction again to obtain the job feature vectors of each job name and / or each cluster, and then use the DBSCAN algorithm for clustering. Repeat this process multiple times until no further clustering can be performed, that is, achieve combined iteration through Isomap and DBSCAN.
[0107] In some embodiments, in each iteration, the job names regarded as "densely connected" under the specific parameter settings of DBSCAN will be extracted and named as "job name clusters", while the remaining ungrouped jobs will be input into the next iteration. During the iteration process, gradually increase the radius parameter of DBSCAN (for example, add 1 to the radius after each iteration) to allow the extraction of job name clusters to become gradually sparser. The iteration process continues until no job names can be grouped, and the job names that cannot be grouped at the end of the iteration are regarded as outliers and removed from the job transition sequence.
[0108] Generally, high-frequency job names tend to have relatively high similarity values among themselves, while less common job names have relatively low similarity values. This makes it difficult to set a single, universal threshold for job name clustering. As Figure 4B shown, in the figure, each job name is only connected to its nearest neighbor. The area surrounded by the dashed line loop in the figure represents the sparse area, and the area surrounded by the solid line loop represents the dense area. Compared with the job names in the dense area, the job names in the sparse area appear frequently and have relatively high mutual similarity values. Therefore, if there are both sparse and dense job name areas, it will pose a very great challenge to set a unified density threshold for clustering.
[0109] To address this challenge, in this embodiment, the density-based clustering algorithm DBSCAN is used iteratively, and job names connected at different density levels are grouped with different parameter configurations in different iterations. As Figure 4B shown, software engineers and software programmers, as well as computer programmers, programmers, analysts, and enhanced programmers are clustered in different iterations.
[0110] In addition, even if the jobs in different density regions have similar job names, due to their intimacy values and are different, and they are also less likely to be grouped into the same cluster. As Figure 4B shown, software engineer is the closest position to computer programmer, so they are linked together. However, they cannot be directly clustered together because they belong to different density region levels. Instead, computer programmers are clustered with programmers, analysts, and enhancement programmers, which are also rare in the career transition sequence.
[0111] To solve this problem, a hierarchical clustering process is designed in this embodiment (that is, first use Isomap for feature extraction, and then use DBSCAN for clustering, and implement combined iteration multiple times), which balances the connection density between computer programmers and software engineers. By clustering computer programmers, programmers, analysts, and enhancement programmers into one cluster, computer programmers are closer to software engineers, thus realizing subsequent clustering.
[0112] S107, calculate the intimacy representing the similarity between each similar position cluster, and use the density-based clustering algorithm to cluster multiple said similar position clusters to obtain different position cluster classes.
[0113] In some embodiments, the same formula as in equations (1) and (2) is used to calculate the intimacy between different similar position clusters. The only difference is that ji and jj in equations (1) and (2) represent position name clusters. Finally, based on the intimacy between different position name clusters, use DBSCAN clustering to further group these position name clusters into the final position cluster classes, that is, by merging position name clusters, similar position names that appear at different frequencies can be effectively clustered to form the final position cluster classes.
[0114] Exemplarily, referring to Figure 3C , clustering using the above clustering method gives: the position names surrounded by each dotted line represent a group of position names; each group surrounded by a dashed line represents a position cluster class, that is, different position names in this class actually correspond to the same position, just different names.
[0115] Usually in the process of career pattern mining, since the career transition sequence dataset stored in the database contains a large number of different position names, a large amount of work is required to encode the career transition sequences, which not only increases the workload, but also reduces the significance of career pattern mining to a certain extent, because a large number of various position names may obscure the truly meaningful career patterns.
[0116] Therefore, in this embodiment, the number of jobs is reduced through a clustering method based on a job matrix. Moreover, the clustering of job names does not rely on additional specific job knowledge, but directly learns the similarity between job names from the career transition sequence, thereby summarizing the complex career transition sequence information into a high-dimensional job matrix (such as a job graph). Then, based on functions and levels (i.e., job status), job names are clustered to create categories of similar jobs, solving two technical challenges: the huge dimension of job names increases the difficulty of clustering, and the imbalance in the frequency of job names in the career transition sequence leads to the problem of inaccurate clustering.
[0117] Embodiment 2
[0118] The present invention also provides another job clustering method based on big data, which includes each step in the above Embodiment 1. The difference is that considering that in the original data obtained from network big data, even for the same job name, different companies may have different naming methods. Therefore, in this embodiment, before performing the above step S101, it further includes the step of preprocessing each job name, so as to reduce noise and further reduce the difficulty of job name clustering.
[0119] Specifically, the step of preprocessing each job name specifically includes:
[0120] S201, construct a standard job name data set.
[0121] In some embodiments, obtain the original data from the Internet, and then use the top pre-set threshold (for example, 5000, specifically, the pre-set threshold can be determined based on the size of the sample in the original data) frequently occurring job names in the original data set as the standard job name data set.
[0122] S202, extract the core information of all job names in all career transition sequences.
[0123] In some embodiments, use more than 30,000 job levels and job function entities from the IPOD job data set to perform exact keyword matching and extract the core information of all job names in the original data set. A job usually consists of three parts: (1) the job name and its level, such as junior, senior, and chief; (2) the core function, such as software developer; (3) redundant information, such as MLCM Global Label Quality Coordinator. The first two parts constitute the core information of the job name. Therefore, a job that only contains these two parts of information is called a standard job name without redundant information. That is, the core information includes: job level and core function.
[0124] S203, use a fuzzy matching package to align the extracted core information with the standard job name data set.
[0125] In some embodiments, the fuzzywuzzy package is used to align the extracted core information with the standard job title dataset. During the fuzzy matching process, job titles with a similarity score higher than 80 are mapped to their corresponding standard job titles, while those lower than the threshold (80) are classified as "unknown".
[0126] Embodiment 3
[0127] Based on the above clustering method, the present invention also provides another clustering method (or called career pattern mining method), which includes all the steps of Embodiment 1 or 2 above. That is, after obtaining the clustering result, it further includes the steps of: merging adjacent different job titles in the same job title cluster into one job title; encoding the job titles in the merged career transition sequence to obtain the sequence length of the career transition sequence; screening out the career transition sequences with a sequence length less than or equal to a preset length threshold from them to obtain several candidate career transition sequences; and using the frequent sequence mining algorithm to perform frequent pattern mining to obtain the frequent sequence pattern.
[0128] See Figure 3D , merge each job title in the same job title cluster C1 (this cluster C1 includes a group C11 with one job title and an adjacent group C12 with multiple different job titles; in the same cluster, each different job title in the same similar group is merged into one job title) into one job title; merge the two job titles and in the same job cluster C2 into one job title (for example, both are called or ); merge each job title in cluster C3 (this cluster C1 includes a group C31 with one job title and a group C32 with multiple different job titles) into one job title. Among them, C1, C2, and C3 are the encodings of the corresponding clusters, and C11, C12, C31, and C32 are the encodings of each job title group under the corresponding clusters.
[0129] Then, re-encode the job titles in the original career transition sequence in Figure 3A :
[0130] For Figure 3A the career transition sequence in and belong to the same cluster C3, that is, they correspond to the same job title. Therefore, after replacing them with the merged job title, the career transition sequence It only includes three job titles (for convenience, cluster codes are used in the figure), that is, the length of the original career transition sequence changes from 4 to 3;
[0131] For Figure 3A the career transition sequence , since and belong to the same cluster C2, that is, they correspond to the same job title. Therefore, after replacing with the merged job title, the career transition sequence only includes two job titles (for convenience, cluster codes are used in the figure), that is, the length of the original career transition sequence changes from 3 to 2;
[0132] For Figure 3A the career transition sequence , since and belong to the same cluster C1, that is, they correspond to the same job title. Therefore, after replacing with the merged job title, the career transition sequence only includes two job titles (for convenience, cluster codes are used in the figure), that is, the length of the original career transition sequence changes from 4 to 3.
[0133] In this embodiment, the frequent sequence mining algorithm is mined using the frequent sequence mining algorithm commonly used in the art, which will not be elaborated here. After clustering the job titles through the above clustering method in this embodiment, not only the workload of clustering before mining is greatly reduced, but also the workload of encoding the career transition sequence is reduced. Moreover, through the above clustering method, the dimension of the job titles is greatly reduced, which is conducive to mining truly meaningful career patterns.
[0134] Embodiment 4
[0135] The following is a detailed description in combination with specific embodiments and drawings.
[0136] The original data for this clustering analysis comes from the LinkedIn platform. Filter the employment profiles with a tenure of 15 years or more and employment profiles related to IT. Finally, N = 155 million LinkedIn career transition sequences in the IT industry are obtained. Among them, each career transition sequence corresponds to an IT practitioner, and each career transition sequence contains multiple job title change records.
[0137] Each position includes information such as the position name, company, and employment duration. The original dataset contains different positions held by millions of IT practitioners. Due to subjective naming conventions of different companies (such as Software Development Engineer, Software Engineer, and SDE), additional descriptive information for specific technologies (such as Machine Learning, C++, Java), specific project, program, or team names (such as Test and Startup Project Coordinator), additional punctuation and spelling mistakes (such as Senior Engineer), mixed job titles (such as Data Scientist / Data Analyst), and so on.
[0138] Such a large scale of position names poses a significant challenge to the effective application of clustering. Therefore, it is necessary to standardize the position names and unify various similar expressions of position names into one format.
[0139] As mentioned above, through preprocessing (i.e., the process of standardizing position names), the original millions of positions are transformed into 4,534 different positions, greatly reducing the order of magnitude of the position names, and the occurrence times of "unknown" positions account for 11.8% of the total occurrence times of all positions.
[0140] After clustering the preprocessed data using the clustering method in Embodiment 1 above, a total of 19 position name clusters are finally obtained. Among them, the top 10 frequently occurring position names in each position name cluster are shown in Table 1-3.
[0141] Table 1 Position Name Cluster: The Top 10 Frequently Occurring Position Names in IT Technical Positions - Service and Support
[0142]
[0143] Table 2 Position Name Cluster: The Top 10 Frequently Occurring Position Names in IT Technical Positions - Software and System Development and Implementation
[0144]
[0145] Table 3 Position Name Cluster: The Top 10 Frequently Occurring Position Names in IT Management Positions
[0146]
[0147] In addition, the top 10 words with the highest tf-idf values in each job title cluster are used to represent the most distinctive keywords in each cluster. The semantic meanings of these keywords are consistent with the meanings of the 10 most frequently occurring job titles in each cluster. This proves the accuracy of the job title clustering results of the present invention. Then, we assign a job label to each job title cluster. We find that the labels of each cluster are significantly different in terms of job functions and job levels. These 19 jobs can be divided into two categories: technical IT jobs and managerial IT jobs. Technical IT jobs can be further divided into two different types according to their basic functions and responsibilities: (1) service and support-related jobs, and (2) software and system development and implementation-related jobs.
[0148] Evaluation of Job Title Clustering Results
[0149] See Figures 5A - 5F , by comparing the job title clustering results using the present invention with different baseline methods, including: a. Job title clustering based on lexical similarity: (1) Mean shift clustering based on the Jaro-winkler distance, (2) Affinity propagation clustering based on the Jaro-winkler distance (Frey and Dueck, 2007), and (3) AlterHier clustering; b. Job title clustering based on semantic meaning: (4) Word2Vec; c. Job title clustering based on the temporal correlation in the career transition sequence: (5) Temporal skeletonization.
[0150] Baseline 1 (Mean shift clustering based on the Jaro-winkler distance) (JW-MS): In the first baseline, the Jaro-winkler algorithm is used to measure the lexical similarity between job titles. The Jaro-winkler string metric has been proven to be superior to other related metrics, including the widely used Levenshtein distance (Cohen et al., 2003), in measuring the edit distance between different strings. Then, the job title distance matrix based on lexical similarity is used as input to implement mean shift clustering to group job titles. Mean shift clustering can be performed automatically when choosing the kernel function k and bandwidth h without having to determine the total number of clusters in advance. Since the choice of the kernel function has little impact on the kernel density estimation and the accuracy of clustering for each point, a Gaussian kernel function is used. In addition, the bandwidth parameter is calculated, where is the standard deviation of the pairwise distances (Dehnad, 1987).
[0151] Baseline 2 (Affinity Propagation Clustering based on Jaro - winkler Distance) (JW - AP): In the second baseline, Affinity Propagation (AP) clustering is implemented. In AP clustering, the number of clusters does not need to be specified. Traditional k - center clustering methods (such as k - means clustering) start from randomly selected centers and iteratively optimize these centers to reduce the squared error between each point and its nearest center. Therefore, the clustering results are very sensitive to the choice of initial centers. AP improves the traditional k - center clustering method by considering all points as potential centers simultaneously and using message passing between different points to find the optimal set of centers and the corresponding clusters. Thus, AP uses the Jaro - winkler similarity metric described above as the similarity between data point pairs.
[0152] Baseline 3 (AlterHier Clustering): In the third baseline, the hierarchical clustering method developed by Oh and Kim (2004) is used, which is used for clustering sequence datasets. Different from the commonly used metric of comparing two sequences by calculating the edit distance, Oh and Kim (2004) define the similarity between two sequences and as , where is the set of all subsequences of length 2 in sequence , and the AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) values are used to determine the optimal number of clusters.
[0153] Baseline 4 (Word2Vec): In the fourth baseline, semantic similarity is utilized to group job names. First, the Word2Vec algorithm is implemented to capture the semantic correlations between different job titles. Specifically, each career transition sequence is regarded as a "sentence", where each job name is regarded as a separate "word". By training Word2Vec on a large corpus of career transition sequences, vectors representing each job name in the career transition sequence "corpus" are obtained to capture the semantic information of each job name. Then, the mean - shift clustering algorithm is implemented to group the job names into different clusters as described above.
[0154] Baseline 5 (Time Skeletonization): Time skeletonization is applied to career path data to group job names. Time skeletonization first summarizes the temporal correlations between different symbols in a sequence into an undirected graph, and then uses the embedding topology of the graph in the metric space to perform clustering of the symbols.
[0155] Such as Figures 5A - 5FAs shown in (the horizontal and vertical coordinates in the figure represent the two dimensions of the vector) and Table 4-7, the clustering results of job names with different baselines are as follows:
[0156] Table 4 Partial results of mean shift clustering based on Jaro - winkler distance
[0157]
[0158] Table 5 Affinity propagation clustering results based on Jaro - winkler distance
[0159]
[0160] Table 6 Clustering results of Word2Vec
[0161]
[0162] Table 7 Clustering results of time skeletonization
[0163]
[0164] Figures 5A - 5E The t - distributed Stochastic Neighbor Embedding (t - SNE) (Van der Maaten and Hinton, 2008) of job name clustering is visualized, where each name title is represented in a two - dimensional space. From the visualization results, it can be observed that the job name clusters obtained by JW - MS, AlterHier, time skeletonization, and Word2Vec are evenly distributed in the embedding space, with few cases of clustering together. In addition, different clusters cannot be clearly separated because they are closely adjacent. This indicates poor clustering results. Only the baseline JW - AP seems to create relatively reasonable job clusters, and some of the clusters can be clearly identified. In contrast, the clusters obtained by the clustering method of the present invention are closely clustered together with a larger cluster spacing, which proves that the clustering method of the present invention has better performance in the job name clustering task.
[0165] In addition, the clustering results of the clustering method of the present invention are evaluated by comparing the top three keywords and common job names of each cluster obtained by different methods.
[0166] Table 8: Semantic annotation of Top3 job clusters under different methods
[0167]
[0168] See Table 8, which only shows the results of the first three job clusters for comparison. In the AlterHier clustering, almost all job names are included in the same cluster. Therefore, only the semantic annotations of job clusters in the other four baseline methods are analyzed. Consistent with the job cluster visualization results, the semantic annotations of job clusters again show that the clustering results of different baseline methods are poor. As for the baseline JW-MS and JW-AP, they cluster job names by evaluating the lexical similarity between job names, and the job names in the same cluster are only similar at the lexical level. Therefore, the job clusters are heterogeneous in terms of function and rank, and thus it is impossible to clearly annotate each job cluster. For example, in the baseline JW-MS, the job names in C13 include lexically similar words such as "technical", "technologist", and "technology", as indicated by the keywords. However, the job names in C13 (such as "technical lead", "technical writer", and "technical support") belong to different functional areas and ranks.
[0169] By learning the semantic meanings of each job cluster, the clustering results of the baseline word2vec have been greatly improved compared with the results considering only lexical similarity. For example, in C4, "architect" is clustered with "senior software engineer", although their lexical overlap is small. However, the differences between some clusters still cannot be well identified. For example, the job names in C3, C6, and C11 are all related to the software development and analysis functional areas. As Figure 5D shown, these three clusters are intertwined and cannot be clearly separated. In addition, due to the heterogeneity of the clustering, there are still some clusters that cannot be clearly annotated. For example, "project manager", "technician", "director", and "customer service representative" are three job names with different job responsibilities and ranks, but they are grouped together in C1. Similarly, from the results of temporal skeletonization, we still cannot clearly annotate different job name clusters. For example, the keywords "consultant", "engineer", and "manager" appear in C1, C2, and C3 at the same time.
[0170] The job clustering method used in Lappas (2020)'s research is based on job description data and is not suitable as a baseline for direct comparison. From their job clustering results, it can be found that job titles are clustered into several groups, each group representing jobs with different functions in the IT field, which is better than other baseline methods. However, in Lappas (2020)'s research, job titles of jobs belonging to the same functional area but different ranks are grouped together, while the clustering method of the present invention can clearly divide these jobs into different groups. For example, "software engineer", "lead software engineer" and "software architect" are grouped into the same cluster in Lappas (2020)'s research, while in the clustering method of the present invention, they are regarded as different types of jobs. Although these three jobs all belong to the software development field, they represent three different ranks. Changes in job functions and ranks are important indicators of career mobility. Therefore, the ability to distinguish different ranks during the job title classification stage is crucial for understanding career mobility patterns.
[0171] Verifying the results of career pattern mining
[0172] To evaluate the performance of career pattern mining, the method of the present invention is compared with a baseline method. In the baseline method, the frequent sequence mining (FSM) algorithm is applied to the sequences encoded by the job title clusters obtained by word2vec and time skeletonization. Note that in this verification experiment, no comparison is made with the clustering method based on lexical similarity because their career clustering results are very poor and cannot capture the similarity of jobs. Therefore, the next step of sequence pattern mining cannot be carried out.
[0173] Table 9 Typical career patterns of lengths 2, 3, and 4
[0174]
[0175] Table 9 shows the top 3 frequent career patterns of lengths 2, 3, and 4 obtained by the method of the present invention and the above two baseline methods, namely baseline method 4 and baseline method 6, respectively. The results show that the baseline algorithms are ineffective: taking baseline method 6 (time skeletonization) as an example, most of the top frequent career patterns are combinations of different orders of C1, C2, and C3. The job title clusters C1, C2, and C3 all belong to jobs with similar keywords "consultant", "manager", and "analyst". The typical career patterns obtained by the method of the present invention show the changes in job functions and ranks during the career development process.
[0176] Example 5
[0177] Based on the above job title clustering method, the present invention also provides a job clustering device based on big data, which includes:
[0178] A database for pre-storing multiple career transition sequences in each different professional field;
[0179] A data acquisition module for extracting job transition information from multiple career transition sequences in any professional field; wherein, the job transition information includes: job title, job co-occurrence frequency, job co-occurrence closeness, job transition duration, and job transition direction;
[0180] An intimacy recognition module for calculating the intimacy characterizing the similarity between job titles in the career transition sequence based on the job transition information obtained by the data acquisition module; and constructing a high-dimensional job matrix based on the intimacy;
[0181] A sparsification module for associating each job title in the high-dimensional job matrix only with the job title having the greatest intimacy with it, to obtain a sparsified high-dimensional job matrix;
[0182] A first clustering module for extracting job features from the high-dimensional job matrix using a manifold learning algorithm to obtain the job features of each job title; and then clustering the job titles in the high-dimensional job matrix using a density-based clustering algorithm according to the job features, to obtain similar job clusters with different density levels; and repeatedly performing job feature extraction and clustering in sequence multiple times until no further clustering can be performed;
[0183] A second clustering module for calculating the intimacy characterizing the similarity between each similar job cluster, and clustering the multiple similar job clusters using a density-based clustering algorithm to obtain different job cluster classes.
[0184] In some embodiments, the intimacy recognition module calculates the intimacy characterizing the similarity between jobs in the career transition sequence using the above formulas (1) and (2).
[0185] In some embodiments, the above database is further used to store a pre-constructed standard job title dataset; correspondingly, the above job clustering device further includes:
[0186] A data preprocessing module for extracting the core information of all job titles in all career transition sequences before the data acquisition module extracts job transition information, and aligning the extracted core information with the pre-constructed standard job title dataset using a fuzzy matching package; wherein, the core information includes: job level and core function.
[0187] In some embodiments, during the fuzzy matching process, the above data preprocessing module maps job names with a similarity greater than a preset threshold to corresponding standard job names, and classifies job names with a similarity less than the preset threshold as "unknown".
[0188] In some other embodiments, the job clustering device further includes:
[0189] An encoding module, configured to merge adjacent different job names in the same similar job cluster into one job name; and encode the job names in the merged career transition sequence to obtain the sequence length of the career transition sequence;
[0190] A frequent pattern mining module, configured to filter out career transition sequences with a sequence length less than or equal to a preset length threshold to obtain a number of candidate career transition sequences; and then perform frequent pattern mining using the frequent sequences according to a preset rule to obtain a frequent sequence pattern that conforms to the preset rule; the preset rule includes: flow within an organization and across organizations, or flow between different positions.
[0191] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0192] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a computer terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0193] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims. All of these are within the protection scope of the present invention.
Claims
1. A job title clustering method based on big data, characterized in that Including the steps: Extracting job transition information based on multiple career transition sequences within a professional field; the job transition information includes: job name, job co-occurrence frequency, job co-occurrence closeness, job transition duration, and job transition direction; Calculating the closeness representing the similarity between each job name in the career transition sequence based on the job transition information, and constructing a high-dimensional job matrix based on the closeness; Associating each job name in the high-dimensional job matrix only with the job name having the maximum closeness to it, obtaining a sparsified high-dimensional job matrix; Extracting job features of each job name by using a manifold learning algorithm for the high-dimensional job matrix; Clustering the job names in the high-dimensional job matrix according to the job features and using a density-based clustering algorithm to obtain similar job clusters with different density levels; Repeating the above job feature extraction and clustering steps until no further clustering can be performed; Calculating the closeness representing the similarity between each similar job cluster, and using a density-based clustering algorithm to cluster multiple similar job clusters to obtain different job name clusters; Wherein, the job co-occurrence closeness refers to the tightness between two job names that co-occur in the same career transition sequence, and the tightness between the two job names is represented by the absolute value of the distance between the two job names in the career transition sequence; The closeness of the similarity between each job name refers to the distance between jobs, and is used to represent the similarity between each job name in the career transition sequence.
2. The method for clustering job titles based on big data according to claim 1, wherein Calculate the intimacy characterizing the similarity between positions in the career transition sequence using the following formula: ; Among them, represents the intimacy between job name and job name when transitioning from job name to job name in the career transition sequence, = , represents the intimacy between job name and job name when transitioning from job name to job name in the career transition sequence, is the total number of career transition sequences, is the th total number of job names in the th career transition sequence, and are the th and th job names in the th career transition sequence, , and , and is the number of years of work required to jump from the th job name to the th job name in the th career transition sequence, is the number of years of work required to jump from the ( - 1)th job name to the th job name in the th career transition sequence, and are non-increasing functions with as parameters respectively.
3. The method for clustering job titles based on big data according to claim 1, wherein Before extracting the job transition information, preprocessing each job name, specifically including: Constructing a standard job name dataset; Extracting the core information of all job names in all career transition sequences; the core information includes: job level and core function; Using a fuzzy matching package to align the extracted core information with the standard job name dataset.
4. The method for clustering job titles based on big data according to claim 3, wherein During the fuzzy matching process, mapping the job names with a similarity greater than a preset threshold to the corresponding standard job names, and classifying the job names with a similarity less than the preset threshold as "unknown".
5. The method for clustering job titles based on big data according to claim 1, characterized in that, Also including the steps: Merging adjacent different job names in the same job name cluster into one job name; Encoding the job names in the merged career transition sequence to obtain the sequence length of the career transition sequence; Screening out the career transition sequences with a sequence length less than or equal to a preset length threshold from them to obtain a number of candidate career transition sequences; Performing frequent pattern mining by using a frequent sequence mining algorithm to obtain a frequent sequence pattern.
6. A job title clustering device based on big data, characterized in that, Including the steps: A database for pre-storing multiple career transition sequences within each different professional field; A data acquisition module for extracting job transition information from multiple career transition sequences within any professional field; the job transition information includes: job name, job co-occurrence frequency, job co-occurrence closeness, job transition duration, and job transition direction; An intimacy recognition module, configured to calculate the intimacy representing the similarity between each job name in the career transition sequence based on the job transition information obtained by the data acquisition module; and construct a high-dimensional job matrix based on the intimacy; A sparsification module, configured to associate each job name in the high-dimensional job matrix only with the job name having the greatest intimacy, so as to obtain a sparsified high-dimensional job matrix; A first clustering module, configured to extract job features of each job name by using a manifold learning algorithm for the high-dimensional job matrix, so as to obtain job features of each job name; then cluster the job names in the high-dimensional job matrix according to the job features and by using a density-based clustering algorithm, so as to obtain similar job clusters with different density levels; and repeatedly perform job feature extraction and clustering in sequence until no further clustering can be performed; A second clustering module, configured to calculate the intimacy representing the similarity between each similar job cluster, and cluster multiple similar job clusters by using a density-based clustering algorithm, so as to obtain different job name clusters; Wherein, the job co-occurrence closeness refers to the tightness between two job names that co-occur in the same career transition sequence, and the absolute value of the distance between the two job names in the career transition sequence is used to represent the tightness between the two job names; The intimacy of the similarity between each job name refers to the distance between jobs, and is used to represent the similarity between each job name in the career transition sequence.
7. The apparatus for clustering job titles based on big data according to claim 6, wherein The intimacy recognition module calculates the intimacy characterizing the similarity between positions in the career transition sequence using the following formula: ; Among them, represents the intimacy between job name and job name when transitioning from job name to job name in the career transition sequence. = , represents the intimacy between job name and job name when transitioning from job name to job name in the career transition sequence. is the total number of career transition sequences, is the total number of job names in the th career transition sequence, and are the th and th job names in the nth career transition sequence, , , and , is the number of years of work required to jump from the th job to the th job, is the number of years of work required to jump from the ( - 1)th job name to the th job name in the nth career transition sequence, and are non-increasing functions with and as parameters respectively.
8. The device for clustering job titles based on big data according to claim 7, characterized in that, The database is further configured to store a pre-constructed standard job name data set; The job clustering device further includes: a data preprocessing module, configured to extract the core information of all job names in all career transition sequences before the data acquisition module extracts the job transition information, and align the extracted core information with the standard job name data set by using a fuzzy matching package; wherein, the core information includes: job level and core function.
9. The apparatus for clustering job titles based on big data according to claim 8, wherein During the fuzzy matching process, the data preprocessing module maps the job names with a similarity greater than a preset threshold to the corresponding standard job names, and classifies the job names with a similarity less than the preset threshold as "unknown".
10. The apparatus for clustering job titles based on big data according to claim 9, wherein It further includes: An encoding module, configured to merge adjacent different job names in the same similar job cluster into one job name; And encode the job names in the merged career transition sequence to obtain the sequence length of the career transition sequence; A frequent pattern mining module, configured to filter out the career transition sequences with a sequence length less than or equal to a preset length threshold to obtain a number of candidate career transition sequences; and perform frequent pattern mining by using the frequent sequences to obtain frequent sequences.
Citation Information
Patent Citations
Job composition and automatic clustering method
CN109829500A
Occupational development planning method based on time sequence knowledge graph
CN115455205A
Occupational development prediction system and method based on big data analysis
CN117371625A
Device and method for name disambiguation clustering
CN102654881A
Method and system for processing talent model information based on talent market
CN110390514A