Electronic device and method for performing data table-based syntax-wise job fit assessment
The electronic device and method address inefficiencies in recruitment systems by using data table-based phrase-unit job suitability evaluation with clustering keywords and semantic role analysis, enhancing accuracy and reducing manual workload in determining job suitability.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MUHAYU INC
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional recruitment systems face significant time and cost inefficiencies due to manual processing of large volumes of applicant documents, and the difficulty in determining job suitability based on contextual meaning, especially for large corporations receiving over 10,000 applicants.
An electronic device and method for data table-based phrase-unit job suitability evaluation using clustering keywords and semantic role analysis to construct a keyword pool, enabling accurate job suitability determination by matching input data with preprocessed documents.
Improves the accuracy of job suitability assessment by reflecting contextual meaning and reducing manual workload, constructing a meaningful keyword dictionary for consistent evaluation across different HR personnel.
Smart Images

Figure KR2025017238_07052026_PF_FP_ABST
Abstract
Description
Electronic device and method for performing data table-based syntax-unit job fit evaluation
[0001] The present disclosure relates to a natural language processing device, and more specifically, to an electronic device and method for performing a data table-based syntax-unit job suitability evaluation.
[0002] Conventional recruitment systems faced the problem of incurring significant time and costs because HR personnel had to manually read applicants' documents, extract necessary information, and organize it. This was particularly problematic in the case of open recruitment by large corporations attracting over 10,000 applicants, where the human resource costs and workload for processing were substantial, and maintaining the consistency of extracted information was difficult depending on the proficiency of the HR personnel.
[0003] Therefore, answer items within the above recruitment documents must be extracted using natural language processing and preprocessed into a document of a certain format so that HR personnel can verify only the meaningful answer items.
[0004] Furthermore, it is necessary to determine whether an applicant is suitable for the job group or position they applied for based on the recruitment documents they submitted. However, even if the documents are preprocessed into a specific format, there was a problem in that it was difficult to determine actual suitability for the job group or position because the contextual meaning could not be reflected by matching only words.
[0005] In conventional patent literature, it is possible to provide feedback on the answer level by determining whether the interviewer satisfies the level of answer desired by the interviewer, but it is not possible to determine the suitability for the job group or job from the answer level.
[0006] The purpose of the embodiments disclosed in this disclosure is to provide an electronic device and method for performing a data table-based phrase-unit job suitability evaluation by extracting valid keywords through clustering keywords and comparing the extracted valid keywords at the phrase unit level according to semantic roles.
[0007] The problems that this disclosure aims to solve are not limited to those mentioned above, and other unmentioned problems will be clearly understood by a person skilled in the art from the description below.
[0008] An electronic device for performing a data table-based phrase-unit job suitability evaluation according to the present disclosure, for achieving the technical problem described above, comprises: a memory storing at least one process for performing a phrase-unit job suitability evaluation operation; and at least one processor for performing the job suitability evaluation operation according to the process. The at least one processor may be configured to construct a standard dataset based on competency elements for each job group or job, extract clustering keywords based on a plurality of clusters configured for each job group or job in the standard dataset, classify and label each text in the standard dataset according to a preset semantic role, construct a keyword pool using the clustering keywords and the labeled text, and evaluate job suitability for input data based on the phrase-unit matching degree with the keyword pool.
[0009] A method for performing a data table-based syntactic unit job suitability evaluation according to the present disclosure for achieving the aforementioned technical task may include: a step of configuring a standard dataset based on competency elements for each job group or job; a step of extracting clustering keywords based on a plurality of clusters configured for each job group or job in the standard dataset; a step of classifying and labeling each text in the standard dataset according to a preset semantic role; a step of constructing a keyword pool using the clustering keywords and the labeled text; and a step of evaluating job suitability for input data based on the syntactic unit matching degree with the keyword pool.
[0010] In addition to this, a computer program stored on a computer-readable recording medium for implementing the present disclosure may be further provided.
[0011] In addition to this, a computer-readable recording medium for recording a computer program for implementing the present disclosure may be further provided.
[0012] According to the aforementioned means for solving the problem of the present disclosure, a meaningful keyword dictionary capable of determining actual job suitability is constructed by labeling phrases containing clustering keywords in a self-introduction letter or job description according to semantic roles, thereby providing the effect of improving the accuracy of job suitability determination.
[0013] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below.
[0014] FIG. 1 is a block diagram briefly illustrating the configuration of an electronic device that performs a data table-based syntax unit job suitability evaluation according to the present disclosure.
[0015] FIG. 2 is a process diagram illustrating the job suitability evaluation process of an electronic device that performs a data table-based syntax unit job suitability evaluation according to the present disclosure.
[0016] FIG. 3 is a block diagram illustrating the process of generating a keyword dictionary for an electronic device that performs a data table-based phrase-unit job suitability evaluation according to the present disclosure.
[0017] FIG. 4 is an exemplary diagram illustrating evaluation elements of an electronic device that performs a data table-based syntax unit job suitability evaluation according to the present disclosure.
[0018] FIG. 5 is an exemplary diagram illustrating a standard dataset of an electronic device that performs a data table-based syntax unit job suitability evaluation according to the present disclosure.
[0019] FIG. 6 is a keyword frequency table generated according to the word frequency analysis within a document of an electronic device that performs a data table-based phrase-unit job suitability evaluation according to the present disclosure.
[0020] Figure 7 is a diagram illustrating the filtering process of clusters generated from the keyword frequency table of Figure 6.
[0021] FIG. 8 is a diagram illustrating the result of a text syntactically segmented by an electronic device that performs a data table-based syntactic unit job suitability evaluation according to the present disclosure.
[0022] FIG. 9 is a table in which an electronic device performing a data table-based phrase-unit job suitability evaluation according to the present disclosure labels corresponding semantic roles for divided phrases.
[0023] FIG. 10 is a data table of a keyword pool generated by an electronic device that performs a data table-based phrase-unit job suitability evaluation according to the present disclosure.
[0024] FIG. 11 is a data table of a document to be analyzed generated by an electronic device that performs a data table-based syntax unit job suitability evaluation according to the present disclosure.
[0025] FIG. 12 is a table showing the weights for each job within the same job group when calculating job suitability in an electronic device that performs a data table-based phrase-unit job suitability evaluation according to the present disclosure.
[0026] Throughout this disclosure, the same reference numerals denote the same components. This disclosure does not describe all elements of the embodiments, and general content in the art to which this disclosure pertains or content that overlaps between embodiments is omitted. The terms 'part, module, component, block' as used in the specification may be implemented in software or hardware, and depending on the embodiments, a plurality of 'parts, modules, components, blocks' may be implemented as a single component, or a single 'part, module, component, block' may include a plurality of components.
[0027] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are directly connected but also cases where they are indirectly connected, and indirect connections include connections made via a wireless communication network.
[0028] Furthermore, when it is stated that a part "includes" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0029] Throughout the specification, when it is stated that a component is located "on" another component, this includes not only cases where a component is in contact with another component, but also cases where another component exists between the two components.
[0030] The terms first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.
[0031] Singular expressions include plural expressions unless there is an obvious exception in the context.
[0032] In each step, identification codes are used for convenience of explanation and do not describe the order of the steps; the steps may be performed differently from the specified order unless a specific order is clearly indicated in the context.
[0033] The operating principles and embodiments of the present disclosure will be described below with reference to the attached drawings.
[0034] In this specification, the term "device according to the present disclosure" includes all various devices capable of performing computational processing and providing results to a user. For example, the device according to the present disclosure may include all of a computer, a server device, and a portable terminal, or may be in the form of any one of these.
[0035] Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.
[0036] The above server device is a server that processes information by communicating with an external device, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.
[0037] The above portable terminal may include, for example, all types of handheld-based wireless communication devices such as PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminals, smartphones, etc., as well as wearable devices such as watches, rings, bracelets, anklets, necklaces, glasses, contact lenses, or head-mounted devices (HMDs).
[0038] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0039] The predefined operating rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined operating rules or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0040] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained by the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0041] According to an exemplary embodiment of the present disclosure, a processor can implement artificial intelligence. Artificial intelligence refers to a machine learning method based on an artificial neural network that enables a machine to learn by mimicking human biological neurons. Methodologies of artificial intelligence can be classified according to the learning method into supervised learning, where input and output data are provided together as training data and the solution (output data) to the problem (input data) is predetermined; unsupervised learning, where only input data is provided without output data and the solution (output data) to the problem (input data) is not predetermined; and reinforcement learning, where a reward is given from an external environment whenever an action is taken from the current state, and learning proceeds in a direction that maximizes such reward. In addition, artificial intelligence methodologies can be classified according to the architecture, which is the structure of the learning model. The architectures of widely used deep learning technologies can be classified into Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Transformers, and Generative Adversarial Networks (GAN).
[0042] The device and system may include an artificial intelligence model. The artificial intelligence model may be a single model or may be implemented as multiple models. The artificial intelligence model may be composed of a neural network (or artificial neural network) and may include statistical learning algorithms in machine learning and cognitive science that mimic biological neurons. A neural network may refer to a model that possesses problem-solving capabilities by having artificial neurons (nodes) that form a network through synaptic connections and change the strength of synaptic connections through learning. The neurons of a neural network may include combinations of weights or biases. A neural network may include one or more layers composed of one or more neurons or nodes. For example, the device may include an input layer, a hidden layer, and an output layer. The neural network constituting the device can infer a result (output) to be predicted from an arbitrary input by changing the weights of the neurons through learning.
[0043] The processor can create neural networks, train or learn neural networks, perform computations based on received input data, generate information signals based on the results of the computation, or retrain neural networks. Neural network models may include, but are not limited to, various types of models such as Convolutional Neural Networks (CNN), Region with Convolutional Neural Networks (R-CNN), Region Proposal Networks (RPN), Recurrent Neural Networks (RNN), Stacking-based Deep Neural Networks (S-DNN), State-Space Dynamic Neural Networks (S-SDNN), Deconvolution Networks, Deep Belief Networks (DBN), Restructured Boltzmann Machines (RBM), Fully Convolutional Networks, Long Short-Term Memory Networks (LSTM), and Classification Networks, such as GoogleNet, AlexNet, and VGG Network. The processor may include one or more processors to perform computations according to neural network models. For example, a neural network is a deep neural network It may include a (Deep Neural Network).
[0044] Neural networks include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), perceptron, multilayer perceptron, FF (Feed Forward), RBF (Radial Basis Function), DFF (Deep Feed Forward), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), AE (Auto Encoder), VAE (Variational Auto) Encoder), DAE (Denoising Auto Encoder), SAE (Sparse Auto Encoder), MC (Markov Chain), HN (Hopfield Network), BM (Boltzmann Machine), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), DCN (Deep Convolutional Network), DN (Deconvolutional Network), DCIGN (Deep Convolutional Inverse Graphics Network), GAN (Generative Adversarial Network), LSM (Liquid State Machine), ELM (Extreme Learning Machine), ESN (Echo It will be understood by a person skilled in the art that any neural network may be included, but is not limited to, State Network, Deep Residual Network, Differential Neural Computer, Neural Turning Machine, Capsule Network, Kohonen Network, and Attention Network.
[0045] According to exemplary embodiments of the present disclosure, the processor comprises a Convolutional Neural Network (CNN) such as GoogleNet, AlexNet, VGG Network, Region with Convolutional Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based Deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolution Network, Deep Belief Network (DBN), Restructured Boltzmann Machine (RBM), Fully Convolutional Network, Long Short-Term Memory (LSTM) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA, Text Analysis, Dialog System, GPT-3, GPT-4 for Natural Language Processing, Visual Analytics, Visual Understanding, Video Synthesis for Vision Processing, Anomaly Detection, Prediction, Time-Series Forecasting, Optimization for ResNet Data Intelligence, Various artificial intelligence structures and algorithms, such as recommendation and data creation, may be used, but are not limited thereto. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.
[0046] FIG. 1 is a block diagram briefly illustrating the configuration of an electronic device (100) that performs a data table-based syntax unit job suitability evaluation according to the present disclosure.
[0047] Referring to FIG. 1, the electronic device (100) according to the present disclosure may include an input / output module (110), a communication module (120), a memory (130), and a processor (140). Hereinafter, the electronic device (100) is an electronic device that performs a data table-based syntax unit job suitability evaluation, and the method thereof is assumed to be implemented through the electronic device (100).
[0048] The input / output module (110) may be various interfaces or connection ports that receive user input or output information to the user. The input / output module (110) may be divided into an input module and an output module.
[0049] The input module receives user input from the user. The input module is for inputting video information (or signals), audio information (or signals), data, or information input by the user, and may include at least one of at least one camera, at least one microphone, and a user input unit. Voice data or image data collected by the input unit may be analyzed and processed into a user control command.
[0050] User input can take various forms, including key input, touch input, and voice input. Examples of input modules capable of receiving such user input include traditional keypads, keyboards, and mice; as well as touch sensors that detect user touch; microphones that receive voice signals; cameras that recognize gestures through image recognition; proximity sensors consisting of light or infrared sensors that detect user approach; motion sensors that recognize user movements using accelerometers or gyroscopes; and all other diverse forms of input means that detect or receive various types of user input. This is a comprehensive concept.
[0051] Here, the touch sensor can be implemented as a piezoelectric or capacitive touch sensor that detects touch through a touch panel or touch film attached to the display panel, or as an optical touch sensor that detects touch by an optical method. In addition, the input module may be implemented in the form of an input interface (USB port, PS / 2 port, etc.) that connects an external input device to receive user input, instead of a device that detects user input itself.
[0052] The output module can output various types of information and provide it to the user. The output module is a comprehensive concept that includes a display for outputting video, a speaker for outputting sound (and / or an amplifier connected thereto), a haptic device for generating vibration, and various other forms of output means. In addition, the output module may be implemented in the form of a port-type output interface that connects the individual output means described above.
[0053] For example, an output module in the form of a display can display text, still images, and videos. The term "display" refers to a broad concept of an image display device that includes all types of devices capable of performing image output functions, such as Liquid Crystal Displays (LCDs), Light Emitting Diode (LED) displays, Organic Light Emitting Diode (OLED) displays, Flat Panel Displays (FPDs), transparent displays, Curved Displays, flexible displays, 1D displays, holographic displays, projectors, and others. Such a display may also take the form of a touch display integrated with the touch sensor of an input module.
[0054] In other words, the input / output module (110) can receive user input or provide output to the user based on a user interface.
[0055] The communication module (120) can communicate with an external device. Accordingly, the device can transmit and receive information with an external device through the communication module. For example, the device can communicate with an external device using the communication module so that information stored and generated within the electric vehicle charging management system is shared. The communication module (120) may include, for example, at least one of a wired communication module, a wireless communication module, a short-range communication module, and a location information module.
[0056] Here, communication, that is, the transmission and reception of data, can be performed via wired or wireless means. To this end, the communication module may be composed of a wired communication module that connects to the Internet, etc., via a Local Area Network (LAN); a mobile communication module that connects to a mobile communication network via a mobile communication base station to transmit and receive data; a short-range communication module that uses a Wireless Local Area Network (WLAN) family communication method such as Wi-Fi or a Wireless Personal Area Network (WPAN) family communication method such as Bluetooth or Zigbee; a satellite communication module that uses a Global Navigation Satellite System (GNSS) such as GPS; or a combination thereof. The wireless communication technology used for communication may include Narrowband Internet of Things (NB-IoT) for low-power communication. In this case, for example, NB-IoT technology may be an example of LPWAN (Low Power Wide Area Network) technology and may be implemented according to standards such as LTE Cat (category) NB1 and / or LTE Cat NB2, but is not limited to the names mentioned above. Additionally, or generally, wireless communication technology implemented in wireless devices according to various embodiments may perform communication based on LTE-M technology. In this case, for example, LTE-M technology may be an example of LPWAN technology and may be referred to by various names such as eMTC (enhanced Machine Type Communication).For example, LTE-M technology may be implemented in at least one of various standards such as 1) LTE CAT 0, 2) LTE Cat M1, 3) LTE Cat M2, 4) LTE non-BL (non-Bandwidth Limited), 5) LTE-MTC, 6) LTE Machine Type Communication, and / or 7) LTE M, and is not limited to the names mentioned above. Additionally or generally, wireless communication technology implemented in wireless devices according to various embodiments may include at least one of ZigBee, Bluetooth, and Low Power Wide Area Network (LPWAN) for low-power communication, and is not limited to the names mentioned above. As an example, ZigBee technology can create personal area networks (PANs) related to small / low-power digital communication based on various standards such as IEEE 802.15.4 and may be referred to by various names.
[0057] The memory (130) can store various types of information. The memory can store data temporarily or semi-permanently. For example, the memory may store an operating system (OS) for operating the first device and / or the second device, data for hosting a website, or data regarding a program or application (e.g., a web application) for generating Braille. In addition, the memory may store modules in the form of computer code as described above.
[0058] Examples of memory (130) may include a hard disk drive (HDD), a solid state drive (SSD), flash memory, ROM (Read-Only Memory), and RAM (Random Access Memory). These memories may be provided as built-in or removable types.
[0059] The processor (140) controls the overall operation of the electronic device (100). To this end, the processor (140) performs computation and processing of various information and can control the operation of the components of the first device and / or the second device.
[0060] The processor (140) may be implemented as a computer or a similar device according to hardware, software, or a combination thereof. Hardware-wise, the processor (140) may be provided in the form of an electronic circuit that processes electrical signals to perform control functions, and software-wise, it may be provided in the form of a program that drives the hardware processor. Meanwhile, unless otherwise specifically mentioned in the following description, the operation of the first device and / or the second device may be interpreted as being performed by the control of the processor (140). That is, the modules may be interpreted as the processor (140) controlling the first device and / or the second device to perform the following operations.
[0061] The processor (140) may be implemented with a memory that stores data for an algorithm or a program that reproduces the algorithm for controlling the operation of components within the device, and at least one sub-processor (not shown) that performs the aforementioned operation using the data stored in the memory. In this case, the memory and the processor may each be implemented as separate chips. Alternatively, the memory and the processor may be implemented as a single chip.
[0062] Additionally, the processor (140) can control one or a combination of the components described above in order to implement various embodiments according to the present disclosure, which will be described in the drawings below, on the device.
[0063] Hereinafter, an electronic device and method for performing a data table-based syntax unit job suitability evaluation according to one embodiment of the present disclosure will be described using FIGS. 2 to 12.
[0064] An electronic device (100) for performing a data table-based phrase-unit job suitability evaluation according to the present disclosure may include a memory (130) in which at least one process for performing a phrase-unit job suitability evaluation operation is stored, and at least one processor (140) for performing the job suitability evaluation operation according to said process.
[0065] As illustrated in FIG. 2, the electronic device (100) according to the present disclosure can extract keywords when recruitment document data, such as self-introduction data, is input as input data, identify valid keywords among the extracted keywords, and calculate a job suitability score based on the valid keywords.
[0066] In one embodiment, the electronic device (100) according to the present disclosure can extract clustering keywords without extracting words one by one to identify them as valid keywords. By clustering multiple words into a single cluster and adopting or deleting them as valid keywords, valid keywords that reflect the meaning of the entire input data can be identified.
[0067] In addition, as an embodiment, the electronic device (100) according to the present disclosure can calculate a job suitability score by analyzing the syntax in expanded syntax units using clustering keywords. Since the overall context may not be reflected when determining matching at the word level, matching can be determined in expanded syntax units based on clustering keywords.
[0068] Therefore, when extracting keywords, a keyword pool to compare with the input data must be built in advance.
[0069] As illustrated in FIG. 3, the at least one processor (140) can configure a standard dataset based on competency elements for each job group or job (S301), extract clustering keywords based on a plurality of clusters configured for each job group or job in the standard dataset (S302), and perform syntactic normalization based on semantic role analysis by classifying and labeling each text in the standard dataset according to a preset semantic role (S303), and can build a keyword pool using the clustering keywords and the labeled text. In other words, it can perform keyword dictionary generation (S304), which is reference data that is compared with input data.
[0070] Additionally, the at least one processor (140) may be configured to evaluate job suitability based on the degree of syntactic matching with the keyword pool for the input data.
[0071] Regarding step (S301), when constructing a standard dataset based on competency elements by job group or job function, the competency elements may include knowledge, skills, attitudes, and performance criteria related to the job group or job function. Job description data according to the job group or job classification system is collected in accordance with the competency unit structure presented in the National Competency Standards (NCS), and the applicant's competency elements required by each job description are extracted through standardization.
[0072] For example, job categories can be classified into Management / Office, Sales / Marketing, Public / Service, Manufacturing / Research, ICT, Design, and Production / Maintenance, and job functions, taking Management / Office as an example, can be classified into Administration / Office Management, Planning, HR, Finance / Accounting / Investment, Purchasing / Logistics, Legal / Auditing, and Others.
[0073] At least one processor (140) can collect and analyze job descriptions including knowledge, skills, attitudes, and performance criteria according to classified job groups and duties, and generate a standard dataset including standard data for each job group and duty.
[0074] Referring to FIG. 4, at least one processor (140) can collect an NCS-based job description (11) as in FIG. 4 and analyze job performance content as a performance criterion, necessary knowledge as knowledge, necessary skills as skills, and job performance attitude as an attitude.
[0075] For example, competency elements such as knowledge, skills, and attitudes refer to the knowledge, skills, and attitudes required to perform the relevant job, while competency elements such as performance criteria refer to an overall description of the job performance, which may include job performance content and evaluation items.
[0076] By analyzing the NCS-based job description (11), standard data (13) for the job can be processed and generated as shown in FIG. 5. The NCS-based job description (11) can be standardized into standard data (13) in a consistent form that includes a recruitment-related title, job group classification, job classification, job performance content, required knowledge, required skills, and job performance attitude.
[0077] Therefore, a standard dataset can be constructed by collecting standardized data (13) for each job group and job function, and statistical analysis techniques can be applied to it to extract useful job function or job group related keywords.
[0078] At least one processor (140) according to one embodiment of the present disclosure may be configured to preprocess by labeling the part of speech of words in a document corresponding to each job in the standard dataset, and to cluster the words into a plurality of clusters based on word frequency analysis in the document corresponding to each job.
[0079] Specifically, morphological analysis can be performed on text included in standard datasets or input data. Morphological analysis is the process of identifying the part of speech of each word and vocabulary, which can divide sentences into morpheme units and map the part of speech of each morpheme.
[0080] You can retain only meaningful words by performing filtering based on parts of speech on morphologically analyzed standard datasets or input data. You can keep words primarily mapped to common nouns and filter out the rest.
[0081] In one embodiment, determiners (MM), adverbs (MA), pronouns (NP), numerals (NR), adjectives (VA), verbs (VV), or roots (XR) may be removed as keywords because they are unclear as keywords for knowledge and technical competency elements.
[0082] In addition, morphemes shorter than 2 syllables can be removed as keywords. Or, noun derivation suffixes (XSN) including ...적, ...화 can be identified as separate tags and integrated with nouns.
[0083] Through the processing rules described above, meaningless words can be removed from the keyword pool, and mis-segmented or over-segmented words can be processed.
[0084] For example, at least one processor (140) can preprocess the sentence “Python programming skills are required for data analysis, and creative and analytical thinking is needed” within a standard dataset or input data through the above processing rule to leave only the words [data, analysis, programming, skills, requirements, creative, analytical, thinking, need].
[0085] At least one processor (140) can generate a keyword frequency table or a keyword visualization table (21) as illustrated in FIG. 6 using preprocessed words of a standard dataset. A keyword frequency table can be generated for each job and competency element, or a keyword visualization can be generated to be used for subsequent keyword extraction.
[0086] Referring to Figure 6, keywords preprocessed in IT development roles—experience, knowledge, communication, related, security, about, network, server, DB, Cloud—along with their indices and the number of times they appear in a standard dataset (#count) can be viewed in a keyword frequency table. Additionally, to allow users to easily verify the information, the data can be visualized and displayed in a table with different sizes or colors depending on the frequency.
[0087] At least one processor (140) can form multiple clusters using a keyword frequency table for one job and competency element.
[0088] In other words, by clustering keywords that have been filtered once based on parts of speech using a keyword frequency table, words that are frequently used in specific job and competency elements, words that are used restrictively, and words that are used universally can be distinguished, and words that are not important in those job and competency elements can be filtered twice.
[0089] Keyword clustering can be performed using the Exploratory Factor Analysis (EFA) or Principal Component Analysis (PCA) methods.
[0090] EFA is an exploratory factor analysis method that analyzes job factors distinguished by the frequency of construct keywords. PCA is a principal component analysis method that analyzes job characteristics correlated with frequency. In other words, it allows for the statistical clustering of keywords reflecting the potential characteristics of a job based on their frequency.
[0091] At least one processor (140) can generate clustering keywords by performing word frequency-based clustering or embedding-based filtering on preprocessed words.
[0092] In one embodiment, the at least one processor (140) may be configured to calculate the TF-IDF (Term Frequency-Inverse Document Frequency) and IDF (Inverse Document Frequency) of each word within the plurality of clusters, remove words within the plurality of clusters that have an IDF lower than a preset threshold value for sparse words, remove the top m clusters containing many words that have a TF / IDF lower than a preset threshold value for job-specific words, and then extract clustering keywords, thereby enabling word frequency-based filtering. In this case, m is a positive integer.
[0093] This allows for a first filtering of meaningless words based on parts of speech for text in a standard dataset, and a second filtering of clusters containing clustering keywords that are mentioned below a certain level within the relevant job field from clusters generated based on statistical techniques including EFA or PCA according to word frequency.
[0094] In other words, by filtering the cluster itself, you can coarsely filter out keywords that are not important in specific job fields and competency elements and are rarely mentioned.
[0095] In one embodiment, the at least one processor (140) may be configured to obtain at least one hint keyword for each job, remove the top n clusters containing many words with embedding vectors that differ from the embedding vector of the hint keyword by more than a preset job threshold value, and then extract a clustering keyword, thereby enabling embedding-based filtering. In this case, n is a positive integer.
[0096] In the word embedding process, keywords can be mapped into a vector space using embedding models such as Word2Vec, GloVe, or BERT to quantify semantic relationships and calculate distances.
[0097] Hint keywords are selected core keywords that are typically used in specific job fields or evaluation factors, and the embedding vector values of the hint keywords can be calculated and stored in advance.
[0098] Embedding-based filtering can be performed to consider even meaningful keywords that are not frequently mentioned and have a low frequency.
[0099] Accordingly, at least one processor (140) can perform word frequency-based filtering and embedding-based filtering serially in sequence.
[0100] An electronic device (100) according to one embodiment of the present disclosure can perform word frequency-based filtering and embedding-based filtering to select clustering keywords that can provide a large amount of information for judging the quality of applicants, and can perform syntactic analysis based on clustering keywords.
[0101] As shown in Fig. 7, multiple clusters are automatically grouped according to their respective features through data mining, and if a cluster is determined to contain clustering keywords with low information content in a job description of a specific job field, that cluster is removed so that only keywords suitable for job evaluation can be finally extracted.
[0102] As shown in the cluster and factor table (23) in FIG. 7, the clusters to be removed can be determined by calculating the association between cluster groups 1, 2, and 3 containing clustering keywords and factors reflecting specific potential characteristics of the job using word frequency-based filtering and embedding-based filtering.
[0103] By using a keyword frequency table or keyword visualization table (21), the clusters and factors (23) can be organized, and word frequency-based filtering and embedding-based filtering can be performed on them to determine the clustering keywords to be removed and the clustering keywords to be used.
[0104] As one embodiment, a stop word dictionary may be stored in a dictionary, and clusters containing keywords in the stop word dictionary may be removed.
[0105] Meanwhile, since a method of simply matching text and keywords cannot reflect contextual information and cannot identify synonyms, an electronic device (100) according to one embodiment of the present disclosure can expand clustering keywords grouped into clusters into phrase units.
[0106] In other words, to reflect semantic elements that are difficult to reflect with keywords alone, a syntactic analysis method can be applied that groups multiple words into a single keyword.
[0107] Due to the nature of job descriptions, individual text units may not be distinguished by punctuation or sentence-ending forms, and there may be numerous instances where they are composed in an ungrammatical format or contain multiple phrases in parallel or nested.
[0108] Therefore, in order to effectively analyze the competency elements inherent in each phrase, it is necessary to classify and organize each phrase according to semantic criteria before applying syntactic analysis methods.
[0109] The above at least one processor (140) may be configured to use a machine learning model to divide each text into phrases according to semantic criteria before classifying and labeling each text in the standard dataset according to a preset semantic role.
[0110] As illustrated in FIG. 8, phrase division can be set up differently for phrase units corresponding to a single semantic role. In the phrase division table (31), 'before division' is listed as a single line without phrase separation, but 'after division' is listed on separate lines considering the meaning of each phrase, so that the overall context can be checked more easily.
[0111] Then, at least one processor (140) may be configured to label corresponding semantic roles for segmented phrases in job-specific text using an artificial intelligence model trained with training data that labels word sequences and semantic roles of said word sequences.
[0112] In this case, the word sequence may refer to part or all of the separated phrases depending on the context.
[0113] For example, 'data processing capability is required for the implementation of artificial intelligence services' may be expressed with slight variations, such as 'requires data processing capability for the implementation of artificial intelligence services', 'data processing capability required for the implementation of artificial intelligence services', 'ability to implement artificial intelligence services through data processing', or 'a person capable of implementing artificial intelligence services by possessing data processing capability'.
[0114] In other words, when expressing the objective of “implementing AI services” and the capability of “data processing ability” in a job description, they may be expressed in the passive voice, such as “required,” or the active voice, such as “requires,” or the order of the object and predicate may be reversed, specific phrases may be omitted, or the subject described may be changed, such as “a capable person.”
[0115] Even when expressing the same purpose and capability, the sentences expressing it may be converted into various forms, but the semantic roles that are meaningful for evaluating job suitability are the same, so at least one processor (140) can define semantic roles and extract or generate phrases corresponding to semantic roles.
[0116] As illustrated in FIG. 9, at least one processor (140) can classify three semantic roles specialized for a self-introduction letter—behavior / experience, competence / characteristics, and purpose / condition—and define descriptions for them.
[0117] In the semantic role definition and example table (33), the range to be expressed by the text corresponding to the semantic role is defined, and the phrase corresponding to the semantic role is extracted from the input data to label the phrase and the semantic role.
[0118] As an example, to prevent the same syntax from being judged differently depending on some word differences, input data can be preprocessed and normalized at the morpheme level, and then semantic role labeling can be performed.
[0119] As an example, even in sentences expressed with some modifications, such as 'data processing capability is required for the implementation of artificial intelligence services', 'requires data processing capability for the implementation of artificial intelligence services', 'data processing capability required for the implementation of artificial intelligence services', 'ability to implement artificial intelligence services through data processing', and 'a person capable of implementing artificial intelligence services by possessing data processing capability', if preprocessed and normalized at the morpheme level and then subjected to semantic role analysis, they can be analyzed identically as 'artificial intelligence service implementation-purpose' and 'data processing capability-competency'.
[0120] In addition, an artificial intelligence model trained with training data labeled with word sequences and semantic roles of said word sequences can perform semantic role analysis functions by training machine learning models such as SVM and CRF, or deep learning models such as LSTM and BERT, to infer semantic roles.
[0121] The above artificial intelligence model can perform supervised learning using word sequences and the semantic roles of the word sequences as training data.
[0122] The above artificial intelligence model can identify input data as a sequence of tokens, output the semantic role of each token to delineate boundaries, or output special tokens that identify semantic roles between tokens.
[0123] As one embodiment, it was described above that semantic role identification is performed after syntactic partitioning, but as another embodiment, semantic role identification may be performed simultaneously with syntactic partitioning.
[0124] Additionally, the at least one processor (140) may be configured to convert the plurality of divided phrases into morpheme units and arrange and output them in parallel when the same semantic role is labeled for the plurality of divided phrases.
[0125] For example, as illustrated in FIG. 9, in the semantic role definition and example table (33), 'in performing work' can be labeled as a purpose / condition and communication ability as a competency / characteristic, and the same behavior / experience can be labeled as 'understanding what another person intended by reading and listening to text and speech' and 'accurately writing or speaking what one intended through text and speech'.
[0126] At this time, phrases labeled as actions / experiences can be parallelized and displayed, and each phrase can be converted into a morpheme unit and output.
[0127] Accordingly, through the process described above, at least one processor (140) can obtain defined semantic roles, labeled phrases, and clustering keywords.
[0128] The above at least one processor (140) may be configured to construct a keyword pool in the form of a data table using the type of document containing text, job, evaluation elements of the job, clustering keywords by job, phrases containing the clustering keywords, and semantic roles labeled in the phrases, and to analyze input data in the form of the data table to evaluate job suitability according to the degree of matching at the phrase unit with the keyword pool.
[0129] As illustrated in FIGS. 10 and 11, a data table (41) of a standard dataset and a data table (43) of input data can be generated. A list of keywords for job and competency elements can be constructed in the form of a data table by dividing the phrases and normalizing them according to semantic roles.
[0130] Through this, it is possible to evaluate job suitability by matching key phrases, semantic roles, and clustering keywords to document types, job functions, and competency elements, thereby considering the overall context rather than simple word matching.
[0131] In addition, regarding the evaluation logic for evaluating job suitability, the electronic device (100) according to the present disclosure constructs a data table in phrase units, so it is not possible to calculate a job suitability score by applying weights to each word.
[0132] Accordingly, the electronic device (100) according to the present disclosure can calculate a job suitability score by applying a scoring method for phrases and applying phrase matching degree and weights in a data table.
[0133] In other words, the at least one processor (140) may be configured to calculate a similarity score in morpheme units for phrases with identical labeled semantic roles by comparing the input data with the data table of the keyword pool, and to calculate a job suitability for any job group by applying a job-specific weight corresponding to the labeled semantic role to the similarity score.
[0134] For example, if the input data is a self-introduction letter, the similarity at the morpheme level can be compared for phrases with identical semantic roles by comparing the self-introduction letter with a data table. A similarity score can be calculated by determining how many morphemes overlap based on the phrases in the data table.
[0135] In this case, the Longest Common Subsequence (LCS) metric can be used as a method to calculate similarity at the morpheme level.
[0136] In addition, since each semantic role may have different importance depending on the job and evaluation factors, different weights according to the semantic role can be assigned to the similarity score.
[0137] For example, when comparing the standard data set and the input data using the LCS metric, LCS[Proactive, Feedback, Reflection],[Proactive, Feedback, Reflection, Attitude]=3, and Score_CompetencyElement=LCS / len[Proactive, Feedback, Reflection, Attitude]=3 / 4.
[0138] When active, feedback, and reflection are keywords for 'competency' among competency elements and the weight for competency / characteristic is defined as 0.9, the phrase matching score can be calculated as 3 / 4 * 0.9 = 0.63.
[0139] The final job suitability score can be calculated based on the matching score of each phrase; in this case, if there are multiple jobs within a single job group, the final job suitability score can be calculated by applying different job weights to each job.
[0140] As shown in Fig. 12, the final job suitability score can be calculated by considering the job to be judged and all similar jobs according to the job weight for each job.
[0141] For example, when a job suitability score of 0.7 is calculated for a 'communication' job and 0.5 for an 'AI / data' job, a specific job suitability score can be obtained as 4.4*0.7+0.28*0.5=3.22 for the 'communication' job and 0.24*0.7+4.4*0.5=2.368 for the 'AI / data' job.
[0142] The final job suitability score can be calculated by adjusting based on the reflection ratio between job functions.
[0143] Meanwhile, the method for performing a data table-based phrase-unit job suitability evaluation according to the present disclosure may be performed by a computing device comprising a memory (130) storing at least one process for performing a phrase-unit job suitability evaluation operation and at least one processor (140) for performing the job suitability evaluation operation according to said process.
[0144] The above method may include the steps of: configuring a standard dataset based on competency elements for each job group or job function; extracting clustering keywords based on a plurality of clusters configured for each job group or job function in the standard dataset; classifying and labeling each text in the standard dataset according to a pre-set semantic role; constructing a keyword pool using the clustering keywords and the labeled text; and evaluating job suitability for input data based on the degree of syntactic matching with the keyword pool.
[0145] Content that overlaps with the above is omitted for the sake of brevity in the specification.
[0146] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium that stores instructions executable by a computer. The instructions may be stored in the form of program code and, when executed by a processor, may generate a program module to perform the operation of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.
[0147] Computer-readable recording media include all types of recording media that store instructions that can be decoded by a computer. Examples include ROM (Read Only Memory), RAM (Random Access Memory), magnetic tape, magnetic disk, flash memory, optical data storage devices, etc.
[0148] As described above, the disclosed embodiments have been explained with reference to the attached drawings. Those skilled in the art will understand that the present disclosure may be practiced in forms different from the disclosed embodiments without changing the technical spirit or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be interpreted restrictively.
Claims
1. Memory storing at least one process for performing a job suitability evaluation operation at the syntax unit level; and It includes at least one processor that performs the job suitability evaluation operation according to the above process; and the at least one processor, Construct a standard dataset based on competency elements by job group or function, and Clustering keywords are extracted based on multiple clusters configured for each job group or job function in the above standard dataset, and In the above standard dataset, each text is labeled by distinguishing it according to a pre-set semantic role, and A keyword pool is constructed using the clustering keywords and the labeled text above, and Configured to evaluate job suitability based on the degree of syntactic matching with the aforementioned keyword pool for input data An electronic device that performs data table-based syntax-unit job fit evaluation.
2. In Paragraph 1, The above-mentioned at least one processor is, Preprocess the words within the documents corresponding to each job in the above standard dataset by labeling their parts of speech, and Configured to cluster the said words into multiple clusters based on word frequency analysis within documents corresponding to each job function. An electronic device that performs data table-based syntax-unit job fit evaluation.
3. In Paragraph 2, The above-mentioned at least one processor is, Calculate the TF-IDF (Term Frequency-Inverse Document Frequency) and IDF (Inverse Document Frequency) of each word within the above plurality of clusters, and Words with an IDF lower than the preset threshold value of sparse words are removed from the aforementioned multiple clusters, and It is configured to extract clustering keywords after removing the top m clusters containing many words with TF / IDF lower than the preset word threshold value by job, where m is a positive integer. An electronic device that performs data table-based syntax-unit job fit evaluation.
4. In Paragraph 2, The above-mentioned at least one processor is, Obtain at least one hint keyword per job, and It is configured to extract clustering keywords after removing the top n clusters containing many words with embedding vectors that differ from the embedding vector of the hint keyword by more than a preset job standard value, where n is a positive integer. An electronic device that performs data table-based syntax-unit job fit evaluation.
5. In Paragraph 1, The above-mentioned at least one processor is, Before classifying and labeling each text in the above standard dataset according to pre-set semantic roles, Configured to split each text into phrases based on semantic criteria using a machine learning model An electronic device that performs data table-based syntax-unit job fit evaluation.
6. In Paragraph 5, The above-mentioned at least one processor is, Configured to label corresponding semantic roles for segmented phrases in job-specific text using an artificial intelligence model trained on training data labeled with word sequences and the semantic roles of said word sequences. An electronic device that performs data table-based syntax-unit job fit evaluation.
7. In Paragraph 6, The above-mentioned at least one processor is, When the same semantic role is labeled for multiple divided phrases, the multiple divided phrases are converted into morpheme units and arranged and output in parallel. An electronic device that performs data table-based syntax-unit job fit evaluation.
8. In Paragraph 1, The above-mentioned at least one processor is, A keyword pool in the form of a data table is constructed using the type of document containing text, job, evaluation factors of the said job, clustering keywords by job, phrases containing the said clustering keywords, and semantic roles labeled in the said phrases, and Configured to analyze input data in the form of the above data table and evaluate job suitability based on the degree of phrase-level matching with the above keyword pool. An electronic device that performs data table-based syntax-unit job fit evaluation.
9. In Paragraph 8, The above-mentioned at least one processor is, By comparing the above input data with the data table of the above keyword pool, a similarity score is calculated in morpheme units for phrases with identical labeled semantic roles, and Configured to calculate job fit for any job group by applying job-specific weights corresponding to semantic roles labeled in the aforementioned similarity score. An electronic device that performs data table-based syntax-unit job fit evaluation.
10. A method for performing a data table-based syntax-unit job suitability evaluation performed by a processor of a device, Step of constructing a standard dataset based on competency elements by job group or function; A step of extracting clustering keywords based on a plurality of clusters configured for each job group or job function in the above standard dataset; A step of classifying and labeling each text in the above standard dataset according to a pre-set semantic role; A step of constructing a keyword pool using the clustering keyword and the labeled text; and A step of evaluating job suitability based on the degree of syntactic matching with the keyword pool for the input data; Method for performing a data table-based syntax-unit job suitability assessment.
Citation Information
Patent Citations
Human resource competence evaluation system and method
KR101235292B1
Method for resizing window area and electronic device for the same
KR1020200087742A
KR20220129713A
KR20220167608A
KR20240113342A