Electronic apparatus and method for extracting detailed information from text-based employment document
Patent Information
- Application Number
- KR1020240082474
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2044-06-25
Smart Images

Figure 112024068535737-PAT00002_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to an apparatus and method for extracting data from a document, and more specifically, to an electronic apparatus and method for extracting detailed information within a text-based recruitment document. Background Technology
[0002] Conventional recruitment systems faced the problem of incurring significant time and costs because HR personnel had to manually read applicants' documents, extract necessary information, and organize it. This was particularly problematic in the case of open recruitment by large corporations attracting over 10,000 applicants, where the human resource costs and workload for processing were substantial, and maintaining the consistency of extracted information was difficult depending on the proficiency of the HR personnel.
[0003] To address this, machine models were used to extract keyword-based information from recruitment documents; however, utilizing keywords required the construction of a prior database, and there were issues where meaningless information was extracted because keywords varied depending on the job category or function in the recruitment market.
[0004] In addition, as text processing became difficult due to the diverse file extensions of recruitment documents, image-based OCR was sometimes performed; however, it was difficult to extract specific information to the desired level due to typos and inaccurate data extraction resulting from the OCR.
[0005] In conventional patent literature, an item extraction model corresponding to the fingerprint was used by utilizing the fingerprint within each recruitment document, but this was also difficult to process if an item extraction model matching each fingerprint was not prepared in advance.
[0006] Therefore, the need has arisen for an invention that extracts specific desired information based on text from recruitment documents written in various formats or various extensions. Prior art literature
[0007] Republic of Korea Registered Publication 10-2601932(2023.11.09) The problem to be solved
[0008] To solve conventional problems, the embodiments disclosed in this disclosure aim to provide an electronic device and method for performing detailed information extraction within text-based recruitment documents, which improves extraction accuracy by labeling overall key areas of unstructured recruitment documents and then extracting specific detailed information.
[0009] The problems that this disclosure aims to solve are not limited to those mentioned above, and other unmentioned problems will be clearly understood by a person skilled in the art from the description below. means of solving the problem
[0010] An electronic device for performing detailed information extraction within a text-based recruitment document according to the present disclosure for achieving the aforementioned technical problem comprises: an artificial intelligence model that receives an unstructured recruitment document as input and outputs a structured recruitment document; a memory storing at least one process for performing a structured recruitment document generation operation to enable information extraction; and at least one processor that performs the structured recruitment document generation operation according to the process. The at least one processor may be configured to divide the text within the input recruitment document into pages, preprocess the divided pages into pre-set units, input the pre-processed pages into the artificial intelligence model to output a page in which area names are labeled in key areas, and extract detailed information based on a pre-set prompt for the output page to output a structured recruitment document.
[0011] At least one processor according to one embodiment of the present disclosure may be configured to extract text from the employment document through a document filter module before dividing the text within the input employment document into pages.
[0012] According to one embodiment of the present disclosure, the recruitment document is written with at least one of a plurality of extensions, and the at least one processor may be configured to extract text by operating a text extraction library corresponding to the extension based on a meta tag of the extension.
[0013] At least one processor according to one embodiment of the present disclosure is configured to post-process the different area name as the first area name when a plurality of area names labeled in the main areas are identical as the first area name and a different area name exists between the plurality of area names, and the first area name may be a self-introduction letter.
[0014] At least one processor according to one embodiment of the present disclosure may be configured to output detailed information among at least one of personal information, educational background information, and self-introduction information from the output page using at least one of a first prompt for extracting personal information, a second prompt for extracting educational background information, and a third prompt for extracting self-introduction information.
[0015] According to one embodiment of the present disclosure, the at least one processor may be configured to output the detailed information in JSON (JavaScript Object Notation) format.
[0016] According to one embodiment of the present disclosure, the at least one processor may be configured to perform preprocessing by giving a greater weight to the subsequent page than to the preceding page for the divided pages and summing the divided pages into three groups.
[0017] According to one embodiment of the present disclosure, the at least one processor may be configured to label area names in the main areas within the three groups according to the extent that it includes information corresponding to at least one of the first prompt, the second prompt, and the third prompt.
[0018] According to one embodiment of the present disclosure, the at least one processor may be configured to provide a standardized recruitment document to a user interface by aggregating a key-value pair in which the labeled area name and the detailed information are matched, or a list containing the key-value pair.
[0019] In addition, a method for extracting detailed information within a text-based recruitment document, performed by a computing device comprising: an artificial intelligence model that receives an unstructured recruitment document according to the present disclosure and outputs a structured recruitment document; a memory; and at least one processor, for achieving the technical problem described above, wherein the method may include: a step of dividing the text within the input recruitment document into page units and preprocessing the divided pages into pre-set units; a step of inputting the pre-processed pages into the artificial intelligence model to output a page in which area names are labeled in key areas; and a step of extracting detailed information based on a pre-set prompt for the output page to output a structured recruitment document.
[0020] In addition to this, a computer program stored on a computer-readable recording medium for implementing the present disclosure may be further provided.
[0021] In addition to this, a computer-readable recording medium for recording a computer program for implementing the present disclosure may be further provided. Effects of the invention
[0022] According to the aforementioned means for solving the problem of the present disclosure, the effect of accurately extracting detailed information about employment documents having various extensions is provided even without a prior database construction process related to the format of the employment documents.
[0023] In addition, according to the aforementioned means for solving the problem of the present disclosure, by training an artificial intelligence model to extract detailed information based on prompts after labeling key areas, it is possible to extract detailed information more accurately and effectively compared to simply inputting the entire recruitment document.
[0024] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing
[0025] FIG. 1 is a block diagram briefly illustrating the configuration of an electronic device for extracting detailed information from text-based recruitment documents according to one embodiment of the present disclosure. FIG. 2 is a process diagram illustrating a detailed information extraction process of an electronic device that performs detailed information extraction within a text-based recruitment document according to one embodiment of the present disclosure. FIG. 3 is a conceptual diagram illustrating a text preprocessing process of an electronic device that performs detailed information extraction within a text-based recruitment document according to one embodiment of the present disclosure. FIG. 4 is a diagram illustrating the input / output process of an artificial intelligence model of an electronic device that performs detailed information extraction within a text-based recruitment document according to one embodiment of the present disclosure. FIGS. 5 to 7 are embodiments illustrating area-specific prompts of an electronic device for extracting detailed information within a text-based recruitment document according to one embodiment of the present disclosure. FIG. 8 is a drawing illustrating labeled text of an electronic device that performs detailed information extraction within a text-based recruitment document according to one embodiment of the present disclosure. FIG. 9 is a table illustrating area-specific detailed information used in an electronic device that performs detailed information extraction within a text-based recruitment document according to one embodiment of the present disclosure. FIG. 10 is a flowchart illustrating the flow of a method for extracting detailed information from text-based recruitment documents according to one embodiment of the present disclosure. Specific details for implementing the invention
[0026] Throughout this disclosure, the same reference numerals denote the same components. This disclosure does not describe all elements of the embodiments, and general content in the art to which this disclosure pertains or content that overlaps between embodiments is omitted. The terms 'part, module, component, block' as used in the specification may be implemented in software or hardware, and depending on the embodiments, a plurality of 'parts, modules, components, blocks' may be implemented as a single component, or a single 'part, module, component, block' may include a plurality of components.
[0027] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are directly connected but also cases where they are indirectly connected, and indirect connections include connections made via a wireless communication network.
[0028] Furthermore, when it is stated that a part "includes" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0029] Throughout the specification, when it is stated that a component is located "on" another component, this includes not only cases where a component is in contact with another component, but also cases where another component exists between the two components.
[0030] The terms first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.
[0031] Singular expressions include plural expressions unless there is an obvious exception in the context.
[0032] In each step, identification codes are used for convenience of explanation and do not describe the order of the steps; the steps may be performed differently from the specified order unless a specific order is clearly indicated in the context.
[0033] The operating principles and embodiments of the present disclosure will be described below with reference to the attached drawings.
[0034] In this specification, the term "device according to the present disclosure" includes all various devices capable of performing computational processing and providing results to a user. For example, the device according to the present disclosure may include all of a computer, a server device, and a portable terminal, or may be in the form of any one of these.
[0035] Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.
[0036] The above server device is a server that processes information by communicating with an external device, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.
[0037] The above portable terminal may include, for example, all types of handheld-based wireless communication devices such as PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminals, smartphones, etc., as well as wearable devices such as watches, rings, bracelets, anklets, necklaces, glasses, contact lenses, or head-mounted devices (HMDs).
[0038] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0039] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0040] The predefined operating rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined operating rules or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0041] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0042] According to an exemplary embodiment of the present disclosure, a processor can implement artificial intelligence. Artificial intelligence refers to a machine learning method based on an artificial neural network that enables a machine to learn by mimicking human biological neurons. Methodologies of artificial intelligence can be classified according to the learning method into supervised learning, where input and output data are provided together as training data and the solution (output data) to the problem (input data) is predetermined; unsupervised learning, where only input data is provided without output data and the solution (output data) to the problem (input data) is not predetermined; and reinforcement learning, where a reward is given from an external environment whenever an action is taken from the current state, and learning proceeds in a direction that maximizes such reward. In addition, artificial intelligence methodologies can be classified according to the architecture, which is the structure of the learning model. The architectures of widely used deep learning technologies can be classified into Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Transformers, and Generative Adversarial Networks (GAN).
[0043] The device and system may include an artificial intelligence model. The artificial intelligence model may be a single model or may be implemented as multiple models. The artificial intelligence model may be composed of a neural network (or artificial neural network) and may include statistical learning algorithms that mimic biological neurons in machine learning and cognitive science. A neural network may refer to a model that possesses problem-solving capabilities by having artificial neurons (nodes) that form a network through synaptic connections and change the strength of synaptic connections through learning. The neurons of a neural network may include combinations of weights or biases. A neural network may include one or more layers composed of one or more neurons or nodes. For example, the device may include an input layer, a hidden layer, and an output layer. The neural network constituting the device can infer a result (output) to be predicted from an arbitrary input by changing the weights of the neurons through learning.
[0044] The processor can create neural networks, train or learn neural networks, perform computations based on received input data, generate information signals based on the results of the computation, or retrain neural networks. Neural network models may include, but are not limited to, various types of models such as Convolutional Neural Networks (CNN), Region with Convolutional Neural Networks (R-CNN), Region Proposal Networks (RPN), Recurrent Neural Networks (RNN), Stacking-based Deep Neural Networks (S-DNN), State-Space Dynamic Neural Networks (S-SDNN), Deconvolution Networks, Deep Belief Networks (DBN), Restructured Boltzmann Machines (RBM), Fully Convolutional Networks, Long Short-Term Memory Networks (LSTM), and Classification Networks, such as GoogleNet, AlexNet, and VGG Network. The processor may include one or more processors to perform computations according to neural network models. For example, a neural network is a deep neural network It may include a (Deep Neural Network).
[0045] Neural networks include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), perceptron, multilayer perceptron, FF (Feed Forward), RBF (Radial Basis Network), DFF (Deep Feed Forward), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), AE (Auto Encoder), VAE (Variational Auto) Encoder), DAE (Denoising Auto Encoder), SAE (Sparse Auto Encoder), MC (Markov Chain), HN (Hopfield Network), BM (Boltzmann Machine), RBM (Restricted Boltzmann Machine), DBN (Depp Belief Network), DCN (Deep Convolutional Network), DN (Deconvolutional Network), DCIGN (Deep Convolutional Inverse Graphics Network), GAN (Generative Adversarial Network), LSM (Liquid State Machine), ELM (Extreme Learning Machine), ESN (Echo It will be understood by a person skilled in the art that any neural network may be included, but is not limited to, State Network, Deep Residual Network, Differential Neural Computer, Neural Turning Machine, Capsule Network, Kohonen Network, and Attention Network.
[0046] According to an exemplary embodiment of the present disclosure, the processor comprises a Convolutional Neural Network (CNN) such as GoogleNet, AlexNet, VGG Network, Region with Convolutional Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based Deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolution Network, Deep Belief Network (DBN), Restructured Boltzmann Machine (RBM), Fully Convolutional Network, Long Short-Term Memory (LSTM) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA, Text Analysis, Dialog System, GPT-3, GPT-4 for Natural Language Processing, Visual Analytics, Visual Understanding, Video Synthesis for Vision Processing, Anomaly Detection, Prediction, Time-Series Forecasting, Optimization for ResNet Data Intelligence, Various artificial intelligence structures and algorithms, such as recommendation and data creation, may be used, but are not limited thereto. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.
[0047] FIG. 1 is a block diagram briefly illustrating the configuration of an electronic device (100) that performs the extraction of detailed information within a text-based recruitment document according to one embodiment of the present disclosure.
[0048] Referring to FIG. 1, the electronic device (100) according to the present disclosure may include an input / output module (110), a communication module (120), a memory (130), and a processor (140). Hereinafter, the electronic device (100) is an electronic device that performs a standardized recruitment document generation operation to enable information extraction, and the method for extracting detailed information within a text-based recruitment document of the electronic device (100) is assumed to be implemented through the electronic device (100) that performs a standardized recruitment document generation operation to enable information extraction.
[0049] The input / output module (110) may be various interfaces or connection ports that receive user input or output information to the user. The input / output module (110) may be divided into an input module and an output module.
[0050] The input module receives user input from the user. The input module is for inputting video information (or signals), audio information (or signals), data, or information input by the user, and may include at least one of at least one camera, at least one microphone, and a user input unit. Voice data or image data collected by the input unit may be analyzed and processed into a user control command.
[0051] User input can take various forms, including key input, touch input, and voice input. Examples of input modules capable of receiving such user input include traditional keypads, keyboards, and mice; as well as touch sensors that detect user touch; microphones that receive voice signals; cameras that recognize gestures through image recognition; proximity sensors consisting of light or infrared sensors that detect user approach; motion sensors that recognize user movements using accelerometers or gyroscopes; and all other diverse forms of input means that detect or receive various types of user input. This is a comprehensive concept.
[0052] Here, the touch sensor can be implemented as a piezoelectric or capacitive touch sensor that detects touch through a touch panel or touch film attached to the display panel, or as an optical touch sensor that detects touch by an optical method. In addition, the input module may be implemented in the form of an input interface (USB port, PS / 2 port, etc.) that connects an external input device to receive user input, instead of a device that detects user input itself.
[0053] The output module can output various types of information and provide it to the user. The output module is a comprehensive concept that includes a display for outputting video, a speaker for outputting sound (and / or an amplifier connected thereto), a haptic device for generating vibration, and various other forms of output means. In addition, the output module may be implemented in the form of a port-type output interface that connects the individual output means described above.
[0054] For example, an output module in the form of a display can display text, still images, and videos. The term "display" refers to a broad concept of an image display device that includes all types of devices capable of performing image output functions, such as Liquid Crystal Displays (LCDs), Light Emitting Diode (LED) displays, Organic Light Emitting Diode (OLED) displays, Flat Panel Displays (FPDs), transparent displays, Curved Displays, flexible displays, 1D displays, holographic displays, projectors, and others. Such a display may also take the form of a touch display integrated with the touch sensor of an input module.
[0055] In other words, the input / output module (110) can receive user input or provide output to the user based on a user interface.
[0056] The communication module (120) can communicate with an external device. Accordingly, the device can transmit and receive information with an external device through the communication module. For example, the device can communicate with an external device using the communication module so that information stored and generated within the electric vehicle charging management system is shared. The communication module (120) may include, for example, at least one of a wired communication module, a wireless communication module, a short-range communication module, and a location information module.
[0057] Here, communication, that is, the transmission and reception of data, can be performed via wired or wireless means. To this end, the communication module may be composed of a wired communication module that connects to the Internet, etc., via a Local Area Network (LAN); a mobile communication module that connects to a mobile communication network via a mobile communication base station to transmit and receive data; a short-range communication module that uses a Wireless Local Area Network (WLAN) family communication method such as Wi-Fi or a Wireless Personal Area Network (WPAN) family communication method such as Bluetooth or Zigbee; a satellite communication module that uses a Global Navigation Satellite System (GNSS) such as GPS; or a combination thereof. The wireless communication technology used for communication may include Narrowband Internet of Things (NB-IoT) for low-power communication. In this case, for example, NB-IoT technology may be an example of LPWAN (Low Power Wide Area Network) technology and may be implemented according to standards such as LTE Cat (category) NB1 and / or LTE Cat NB2, but is not limited to the names mentioned above. Additionally, or generally, wireless communication technology implemented in wireless devices according to various embodiments may perform communication based on LTE-M technology. In this case, for example, LTE-M technology may be an example of LPWAN technology and may be referred to by various names such as eMTC (enhanced Machine Type Communication).For example, LTE-M technology may be implemented in at least one of various standards such as 1) LTE CAT 0, 2) LTE Cat M1, 3) LTE Cat M2, 4) LTE non-BL (non-Bandwidth Limited), 5) LTE-MTC, 6) LTE Machine Type Communication, and / or 7) LTE M, and is not limited to the names mentioned above. Additionally or generally, wireless communication technology implemented in wireless devices according to various embodiments may include at least one of ZigBee, Bluetooth, and Low Power Wide Area Network (LPWAN) for low-power communication, and is not limited to the names mentioned above. As an example, ZigBee technology can create personal area networks (PANs) related to small / low-power digital communication based on various standards such as IEEE 802.15.4 and may be referred to by various names.
[0058] The wired communication module may include various wired communication modules such as a Local Area Network (LAN) module, a Wide Area Network (WAN) module, or a Value Added Network (VAN) module, as well as various cable communication modules such as USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface), DVI (Digital Visual Interface), RS-232 (recommended standard 232), power line communication, or POTS (plain old telephone service).
[0059] In addition to Wi-Fi modules and WiBro (Wireless broadband) modules, the wireless communication module may include wireless communication modules that support various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), LTE (Long Term Evolution), 4G, 5G, and 6G.
[0060] The wireless communication module may include a wireless communication interface comprising an antenna and a transmitter that transmit a signal. Additionally, the wireless communication module may further include a signal conversion module that modulates a digital control signal output from the control unit through the wireless communication interface into an analog wireless signal under the control of the control unit.
[0061] The wireless communication module may include a wireless communication interface comprising an antenna and a receiver for receiving a signal. Additionally, the wireless communication module may further include a signal conversion module for demodulating an analog wireless signal received through the wireless communication interface into a digital control signal.
[0062] A short-range communication module is for short-range communication and can support short-range communication by using at least one of Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies.
[0063] The location information module is a module for obtaining the location (or current location) of the electronic device (10) according to the present disclosure, and representative examples thereof include a GPS (Global Positioning System) module or a WiFi (Wireless Fidelity) module. For example, if a GPS module is utilized, the location of the electronic device can be obtained using signals sent from GPS satellites. As another example, if a WiFi module is utilized, the location of the electronic device (10) can be obtained based on information from a Wireless Access Point (AP) that transmits or receives wireless signals from the WiFi module. If necessary, the location information module may perform any of the functions of other modules of the communication module to obtain data regarding the location of the device, either substituted or additionally. The location information module is a module used to obtain the location (or current location) of the device, and is not limited to a module that directly calculates or obtains the location of the device.
[0064] The memory (130) can store various types of information. The memory can store data temporarily or semi-permanently. For example, the memory may store an operating system (OS) for operating the first device and / or the second device, data for hosting a website, or data regarding a program or application (e.g., a web application) for generating Braille. In addition, the memory may store modules in the form of computer code as described above.
[0065] Examples of memory (130) may include a hard disk drive (HDD), a solid state drive (SSD), flash memory, ROM (Read-Only Memory), and RAM (Random Access Memory). These memories may be provided as built-in or removable types.
[0066] The processor (140) controls the overall operation of the electronic device (100). To this end, the processor (140) performs computation and processing of various information and can control the operation of the components of the first device and / or the second device.
[0067] The processor (140) may be implemented as a computer or a similar device according to hardware, software, or a combination thereof. Hardware-wise, the processor (140) may be provided in the form of an electronic circuit that processes electrical signals to perform control functions, and software-wise, it may be provided in the form of a program that drives the hardware processor. Meanwhile, unless otherwise specifically mentioned in the following description, the operation of the first device and / or the second device may be interpreted as being performed by the control of the processor (140). That is, the modules may be interpreted as the processor (140) controlling the first device and / or the second device to perform the following operations.
[0068] The processor (140) may be implemented with a memory that stores data for an algorithm or a program that reproduces the algorithm for controlling the operation of components within the device, and at least one sub-processor (not shown) that performs the aforementioned operation using the data stored in the memory. In this case, the memory and the processor may each be implemented as separate chips. Alternatively, the memory and the processor may be implemented as a single chip.
[0069] Additionally, the processor (140) may control one or a combination of the components described above in order to implement various embodiments according to the present disclosure, which will be described in FIGS. 2 to 10 below, on the device.
[0070] FIG. 2 is a process diagram illustrating a detailed information extraction process of an electronic device (100) that performs detailed information extraction within a text-based recruitment document according to one embodiment of the present disclosure. Hereinafter, an electronic device (100) and a method according to one embodiment of the present disclosure will be described with reference to FIG. 3 to FIG. 10.
[0071] An electronic device (100) for performing detailed information extraction within a text-based recruitment document according to the present disclosure may include an artificial intelligence model (150, see FIG. 4) that receives an unstructured recruitment document and outputs a structured recruitment document, a memory (130) that stores at least one process for performing a structured recruitment document generation operation to enable information extraction, and at least one processor (140) that performs the structured recruitment document generation operation according to the process.
[0072] In one embodiment, the at least one processor (140) may be configured to divide text within an input recruitment document into pages, preprocess the divided pages into pre-set units, input the pre-processed pages into an artificial intelligence model to output a page with area names labeled in key areas, and extract detailed information based on a pre-set prompt for the output page to output a standardized recruitment document.
[0073] As illustrated in FIG. 2, at least one processor (140) can generate and output a recruitment document as a standardized recruitment document by performing text extraction (210), extracted text preprocessing (220), major area labeling (230), extraction of detailed information by major area (240), and aggregation of extracted detailed information (250).
[0074] In relation to text extraction (210), at least one processor (140) according to one embodiment of the present disclosure may be configured to extract text from the employment document through a document filter module.
[0075] The document filter module can extract text from recruitment documents by performing document filtering.
[0076] According to one embodiment of the present disclosure, the recruitment document is written with at least one of a plurality of extensions, and the at least one processor (140) may be configured to extract text by operating a text extraction library corresponding to the extension based on a meta tag of the extension.
[0077] In one embodiment, recruitment documents may be created with at least one of various extensions, such as hwp files, doc files, and pdf files. Each extension has text corresponding to that extension stored in a meta tag, and the stored text can be extracted by refining the program or library code corresponding to each extension.
[0078] Therefore, when recruitment documents generated with multiple extensions are input, the recruitment documents can be classified according to each extension, and a text extraction program or library corresponding to the extension can be operated to extract the text.
[0079] If there is no fixed format for recruitment documents, they may be entered with various extensions; however, the text can be extracted into a single format without the need to build a database in advance.
[0080] Regarding the extracted text preprocessing (220), the extracted text can be preprocessed into a form and size that can be input into the artificial intelligence model (150) thereafter.
[0081] As an example, the entire text extracted from the recruitment document can be structured and output page by page. If text is extracted from the entire page of the recruitment document in sentence units, the labeling or classification processing speed of the artificial intelligence model (150) may be significantly reduced. Furthermore, since recruitment documents are often divided into specific areas on each page and have a unified topic on each page, the text can be separated by page.
[0082] In other words, you can organize text by separating it into page units.
[0083] In addition, as an embodiment, the recruitment documents can be aggregated in specific units, taking into account the size of the text separated into pages that can be input into the artificial intelligence model (150).
[0084] As one example, the recruitment document can generally be divided into personal information, educational background, and a self-introduction, with a volume of 4 to 6 pages. Accordingly, the entire page can be divided into three sections.
[0085] At this time, the total number of pages can be divided by 3, divided into pages corresponding to the quotient, and then the remaining number of pages can be added up to divide it into 3 sections.
[0086] As shown in FIG. 3, if there are a total of 4 employment documents, page 1 can be preprocessed by grouping it into the first group (301), page 2 into the second group (302), and pages 3 and 4 into the third group (303) as 1, which is the 4 / 3 share.
[0087] Since key sections requiring more pages, such as personal statements, are more likely to be located on the later pages of the recruitment documents, more weight is assigned to the later pages than the earlier ones to sum them up as a single unit.
[0088] Recruitment documents can be preprocessed by splitting them into pages by topic to enable more efficient input when extracting detailed information or labeling based on prompts.
[0089] Accordingly, the at least one processor according to one embodiment of the present disclosure may be configured to perform preprocessing by giving a greater weight to the subsequent page than to the preceding page for the divided pages and summing the divided pages into three groups.
[0090] Then, major area labeling (230) and extraction of detailed information by major area (240) can be performed.
[0091] If the artificial intelligence model (150) is to extract detailed information all at once, the text volume is large, making it difficult to accurately extract the desired information. Therefore, at least one processor (140) according to one embodiment of the present disclosure can label the preprocessed first group (301), second group (302), and third group (303) with major area names to indicate which area they correspond to, and then extract detailed information for each labeled area.
[0092] For example, the first group (301) may be labeled as personal information and educational background, the second group (302) as career information and specific details, and the third group (303) as a self-introduction letter.
[0093] As an example, detailed information corresponding to a region can be extracted based on a corresponding prompt for each group labeled with a region name.
[0094] Alternatively, as an example, a group labeled with two or more area names, for example, a first group (301) and a second group (302), can be further subdivided by the number of area names to extract detailed information.
[0095] For example, if the first group (301) is labeled as personal information and educational background, the second group (302) as career information and special notes, and the third group (303) as a self-introduction letter, then the first group is divided into personal information, the first group into educational background, the second group into career information, the second group into special notes, and the third group into a self-introduction letter, and then detailed information can be extracted for the five groups based on separate prompts.
[0096] In one embodiment, at least one processor (140) according to one embodiment of the present disclosure is configured to post-process the different area name as the first area name when a plurality of area names labeled in the main areas are identical as the first area name and there exists one different area name between the plurality of area names, and the first area name may be a self-introduction letter.
[0097] For example, if the labeled sections are sequentially Personal Statement, Personal Statement, Education, Personal Statement, since it is uncommon for Education to be placed between Personal Statements, the Education section can also be considered a Personal Statement and labeled as such.
[0098] Additionally, specifically, as illustrated in FIG. 4, the artificial intelligence model (150) can receive a preprocessed page (12) as input and output a labeled page (22).
[0099] In one embodiment, the artificial intelligence model (150) can label a region based on organized instructions regarding what information a specific text region contains within a preprocessed page (12).
[0100] Alternatively, as an example, at least one processor (140) may label areas into each major area based on instructions organized on what information a specific text area contains within a preprocessed page (12) based on an algorithm.
[0101] Additionally, the artificial intelligence model (150) can receive a prompt (14) and extract detailed information (24).
[0102] Detail information (24) can be extracted through instructions written to extract pre-set details from the labeled page (22).
[0103] According to one embodiment of the present disclosure, at least one processor (140) may be configured to output detailed information among at least one of personal information, educational background information, and self-introduction information from the output page using at least one of a first prompt for extracting personal information, a second prompt for extracting educational background information, and a third prompt for extracting self-introduction information.
[0104] Referring to Fig. 5, detailed information including personal information, Korean name, English version of Korean name, date of birth, gender, address, mobile phone number, (landline) phone number, and email address can be extracted within a major area labeled as personal information based on a first prompt for extracting personal information.
[0105] Based on the first prompt, detailed information can be extracted within the main area where the area name is personal information according to the conditional statement.
[0106] In one embodiment, at least one processor (140) may be configured to output the detailed information in JSON (JavaScript Object Notation) format.
[0107] JSON format is easily compatible with data types such as lists and dictionaries in various programming languages, making it advantageous for secondary processing and allowing text to be converted into JSON format.
[0108] As shown in FIG. 5, personal information, Korean name, English version of Korean name, date of birth, gender, address, mobile phone number, (landline) phone number, and email address extracted from the prompt defined by personalInfo (51), nameKor (52), nameEng (53), birth (54), gender (55), address (56), mobile (57), number (58), and email (59) can be stored by matching them in pairs.
[0109] If there is no corresponding detailed information, you can enter an empty value " " in the field where the extracted value is entered. For example, if there is no (landline) phone number, you can enter " ". Alternatively, if only the month and day exist in the date of birth, you can omit the year and enter only the month and day as is.
[0110] In addition, the extracted detailed information can be entered as is based on the first prompt without modification.
[0111] Referring to Fig. 6, detailed information including educational institution classification, school name, major name, start time, end time, and graduation classification can be extracted within the main area labeled as educational information based on the second prompt for extracting educational information.
[0112] Based on the second prompt, detailed information can be extracted within the main area where the area name is educational background information according to the conditional statement.
[0113] In one embodiment, at least one processor (140) may be configured to output the detailed information in JSON (JavaScript Object Notation) format.
[0114] JSON format is easily compatible with data types such as lists and dictionaries in various programming languages, making it advantageous for secondary processing and allowing text to be converted into JSON format.
[0115] As shown in FIG. 6, detailed information extracted from the prompt defined by institution (62), name (63), major (64), startPeriod (64), endPeriod (66), and gradCategory (67), such as the educational institution type, school name, major name, start time, end time, and graduation type, can be stored by matching them in pairs.
[0116] As an example, the institution (62), which is an educational institution, may have name (63), major (64), startPeriod (65), endPeriod (66), and gradCategory (67) written separately depending on whether it is a graduate school, university, high school, etc.
[0117] Alternatively, since multiple admissions to the graduate school are possible, if detailed information about the graduate school is extracted in institution (62), name (63), major (64), startPeriod (64), endPeriod (66), and gradCategory (67) may be entered for each graduate school.
[0118] Therefore, educational background information can be generated as multiple JSON lists. This involves managing pairs of prompts and details of the same format as lists.
[0119] As an example, if there is only one date information in the educational background information, that information is entered into endPeriod (64), and startPeriod (65) can be entered as “ ”.
[0120] Alternatively, as an example, if endPeriod (64) is unclear, gradCategory (67) may be entered as a blank “ “.
[0121] As shown above, detailed information can be extracted in the form of a conditional statement based on the second prompt.
[0122] Referring to Fig. 7, detailed information including the question of the self-introduction question and the answer to the self-introduction question can be extracted within the main area labeled as self-introduction information based on the third prompt for extracting the self-introduction.
[0123] Based on the third prompt, detailed information can be extracted within the main area named "Self-Introduction" according to the conditional statement.
[0124] In one embodiment, at least one processor (140) may be configured to output the detailed information in JSON (JavaScript Object Notation) format.
[0125] JSON format is easily compatible with data types such as lists and dictionaries in various programming languages, making it advantageous for secondary processing and allowing text to be converted into JSON format.
[0126] As shown in FIG. 7, the questions and answers of the self-introduction questions, which are detailed information extracted from the prompts defined as question (72) and answer (73), can be matched and stored as pairs.
[0127] The questions and answers in personal statement items are particularly long and numerous, and can be written in forms that vary widely depending on the recruiter. Therefore, the extracted detailed information can be entered exactly as is, without modification, based on a third-party prompt.
[0128] In addition, as an example, the questions of the self-introduction section can be extracted based on rules such as numbering the beginning of the sentence, enclosing the sentence in parentheses, or adding a new line.
[0129] Additionally, as an example, if there is no distinction between sentences within the main area of the self-introduction letter or the distinction is unclear, question (72) can input “ “ and all extracted sentences can be entered into answer (73).
[0130] In addition, to extract detailed information as described above, at least one processor (1400) can apply four tags to the content of the input recruitment document text.
[0131] As an example, information that does not belong to any other tag or information that is difficult to determine may be attached as an UNK tag, if at least four of the applicant's name, date of birth, gender, address, phone number, and email address can be extracted from the input, an PER tag, if at least five of the applicant's educational institution, school name, major name, admission date, graduation date, degree, and graduation status can be extracted from the input, an EDU tag, and if 'self-introduction letter' is specified in the input, an INT tag may be attached.
[0132] Each tag is evaluated independently, so PER tags and EDU tags may be attached together. In that case, PER tags and EDU tags can be written as a list and output.
[0133] UNK cannot be used with other tags, and the UNK tag can only be attached if it contains information that does not belong to the other three tags.
[0134] It returns a list of appropriate tags in Python list format based on the text content, allowing you to check the area name based on the tag values.
[0135] PER tags can be identified as personal information, EDU tags as educational background information, and INT tags as personal statement information. UNK tags can be identified as other information.
[0136] Accordingly, the at least one processor (140) according to one embodiment of the present disclosure may be configured to label area names in the main areas within the three groups according to the extent that it includes information corresponding to at least one of the first prompt, the second prompt, and the third prompt.
[0137] As described above, in order to label the domain name, an artificial intelligence model (150) must be trained.
[0138] As shown in FIG. 8, for text listed without separate classification, major areas can be labeled as ==Other==(81), ==Personal Information==(82), and ==Educational Background==(83), and an artificial intelligence model (150) can be trained based on the labeled results. Through this, text segmentation and labeling can be learned simultaneously.
[0139] Alternatively, as another embodiment, each major area can be distinguished by the mark == without separate classification, and an artificial intelligence model (150) can be trained using this as training data. Through this, it can learn where to insert the division symbol. It can learn the process of dividing text arranged in a line and organizing it so that it can be identified by the user.
[0140] Meanwhile, as illustrated in FIG. 9, each area and the detailed information items included in the area can be organized into a table (90). Based on the table (90), a prompt can be created to train an artificial intelligence model (150) so that area labeling and detailed information extraction are possible.
[0141] According to one embodiment of the present disclosure, the at least one processor (140) may be configured to provide a standardized recruitment document to a user interface by aggregating a key-value pair in which the labeled area name and the detailed information are matched, or a list containing the key-value pair.
[0142] Detailed information is extracted to create standardized recruitment documents, and talent matching services can be provided to companies to enable them to hire suitable candidates based on these documents.
[0143] By organizing information into a standardized format and allowing users to input their written text exactly as it is, it enables HR managers to smoothly compare candidates and evaluate them accurately.
[0144] FIG. 10 is a flowchart illustrating the flow of a method for extracting detailed information from text-based recruitment documents according to one embodiment of the present disclosure.
[0145] As illustrated in FIG. 10, the method may include the steps of: dividing text within an input recruitment document into pages and preprocessing the divided pages into pre-set units (S1100); inputting the pre-processed pages into an artificial intelligence model to output a page with area names labeled in key areas (S1120); and extracting detailed information based on a pre-set prompt for the output page to output a standardized recruitment document (S1130).
[0146] Content that overlaps with the above is omitted for the sake of brevity in the specification.
[0147] Therefore, by efficiently inputting preprocessed data into artificial intelligence models such as Chat GPT, and by labeling pages by region and extracting detailed information by region, it is possible to optimize the input data as much as possible to enable efficient extraction of detailed information.
[0148] Extracted detailed information can be converted into JSON format and stored as a list to efficiently manage and process detailed information by area.
[0149] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium that stores instructions executable by a computer. The instructions may be stored in the form of program code and, when executed by a processor, may generate a program module to perform the operation of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.
[0150] Computer-readable recording media include all types of recording media that store instructions that can be decoded by a computer. Examples include ROM (Read Only Memory), RAM (Random Access Memory), magnetic tape, magnetic disk, flash memory, optical data storage devices, etc.
[0151] As described above, the disclosed embodiments have been explained with reference to the attached drawings. Those skilled in the art will understand that the present disclosure may be practiced in forms different from the disclosed embodiments without changing the technical spirit or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be interpreted restrictively. Explanation of the symbols
[0152] 100: Electronic device 110: I / O module 120: Communication module 130: Memory 140: Processor 150: Artificial Intelligence Model
Claims
Claim 1 Memory storing at least one process for performing a standardized recruitment document generation operation to enable information extraction; An electronic device for performing detailed information extraction within a text-based recruitment document, comprising: at least one processor that performs the operation of generating the standardized recruitment document according to the above process; wherein the at least one processor divides the text within the input recruitment document into page units and preprocesses the divided pages into pre-set units, inputs the preprocessed pages into an artificial intelligence model to output a page in which area names are labeled in the main areas, and extracts detailed information based on a pre-set prompt for the output page to output a standardized recruitment document, wherein, prior to dividing the text within the input recruitment document into page units, the processor is configured to extract text from the recruitment document through a document filter module, wherein the recruitment document is written with at least one of a plurality of extensions, and the at least one processor operates a text extraction library corresponding to the extension based on the meta tag of the extension to extract text, wherein if a plurality of area names labeled in the main areas are identical as a first area name and a different area name exists between the plurality of area names, the processor is configured to post-process the different area name as the first area name, and wherein the first area name is a self-introduction letter. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 An electronic device for performing detailed information extraction within a text-based recruitment document, wherein the at least one processor is configured to output detailed information among at least one of personal information, educational background information, and self-introduction information from the output page using at least one of a first prompt for extracting personal information, a second prompt for extracting educational background information, and a third prompt for extracting self-introduction information. Claim 6 In claim 5, the at least one processor is an electronic device that performs the extraction of detailed information within a text-based recruitment document configured to output the detailed information in JSON (JavaScript Object Notation) format. Claim 7 delete Claim 8 An electronic device for performing detailed information extraction within text-based recruitment documents, wherein the at least one processor is configured to label area names in key areas within three groups according to the extent that they contain information corresponding to at least one of the first prompt, the second prompt, and the third prompt. Claim 9 In claim 8, the at least one processor is an electronic device configured to perform detailed information extraction within text-based recruitment documents, wherein the processor aggregates a key-value pair in which the labeled area name and the detailed information are matched, or a list containing the key-value pairs, and provides the output structured recruitment documents to a user interface. Claim 10 A method for extracting detailed information within a text-based recruitment document, performed by a computing device comprising: memory; and at least one processor; wherein the method comprises: a step of dividing text within an input recruitment document into page units and preprocessing the divided pages into pre-set units; a step of inputting the preprocessed pages into an artificial intelligence model to output a page in which area names are labeled in key areas; and a step of extracting detailed information based on a pre-set prompt for the output page to output a structured recruitment document; wherein the preprocessing step comprises: a step of extracting text from the recruitment document through a document filter module before dividing the text within the input recruitment document into page units; wherein the recruitment document is written with at least one of a plurality of extensions; and wherein the method comprises: a step of extracting text by operating a text extraction library corresponding to the extension based on a meta tag of the extension. A method for extracting detailed information from text-based recruitment documents, further comprising: a step of post-processing a different area name as the first area name when a plurality of area names labeled in the above major areas are identical to the first area name and a different area name exists between the plurality of area names, wherein the first area name is a self-introduction letter.
Citation Information
Patent Citations
Automated Textual analysis technology for unstructured big data mining in the form of compound documents
KR1020210085306A
System and method for extracting data from document for each company using fingerprints and machine learning
KR1020230066757A
Deep learning-based resume information structuring method and system
CN113220768A
Document data providing device and program thereof
JP2013065153A
Information processing device, information processing method, and program
JP7349219B1