Information processing device, method, program, and system

The system addresses the challenge of incomplete or incorrect resume data extraction by using a generative model to organize and categorize work history, enhancing data accuracy and reducing user input burden.

JP7791528B2Active Publication Date: 2025-12-24FINDY INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023222131
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-12-24
Estimated Expiration
2043-12-28

AI Technical Summary

Technical Problem

Existing systems struggle with accurately extracting work history information from resumes due to inappropriate term usage or font issues, leading to incomplete or incorrect data extraction.

Method used

A system that utilizes a generative model to organize and categorize resume data by generating structured text data based on prompts, reducing the user's burden of inputting work history by automatically extracting and organizing information into predefined categories.

Benefits of technology

The system effectively reduces the user's input burden by automatically structuring and categorizing work history information, improving data extraction accuracy and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007791528000001
    Figure 0007791528000001
  • Figure 0007791528000002
    Figure 0007791528000002
  • Figure 0007791528000003
    Figure 0007791528000003
Patent Text Reader

Abstract

To provide a technique capable of reducing an input burden of a work history.SOLUTION: A program causes a computer to function as means of: acquiring first text data on the basis of work history data representing a work history of a target person; generating second text data obtained by organizing the first text data according to a sentence structure; and acquiring information of the target person organized by category name relating to the work history by providing a generation model with a prompt based on the second text data and the category name relating to the work history.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, a method, a program, and a system. [Background technology]

[0002] When job seekers use job search support services such as career change support services and personnel matching services, they are sometimes asked to enter their work history. If the burden of this input work could be reduced, it could potentially increase the number of users of job search support services.

[0003] Patent document 1 describes a system that analyzes resumes and extracts data corresponding to database fields such as contact information, work history, and educational history. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] U.S. Patent No. 8,600,931 Summary of the Invention [Problem to be solved by the invention]

[0005] In the technical idea of ​​Patent Document 1, terms that match a database of known resume terms or terms displayed in a special font are considered as field names, and field data associated with or close to the field name is extracted. In other words, with this technical idea, if the terms and font used in the resume are not appropriate, necessary information may not be extracted or inappropriate information may be extracted.

[0006] An object of the present disclosure is to provide a technology that can reduce the burden of inputting work history. [Means for solving the problem]

[0007] A program according to one embodiment of the present disclosure causes a computer to function as a means for acquiring first text data based on resume data representing a subject's work history, a means for generating second text data by organizing the first text data according to the structure of the sentence, and a means for acquiring information about the subject organized by category names related to their work history by providing a generative model with prompts based on the second text data and category names related to their work history. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram showing a configuration of an information processing system according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram showing the configuration of a client device according to the present embodiment. [Figure 3] FIG. 2 is a block diagram showing the configuration of a server according to the present embodiment. [Figure 4] FIG. 1 is an explanatory diagram of one aspect of the present embodiment. [Figure 5] 10 is a flowchart of a work history extraction process according to the present embodiment. [Figure 6] FIG. 10 is a diagram showing a structured document acquired in the work history extraction process of the present embodiment. [Figure 7] 10A and 10B are diagrams illustrating an example of a screen displayed in the work history extraction process of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. In the drawings for explaining the embodiment, the same components are generally designated by the same reference numerals, and repeated description thereof will be omitted.

[0010] (1) Information processing system configuration The configuration of the information processing system will now be described with reference to Fig. 1, which is a block diagram showing the configuration of the information processing system according to this embodiment.

[0011] As shown in FIG. 1, the information processing system 1 includes a client device 10 and a server 30. The client device 10 and the server 30 are connected via a network (for example, the Internet or an intranet) NW.

[0012] The client device 10 is an example of an information processing device that transmits a request to the server 30. The client device 10 is, for example, a smartphone, a tablet terminal, or a personal computer.

[0013] The user of the client device 10 is typically, but is not limited to, a job seeker. In this specification, a job seeker includes a person who is looking for a new job or (re)looking for employment, a person who is interested in these activities, or a person who intends to seek employment. For example, a job seeker may include a person who registers his or her own information on a job search support service or prepares to register (e.g., uploads a resume file or a curriculum vitae file, or enters his or her own work history). The user of the client device 10 may also be a person who enters work history information on behalf of the person (an input agent).

[0014] The server 30 is an example of an information processing device that provides the client device 10 with a response in response to a request sent from the client device 10. The server 30 is, for example, a server computer. The server 30 can also manage job search support services such as a job change support service and a human resources matching service.

[0015] (1-1) Client device configuration The configuration of the client device will now be described with reference to Fig. 2, which is a block diagram showing the configuration of the client device of this embodiment.

[0016] 2, the client device 10 includes a storage device 11, a processor 12, an input / output interface 13, and a communication interface 14. The client device 10 is connected to a display 21.

[0017] The storage device 11 is configured to store programs and data, and is, for example, a combination of a read-only memory (ROM), a random access memory (RAM), and a storage (for example, a flash memory or a hard disk).

[0018] The programs include, for example, the following programs: OS (Operating System) programs Applications that process information (e.g., web browsers)

[0019] The data includes, for example, the following data: Databases referenced in information processing Data obtained by performing information processing (i.e., the results of performing information processing)

[0020] The processor 12 is a computer that implements the functions of the client device 10 by running a program stored in the storage device 11. The processor 12 is, for example, at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Gate Array)

[0021] The input / output interface 13 is configured to acquire information (for example, a user's instruction) from an input device connected to the client device 10, and to output information (for example, an image) to an output device connected to the client device 10.

[0022] The input device is, for example, a keyboard, a pointing device, a touch panel, or a combination thereof. The output device is, for example, a display 21, a speaker, or a combination thereof.

[0023] The communication interface 14 is configured to control communication between the client device 10 and an external device (eg, a server 30).

[0024] The display 21 is configured to display an image (a still image or a moving image). The display 21 is, for example, a liquid crystal display or an organic EL display.

[0025] (1-2) Server configuration The configuration of the server will now be described with reference to Fig. 3, which is a block diagram showing the configuration of the server according to this embodiment.

[0026] As shown in FIG. 3, the server 30 includes a storage device 31, a processor 32, an input / output interface 33, and a communication interface .

[0027] The storage device 31 is configured to store programs and data, and is, for example, a combination of ROM, RAM, and storage (for example, flash memory or a hard disk).

[0028] The programs include, for example, the following programs: OS programs Application programs that perform information processing

[0029] The data includes, for example, the following data: Databases referenced in information processing - Results of information processing

[0030] The processor 32 is a computer that implements the functions of the server 30 by running a program stored in the storage device 31. The processor 32 is, for example, at least one of the following: ·CPU GPU ASIC FPGA

[0031] The input / output interface 33 is configured to acquire information (for example, a user's instruction) from an input device connected to the server 30, and to output information (for example, an image) to an output device connected to the server 30.

[0032] The input device is, for example, a keyboard, a pointing device, a touch panel, or a combination thereof. The output device is, for example, a display.

[0033] The communication interface 34 is configured to control communications between the server 30 and an external device (eg, the client device 10).

[0034] (2) One aspect of the embodiment An example of this embodiment will now be described with reference to Fig. 4, which is an explanatory diagram of this example.

[0035] As shown in FIG. 4, user US1 uses client device 10 to transmit (upload) resume data (file) relating to the subject's work history to server 30. The subject may be user US1 himself or herself, or a different person. The resume data may be data created by, for example, scanning or photographing a paper resume to create an electronic file, or may be an electronic file created by editing work history information using word processing software. For example, it is assumed that documents or electronic files created by the subject in the past for job hunting purposes will be used as resume data.

[0036] The server 30 acquires (extracts) text data based on the acquired resume data, and then organizes the text data according to the structure of the sentences.

[0037] The server 30 generates a prompt based on the organized text data and a predetermined category name related to work history. The server 30 provides the generated prompt to the generative model GM2. The generative model GM2 generates text data in which the subject's information (work history information) is organized by this category name, and outputs the text data.

[0038] The server 30 extracts information about the subject by category name (elements of work history information by category name) from the text data generated by the generative model GM2. The server 30 sets the extracted information in the input field of the corresponding category, and presents an input screen for work history to the user via the client device 10.

[0039] In this way, according to this embodiment, the user US1 can avoid inputting at least a part of the work history information of the subject by uploading, for example, already created resume data. In other words, the burden of inputting work history can be reduced.

[0040] (3) Information processing The information processing of this embodiment will be described.

[0041] (3-1) Work history extraction process The work history extraction process of this embodiment will be described. Fig. 5 is a flowchart of the work history extraction process of this embodiment. Fig. 6 is a diagram showing a structured document acquired in the work history extraction process of this embodiment. Fig. 7 is a diagram showing an example of a screen displayed in the work history extraction process of this embodiment.

[0042] The work history extraction process of this embodiment can be initiated, for example, by a user of a client device 10 uploading work history data (for example, a PDF file or other format document file) to the server 30 via an interface displayed on the display 21 of the client device 10.

[0043] Before uploading the resume data, the client device 10 may present a message to warn the user that the resume data is to be input into the generative model, and may then upload the resume data only after receiving explicit consent from the user.

[0044] As shown in FIG. 6, the server 30 acquires curriculum vitae data (S130). Specifically, the server 30 receives the resume data from the client device 10 and stores it in the storage device 31. The server 30 may convert the file format of the resume data.

[0045] After step S130, the server 30 extracts text data (S131). Specifically, server 30 performs character recognition processing on the resume data acquired in step S130 to acquire text data (an example of "first text data") included in the resume data and corresponding text coordinate information. For example, the text coordinate information is the two-dimensional coordinates (x-y coordinates) of pixels representing the corresponding text in the layout of the resume data. For another example, the text coordinate information may be the position of the corresponding text in the layout of the resume data (the page number, the number of lines, the number of characters from the beginning of the line, or a combination thereof). Furthermore, if the resume data includes data representing a table listing the subject's work history, server 30 may extract coordinate information corresponding to the rows, columns, or cells that make up the table. Alternatively, if text data is embedded in the resume data acquired in step S130, the server 30 may use the text data as the extraction result.

[0046] After step S131, the server 30 organizes the text data (S132). Specifically, based on the text data and text coordinate information extracted in step S131, the server 30 organizes the text data according to the structure of the sentence, thereby obtaining text data (an example of "second text data") that is more suitable for subsequent processing (particularly natural language processing). Here, organizing the text data may include, for example, dividing one or more blocks of text included in the text data into a larger number of blocks of text, integrating multiple blocks of text included in the text data into a smaller number of blocks of text, changing the order of text included in the text data, summarizing the text included in the text data, or a combination thereof. As an example, the server 30 estimates the position of a block in a sentence by providing a model input based on the text data and text coordinate information extracted in step S131 to a block estimation model. The position of a block in a sentence may be, for example, a paragraph position, a line break position, or a combination thereof. As a segmentation estimation model, a trained model can be used that has acquired the ability to estimate segmentation positions from text data and corresponding text coordinate information by learning the structure of a large number of documents (preferably resumes) (for example, the relationship between text data obtained from a large number of documents and corresponding text coordinate information and segmentation positions (correct answer data)).

[0047] The server 30 chunks the text data according to the estimated delimiter positions (e.g., paragraph positions). The server 30 also adds line break data to the text data according to the estimated delimiter positions (e.g., line break positions). Note that instead of using a delimiter estimation model, it is also possible to organize the text data on a rule basis.

[0048] The server 30 may also organize the text data based on coordinate information corresponding to rows, columns, or cells that make up the table listing the subject's work history. For example, the server 30 may chunk the text data by row, column, or cell, or add line break data.

[0049] Depending on the format of the resume data acquired in step S130, text data organized according to delimiter positions may be extracted in step S131. In this case, step S132 may be omitted.

[0050] Before or after step S132, or in parallel with step S132, the server 30 may determine the format requirements of the resume data. Specifically, the server 30 determines whether the resume data acquired in step S130 or the text data acquired in step S131 or S132 meets the format requirements. As a first example, the server 30 determines that the data to be judged does not meet the format requirements if it does not have a predetermined required item, such as a header such as "resume." As a second example, the server 30 provides a model input based on the data to be judged to a format determination model, and determines that the data does not meet the format requirements if a determination result is obtained that the data is not a resume. A trained model that has acquired the ability to distinguish resumes from other documents by learning the structural characteristics of many resumes can be used as the format determination model. If the server 30 determines that the data to be judged does not meet the format requirements, it may present an error screen to the user via the client device 10 and terminate the work history extraction process.

[0051] After step S132, the server 30 acquires the structured document (S133). Specifically, the server 30 obtains text data (an example of "third text data") that is a structured document based on the text data obtained in step S132 by providing the generative model with a prompt including the text data. The generative model may be a large-scale language model trained with a large amount of text data, or a model obtained by transfer learning or fine-tuning the large-scale language model. The generative model may also be constructed in a system external to the information processing system 1 (e.g., a cloud environment).

[0052] An example of a structured document is shown in Figure 6. A structured document is a document that contains text that divides the information contained in a resume into personal information and company-related information. Each piece of information can be written in chronological order. In addition to personal and company-related information, a structured document may also include information on the creation date. Personal information can include, for example, the subject's name, qualifications, self-promotion, self-improvement, or a combination thereof. Company-related information can include the subject's achievements and efforts, role, length of employment, job summary, or a combination thereof for each organization or company to which the subject has previously belonged.

[0053] After step S133, the server 30 acquires information organized by category name (S134). Specifically, the server 30 provides the generative model with a prompt including the text data (structured document) obtained in step S133, a predetermined category name related to work history, and instructions for causing the generative model to extract information about the subject. As a result, the server 30 obtains text data (an example of "fourth text data") from which information about the subject organized by category name (elements of work history information by category name) can be extracted. As an example, the server 30 may extract work history information for each category name from the text data and provide the generative model with a prompt instructing the generative model to output the extracted work history information as text data in a format linked to the corresponding category name (e.g., a format such as "Category name A: work history information α, ..."). The server 30 may then extract information about the subject by category name related to work history from the text data obtained from the generative model.

[0054] The predetermined category names may include, for example, company name, project name, role, job type, project period, technology used, project details, or a combination thereof. The predetermined category names are determined to correspond to the input items on the work history input screen of the job search support service provided by the server 30.

[0055] In addition, when the server 30 extracts text data from resume data in PDF file format in step S131, in this step S134, the prompt may include information to inform the generative model that the text data (structured document) has been extracted from a PDF file and therefore may be out of format.

[0056] Furthermore, the server 30 may optionally modify the text data obtained from the generative model in this step S134. Specifically, the server 30 provides the generative model with a prompt including the text data (structured document) obtained in step S133, a predetermined category name related to work history, the text data obtained from the generative model in step S134, and instructions for the generative model to compare these text data and correct any deficiencies or excesses. As a result, the server 30 obtains text data (an example of "fifth text data") from which information about the subject organized by category name related to work history can be extracted. Here, "excess or deficiency" may include, for example, elements of work history information included in the text data obtained from the generative model in step S134 that are not based on the text data obtained in step S133 (i.e., there is no corresponding description). The server 30 may then extract information about the subject by category name related to work history from the text data obtained from the generative model.

[0057] Furthermore, the server 30 may optionally summarize the text data obtained from the generative model in this step S134 (which may include the above-mentioned corrected text data). Specifically, the server 30 provides the generative model with a prompt including the text data obtained from the generative model in step S134 and instructions for the generative model to summarize the text data. As a result, the server 30 obtains a summary of the text data. The server 30 may then extract information about the subject from the summary results by category name related to work history. In step S134, whether to summarize may be determined depending on the volume of the text data obtained from the generative model. For example, summarization may be performed if the overall volume of text or the volume of text for a specific category name exceeds a threshold, and summarization may be omitted if the volume does not exceed a threshold.

[0058] After step S134, the server 30 presents an input screen (S135). Specifically, the server 30 presents a screen for accepting input of the subject's work history (hereinafter referred to as "input screen") in a state in which the subject's information, organized by category name related to work history, is stored in each corresponding object. For example, for each object (e.g., text box) placed on the input screen, the server 30 extracts information on the category name corresponding to the input item of the object from the text data obtained in step S134 and sets it in the object. Then, the server 30 transmits information for displaying the input screen to the client device 10. If the input items of the object and the category name are different, the server 30 may instruct the generative model to regenerate the text data. The prompt instructing the generative model to regenerate may include an instruction to associate a category name that does not contain information corresponding to the original resume data (the text data (structured document) obtained in step S133) with a blank. The server 30 then sets information in the object based on the text data obtained by regeneration.

[0059] The client device 10 displays an input screen on the display 21 based on information from the server 30. An example of the input screen is shown in Fig. 7. The input screen in Fig. 7 includes objects J20 to J28. Object J20 has a one-to-one correspondence with object J21 and represents the input item (category name) assigned to the corresponding object J21.

[0060] The object J21 accepts input of information (work history information) about the target person for the assigned input fields. Note that information may already be set in the object J21 by presenting the input screen described above (S135). If the information set in the object J21 is insufficient or incorrect, or if the object J21 is blank, the user can edit the information set in the object J21 or add appropriate information. Objects J20 to J21 are arranged on the input screen for each input item.

[0061] The object J22 accepts a user instruction to add an object (e.g., objects J20 to J21) for accepting further input of work history. When the object J22 is selected, the client device 10 adds the object for accepting further input of work history to the input screen.

[0062] The object J23 accepts a user instruction to return to the previous screen from the input screen. When the object J23 is selected, the client device 10 transitions from the input screen to the previous screen.

[0063] Object J24 accepts a user instruction to save work history information. When object J24 is selected, the client device 10 creates a resume file based on the input values ​​of each object and saves it in the storage device 11. The resume file can be used as a resume in a format different from the resume uploaded by the user. In response to the selection of object J24, the client device 10 may display a message warning the user that the saved resume file may contain errors, and may further receive explicit approval from the user before saving the resume file on the client device 10.

[0064] Object J25 accepts user instructions for registering work history information with the job change support service. When object J25 is selected, client device 10 transmits the input values ​​of each object to server 30. Server 30 registers the input values ​​of each object in a database (not shown) in association with information identifying the target person.

[0065] The object J26 accepts a user instruction to redo the extraction of work history information. When the object J26 is selected, the client device 10 displays a message on the display 21 prompting the user to re-upload the resume data. The server 30 then executes the work history extraction process again based on the re-uploaded resume data. However, when redoing the work history extraction process, the server 30 may provide a prompt to the generative model including an instruction to cause the generative model to generate text different from that of the previously executed work history extraction process. The client device 10 may also accept a user request to provide more detailed information about a specific work history or to provide a thinner (summarized) information about a specific work history, and transmit the received information to the server 30. The server 30 may then reflect instructions in the prompt to cause the generative model to generate text data in accordance with the user's request. The client device 10 may also accept a user's indication of missing parts of the extraction (text about a specific category or a specific work history) and transmit the received information to the server 30. The server 30 may then reflect in the prompt an instruction to cause the generative model to generate text data including information on the portion pointed out by the user.

[0066] Object J27 accepts a user instruction to reserve an appointment. When object J27 is selected, the client device 10 may launch another program to start reserving the appointment.

[0067] The object J28 receives a user instruction to request a resume creation service. When the object J28 is selected, the client device 10 may launch another program and start the resume creation service request.

[0068] (4) Summary As described above, the server 30 of this embodiment acquires first text data based on resume data representing the work history of a subject, and generates second text data by organizing the first text data according to the structure of the sentence. The server 30 acquires information about the subject organized by category name related to work history by providing a generative model with prompts based on the second text data and category names related to work history. This allows the server 30 to acquire information about the subject organized by category name related to work history without placing the burden on the user of organizing and inputting work history information according to a format required by the system. In other words, the burden of inputting work history can be reduced.

[0069] The server 30 may generate second text data based on the first text data and the coordinate information corresponding to the first text data. This allows the second text data to be obtained in which each piece of text constituting the first text data is appropriately organized according to its position within the document, thereby improving the quality of the information that is ultimately extracted.

[0070] The server 30 may acquire the second text data by determining the structural delimiter positions of the sentence based on the first text data and the coordinate information corresponding to the first text data, and chunking the first text data based on the delimiter positions. This allows the second text data to be acquired in which each text constituting the first text data is appropriately organized according to the structural delimiter positions of the sentence, thereby improving the quality of the information finally extracted.

[0071] The resume data may include data representing a table listing the subject's work history. The server 30 may generate second text data based on coordinate information corresponding to the rows, columns, or cells that make up the table, the first text data, and the coordinate information corresponding to the first text data. This allows the second text data to be obtained in which each piece of text that makes up the first text data is appropriately organized according to the structure of the table included in the resume data, thereby improving the quality of the ultimately extracted information.

[0072] The server 30 may execute a first phase of processing to acquire third text data, which is a structured document based on the second text data, by providing a prompt including the second text data to the generative model, and a second phase of processing to acquire fourth text data from which information about the subject organized by category names related to work history can be extracted by providing a prompt including the third text data, category names related to work history, and instructions for the generative model to extract information about the subject by the category names. This allows the processing using the generative model to be performed in stages, including a process for generating a structured document and a process for organizing information in the structured document by category, thereby improving the quality of the information finally extracted.

[0073] In the second phase of processing, the server 30 may obtain the fourth text data by providing the generative model with a prompt that further includes information informing the generative model that the third text data may be out of format because it has been extracted from a PDF file. This makes it easier to obtain appropriate fourth text data even if the third text data is out of format.

[0074] The server 30 may further execute a third phase of processing to acquire fifth text data from which information about the subject organized by category name related to work history can be extracted, by comparing the third text data, the category names related to work history, the fourth text data, and the third text data with the generative model, and providing a prompt including an instruction to correct any deficiencies or excesses in the fourth text data to the generative model. This allows the fifth text data to be obtained by correcting any deficiencies in the fourth text data, thereby improving the quality of the information finally extracted.

[0075] The server 30 may obtain a summary result of the fourth text data by providing the generative model with a prompt based on the fourth text data and instructions for causing the generative model to summarize the fourth text data, and extract information about the subject by category name related to work history from the summary result. This allows for a summary result that summarizes the main points of the fourth text data, thereby improving the quality of the information finally extracted.

[0076] The server 30 may present a screen for accepting input of the subject's work history in a state in which the subject's information is organized by category name related to work history and stored in the corresponding object, thereby reducing the input burden on the user.

[0077] The server 30 may receive an instruction from the user to regenerate at least one of the pieces of information stored in any of the objects, so that if the information is not extracted properly, the user can try again until he or she is satisfied.

[0078] The server 30 may determine whether the resume data meets the format requirements, and may display an error screen if it is determined that the resume data does not meet the format requirements. This can, for example, quickly notify the user that incorrect data has been uploaded, or avoid the burden on the server 30 and the generative model caused by processing inappropriate data.

[0079] (5) Other variations The storage device 11 may be connected to the client device 10 via a network NW. The display 21 may be integrated with the client device 10. The storage device 31 may be connected to the server 30 via the network NW.

[0080] Each step of the above information processing can be executed by either the client device 10 or the server 30. In the above explanation, an example in which each step in each process is executed in a specific order is shown, but the execution order of each step is not limited to the example explained above as long as there is no dependency between the steps.

[0081] The above description has shown an example in which the server 30 performs information processing using a generative model (steps S133 to S134). If an error occurs in such information processing, the server 30 may present an error screen to the user via the client device 10. The error screen may include information indicating the cause of the error (for example, the server 30 or the generative model) along with the reason for the error. The manner in which the information is presented may differ depending on the cause of the error. Furthermore, if an error occurs, information may be presented to prompt the user to re-upload their resume data.

[0082] Although the embodiments of the present invention have been described in detail above, the scope of the present invention is not limited to the above-described embodiments. Furthermore, the above-described embodiments can be improved or modified in various ways without departing from the spirit of the present invention. Furthermore, the above-described embodiments and modifications can be combined.

[0083] (6) Supplementary notes The matters explained in the embodiment and the modified examples are additionally noted below.

[0084] (Appendix 1) A computer (30), A means for acquiring first text data based on resume data representing the work history of a subject (S131); A means for generating second text data by organizing the first text data according to the structure of the sentence (S132); a means for acquiring information about the subject organized by category names related to work history by providing prompts based on the second text data and category names related to work history to the generation model (S133 to S134); A program that functions as a

[0085] (Appendix 2) the means for generating the second text data generates the second text data based on the first text data and coordinate information corresponding to the first text data; The program described in Appendix 1.

[0086] (Appendix 3) the means for generating the second text data determines structural delimiter positions of the sentence based on the first text data and coordinate information corresponding to the first text data, and chunks the first text data based on the delimiter positions to obtain the second text data; The program described in Appendix 2.

[0087] (Appendix 4) The resume data includes data representing a table listing the subject's work history; the means for generating the second text data generates the second text data based on coordinate information corresponding to rows, columns, or cells constituting the table, the first text data, and the coordinate information corresponding to the first text data; The program described in Appendix 1.

[0088] (Appendix 5) The means to obtain information on subjects organized by category name regarding work history is as follows: a first phase process of obtaining third text data, which is a structured document based on the second text data, by providing a prompt including the second text data to a generative model; A second phase process of obtaining fourth text data from which information on subjects organized by category names related to work history can be extracted by providing the third text data, category names related to work history, and prompts including instructions for extracting information on subjects by the category names to the generative model; To execute The program described in Appendix 1.

[0089] (Appendix 6) The means for obtaining the subject's information organized by category name related to work history is to obtain the fourth text data by providing the generative model with a prompt containing further information to inform the generative model that the third text data is extracted from a PDF file and therefore may be out of format. The program described in Appendix 5.

[0090] (Appendix 7) The means for acquiring information about the subject organized by category name related to work history further executes a third phase of processing to acquire fifth text data from which information about the subject organized by category name related to work history can be extracted by comparing the third text data, the category name related to work history, the fourth text data, and the third text data with the fourth text data, and providing a prompt to the generative model including an instruction to correct any excess or deficiency in the third text data and the fourth text data. The program described in Appendix 6.

[0091] (Appendix 8) The means to obtain information on subjects organized by category name regarding work history is as follows: obtaining a summary result of the fourth text data by providing the generative model with fourth text data and prompts based on instructions for causing the generative model to summarize the fourth text data, and extracting information about the subject by category name related to work history from the summary result; 1. The program described in Appendix 7.

[0092] (Appendix 9) causing the computer to function as a means (S135) for presenting a screen for accepting input of the subject's work history in a state in which the subject's information organized by category name related to work history is stored in each corresponding object; The program described in Appendix 1.

[0093] (Appendix 10) causing the computer to function as a means for receiving instructions from a user to reproduce at least one of the pieces of information stored in any of the objects; 10. The program described in Appendix 9.

[0094] (Appendix 11) Computer, A means for determining whether the resume data meets format requirements; A means for displaying an error screen when it is determined that the resume data does not meet format requirements; To function as, The program described in Appendix 1.

[0095] (Appendix 12) A means for acquiring first text data based on resume data representing the work history of a subject (S131); A means for generating second text data by organizing the first text data according to the structure of the sentence (S132); A means for acquiring information about a subject organized by category names related to work history by providing a prompt based on the second text data and category names related to work history to a generation model (S133 to S134); An information processing device (30) comprising:

[0096] (Appendix 13) A computer (30) A step (S131) ​​of acquiring first text data based on resume data representing the work history of a subject; A step (S132) of generating second text data by organizing the first text data according to the structure of the sentence; a step (S133 to S134) of acquiring information on the subject organized by category names related to work history by providing prompts based on the second text data and category names related to work history to the generation model; How to do it.

[0097] (Appendix 14) A system (1) including a plurality of information processing devices (10, 30), A means for acquiring first text data based on resume data representing the work history of a subject (S131); A means for generating second text data by organizing the first text data according to the structure of the sentence (S132); A means for acquiring information about a subject organized by category names related to work history by providing a prompt based on the second text data and category names related to work history to a generation model (S133 to S134); A system comprising: [Explanation of symbols]

[0098] 1: Information processing system 10: Client device 11:Storage device 12: Processor 13: Input / output interface 14: Communication interface 21: Display 30: Server 31:Storage device 32: Processor 33: Input / output interface 34: Communication interface

Claims

1. Computer, means for acquiring first text data which is text data relating to the work history of the subject included in curriculum vitae data representing the work history of the subject; a means for generating a prompt by including a structured document output in response to providing information based on the first text data to a generative model and a category name in an instruction to extract the work history information of the subject by category name, and providing the prompt to the generative model to acquire the information of the subject by the category name; A program that functions as a

2. The computer means for dividing the first text data at sentence delimiters to generate second text data consisting of a plurality of delimited sentences; It functions as the means for generating the second text data determines sentence division positions in the first text data based on a plurality of text data included in the first text data and coordinate information in a layout of the resume data for each of the plurality of text data, and divides the first text data at the division positions in the sentences to generate the second text data consisting of a plurality of divided sentences; the structured document is a structured document based on the second text data; The program according to claim 1.

3. The means for generating the second text data determines structural division positions of sentences in the first text data based on a plurality of text data included in the first text data and coordinate information in the layout of the resume data of each of the plurality of text data, and chunks the first text data based on the division positions, thereby obtaining the second text data consisting of a plurality of divided sentences. The program according to claim 2.

4. The computer means for dividing the first text data at sentence delimiters to generate second text data consisting of a plurality of delimited sentences; It functions as the resume data includes data representing a table listing the subject's work history; the means for generating the second text data chunks the first text data in units of rows, columns, or cells that make up the table, based on a plurality of text data corresponding to the rows, columns, or cells that make up the table and coordinate information for each of the plurality of text data, thereby generating the second text data consisting of a plurality of delimited sentences; the structured document is a structured document based on the second text data; The program according to claim 1.

5. Computer, means for acquiring first text data which is text data relating to the work history of the subject included in curriculum vitae data representing the work history of the subject; A program that functions as a means for acquiring information about the subject by category name regarding work history, The means for acquiring information of the subject by category name includes: a first phase of processing in which a first prompt is generated by including the first text data in an instruction to generate a structured document in which information included in the resume data is divided into personal information and company information, and the first prompt is provided to a generation model to obtain third text data, which is a structured document based on the first text data; a second phase process of generating a second prompt by including the third text data and a category name related to work history in an instruction to extract information about the subject by category name, and providing the second prompt to a generation model to obtain fourth text data from which information about the subject can be extracted by category name related to work history; and a process of acquiring information about the subject by category name based on the fourth text data.

6. The means for acquiring the information of the subject by category name generates the second prompt by further including information informing the generative model that the third text data may be out of format because it has been extracted from a PDF file in the second phase of processing, and acquires the fourth text data by providing the second prompt to the generative model. The program according to claim 5.

7. The means for acquiring the information of the subject by category name generates a third prompt by comparing the third text data, the category name related to the work history, the fourth text data, and the generative model, and including an instruction to correct any excess or deficiency in the third text data and to provide the third prompt to the generative model, thereby further executing a third phase of processing to acquire fifth text data from which the information of the subject can be extracted by the category name. The program according to claim 6.

8. The means for acquiring information of the subject by category name includes: generating a fourth prompt by including the fourth text data in an instruction to summarize the fourth text data, obtaining a summary result of the fourth text data by providing the fourth prompt to the generation model, and further performing a process of extracting information about the subject from the summary result by category name related to the work history. The program according to claim 7.

9. causing the computer to function as a means for presenting a screen for accepting input of the subject's work history by category name, in a state in which the subject's information is stored in corresponding objects by category name regarding the work history; The program according to claim 1.

10. The computer a means for receiving an instruction from a user to regenerate at least one piece of information stored in any of the objects; The means for acquiring information of the subject by category name includes: generating a prompt by including a structured document based on the first text data and a category name in the instruction to regenerate, and providing the prompt to a generative model to execute a process of re-obtaining at least one of the pieces of information stored in any of the objects; The program according to claim 9.

11. means for acquiring first text data which is text data relating to the work history of the subject included in curriculum vitae data representing the work history of the subject; a means for generating a prompt by including a structured document output in response to providing information based on the first text data to a generative model and a category name in an instruction to extract the work history information of the subject by category name, and providing the prompt to the generative model to acquire the information of the subject by the category name; An information processing device comprising:

12. The computer A step of acquiring first text data which is text data relating to the work history of the subject included in curriculum vitae data representing the work history of the subject; generating a prompt by including a structured document output in response to providing information based on the first text data to a generative model and a category name in an instruction to extract work history information of the subject by category name, and providing the prompt to a generative model to acquire information of the subject by the category name; How to do it.

13. means for acquiring first text data which is text data relating to the work history of the subject included in curriculum vitae data representing the work history of the subject; a means for generating a prompt by including a structured document output in response to providing information based on the first text data to a generative model and a category name in an instruction to extract the work history information of the subject by category name, and providing the prompt to the generative model to acquire the information of the subject by the category name; A system comprising:

Citation Information

Patent Citations

  • Apparatuses, methods and systems for automated online data submission

    US8600931B1