Computer program, information processing method, and information processing device
A language model-based system extracts and integrates labels from patent text to address the lack of effective label generation for patent documents, improving searchability by generating concise problem labels.
Patent Information
- Application Number
- JP2024018757
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-09
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-02-09
AI Technical Summary
Existing technologies lack effective methods for generating additional information such as tags or labels for text information in patent documents, particularly for identifying and summarizing the problems described in patent documents.
A computer program and information processing device utilize a language model that has undergone machine learning to extract sentences containing specified terms from patent text, generate labels indicating the content of these sentences, and integrate similar labels based on frequency and similarity, supporting the generation of concise problem labels for patent documents.
Enables efficient generation of concise problem labels for patent documents, improving searchability and usability by summarizing and merging similar labels, thereby enhancing the search and retrieval of patent information.
Smart Images

Figure 2025122975000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a computer program, an information processing method, and an information processing device. [Background technology]
[0002] Patent document 1 describes a proposal device that extracts multiple terms from multiple patent documents, obtains a score regarding the importance of the term in a patent classification based on the frequency of occurrence of each of the multiple terms, selects at least one of the multiple terms based on the score of each of the multiple terms, and proposes a revision of the patent classification based on the selected at least one term. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-48781 Summary of the Invention [Problem to be solved by the invention]
[0004] An object of the present disclosure is to provide a computer program, an information processing method, and an information processing device that are expected to support the generation of additional information such as tags or labels for text information such as patent documents. [Means for solving the problem]
[0005] A computer program according to one embodiment causes a computer to acquire text information relating to a patent, extract sentences containing specified terms from the acquired text information, and, based on the extracted sentences, generate one or more labels indicating the content of the sentences using a language model that has undergone machine learning in advance.
[0006] In one embodiment of the computer program, the label indicates the content of the assignment described in the text information.
[0007] A computer program according to one embodiment extracts sentences containing the predetermined term from English text information.
[0008] In one embodiment, the computer program extracts, when the text information is in a language other than English, a sentence containing the predetermined term from the text information obtained by translating the text information into English.
[0009] A computer program according to one embodiment acquires multiple pieces of text information relating to multiple patents, generates one or more labels for each of the multiple pieces of text information, and integrates the multiple labels generated for the multiple pieces of text information.
[0010] A computer program according to one embodiment calculates the number of occurrences of each label included in the generated plurality of labels, determines a reference label from the plurality of labels based on the calculated number of occurrences, calculates a similarity between the determined reference label and labels other than the reference label, and merges the labels other than the reference label into the reference label based on the calculated similarity.
[0011] In one embodiment of the computer program, the text information is text relating to the abstract, background, or problem of a patent document.
[0012] A computer program according to one embodiment extracts sentences containing the predetermined term, and then removes unnecessary symbols from the sentences.
[0013] A computer program according to an embodiment displays the extracted sentences and the labels generated based on the sentences on a display unit in association with each other.
[0014] An information processing method according to one embodiment includes an acquisition step in which an information processing device acquires text information relating to a patent, an extraction step in which a sentence containing a predetermined term is extracted from the acquired text information, and a generation step in which, based on the extracted text, a language model that has undergone machine learning in advance is used to generate one or more labels indicating the content of the text.
[0015] An information processing device according to one embodiment includes a processing unit that acquires text information related to a patent, extracts sentences containing predetermined terms from the acquired text information, and generates one or more labels indicating the content of the sentences based on the extracted sentences using a language model that has undergone machine learning in advance.
[0016] In one embodiment, it is expected that the generation of additional information such as tags or labels for text information such as patent documents will be supported. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a schematic diagram illustrating an overview of an information processing system according to an embodiment of the present invention. [Figure 2] 1 is a block diagram showing an example of a configuration of an information processing device according to an embodiment of the present invention; [Figure 3] FIG. 2 is a block diagram showing the configuration of a terminal device according to the present embodiment. [Figure 4] 10 is a flowchart illustrating an example of the procedure of a task label generation process performed by an information processing device according to the present embodiment. [Figure 5] 10 is a flowchart illustrating an example of the procedure of a task label generation process performed by an information processing device according to the present embodiment. [Figure 6] FIG. 2 is a schematic diagram illustrating an example of input and output information of a large-scale language model. [Figure 7] 10 is a flowchart illustrating an example of a procedure for generating subject labels for a plurality of patent documents performed by an information processing device according to the present embodiment. [Figure 8] 10 is a flowchart illustrating an example of a procedure of a process performed by a terminal device according to the present embodiment. [Figure 9] FIG. 10 is a schematic diagram showing an example of label information displayed by a terminal device. [Figure 10] FIG. 10 is a schematic diagram illustrating a configuration of an information processing system according to a modified example. DETAILED DESCRIPTION OF THE INVENTION
[0018] Specific examples of the information processing system according to the present embodiment will be described below with reference to the drawings. The present technology is not limited to these examples, but is defined by the claims, and is intended to include all modifications within the meaning and scope of the claims.
[0019] <System Overview> FIG. 1 is a schematic diagram for explaining an overview of an information processing system according to this embodiment. The information processing system according to this embodiment is a system in which an information processing device 1 generates additional information such as tags or labels (hereinafter referred to as problem labels) related to problems of inventions described in text information of patent documents. In this embodiment, tags or labels are information such as words, word combinations, or sentences (shorter than the original sentences) that succinctly express the content of the text information, and are, for example, information on character strings of several to several tens of characters. In this embodiment, problem labels are information such as words, word combinations, or sentences that succinctly express the problems of inventions described in text information such as patent documents. In this embodiment, patent documents are documents that publish the contents of patent applications, utility model registration applications, etc., and may include published patent gazettes, patent gazettes, utility model gazettes, etc.
[0020] In this example, the information processing system is configured to include an information processing device 1, a terminal device 3, and a patent management server device 5. The terminal device 3 in this embodiment is a device used by a user, and can be configured using a general-purpose information processing device such as a personal computer, a smartphone, or a tablet terminal. The user uses the terminal device 3 to upload data (files) of patent documents for which the user wishes to generate a challenge label to the information processing device 1. Alternatively, the user may notify the information processing device 1 of identification information (application number, publication number, publication number, announcement number, patent number, etc.) of the patent document for which the user wishes to generate a challenge label.
[0021] The information processing device 1 according to this embodiment can be configured by installing a computer program according to this embodiment in a general-purpose information processing device such as a server computer or a personal computer. The information processing device 1 acquires patent document data from a terminal device 3 and generates a problem label indicating the problem of the invention from the content described in the text information of the acquired patent document. If the terminal device 3 acquires patent document identification information instead of patent document data, the information processing device 1 acquires the corresponding patent document data from a patent management server device 5 managed and operated by the Japan Patent Office or the like. The information processing device 1 requests the patent management server device 5 to transmit the document by specifying the identification information acquired from the terminal device 3. In response to this document request, the patent management server device 5 reads the patent document data stored in a patent document DB (database) 6 and transmits it to the information processing device 1. The information processing device 1 receives patent document data from the patent management server device 5 and generates a problem label for the patent document.
[0022] In this embodiment, the task label generated by the information processing device 1 is information on a string of characters expressed as a word or a combination of words of several to a dozen characters in Japanese, such as "reducing power consumption," "maintaining comfort," or "energy saving." The task label does not have to be in Japanese, and can be generated in various languages such as English or Chinese.
[0023] The text information of patent documents for which the information processing device 1 generates issue labels does not have to be in Japanese, and may be written in various languages such as English or Chinese. However, in this embodiment, the information processing device 1 basically processes in English, and for patent documents written in languages other than English, it translates them into English in advance and uses the translated English text information in subsequent processing. However, the information processing device 1 may perform processing in a language other than English, and the processing language may be Japanese, Chinese, etc.
[0024] In this embodiment, the information processing device 1 extracts sentence information corresponding to items such as "background" or "task" from all sentence information (text information) included in the obtained patent documents. The information processing device 1 extracts one or more sentences containing predetermined terms (such as "not" or "difficult") from the extracted sentence information such as the task as the task sentence. The information processing device 1 generates a task label based on the extracted one or more task sentences.
[0025] The information processing device 1 according to this embodiment uses a large language model (LLM) that has undergone machine learning in advance to generate issue labels for patent documents. The large-scale language model may employ a learning model such as a Transformer equipped with an attention mechanism in a large-scale neural network, BERT (Bidirectional Encoder Representations from Transformers), or GPT (Generative Pre-trained Transformer). The large-scale language model used by the information processing device 1 may be a widely available general-purpose large-scale language model, or may be one that has been trained on information related to patents, etc., through fine tuning.
[0026] The information processing device 1 inputs one or more task sentences containing predetermined terms extracted from patent documents into a large-scale language model, and also inputs instructions (prompts) to generate one or more task labels representing the content of the task contained in the task sentences into the large-scale language model. The information processing device 1 acquires the text information output by the large-scale language model in response to these inputs, and extracts one or more task labels contained in the acquired text information, thereby enabling the information processing device 1 to generate task labels for patent documents.
[0027] In the information processing system according to this embodiment, a user can specify multiple (e.g., 100 or more) patent documents at once and request the generation of problem labels. In this case, the information processing device 1 generates problem labels for each of the multiple patent documents using the above-described process, and then merges similar problem labels. Problem labels generated by a large-scale language model may express the same or similar problem differently depending on the wording used in each patent document. Therefore, the information processing device 1 merges similar problem labels based on the similarity and frequency of occurrence of the multiple problem labels generated by the large-scale language model. This allows the information processing device 1 to improve the user's convenience when searching for patent documents based on problem labels.
[0028] The information processing device 1 transmits information about the challenge label generated for the patent document provided by the terminal device 3 to the requesting terminal device. The terminal device 3 receives the information from the information processing device 1 and displays the information about the challenge label for the patent document specified by the user. At this time, the information processing device 1 transmits information about the challenge sentence (one or more challenge sentences input to the large-scale language model) used to generate the challenge label to the terminal device 3, and the terminal device 3 may display this information as the sentence that served as the basis together with the challenge label.
[0029] <Device configuration> 2 is a block diagram showing an example of the configuration of an information processing device 1 according to this embodiment. The information processing device 1 according to this embodiment can be realized by installing a predetermined application program or the like in a general-purpose information processing device such as a personal computer or a server computer. The information processing device 1 according to this embodiment is configured to include a processing unit (processor) 11, a memory unit (storage) 12, and a communication unit (transceiver) 13. In this embodiment, the processing will be described as being performed by one information processing device 1, but the processing of the information processing device 1 may be distributed among a plurality of devices.
[0030] The processing unit 11 is configured using an arithmetic processing device such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit) or a quantum processor, a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The processing unit 11 reads and executes a program 12a stored in the storage unit 12 to perform various processes such as a process of acquiring patent documents and a process of generating problem labels from text information included in the patent documents.
[0031] The storage unit 12 is configured using a large-capacity storage device such as a hard disk or an SSD (Solid State Drive). The storage unit 12 stores various programs executed by the processing unit 11 and various data required for the processing of the processing unit 11. In this embodiment, the storage unit 12 stores a program 12a executed by the processing unit 11. The storage unit 12 is provided with a model information storage unit 12b that stores information related to a trained learning model used in the processing performed by the information processing device 1. The storage unit 12 is provided with a term storage unit 12c that stores predetermined terms used when extracting task sentences related to tasks from text information included in patent documents.
[0032] In this embodiment, the program (computer program, program product) 12a is provided in a form recorded on a recording medium 99 such as a memory card or an optical disc, and the information processing device 1 reads the program 12a from the recording medium 99 and stores it in the storage unit 12. However, the program 12a may also be written to the storage unit 12 during the manufacturing stage of the information processing device 1. The program 12a may be distributed by a remote server device or the like and acquired by the information processing device 1 via communication. The program 12a may be recorded on the recording medium 99 and read by a writing device and written to the storage unit 12 of the information processing device 1. The program 12a may be provided in a form distributed via a network, or may be provided in a form recorded on the recording medium 99.
[0033] The model information storage unit 12b stores information about a learning model that has been previously subjected to machine learning. The information about the learning model may include information indicating the configuration of the learning model and information such as the values of internal parameters determined by machine learning. In this embodiment, the model information storage unit 12b stores information about a large-scale language model for generating a problem label for a problem in a patent document based on one or more problem sentences extracted from the patent document. The model information storage unit 12b stores information about a learning model that translates text information included in a patent document into English. The learning model that generates the problem label and the learning model that translates into English may be the same learning model.
[0034] In this embodiment, information about the learning model is stored in the information processing device 1, and processing using the learning model is performed by the information processing device 1, but this is not limited to this. Information about the learning model may be stored in a device different from the information processing device 1, and this device may perform processing using the learning model, and the information processing device 1 may acquire the processing results from this device. Machine learning processing of the learning model may be performed by the information processing device 1, or may be performed by a device different from the information processing device 1.
[0035] The terminology storage unit 12c stores one or more terms predetermined by a designer, administrator, or the like of the information processing system according to this embodiment. In this embodiment, the terminology storage unit 12c stores in advance English words such as "but," "not," "important," "problem," "issue," "however," "can't," "isn't," and "didn't." The above terms are extracted by the developer of the information processing system according to this embodiment based on the results of collecting and investigating sentences contained in patent documents, and are terms that are often included in sentences related to problems. However, the above terms are merely examples and are not limited to these.
[0036] The communication unit 13 transmits and receives data to and from devices such as the terminal device 3 and the patent management server device 5 via a wired or wireless network N. In this embodiment, the information processing device 1 can acquire patent documents or patent document identification information, etc. from the terminal device 3 and transmit information about the assignment label generated for the patent document to the terminal device 3 by the communication unit 13 communicating with one or more terminal devices 3. The information processing device 1 can request the patent management server device 5 to transmit patent document data corresponding to the patent document identification information acquired from the terminal device 3 and receive the patent document data transmitted from the patent management server device 5 in response to the request by the communication unit 13 communicating with the patent management server device 5. The communication unit 13 transmits data provided by the processing unit 11 to other devices, receives data from other devices, and provides the received data to the processing unit 11.
[0037] The storage unit 12 may be an external storage device connected to the information processing device 1. The information processing device 1 may be a multi-computer including multiple computers, or may be a virtual machine virtually constructed by software. The information processing device 1 is not limited to the above configuration, and may include a reading unit that reads information stored in a portable storage medium, an input unit that accepts operation input, or a display unit that displays images.
[0038] In the information processing device 1 according to this embodiment, the processing unit 11 reads and executes the program 12a stored in the storage unit 12, whereby a patent document acquisition unit 11a, a problem sentence extraction unit 11b, a problem label generation unit 11c, a problem label integration unit 11d, a label information transmission processing unit 11e, etc. are realized as software functional units in the processing unit 11. In this figure, functional units related to the generation processing of problem labels for patent documents are shown as functional units of the processing unit 11, and functional units related to other processing are not shown.
[0039] The patent document acquisition unit 11a performs a process of acquiring data on patent documents that are the subject of generation of a problem label. The patent document acquisition unit 11a communicates with the terminal device 3 via the communication unit 13 and acquires the patent document data by receiving the patent document data transmitted from the terminal device 3. When the patent document acquisition unit 11a receives identification information for a patent document from the terminal device 3, it communicates with the patent management server device 5 via the communication unit 13 and requests the patent management server device 5 to transmit data on the patent document corresponding to the identification information acquired from the terminal device 3. The patent document acquisition unit 11a acquires the patent document data by receiving data transmitted by the patent management server device 5 in response to this request via the communication unit 13. The patent document acquisition unit 11a stores the acquired patent document data in the memory unit 12.
[0040] In this embodiment, the patent document data acquired by the patent document acquisition unit 11a may be data in a format including text information and image information in one file, such as PDF (Portable Document Format), or may be data divided into multiple files, such as text data and image data in JPEG (Joint Photographic Experts Group) or GIF (Graphics Interchange Format). In this embodiment, the patent document acquisition unit 11a acquires at least the text information of the patent document, and does not necessarily acquire image data of the patent document.
[0041] The subject sentence extraction unit 11b performs a process of extracting subject sentences included in patent documents acquired by the patent document acquisition unit 11a. First, the subject sentence extraction unit 11b extracts sentence information (text information) included in the patent document data, and if this sentence information is written in a language other than English, translates it into English using a learning model for translation. At this time, the subject sentence extraction unit 11b may perform the translation using a learning model for translation stored in the model information storage unit 12b, or may send the sentence information to another device equipped with a learning model for translation to request translation into English, and then acquire the sentence information translated into English.
[0042] The problem sentence extraction unit 11b extracts problem sentences from the entire text information extracted from the patent document, focusing on text information described in sections such as "Background Art" and "Technical Problem." The problem sentence extraction unit 11b first extracts, as problem sentences, sentences containing at least one predetermined term stored in the term storage unit 12c from among the sentences or sentences described in the "Problem to be Solved by the Invention" section. If there is no sentence containing the predetermined term in the "Problem to be Solved by the Invention," the problem sentence extraction unit 11b extracts, as problem sentences, sentences containing the predetermined term from among the sentences or sentences described in the "Background Art" section. If there is no sentence containing the predetermined term in the "Background Art," the problem sentence extraction unit 11b extracts the sentence in the "Summary" of the patent document as problem sentences. The problem sentence extraction unit 11b may perform a process to remove unnecessary symbols, etc. from the extracted problem sentences.
[0043] The problem label generation unit 11c performs a process of generating problem labels indicating the content of the problem described in the patent document based on the problem sentence extracted by the problem sentence extraction unit 11b. In this embodiment, the problem label generation unit 11c generates problem labels based on the problem sentence using a large-scale language model stored in the model information storage unit 12b. The problem label generation unit 11c inputs the extracted problem sentence to the large-scale language model and issues an instruction (prompt) to the large-scale language model to generate one or more problem labels indicating the content of the problem contained in the problem sentence. The problem label generation unit 11c acquires sentence information (text data) output by the large-scale language model in accordance with the above instruction, and generates one or more problem labels for the patent document by extracting one or more problem labels contained in the acquired sentence information.
[0044] The problem label integrating unit 11d integrates similar problem labels generated for multiple patent documents. In this embodiment, the problem label integrating unit 11d integrates problem labels by modifying problem labels with low occurrence counts into similar problem labels with high occurrence counts based on the frequency of occurrence of each problem label and the similarity between the problem labels. For example, if 100 patent documents are provided from the terminal device 3 as the subject for generating problem labels, and three problem labels are generated for each patent document, resulting in 300 problem labels, the problem label integrating unit 11d calculates the frequency of occurrence of each problem label (the number of times the same problem label is included among the 300 labels), and sets the problem label with the highest frequency of occurrence (e.g., the top 1% or 10%) as the reference label. The problem label integrating unit 11d calculates the similarity between the reference label and problem labels other than the reference label, and modifies the problem labels other than the reference label into the most similar reference label. At this time, the problem label integration unit 11d does not need to modify the problem labels other than the reference label for which there is no reference label whose similarity exceeds the threshold value. This allows the problem label integration unit 11d to integrate the 300 problem labels into a predetermined number or a number of problem labels close to this predetermined number. The numbers 300 and the top 1% used in the above explanation are examples and are not limited to these. If the number of patent documents provided by the terminal device 3 is small, for example, one to several, the problem labels do not need to be integrated by the problem label integration unit 11d.
[0045] The label information transmission processing unit 11e performs processing to transmit information about one or more task labels generated for a patent document to the terminal device 3 that requested the generation of the task labels. The label information transmission processing unit 11e acquires the task labels generated for the patent document by the task label generation unit 11c, or the task labels integrated by the task label integration unit 11d as necessary, and transmits information associating the identification information of the patent document with the generated one or more task labels to the terminal device 3 as the result of generating the task labels. The label information transmission processing unit 11e may transmit information associating the task sentences extracted by the task sentence extraction unit 11b with the identification information and task labels of the patent document to the terminal device 3.
[0046] 3 is a block diagram showing the configuration of a terminal device 3 according to this embodiment. The terminal device 3 according to this embodiment is configured to include a processing unit (processor) 31, a memory unit (storage) 32, a communication unit (transceiver) 33, a display unit (display) 34, and an operation unit 35. The terminal device 3 is a device used by a user according to the patent document, and can be configured using an information processing device such as a personal computer, a smartphone, or a tablet terminal device.
[0047] The processing unit 31 is configured using an arithmetic processing unit such as a CPU or an MPU, a ROM, etc. The processing unit 31 reads and executes a program 32a stored in the storage unit 32, thereby performing various processes such as a process of accepting input of information related to patent documents for which a problem label is to be generated, and a process of displaying information related to the problem label generated for the patent document.
[0048] The storage unit 32 is configured using a non-volatile memory element such as a flash memory or a storage device such as a hard disk. The storage unit 32 stores various programs executed by the processing unit 31 and various data required for processing by the processing unit 31. In this embodiment, the storage unit 32 stores the program 32a executed by the processing unit 31. In this embodiment, the program 32a is distributed by a remote server device or the like, and the terminal device 3 acquires the program 32a via communication and stores it in the storage unit 32. However, the program 32a may also be written to the storage unit 32 during the manufacturing stage of the terminal device 3. The program 32a may be read by the terminal device 3 from a recording medium 98 such as a memory card or an optical disk and stored in the storage unit 32. The program 32a may also be read from the recording medium 98 by a writing device and written to the storage unit 32 of the terminal device 3. The program 32a may be provided in the form of distribution via a network or in the form of being recorded on the recording medium 98.
[0049] The communication unit 33 communicates with various devices via a network N including a mobile phone communication network, a wireless LAN, the Internet, etc. In this embodiment, the communication unit 33 communicates with the information processing device 1 via the network N. The communication unit 33 transmits data provided by the processing unit 31 to other devices, and provides data received from other devices to the processing unit 31.
[0050] The display unit 34 is configured using a liquid crystal display or the like, and displays various images, characters, etc. based on processing by the processing unit 31. The operation unit 35 accepts user operations and notifies the processing unit 31 of the accepted operations. The operation unit 35 accepts user operations via input devices such as mechanical buttons or a touch panel provided on the surface of the display unit 34. The operation unit 35 may be input devices such as a mouse and a keyboard, and these input devices may be configured to be detachable from the terminal device 3.
[0051] In the terminal device 3 according to this embodiment, the processing unit 31 reads and executes the program 32a stored in the storage unit 32, thereby realizing the patent document acquisition unit 31a, the display processing unit 31b, etc. as software functional units in the processing unit 31. The program 32a may be a program dedicated to the information processing system according to this embodiment, or may be a general-purpose program such as an internet browser or a web browser.
[0052] The patent document acquisition unit 31a performs a process of acquiring information about one or more patent documents for which a challenge label is to be generated. If the user already has the patent document data (file), the patent document acquisition unit 31a acquires the patent document by accepting a data selection operation from the user, and transmits the acquired patent document data to the information processing device 1 to request the generation of a challenge label. If the user does not have the patent document data, the patent document acquisition unit 31a acquires information about the patent document by accepting input of patent document identification information (application number, publication number, publication number, announcement number, patent number, etc.) from the user, and transmits the acquired patent document identification information to the information processing device 1 to request the generation of a challenge label.
[0053] The display processing unit 31b performs processing to display various characters, images, etc. on the display unit 34. In this embodiment, the display processing unit 31b displays information about the assignment labels generated for patent documents on the display unit 34. The display processing unit 31b receives information about the assignment labels sent by the information processing device 1 in response to a request from the patent document acquisition unit 31a via the communication unit 13, and based on the received information, displays a list of the identification information of the patent document specified by the user, one or more assignment labels generated by the information processing device 1 for this patent document, and the assignment sentences that served as the basis for generating these assignment labels, in association with each other.
[0054] <Issue label generation process> 4 and 5 are flowcharts showing an example of the procedure of the challenge label generation process performed by the information processing device 1 according to this embodiment. The patent document acquisition unit 11a of the processing unit 11 of the information processing device 1 according to this embodiment receives information transmitted from the terminal device 3 via the communication unit 13, and thereby determines whether or not a request for generating a challenge label for a patent document has been received from the terminal device 3 (step S1). If a request for generating a challenge label has not been received (S1: NO), the patent document acquisition unit 11a waits until a request for generating a challenge label is received from the terminal device 3.
[0055] When a request for generating a problem label is received (S1: YES), the patent document acquisition unit 11a determines whether or not it has received identification information of the patent document for which the problem label is to be generated together with the request from the terminal device 3 (step S2). When it has not received the identification information (step S2: NO), that is, when it has received patent document data (file) from the terminal device 3, the patent document acquisition unit 11a stores the patent document data acquired from the terminal device 3 in the storage unit 12 (step S5).
[0056] If the identification information of the target patent document is received (S2: YES), the patent document acquisition unit 11a communicates with the patent management server device 5 via the communication unit 13 and requests the patent management server device 5 to transmit data on the patent document corresponding to the identification information provided by the terminal device 3 (step S3). In response to this request, the patent management server device 5 reads the data on the patent document corresponding to the requested identification information from the patent document DB 6 and transmits it to the information processing device 1. The patent document acquisition unit 11a of the information processing device 1 receives the data on the patent document transmitted from the patent management server device 5 via the communication unit 13 (step S4). The patent document acquisition unit 11a stores the data on the patent document acquired from the patent management server device 5 in the memory unit 12 (step S5).
[0057] Next, the challenge text extraction unit 11b of the processing unit 11 determines whether or not the patent document is written in English based on the patent document data stored in the storage unit 12 in step S5 (step S6). The challenge text extraction unit 11b can determine whether or not the patent document is written in English based on whether or not the text included in the patent document satisfies predetermined conditions for determining English (such as whether the characters included in the patent document are alphabetic). The challenge text extraction unit 11b may input the text of the patent document into a large-scale language model to determine whether or not it is in English. Any method may be used to determine whether or not a patent document is in English.
[0058] If the patent document is not written in English (S6: NO), the problem sentence extraction unit 11b translates the patent document into English (step S7) and proceeds to step S8. The problem sentence extraction unit 11b may translate the text of the patent document into English using a large-scale language model, or may translate the text of the patent document into English using a server device that provides a translation service. The technology for translating text from a language other than English into English is an existing technology, so a detailed explanation will be omitted. The problem sentence extraction unit 11b may use any method to translate the text of the patent document into English. If the patent document is in English (S6: YES), the problem sentence extraction unit 11b proceeds to step S8 without translating the patent document.
[0059] Next, the problem sentence extraction unit 11b determines whether or not there is a sentence containing the predetermined term stored in the term storage unit 12c in the text information of the "problem related to the invention" included in the patent document (step S8). The problem sentence extraction unit 11b can determine which part of the patent document corresponds to the "problem related to the invention" by searching for item names enclosed in parentheses or item names emphasized by changing the font size in the patent document. If there is no text information corresponding to the "problem related to the invention" in the patent document, the problem sentence extraction unit 11b may determine in step S8 that there is no sentence containing the predetermined term in the "problem related to the invention." If there is a sentence containing the predetermined term in the "problem related to the invention" (S8: YES), the problem sentence extraction unit 11b extracts one or more sentences containing the predetermined term from the text information of the "problem related to the invention" as the problem sentence (step S9), and proceeds to step S13.
[0060] If there is no sentence containing a predetermined term in the "Problem Related to the Invention" (S8: NO), the problem sentence extraction unit 11b determines whether there is a sentence containing a predetermined term stored in the term storage unit 12c in the text information of the "background art" included in the patent document (step S10). The problem sentence extraction unit 11b can determine which part of the patent document corresponds to the "background art" by searching for item names enclosed in parentheses or item names emphasized by changing the font size in the patent document. If there is no text information corresponding to the "background art" in the patent document, the problem sentence extraction unit 11b may determine in step S10 that there is no sentence containing a predetermined term in the "background art." If there is a sentence containing a predetermined term in the "background art" (S10: YES), the problem sentence extraction unit 11b extracts one or more sentences containing a predetermined term from the text information of the "background art" as the problem sentence (step S11), and proceeds to step S13.
[0061] If there is no sentence containing the predetermined term in the "Background Art" (S10: NO), the subject sentence extraction unit 11b extracts the "abstract" sentence contained in the patent document as the subject sentence (step S12), and proceeds to step S13. The subject sentence extraction unit 11b can determine which part of the patent document corresponds to the "abstract" by searching for item names enclosed in parentheses or item names emphasized by changing the font size in the patent document.
[0062] Next, the problem sentence extraction unit 11b formats the problem sentences extracted from the patent document in step S9, S11, or S12 (step S13). When the problem sentence extraction unit 11b extracts multiple problem sentences from one patent document, it concatenates these multiple problem sentences into one problem sentence. The sentences in the patent document may include sentences such as "~ device 10" or "~ unit 20" that add symbols in the drawings to the end of the names of the components of the invention to clarify the correspondence. The problem sentence extraction unit 11b removes the symbols in such sentences and corrects "~ device 10" to "~ device". The problem sentence extraction unit 11b performs this sentence formatting in step S13. However, the sentence formatting performed by the problem sentence extraction unit 11b may include any process, not limited to combining multiple sentences and removing symbols.
[0063] Next, the problem label generation unit 11c of the processing unit 11 generates a problem label for the patent document based on the problem sentence formatted in step S13, using the large-scale language model stored in the model information storage unit 12b (step S14). The problem label generation unit 11c inputs the problem sentence extracted from the patent document to the large-scale language model, and also inputs an instruction (prompt) to generate one or more problem labels representing the content of the problem contained in the problem sentence to the large-scale language model.
[0064] FIG. 6 is a schematic diagram showing an example of input and output information of a large-scale language model. In this example, the problem label generation unit 11c Please extract labels for the issues from the patent text below. For example, an issue label might look like this: - Cost of production - Improved yield - Prevents dirt *When extracting, be sure to include "no" *There must be only one "の" in each issue label. *Please limit the number of assignment labels to three. *No words that are difficult for humans to interpret ", which instructs the generation of a problem label, and the English problem sentence, "However, the heart pump require ...", are input to the large-scale language model. In response to this input, the large-scale language model generates and outputs three problem labels: "reducing power consumption," "maintaining comfort," and "energy saving." The problem label generation unit 11c inputs the above information and generates problem labels for the patent document by obtaining one or more problem labels output by the large-scale language model in response to this.
[0065] In this example, the task label generation unit 11c inputs an English task sentence to the large-scale language model. However, if the original patent document is written in a language other than English, a sentence in the original language corresponding to the extracted English task sentence may be input to the large-scale language model. In this example, the large-scale language model generates task labels in Japanese, but this is not limited thereto, and task labels may be generated in languages such as English or Chinese. The task label generation unit 11c can accept a language selection from the user of the terminal device 3 and instruct the large-scale language model to generate task labels in the selected language.
[0066] Next, the label information transmission processing unit 11e of the processing unit 11 transmits information about the generated assignment label to the terminal device 3 that requested the generation of the assignment label (step S15), and terminates the processing. The label information transmission processing unit 11e transmits label information to the terminal device 3, including the identification information of the patent document, one or more assignment labels generated for the patent document, and the assignment text used to generate the assignment label. When a patent document in a language other than English is translated into English, the label information transmission processing unit 11e may extract from the original patent document a non-English assignment text corresponding to the English assignment text used to generate the assignment label, and transmit the assignment text in the original language together with the assignment label to the terminal device 3. The terminal device 3 that receives the label information from the information processing device 1 displays the identification information of the patent document, the assignment label, and the assignment text in association with each other. The terminal device 3 can provide the user with the assignment text corresponding to the assignment label as the text in the patent document that is the basis for generating the assignment label.
[0067] FIG. 7 is a flowchart showing an example of the procedure of the problem label generation process for multiple patent documents performed by the information processing device 1 according to this embodiment. When the processing unit 11 of the information processing device 1 according to this embodiment is requested by the terminal device 3 to generate problem labels for multiple patent documents, it appropriately selects one patent document from the multiple patent documents and performs problem label generation process for this one patent document (step S21). In step S21, the processing unit 11 generates one or more problem labels for one patent document according to the procedure shown in the flowcharts of FIGS. 4 and 5. The processing unit 11 determines whether the generation of problem labels has been completed for all patent documents requested by the terminal device (step S22). If the generation of problem labels has not been completed for all patent documents (S22: NO), the processing unit 11 returns to step S21 and appropriately selects unprocessed patent documents and repeatedly generates problem labels.
[0068] When the generation of problem labels for all patent documents has been completed (S22: YES), the problem label integration unit 11d of the processing unit 11 calculates the number of occurrences of each problem label (i.e., how many of each problem label are included in the multiple problem labels) based on the multiple problem labels generated for multiple patent documents (step S23).Based on the calculated number of occurrences of each problem label, the problem label integration unit 11d, for example, sorts the multiple problem labels in order of the number of occurrences, and determines the problem labels with the highest number of occurrences up to a predetermined percentage (1% or 10%, etc.) as the reference labels (step S24).
[0069] Next, the challenge label integrating unit 11d calculates the similarity between the reference label and challenge labels other than the reference label (step S25). At this time, the challenge label integrating unit 11d selects pairs of the reference label and challenge labels other than the reference label and calculates the similarity, and calculates the similarity for all pairs of labels by repeatedly selecting pairs of labels and calculating the similarity. The challenge label integrating unit 11d can calculate the similarity between the two labels by converting the reference label and challenge labels other than the reference label into feature vectors using a pre-generated encoder and calculating the cosine similarity between the two vectors. However, the similarity between the two labels is not limited to being calculated by the cosine similarity of the vectors and may be calculated by any method.
[0070] The task label integrating unit 11d integrates multiple task labels based on the similarity between the reference label and the task labels other than the reference label calculated in step S25 (step S26). The task label integrating unit 11d selects combinations where the similarity between the reference label and the task labels other than the reference label exceeds a threshold, and integrates the task labels by changing the task labels other than the reference label in this combination to the reference label. Task labels other than the reference label where there is no reference label whose similarity exceeds the threshold do not need to be integrated. The above procedure for integrating task labels is an example and is not limited to this. The task label integrating unit 11d may calculate the similarity between two reference labels, select a combination of reference labels whose similarity exceeds a threshold, and integrate the base labels by changing the base label with a lower frequency of occurrence in this combination to the base label with a higher frequency of occurrence.
[0071] The label information transmission processing unit 11e of the processing unit 11 transmits label information that associates identification information for multiple patent documents, one or more generated task labels, and the task sentences used to generate the task labels to the terminal device 3 (step S27), and then terminates the processing.
[0072] 8 is a flowchart showing an example of a processing procedure performed by the terminal device 3 according to this embodiment. The patent document acquisition unit 31a of the processing unit 31 of the terminal device 3 according to this embodiment accepts input of a patent document for which a problem label is to be generated, based on a user's operation on the operation unit 35 (step S31). At this time, the patent document acquisition unit 31a may display a file selection screen or the like to accept the selection of a patent document file, or may accept input of identification information for the patent document. The patent document acquisition unit 31a communicates with the information processing device 1 via the communication unit 33, and requests the information processing device 1 to generate a problem label for the patent document acquired in step S31 (step S32).
[0073] The display processing unit 31b of the processing unit 31 receives the label information sent by the information processing device 1 in response to the request made in step S32 (step S33). The label information received from the information processing device 1 includes, in association with each other, identification information of the patent document, one or more problem labels generated for this patent document, and sentences in the patent document that served as the basis for generating the problem labels. The display processing unit 31b displays the label information received in step S33 on the display unit 34 (step S34), and ends the processing.
[0074] FIG. 9 is a schematic diagram showing an example of label information displayed by the terminal device 3. The terminal device 3 displays a character string indicating the title "Problem Label Generation Results for Patent Documents" at the top of the screen of the display unit 34, and below this title string, displays a list of label information acquired from the information processing device 1 in a table format. In this example, the terminal device 3 displays a table with the following fields: "Patent Document," "Problem Label," and "Reason Text." The "Patent Document" field contains identification information for the patent document entered by the user as the target for generating a problem label. Publication numbers of patent publications such as "JP Patent Publication No. 2023-12345," "JP Patent Publication No. 2022-234567," and "JP Patent Publication No. 2020-001122" may be displayed. The "Problem Label" field contains problem labels, such as "Power Consumption Reduction," "Comfort Maintenance," and "Energy Conservation," generated by the information processing device 1 for the patent document. In this example, three problem labels are generated for one patent document. The "basis sentence" is a sentence such as "However, there is a problem that..." extracted from a patent document, and is the sentence used to generate the issue label.
[0075] The display format of the label information by the terminal device 3 is not limited to that shown in Fig. 9, and any display format may be adopted. The values in the table shown in Fig. 9 are merely examples and are not limiting.
[0076] <Modification> 10 is a schematic diagram for explaining the configuration of an information processing system according to a modified example. In the information processing system according to the modified example, the information processing device 1 communicates with the patent management server device 5 at a predetermined cycle, such as once a day or once a week, to acquire patent documents stored in the patent document DB 6. At this time, the information processing device 1 acquires from the patent management server device 5 patent documents that meet conditions set by an administrator of the information processing system or the like and that have not been acquired before. Various conditions, such as the technical field or the name of the applicant, can be specified as the conditions for acquiring patent documents.
[0077] The information processing device 1 generates issue labels using a large-scale language model for one or more patent documents acquired from the patent management server device 5. The information processing device 1 according to the modified example is equipped with a patent document DB2. The information processing device 1 assigns one or more issue labels generated for a patent document (or identification information for this patent document) acquired from the patent management server device 5 and stores the patent document DB2. This allows the information processing device 1 to collect necessary patent documents from the patent documents stored in the patent document DB6 of the patent management server device 5 and accumulate the information in its own patent document DB2.
[0078] A user of the information processing system according to the modified example can access the information processing device 1 using his / her own terminal device 3 and search, acquire, view, etc., information stored in the patent document DB2. The terminal device 3 according to the modified example accepts input of the wording of the subject label as a search condition from the user and issues a search request for patent documents to the information processing device 1 using the subject label entered by the user as the search condition. Upon receiving the search request from the terminal device 3, the information processing device 1 extracts patent documents from the patent document DB2 that have been assigned a subject label that matches or is similar to the subject label specified as the search condition. The information processing device 1 transmits the patent documents (or their identification information) extracted from the patent document DB2 to the terminal device 3 that originated the search request. The terminal device 3 receives information on the search results for the search request from the information processing device 1 and displays information on the patent documents corresponding to the subject label entered as the search condition on the display unit 34.
[0079] <Summary> In the information processing system according to the present embodiment, the information processing device 1 acquires text information (patent documents) related to patents, extracts sentences containing predetermined terms from the acquired patent documents as challenge sentences, and generates one or more challenge labels indicating the content of the challenge sentences using a large-scale language model that has undergone machine learning based on the extracted challenge sentences. By extracting challenge sentences containing predetermined terms in advance, the information processing device 1 is expected to generate more accurate challenge labels compared to when the entire patent document is input into a large-scale language model to generate challenge labels. As a result, the information processing system according to the present embodiment is expected to reduce the processing load associated with generating challenge labels. Therefore, the information processing system according to the present embodiment is expected to support the generation of challenge labels for patent documents.
[0080] In the information processing system according to this embodiment, the information processing device 1 translates patent documents in languages other than English into English and extracts task sentences containing predetermined terms from the English patent documents. While Japanese sentences may contain ambiguous expressions, English sentences are clearer than Japanese sentences. Therefore, the information processing system according to this embodiment is expected to improve the accuracy of task sentence extraction by extracting task sentences based on English sentences.
[0081] In the information processing system according to this embodiment, the information processing device 1 acquires multiple patent documents, generates one or more challenge labels for each of them, and integrates the multiple challenge labels generated for the multiple patent documents. The information processing device 1 calculates the number of occurrences of each of the multiple challenge labels and determines a reference label from among the multiple challenge labels based on the calculated number of occurrences. The information processing device 1 calculates the similarity between each of the reference labels and challenge labels other than the reference label, and integrates the challenge labels other than the reference label into the reference label based on the calculated similarity. This allows the information processing system according to this embodiment to integrate challenge labels that are not identical but have similar content, and is expected to reduce the number of types of challenge labels generated for multiple patent documents.
[0082] In the information processing system according to this embodiment, the information processing device 1 extracts problem sentences from text information relating to the abstract, background, or problem of a patent document. As a result, the information processing system according to this embodiment is expected to process passages that are likely to contain a description relating to the problem of the invention.
[0083] In the information processing system according to this embodiment, the information processing device 1 removes unnecessary symbols contained in the extracted challenge text. As a result, the information processing system according to this embodiment can remove symbols such as signs that are often contained in the text of patent documents in advance, and input challenge text that does not contain unnecessary information to the large-scale language model, which is expected to improve the accuracy of challenge labels generated by the large-scale language model.
[0084] In the information processing system according to this embodiment, the information processing device 1 transmits information including the problem label generated for the patent document and the problem sentence used to generate the problem label to the terminal device 3, and causes the terminal device 3 to display the problem label and the problem sentence in association with each other. As a result, the information processing system according to this embodiment is expected to provide the user of the terminal device 3 with the problem label of the patent document together with information on the sentence that is the basis for this problem label.
[0085] In this embodiment, the information processing device 1 generates a problem label indicating the content of the problem of the invention described in the patent document, but this is not limited to this. The information processing device 1 may also generate a purpose label indicating the purpose of the invention described in the patent document or an effect label indicating the effect. The information processing system according to this embodiment includes three types of devices: the information processing device 1, the terminal device 3, and the patent management server device 5, but this is not limited to this. The user may directly operate the information processing device 1 rather than requesting the information processing device 1 to generate a problem label via the terminal device 3. The information processing device 1 and the patent management server device 5 may be a single device, or the information processing device 1, the terminal device 3, and the patent management server device 5 may be a single device. The number of devices included in the information processing system and their respective roles may be determined appropriately by the system designer, etc.
[0086] Although the embodiments have been described above, it will be understood that various changes in form and details can be made without departing from the spirit and scope of the claims. [Explanation of symbols]
[0087] 1. Information processing equipment (computer) 2 Patent document DB 3 Terminal Devices 5 Patent management server device 6 Patent document DB 11 Processing section 11a Patent Document Acquisition Department 11b Assigned text extraction part 11c Issue label generation section 11d Issue Label Integration Section 11e Label information transmission processing unit 12 Storage section 12a Program (computer program) 12b Model information storage section 12c Terminology storage 13 Communications Department 31 Processing section 31a Patent Document Acquisition Department 31b Display processing section 32 Storage section 32a Program 33 Communications Department 34 Display section 35 Control section 98,99 Recording media N Network
Claims
1. On the computer, Obtaining patent-related text information, extracting sentences containing predetermined terms from the acquired sentence information; Based on the extracted sentences, one or more labels indicating the content of the sentences are generated using a language model that has been machine-learned in advance. A computer program that executes a process.
2. The label indicates the content of the assignment described in the text information.
2. The computer program of claim 1.
3. extracting sentences containing the predetermined term from English sentence information; 3. A computer program according to claim 1 or claim 2.
4. If the text information is in a language other than English, extracting sentences containing the predetermined term from text information obtained by translating the text information into English.
4. A computer program according to claim 3.
5. Acquire multiple pieces of text information related to multiple patents, generating one or more labels for each of the plurality of pieces of text information; Integrating the plurality of labels generated for the plurality of pieces of text information; 3. A computer program according to claim 1 or claim 2.
6. Calculating the number of occurrences of each label included in the generated plurality of labels; determining a reference label from among the plurality of labels based on the calculated number of occurrences; Calculating the similarity between the determined reference label and a label other than the reference label; Integrating labels other than the reference label into the reference label based on the calculated similarity.
6. A computer program according to claim 5.
7. The text information is a text relating to the abstract, background, or problem of the patent document.
3. A computer program according to claim 1 or claim 2.
8. extracting a sentence containing the predetermined term, and then removing unnecessary symbols from the sentence; 3. A computer program according to claim 1 or claim 2.
9. The extracted sentences are associated with labels generated based on the sentences and displayed on a display unit.
3. A computer program according to claim 1 or claim 2.
10. The information processing device an acquisition step of acquiring text information related to a patent; an extraction step of extracting sentences including predetermined terms from the acquired sentence information; a generation step of generating one or more labels indicating the content of the sentence based on the extracted sentence, using a language model that has been machine-learned in advance; An information processing method, including:
11. a processing unit; The processing unit Obtaining patent-related text information, extracting sentences containing predetermined terms from the acquired sentence information; generating one or more labels indicating the content of the extracted sentences using a language model that has been machine-learned in advance; Information processing device.
Citation Information
Patent Citations
Article automatic generation system and method based on natural language processing and image algorithm
CN111428472A
Text processing method and device, computer equipment and storage medium
CN115859176A
Text processing method and system and electronic equipment
CN115906858A
Multi-document fusion deep learning title generation method based on continuous learning
CN116842934A
Document retrieval system, document retrieval method and document retrieval program
JP2007058706A