Computer program, information processing method, and information processing device
A computer program and information processing device use a machine-learned language model to extract and integrate labels from patent documents, addressing the lack of efficient label generation for patent text, enhancing searchability by summarizing technical challenges across documents.
Patent Information
- Application Number
- PCT/JP2025/003975
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-09
- Filing Date
- 2025-02-06
- Publication Date
- 2025-08-14
AI Technical Summary
Existing systems lack efficient methods for generating additional information such as tags or labels for text information in patent documents, particularly for identifying and summarizing the technical problems and challenges within patent documents.
A computer program and information processing device utilize a large language model that has undergone machine learning to extract sentences containing specified terms from patent documents, generate labels indicating the content of these sentences, and integrate similar labels based on frequency and similarity, supporting the generation of concise problem labels for patent documents.
This approach enables the efficient generation of concise problem labels for multiple patent documents, improving searchability and user convenience by summarizing the technical challenges across documents, even when phrased differently.
Smart Images

Figure JP2025003975_14082025_PF_FP_ABST
Abstract
Description
Computer program, information processing method, and information processing device
[0001] The present disclosure relates to a computer program, an information processing method, and an information processing device.
[0002] Patent document 1 describes a proposal device that extracts multiple terms from multiple patent documents, obtains a score regarding the importance of the term in a patent classification based on the frequency of occurrence of each of the multiple terms, selects at least one of the multiple terms based on the score of each of the multiple terms, and proposes a revision of the patent classification based on the selected at least one term.
[0003] Japanese Patent Application Laid-Open No. 2022-48781
[0004] An object of the present disclosure is to provide a computer program, an information processing method, and an information processing device that are expected to support the generation of additional information such as tags or labels for text information such as patent documents.
[0005] A computer program according to one embodiment causes a computer to acquire text information relating to a patent, extract sentences containing specified terms from the acquired text information, and, based on the extracted sentences, generate one or more labels indicating the content of the sentences using a language model that has undergone machine learning in advance.
[0006] In one embodiment of the computer program, the label indicates the content of the assignment described in the text information.
[0007] A computer program according to one embodiment extracts sentences containing the predetermined term from English text information.
[0008] In one embodiment, the computer program extracts, when the text information is in a language other than English, a sentence containing the predetermined term from the text information obtained by translating the text information into English.
[0009] A computer program in one embodiment acquires multiple pieces of text information related to multiple patents, generates one or more labels for each of the multiple pieces of text information, and integrates the multiple labels generated for the multiple pieces of text information.
[0010] A computer program according to one embodiment calculates the number of occurrences of each label included in the generated plurality of labels, determines a reference label from the plurality of labels based on the calculated number of occurrences, calculates a similarity between the determined reference label and labels other than the reference label, and merges the labels other than the reference label into the reference label based on the calculated similarity.
[0011] In one embodiment of the computer program, the text information is text relating to the abstract, background, or problem of a patent document.
[0012] A computer program according to one embodiment extracts sentences containing the predetermined term, and then removes unnecessary symbols from the sentences.
[0013] A computer program according to an embodiment displays the extracted sentences and the labels generated based on the sentences on a display unit in association with each other.
[0014] An information processing method according to one embodiment includes an acquisition step in which an information processing device acquires text information relating to a patent, an extraction step in which a sentence containing a predetermined term is extracted from the acquired text information, and a generation step in which, based on the extracted text, a language model that has undergone machine learning in advance is used to generate one or more labels indicating the content of the sentence.
[0015] An information processing device according to one embodiment includes a processing unit that acquires text information related to a patent, extracts sentences containing predetermined terms from the acquired text information, and generates one or more labels indicating the content of the sentences based on the extracted sentences using a language model that has undergone machine learning in advance.
[0016] In one embodiment, it is expected that the generation of additional information such as tags or labels for text information such as patent documents will be supported.
[0017] FIG. 1 is a schematic diagram for explaining an overview of an information processing system according to the present embodiment. FIG. 2 is a block diagram showing an example of the configuration of an information processing device according to the present embodiment. FIG. 3 is a block diagram showing the configuration of a terminal device according to the present embodiment. FIG. 4 is a flowchart showing an example of the procedure of a challenge label generation process performed by an information processing device according to the present embodiment. FIG. 5 is a flowchart showing an example of the procedure of a challenge label generation process performed by an information processing device according to the present embodiment. FIG. 6 is a schematic diagram showing an example of input / output information of a large-scale language model. FIG. 7 is a flowchart showing an example of the procedure of a challenge label generation process for a plurality of patent documents performed by an information processing device according to the present embodiment. FIG. 8 is a schematic diagram showing an example of a display of label information by a terminal device. FIG. 9 is a schematic diagram for explaining the configuration of an information processing system according to a modified example.
[0018] Specific examples of the information processing system according to the present embodiment will be described below with reference to the drawings. The present technology is not limited to these examples, but is defined by the claims, and is intended to include all modifications within the meaning and scope of the claims.
[0019] <System Overview> FIG. 1 is a schematic diagram illustrating an overview of an information processing system according to this embodiment. The information processing system according to this embodiment is a system in which an information processing device 1 generates additional information, such as tags or labels (hereinafter referred to as "problem labels"), related to the problem of an invention described in text information of a patent document. In this embodiment, the tags or labels are information such as words, word combinations, or sentences (shorter than the original text) that succinctly express the content of the text information, for example, character strings of several to several dozen characters. In this embodiment, the problem labels are information such as words, word combinations, or sentences that succinctly express the problem of an invention described in text information such as a patent document. In this embodiment, patent documents are documents that publish the contents of patent applications, utility model registration applications, etc., and may include published patent gazettes, patent publications, utility model publications, etc.
[0020] In this example, the information processing system is configured to include an information processing device 1, a terminal device 3, and a patent management server device 5. The terminal device 3 in this embodiment is a device used by a user and can be configured using a general-purpose information processing device such as a personal computer, a smartphone, or a tablet terminal. The user uses the terminal device 3 to upload data (files) of patent documents for which the user wishes to generate challenge labels to the information processing device 1. Alternatively, the user may notify the information processing device 1 of identification information (application number, publication number, publication number, announcement number, patent number, etc.) of the patent documents for which the user wishes to generate challenge labels.
[0021] The information processing device 1 according to this embodiment can be configured by installing a computer program according to this embodiment on a general-purpose information processing device such as a server computer or a personal computer. The information processing device 1 acquires patent document data from a terminal device 3 and generates a problem label indicating the problem of the invention from the content described in the text information of the acquired patent document. If the terminal device 3 acquires patent document identification information rather than patent document data, the information processing device 1 acquires the corresponding patent document data from a patent management server device 5 managed and operated by the Japan Patent Office or the like. The information processing device 1 requests the patent management server device 5 to transmit the document by specifying the identification information acquired from the terminal device 3. In response to this document request, the patent management server device 5 reads the patent document data stored in a patent document database 6 and transmits it to the information processing device 1. The information processing device 1 receives patent document data from the patent management server device 5 and generates a problem label for the patent document.
[0022] In this embodiment, the task label generated by the information processing device 1 is information on a string of characters expressed as a word or a combination of words of several to a dozen characters in Japanese, such as "reducing power consumption," "maintaining comfort," or "energy saving." The task label does not have to be in Japanese, and can be generated in various languages such as English or Chinese.
[0023] The text information of patent documents for which the information processing device 1 generates issue labels does not have to be in Japanese, and may be written in various languages such as English or Chinese. However, in this embodiment, the information processing device 1 basically processes in English, and for patent documents written in languages other than English, it translates them into English in advance and uses the translated English text information in subsequent processing. However, the information processing device 1 may perform processing in a language other than English, and the processing language may be Japanese, Chinese, etc.
[0024] In this embodiment, the information processing device 1 extracts sentence information corresponding to items such as "background" or "task" from all sentence information (text information) included in the obtained patent documents. The information processing device 1 extracts one or more sentences containing predetermined terms (such as "not" or "difficult") from the extracted sentence information such as the task as the task sentence. The information processing device 1 generates a task label based on the extracted one or more task sentences.
[0025] The information processing device 1 according to the present embodiment uses a large language model (LLM) that has undergone machine learning in advance to generate issue labels for patent documents. The large language model may employ a learning model such as a Transformer equipped with an attention mechanism in a large-scale neural network, a Bidirectional Encoder Representations from Transformers (BERT), or a Generative Pre-trained Transformer (GPT). The large-scale language model used by the information processing device 1 may be a widely available general-purpose large-scale language model, or may be one that has been trained on information related to patents and the like through fine tuning.
[0026] The information processing device 1 inputs one or more task sentences containing predetermined terms extracted from patent documents into a large-scale language model, and also inputs instructions (prompts) to generate one or more task labels representing the content of the task contained in the task sentences into the large-scale language model. The information processing device 1 acquires the text information output by the large-scale language model in response to these inputs, and by extracting one or more task labels contained in the acquired text information, the information processing device 1 can generate task labels for the patent documents.
[0027] In the information processing system according to this embodiment, a user can specify multiple (e.g., 100 or more) patent documents at once and request the generation of problem labels. In this case, the information processing device 1 generates problem labels for each of the multiple patent documents using the above-described process, and then merges similar problem labels. Problem labels generated by a large-scale language model may express the same or similar problem differently depending on the wording used in each patent document. Therefore, the information processing device 1 merges similar problem labels based on the similarity and frequency of occurrence of the multiple problem labels generated by the large-scale language model. This allows the information processing device 1 to improve the user's convenience when searching for patent documents based on problem labels.
[0028] The information processing device 1 transmits, to the requesting terminal device, information on the challenge label generated for the patent document provided by the terminal device 3. The terminal device 3 receives the information from the information processing device 1 and displays the information on the challenge label for the patent document specified by the user. At this time, the information processing device 1 transmits information on the challenge sentence (one or more challenge sentences input to the large-scale language model) used to generate the challenge label to the terminal device 3, and the terminal device 3 may display this information together with the challenge label as the sentence that served as the basis.
[0029] <Device Configuration> Fig. 2 is a block diagram showing an example of the configuration of an information processing device 1 according to this embodiment. The information processing device 1 according to this embodiment can be realized by installing a predetermined application program or the like in a general-purpose information processing device such as a personal computer or a server computer. The information processing device 1 according to this embodiment is configured to include a processing unit (processor) 11, a memory unit (storage) 12, and a communication unit (transceiver) 13. In this embodiment, the processing will be described as being performed by a single information processing device 1, but the processing of the information processing device 1 may be distributed among multiple devices.
[0030] The processing unit 11 is configured using an arithmetic processing device such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit) or a quantum processor, a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The processing unit 11 reads and executes a program 12a stored in the storage unit 12 to perform various processes such as a process of acquiring patent documents and a process of generating problem labels from text information included in the patent documents.
[0031] The storage unit 12 is configured using a large-capacity storage device such as a hard disk or an SSD (Solid State Drive). The storage unit 12 stores various programs executed by the processing unit 11 and various data required for the processing of the processing unit 11. In this embodiment, the storage unit 12 stores a program 12a executed by the processing unit 11. The storage unit 12 is provided with a model information storage unit 12b that stores information related to a trained learning model used in the processing performed by the information processing device 1. The storage unit 12 is provided with a term storage unit 12c that stores predetermined terms used when extracting task sentences related to tasks from text information included in patent documents.
[0032] In this embodiment, the program (computer program, program product) 12a is provided in a form recorded on a recording medium 99 such as a memory card or an optical disc, and the information processing device 1 reads the program 12a from the recording medium 99 and stores it in the storage unit 12. However, the program 12a may also be written to the storage unit 12 during the manufacturing stage of the information processing device 1. The program 12a may be distributed by a remote server device or the like and acquired by the information processing device 1 via communication. The program 12a may be read from the recording medium 99 by a writing device and written to the storage unit 12 of the information processing device 1. The program 12a may be provided in a form distributed via a network or provided in a form recorded on the recording medium 99.
[0033] The model information storage unit 12b stores information about a learning model that has been previously subjected to machine learning. The information about the learning model may include information indicating the configuration of the learning model and information such as the values of internal parameters determined by machine learning. In this embodiment, the model information storage unit 12b stores information about a large-scale language model for generating a task label for a task in a patent document based on one or more task sentences extracted from the patent document. The model information storage unit 12b stores information about a learning model that translates text information included in a patent document into English. The learning model that generates the task label and the learning model that translates into English may be the same learning model.
[0034] In this embodiment, information about the learning model is stored in the information processing device 1, and processing using the learning model is performed by the information processing device 1, but this is not limited to this. Information about the learning model may be stored in a device different from the information processing device 1, and this device may perform processing using the learning model, and the information processing device 1 may acquire the processing results from this device. Machine learning processing of the learning model may be performed by the information processing device 1 or by a device different from the information processing device 1.
[0035] The terminology storage unit 12c stores one or more terms predetermined by a designer, administrator, or the like of the information processing system according to this embodiment. In this embodiment, the terminology storage unit 12c stores in advance English words such as "but," "not," "important," "problem," "issue," "however," "can't," "isn't," and "didn't." The above terms are extracted by the developer of the information processing system according to this embodiment based on the results of collecting and investigating sentences contained in patent documents, and are terms that are often included in sentences related to problems. However, the above terms are merely examples and are not limited to these.
[0036] The communication unit 13 transmits and receives data to and from devices such as the terminal device 3 and the patent management server device 5 via a wired or wireless network N. In this embodiment, the information processing device 1 acquires patent documents or patent document identification information, etc. from the terminal device 3 by communicating with one or more terminal devices 3 via the communication unit 13, and can transmit information regarding the assignment label generated for the patent document to the terminal device 3. The information processing device 1 can request the patent management server device 5 to transmit patent document data corresponding to the patent document identification information acquired from the terminal device 3 by communicating with the patent management server device 5 via the communication unit 13, and can receive the patent document data transmitted from the patent management server device 5 in response to the request. The communication unit 13 transmits data provided by the processing unit 11 to other devices, receives data from other devices, and provides the received data to the processing unit 11.
[0037] The storage unit 12 may be an external storage device connected to the information processing device 1. The information processing device 1 may be a multi-computer including multiple computers, or may be a virtual machine virtually constructed by software. The information processing device 1 is not limited to the above configuration, and may include a reading unit that reads information stored in a portable storage medium, an input unit that accepts operation input, or a display unit that displays images.
[0038] In the information processing device 1 according to this embodiment, the processing unit 11 reads and executes the program 12a stored in the storage unit 12, whereby a patent document acquisition unit 11a, a problem sentence extraction unit 11b, a problem label generation unit 11c, a problem label integration unit 11d, a label information transmission processing unit 11e, etc. are realized as software functional units in the processing unit 11. In this figure, functional units related to the generation processing of problem labels for patent documents are shown as functional units of the processing unit 11, and functional units related to other processing are not shown.
[0039] The patent document acquisition unit 11a performs a process of acquiring data on patent documents that are the subject of generation of a challenge label. The patent document acquisition unit 11a communicates with the terminal device 3 via the communication unit 13 and acquires the patent document data by receiving the patent document data transmitted from the terminal device 3. When the patent document acquisition unit 11a receives identification information for a patent document from the terminal device 3, it communicates with the patent management server device 5 via the communication unit 13 and requests the patent management server device 5 to transmit data on the patent document corresponding to the identification information acquired from the terminal device 3. The patent document acquisition unit 11a acquires the patent document data by receiving data transmitted by the patent management server device 5 in response to this request via the communication unit 13. The patent document acquisition unit 11a stores the acquired patent document data in the memory unit 12.
[0040] In this embodiment, the patent document data acquired by the patent document acquisition unit 11a may be data in a format that includes text information and image information in one file, such as PDF (Portable Document Format), or may be data divided into multiple files, such as text data and image data in JPEG (Joint Photographic Experts Group) or GIF (Graphics Interchange Format). In this embodiment, the patent document acquisition unit 11a acquires at least the text information of the patent document, but does not necessarily acquire image data of the patent document.
[0041] The challenge sentence extraction unit 11b performs a process of extracting challenge sentences included in patent documents acquired by the patent document acquisition unit 11a. First, the challenge sentence extraction unit 11b extracts sentence information (text information) included in the patent document data, and if this sentence information is written in a language other than English, translates it into English using a translation learning model. At this time, the challenge sentence extraction unit 11b may perform the translation using a translation learning model stored in the model information storage unit 12b, or may send the sentence information to another device equipped with a translation learning model to request translation into English, and then acquire the sentence information translated into English.
[0042] The problem sentence extraction unit 11b extracts problem sentences from the entire text information extracted from the patent document, focusing on text information described in sections such as "Background Art" and "Technical Problem." The problem sentence extraction unit 11b first extracts, as problem sentences, sentences containing at least one predetermined term stored in the term storage unit 12c from among the sentences or sentences described in the "Problem to be Solved by the Invention" section. If there is no sentence containing the predetermined term in the "Problem to be Solved by the Invention," the problem sentence extraction unit 11b extracts, as problem sentences, sentences containing the predetermined term from among the sentences or sentences described in the "Background Art" section. If there is no sentence containing the predetermined term in the "Background Art," the problem sentence extraction unit 11b extracts the sentence in the "Summary" of the patent document as problem sentences. The problem sentence extraction unit 11b may perform a process to remove unnecessary symbols, etc. from the extracted problem sentences.
[0043] The problem label generation unit 11c performs a process of generating problem labels indicating the content of the problem described in the patent document based on the problem sentence extracted by the problem sentence extraction unit 11b. In this embodiment, the problem label generation unit 11c generates problem labels based on the problem sentence using a large-scale language model stored in the model information storage unit 12b. The problem label generation unit 11c inputs the extracted problem sentence to the large-scale language model and provides an instruction (prompt) to the large-scale language model to generate one or more problem labels indicating the content of the problem contained in the problem sentence. The problem label generation unit 11c acquires sentence information (text data) output by the large-scale language model in accordance with the above instruction, and extracts one or more problem labels contained in the acquired sentence information, thereby generating one or more problem labels for the patent document.
[0044] The problem label integration unit 11d performs a process of integrating similar problem labels generated for multiple patent documents. In this embodiment, the problem label integration unit 11d integrates problem labels by modifying problem labels with low occurrence counts into similar problem labels with high occurrence counts based on the number of occurrences of each problem label and the similarity between the problem labels. For example, if 100 patent documents are provided from the terminal device 3 as the subject for generating problem labels, and three problem labels are generated per patent document, resulting in 300 problem labels, the problem label integration unit 11d calculates the number of occurrences of each problem label (the number of instances of the same problem label among the 300 labels) and uses a predetermined number of problem labels with the highest occurrence counts (e.g., the top 1% or 10%) as reference labels. The problem label integration unit 11d calculates the similarity between the reference label and problem labels other than the reference label and modifies the problem labels other than the reference label into the most similar reference label. At this time, the problem label integration unit 11d does not need to modify problem labels other than the reference label for which there is no reference label whose similarity exceeds the threshold. This allows the problem label integration unit 11d to integrate 300 problem labels into a predetermined number or a number close to this predetermined number. The numbers 300 and the top 1% used in the above explanation are examples and are not limited to these. If the number of patent documents provided by the terminal device 3 is small, for example, one to several, the problem labels do not need to be integrated by the problem label integration unit 11d.
[0045] The label information transmission processing unit 11e performs processing to transmit information regarding one or more task labels generated for the patent document to the terminal device 3 that requested the generation of the task labels. The label information transmission processing unit 11e acquires the task labels generated for the patent document by the task label generation unit 11c, or the task labels integrated by the task label integration unit 11d as necessary, and transmits information associating the identification information of the patent document with the one or more generated task labels to the terminal device 3 as the generation result of the task labels. The label information transmission processing unit 11e may transmit information associating the task sentences extracted by the task sentence extraction unit 11b with the identification information and task labels of the patent document to the terminal device 3.
[0046] 3 is a block diagram showing the configuration of a terminal device 3 according to this embodiment. The terminal device 3 according to this embodiment is configured to include a processing unit (processor) 31, a memory unit (storage) 32, a communication unit (transceiver) 33, a display unit (display) 34, and an operation unit 35. The terminal device 3 is a device used by a user according to the patent document, and may be configured using an information processing device such as a personal computer, a smartphone, or a tablet terminal device.
[0047] The processing unit 31 is configured using a processing unit such as a CPU or an MPU, a ROM, etc. The processing unit 31 reads and executes a program 32a stored in the storage unit 32, thereby performing various processes such as a process of accepting input of information related to patent documents for which a challenge label is to be generated, and a process of displaying information related to the challenge label generated for the patent document.
[0048] The storage unit 32 is configured using a non-volatile memory element such as a flash memory or a storage device such as a hard disk. The storage unit 32 stores various programs executed by the processing unit 31 and various data required for processing by the processing unit 31. In this embodiment, the storage unit 32 stores the program 32a executed by the processing unit 31. In this embodiment, the program 32a is distributed by a remote server device or the like, and the terminal device 3 acquires it via communication and stores it in the storage unit 32. However, the program 32a may also be written to the storage unit 32 during the manufacturing stage of the terminal device 3. The program 32a may be recorded on a recording medium 98 such as a memory card or an optical disk, and the terminal device 3 reads the program 32a and stores it in the storage unit 32. The program 32a may be recorded on the recording medium 98 and read by a writing device and written to the storage unit 32 of the terminal device 3. The program 32a may be provided in the form of distribution via a network or in the form of being recorded on the recording medium 98.
[0049] The communication unit 33 communicates with various devices via a network N including a mobile phone communication network, a wireless LAN, the Internet, etc. In this embodiment, the communication unit 33 communicates with the information processing device 1 via the network N. The communication unit 33 transmits data provided by the processing unit 31 to other devices, and provides data received from other devices to the processing unit 31.
[0050] The display unit 34 is configured using a liquid crystal display or the like, and displays various images, characters, etc. based on the processing of the processing unit 31. The operation unit 35 accepts user operations and notifies the processing unit 31 of the accepted operations. The operation unit 35 accepts user operations via input devices such as mechanical buttons or a touch panel provided on the surface of the display unit 34. The operation unit 35 may be input devices such as a mouse and a keyboard, and these input devices may be configured to be detachable from the terminal device 3.
[0051] In the terminal device 3 according to this embodiment, the processing unit 31 reads and executes the program 32a stored in the storage unit 32, thereby realizing the patent document acquisition unit 31a, the display processing unit 31b, etc. as software functional units in the processing unit 31. The program 32a may be a program dedicated to the information processing system according to this embodiment, or may be a general-purpose program such as an internet browser or a web browser.
[0052] The patent document acquisition unit 31a performs a process of acquiring information about one or more patent documents for which a challenge label is to be generated. If the user already has patent document data (files), the patent document acquisition unit 31a acquires the patent document by accepting a data selection operation from the user, and transmits the acquired patent document data to the information processing device 1 to request the generation of a challenge label. If the user does not have patent document data, the patent document acquisition unit 31a acquires information about the patent document by accepting input of patent document identification information (application number, publication number, publication number, announcement number, patent number, etc.) from the user, and transmits the acquired patent document identification information to the information processing device 1 to request the generation of a challenge label.
[0053] The display processing unit 31b performs processing to display various characters, images, etc. on the display unit 34. In this embodiment, the display processing unit 31b displays information regarding the assignment labels generated for the patent documents on the display unit 34. The display processing unit 31b receives information regarding the assignment labels sent by the information processing device 1 in response to a request from the patent document acquisition unit 31a via the communication unit 13, and based on the received information, displays a list of the identification information of the patent document specified by the user, one or more assignment labels generated by the information processing device 1 for this patent document, and the assignment sentences that served as the basis for generating these assignment labels, in association with each other.
[0054] 4 and 5 are flowcharts showing an example of the procedure of the task label generation process performed by the information processing device 1 according to this embodiment. The patent document acquisition unit 11a of the processing unit 11 of the information processing device 1 according to this embodiment receives information transmitted from the terminal device 3 via the communication unit 13, and thereby determines whether or not a request for generating a task label for a patent document has been received from the terminal device 3 (step S1). If a request for generating a task label has not been received (S1: NO), the patent document acquisition unit 11a waits until a request for generating a task label is received from the terminal device 3.
[0055] When a request for generating a challenge label is received (S1: YES), the patent document acquisition unit 11a determines whether or not it has received the identification information of the patent document for which the challenge label is to be generated together with the request from the terminal device 3 (step S2). When the identification information has not been received (step S2: NO), that is, when the patent document data (file) has been received from the terminal device 3, the patent document acquisition unit 11a stores the patent document data acquired from the terminal device 3 in the storage unit 12 (step S5).
[0056] If the identification information of the target patent document is received (S2: YES), the patent document acquisition unit 11a communicates with the patent management server device 5 via the communication unit 13 and requests the patent management server device 5 to transmit data on the patent document corresponding to the identification information provided by the terminal device 3 (step S3). In response to this request, the patent management server device 5 reads the patent document data of the requested identification information from the patent document DB 6 and transmits it to the information processing device 1. The patent document acquisition unit 11a of the information processing device 1 receives the patent document data transmitted from the patent management server device 5 via the communication unit 13 (step S4). The patent document acquisition unit 11a stores the patent document data acquired from the patent management server device 5 in the memory unit 12 (step S5).
[0057] Next, the challenge sentence extraction unit 11b of the processing unit 11 determines whether the patent document is written in English based on the patent document data stored in the storage unit 12 in step S5 (step S6). The challenge sentence extraction unit 11b can determine whether the patent document is written in English based on whether the sentence included in the patent document satisfies predetermined conditions for determining English (such as whether the characters included in the patent document are alphabetic). The challenge sentence extraction unit 11b may input the sentence of the patent document into a large-scale language model to determine whether it is in English. Any method may be used to determine whether a patent document is in English.
[0058] If the patent document is not written in English (S6: NO), the problem sentence extraction unit 11b translates the patent document into English (step S7) and proceeds to step S8. The problem sentence extraction unit 11b may translate the patent document into English using a large-scale language model, or may translate the patent document into English using a server device that provides a translation service. Since technology for translating non-English text into English is an existing technology, detailed explanation will be omitted. The problem sentence extraction unit 11b may use any method to translate the patent document into English. If the patent document is in English (S6: YES), the problem sentence extraction unit 11b proceeds to step S8 without translating the patent document.
[0059] Next, the problem sentence extraction unit 11b determines whether or not the text information for the "problem related to the invention" included in the patent document contains a sentence containing the predetermined term stored in the term storage unit 12c (step S8). The problem sentence extraction unit 11b can determine which part of the patent document corresponds to the "problem related to the invention" by searching for item names enclosed in parentheses or item names highlighted by changing the font size in the patent document. If the patent document does not contain text information corresponding to the "problem related to the invention," the problem sentence extraction unit 11b may determine in step S8 that the "problem related to the invention" does not contain a sentence containing the predetermined term. If the "problem related to the invention" contains a sentence containing the predetermined term (S8: YES), the problem sentence extraction unit 11b extracts one or more sentences containing the predetermined term from the text information for the "problem related to the invention" as problem sentences (step S9), and proceeds to step S13.
[0060] If there is no sentence containing a predetermined term in the "Problem Related to the Invention" (S8: NO), the problem sentence extraction unit 11b determines whether there is a sentence containing a predetermined term stored in the term storage unit 12c in the text information of the "background art" included in the patent document (step S10). The problem sentence extraction unit 11b can determine which part of the patent document corresponds to the "background art" by searching for item names enclosed in parentheses or item names highlighted by changing the font size in the patent document. If there is no text information corresponding to the "background art" in the patent document, the problem sentence extraction unit 11b may determine in step S10 that there is no sentence containing a predetermined term in the "background art." If there is a sentence containing a predetermined term in the "background art" (S10: YES), the problem sentence extraction unit 11b extracts one or more sentences containing a predetermined term from the text information of the "background art" as the problem sentence (step S11) and proceeds to step S13.
[0061] If there is no sentence containing the predetermined term in the "Background Art" (S10: NO), the subject sentence extraction unit 11b extracts the "abstract" sentence contained in the patent document as the subject sentence (step S12), and proceeds to step S13. The subject sentence extraction unit 11b can determine which part of the patent document corresponds to the "abstract" by searching for item names enclosed in parentheses or item names emphasized by changing the font size in the patent document.
[0062] Next, the problem sentence extraction unit 11b formats the problem sentences extracted from the patent document in step S9, S11, or S12 (step S13). When the problem sentence extraction unit 11b extracts multiple problem sentences from a single patent document, it concatenates these multiple problem sentences into a single problem sentence. The text of a patent document may include sentences such as "~device 10" or "~unit 20" that add symbols in drawings to the end of the names of components of the invention to clarify the correspondence. The problem sentence extraction unit 11b removes the symbols in such sentences and corrects "~device 10" to "~device." The problem sentence extraction unit 11b performs this text formatting in step S13. However, the text formatting performed by the problem sentence extraction unit 11b may include any processing, not limited to combining multiple sentences and removing symbols.
[0063] Next, the problem label generation unit 11c of the processing unit 11 generates problem labels for the patent documents based on the problem sentences shaped in step S13, using the large-scale language model stored in the model information storage unit 12b (step S14). The problem label generation unit 11c inputs the problem sentences extracted from the patent documents into the large-scale language model, and also inputs instructions (prompts) to the large-scale language model to generate one or more problem labels representing the content of the problem contained in the problem sentences.
[0064] FIG. 6 is a schematic diagram illustrating an example of input and output information for a large-scale language model. In this example, the problem label generation unit 11c inputs the following sentence, which instructs the generation of problem labels: "Please extract problem labels from the following patent text. For example, problem labels are expressed as follows: - manufacturing cost - improved yield - prevention of contamination *When extracting, be sure to include "no" *Make sure that each "no" is used only once in the problem label *Make sure that the problem label length is three or less *Do not include words that are difficult for humans to interpret" into the large-scale language model, along with the English problem text "However, the heart pump require..." In response to this input, the large-scale language model generates and outputs three problem labels: "reducing power consumption," "maintaining comfort," and "energy conservation." The problem label generation unit 11c inputs the above information and obtains one or more problem labels output by the large-scale language model in response to the input, thereby generating problem labels for the patent document.
[0065] In this example, the task label generation unit 11c inputs an English task sentence to the large-scale language model, but if the original patent document is written in a language other than English, a sentence in the original language corresponding to the extracted English task sentence may be input to the large-scale language model. In this example, the large-scale language model generates task labels in Japanese, but this is not limited to this, and task labels may be generated in languages such as English or Chinese. The task label generation unit 11c can accept a language selection from the user of the terminal device 3 and instruct the large-scale language model to generate task labels in the selected language.
[0066] Next, the label information transmission processing unit 11e of the processing unit 11 transmits information about the generated assignment label to the terminal device 3 that requested the generation of the assignment label (step S15), and terminates the processing. The label information transmission processing unit 11e transmits label information to the terminal device 3, including the identification information of the patent document, one or more assignment labels generated for the patent document, and the assignment text used to generate the assignment label. When a patent document in a language other than English is translated into English, the label information transmission processing unit 11e may extract from the original patent document a non-English assignment text corresponding to the English assignment text used to generate the assignment label, and transmit the assignment text in the original language together with the assignment label to the terminal device 3. The terminal device 3 that receives the label information from the information processing device 1 displays the identification information of the patent document, the assignment label, and the assignment text in association with each other. The terminal device 3 can provide the user with the assignment text corresponding to the assignment label as the text in the patent document that served as the basis for generating the assignment label.
[0067] FIG. 7 is a flowchart showing an example of the procedure for generating problem labels for multiple patent documents performed by the information processing device 1 according to this embodiment. When the terminal device 3 requests the processing unit 11 of the information processing device 1 according to this embodiment to generate problem labels for multiple patent documents, the processing unit 11 appropriately selects one patent document from the multiple patent documents and performs the problem label generation process for this single patent document (step S21). In step S21, the processing unit 11 generates one or more problem labels for each patent document according to the procedure shown in the flowcharts of FIGS. 4 and 5. The processing unit 11 determines whether the generation of problem labels for all patent documents requested by the terminal device has been completed (step S22). If the generation of problem labels for all patent documents has not been completed (S22: NO), the processing unit 11 returns to step S21, appropriately selects unprocessed patent documents, and repeatedly generates problem labels.
[0068] When the generation of problem labels for all patent documents has been completed (S22: YES), the problem label integration unit 11d of the processing unit 11 calculates the number of occurrences of each problem label (i.e., how many of each problem label are included in the multiple problem labels) based on the multiple problem labels generated for the multiple patent documents (step S23).Based on the calculated number of occurrences of each problem label, the problem label integration unit 11d, for example, sorts the multiple problem labels in order of number of occurrences, and determines the problem labels with the highest number of occurrences up to a predetermined percentage (1% or 10%, etc.) as reference labels (step S24).
[0069] Next, the challenge label integrating unit 11d calculates the similarity between the reference label and challenge labels other than the reference label (step S25). At this time, the challenge label integrating unit 11d selects pairs of the reference label and challenge labels other than the reference label and calculates the similarity, and calculates the similarity for all pairs of labels by repeatedly selecting pairs of labels and calculating the similarity. The challenge label integrating unit 11d can calculate the similarity between the two labels by converting the reference label and challenge labels other than the reference label into feature vectors using a pre-generated encoder and calculating the cosine similarity between the two vectors. However, the similarity between the two labels is not limited to being calculated by the cosine similarity of the vectors and may be calculated by any method.
[0070] The task label integrating unit 11d integrates multiple task labels based on the similarity between the reference label and the task labels other than the reference label calculated in step S25 (step S26). The task label integrating unit 11d selects combinations where the similarity between the reference label and the task labels other than the reference label exceeds a threshold, and integrates the task labels by changing the task labels other than the reference label in this combination to the reference label. Integration is not required for task labels other than the reference label where there is no reference label whose similarity exceeds the threshold. The above procedure for integrating task labels is an example and is not limited to this. The task label integrating unit 11d may calculate the similarity between two reference labels, select a combination of reference labels whose similarity exceeds a threshold, and integrate the reference labels by changing the reference label with a lower frequency of occurrence in this combination to the reference label with a higher frequency of occurrence.
[0071] The label information transmission processing unit 11e of the processing unit 11 transmits label information that associates identification information for multiple patent documents, one or more generated task labels, and the task text used to generate the task labels to the terminal device 3 (step S27), and then terminates the processing.
[0072] 8 is a flowchart showing an example of a processing procedure performed by the terminal device 3 according to this embodiment. The patent document acquisition unit 31a of the processing unit 31 of the terminal device 3 according to this embodiment accepts input of a patent document for which a challenge label is to be generated based on a user's operation on the operation unit 35 (step S31). At this time, the patent document acquisition unit 31a may display a file selection screen or the like to accept the selection of a patent document file, or may accept input of identification information for the patent document. The patent document acquisition unit 31a communicates with the information processing device 1 via the communication unit 33 and requests the information processing device 1 to generate a challenge label for the patent document acquired in step S31 (step S32).
[0073] The display processing unit 31b of the processing unit 31 receives the label information sent by the information processing device 1 in response to the request made in step S32 (step S33). The label information received from the information processing device 1 includes, in association with each other, the identification information of the patent document, one or more problem labels generated for the patent document, and the text in the patent document that served as the basis for generating the problem labels. The display processing unit 31b displays the label information received in step S33 on the display unit 34 (step S34), and ends the processing.
[0074] FIG. 9 is a schematic diagram showing an example of label information displayed by the terminal device 3. The terminal device 3 displays a string indicating the title "Problem Label Generation Results for Patent Documents" at the top of the screen of the display unit 34, and below this title string, displays a list of label information acquired from the information processing device 1 in table format. In this example, the terminal device 3 displays a table with the following fields: "Patent Document," "Problem Label," and "Reasoning Text." The "Patent Document" field contains identification information for the patent document entered by the user as the target for generating a problem label, and may display information on the publication number of the publication, such as "JP Patent Publication No. 2023-12345," "JP Patent Publication No. 2022-234567," and "JP Patent Publication No. 2020-001122." The "Problem Label" field contains problem labels, such as "Power Consumption Reduction," "Maintaining Comfort," and "Energy Conservation," generated by the information processing device 1 for the patent document. In this example, three problem labels are generated for one patent document. The "basis sentence" is a sentence such as "However, there is a problem that..." extracted from a patent document, and is the sentence used to generate the issue label.
[0075] The display format of the label information by the terminal device 3 is not limited to that shown in Fig. 9, and any display format may be adopted. The values in the table shown in Fig. 9 are merely examples and are not limiting.
[0076] <Modification> Fig. 10 is a schematic diagram for explaining the configuration of an information processing system according to a modification. In the information processing system according to the modification, the information processing device 1 communicates with the patent management server device 5 at a predetermined cycle, such as once a day or once a week, to acquire patent documents stored in the patent document DB 6. At this time, the information processing device 1 acquires from the patent management server device 5 patent documents that meet conditions set by an administrator of the information processing system or the like and that have not been acquired before. Various conditions, such as technical field or applicant name, can be specified as the conditions for acquiring patent documents.
[0077] The information processing device 1 generates issue labels using a large-scale language model for one or more patent documents acquired from the patent management server device 5. The information processing device 1 according to the modified example is equipped with a patent document DB 2. The information processing device 1 assigns one or more issue labels generated for a patent document acquired from the patent management server device 5 (or identification information for the patent document) and stores the patent document in the patent document DB 2. This allows the information processing device 1 to collect necessary patent documents from the patent documents stored in the patent document DB 6 of the patent management server device 5 and accumulate the information in its own patent document DB 2.
[0078] A user of an information processing system according to a modified example can access the information processing device 1 using his or her own terminal device 3 and search, acquire, and view information stored in the patent document DB 2. The terminal device 3 according to a modified example accepts input of the wording of a subject label as a search condition from the user and issues a search request for patent documents to the information processing device 1 using the subject label entered by the user as a search condition. Upon receiving a search request from the terminal device 3, the information processing device 1 extracts patent documents from the patent document DB 2 that have been assigned a subject label that matches or is similar to the subject label specified as a search condition. The information processing device 1 transmits the patent documents (or their identification information) extracted from the patent document DB 2 to the terminal device 3 that issued the search request. The terminal device 3 receives information on the search results for the search request from the information processing device 1 and displays information on the patent documents corresponding to the subject label entered as a search condition on the display unit 34.
[0079] <Summary> In the information processing system according to the present embodiment, the information processing device 1 acquires text information related to patents (patent documents), extracts sentences containing predetermined terms from the acquired patent documents as challenge sentences, and generates one or more challenge labels indicating the content of the challenge sentences using a large-scale language model that has undergone machine learning based on the extracted challenge sentences. By extracting challenge sentences containing predetermined terms in advance, the information processing device 1 is expected to generate more accurate challenge labels compared to when the entire patent document is input into a large-scale language model to generate challenge labels. As a result, the information processing system according to the present embodiment is expected to reduce the processing load associated with generating challenge labels. Therefore, the information processing system according to the present embodiment is expected to support the generation of challenge labels for patent documents.
[0080] In the information processing system according to this embodiment, the information processing device 1 translates patent documents in languages other than English into English and extracts task sentences containing predetermined terms from the English patent documents. While Japanese sentences may contain ambiguous expressions, English sentences are clearer than Japanese sentences. Therefore, the information processing system according to this embodiment is expected to improve the accuracy of task sentence extraction by extracting task sentences based on English sentences.
[0081] In the information processing system according to this embodiment, the information processing device 1 acquires multiple patent documents, generates one or more challenge labels for each of the multiple patent documents, and integrates the multiple challenge labels generated for the multiple patent documents. The information processing device 1 calculates the number of occurrences of each of the multiple challenge labels and determines a reference label from among the multiple challenge labels based on the calculated number of occurrences. The information processing device 1 calculates the similarity between each of the reference labels and challenge labels other than the reference label, and integrates the challenge labels other than the reference label into the reference label based on the calculated similarity. This allows the information processing system according to this embodiment to integrate challenge labels that are not identical but have similar content, and is expected to reduce the number of types of challenge labels generated for multiple patent documents.
[0082] In the information processing system according to this embodiment, the information processing device 1 extracts problem sentences from text information relating to the abstract, background, or problem of a patent document. As a result, the information processing system according to this embodiment is expected to process portions that are likely to contain descriptions relating to the problem of the invention.
[0083] In the information processing system according to this embodiment, the information processing device 1 removes unnecessary symbols from the extracted challenge text. As a result, the information processing system according to this embodiment can remove symbols such as signs that are often included in the text of patent documents in advance, and input challenge text that does not contain unnecessary information to the large-scale language model, which is expected to improve the accuracy of the challenge labels generated by the large-scale language model.
[0084] In the information processing system according to this embodiment, the information processing device 1 transmits information including the problem label generated for the patent document and the problem text used to generate the problem label to the terminal device 3, and causes the terminal device 3 to display the problem label and the problem text in association with each other. As a result, the information processing system according to this embodiment is expected to provide the user of the terminal device 3 with the problem label of the patent document and information on the text that is the basis for this problem label.
[0085] In this embodiment, the information processing device 1 generates a problem label indicating the content of the problem of the invention described in the patent document, but this is not limited to this. The information processing device 1 may also generate a purpose label indicating the purpose of the invention described in the patent document or an effect label indicating the effect. The information processing system according to this embodiment includes three types of devices: the information processing device 1, the terminal device 3, and the patent management server device 5, but this is not limited to this. The user may directly operate the information processing device 1 rather than requesting the information processing device 1 to generate a problem label via the terminal device 3. The information processing device 1 and the patent management server device 5 may be a single device, or the information processing device 1, the terminal device 3, and the patent management server device 5 may be a single device. The number of devices included in the information processing system and their respective roles may be determined appropriately by the system designer, etc.
[0086] Although the embodiments have been described above, it will be understood that various changes in form and details can be made without departing from the spirit and scope of the claims.
[0087] REFERENCE SIGNS LIST 1 Information processing device (computer) 2 Patent document DB 3 Terminal device 5 Patent management server device 6 Patent document DB 11 Processing unit 11a Patent document acquisition unit 11b Problem sentence extraction unit 11c Problem label generation unit 11d Problem label integration unit 11e Label information transmission processing unit 12 Storage unit 12a Program (computer program) 12b Model information storage unit 12c Terminology storage unit 13 Communication unit 31 Processing unit 31a Patent document acquisition unit 31b Display processing unit 32 Storage unit 32a Program 33 Communication unit 34 Display unit 35 Operation unit 98, 99 Recording medium N Network
Claims
1. A computer program that causes a computer to execute the following processes: acquire text information related to a patent; extract sentences containing predetermined terms from the acquired text information; and, based on the extracted sentences, generate one or more labels indicating the content of the sentences using a language model that has been trained by machine learning in advance.
2. The computer program according to claim 1, wherein the label indicates the content of the assignment described in the text information.
3. A computer program according to claim 1 or claim 2, which extracts sentences containing the predetermined term from English text information.
4. The computer program according to claim 3, wherein, when the text information is in a language other than English, a sentence containing the predetermined term is extracted from the text information obtained by translating the text information into English.
5. A computer program as claimed in claim 1 or claim 2, which acquires multiple pieces of text information relating to multiple patents, generates one or more labels for each of the multiple pieces of text information, and integrates the multiple labels generated for the multiple pieces of text information.
6. The computer program according to claim 5, further comprising: calculating the number of times each label included in the generated plurality of labels appears; determining a reference label from among the plurality of labels based on the calculated number of times each label appears; calculating a similarity between the determined reference label and labels other than the reference label; and integrating the labels other than the reference label into the reference label based on the calculated similarity.
7. A computer program as claimed in claim 1 or claim 2, wherein the text information is text relating to the abstract, background or problem of a patent document.
8. The computer program according to claim 1 or 2, further comprising the steps of: extracting a sentence containing the predetermined term; and then removing unnecessary symbols from the sentence.
9. A computer program according to claim 1 or 2, which displays the extracted sentences and the labels generated based on the sentences on a display unit in association with each other.
10. An information processing method comprising: an acquisition step in which an information processing device acquires text information related to a patent; an extraction step in which a sentence containing a predetermined term is extracted from the acquired text information; and a generation step in which, based on the extracted text, a language model that has been previously machine-learned is used to generate one or more labels indicating the content of the text.
11. An information processing device comprising a processing unit that acquires text information related to a patent, extracts sentences containing specified terms from the acquired text information, and generates one or more labels indicating the content of the sentences based on the extracted sentences using a language model that has been previously machine-learned.
Citation Information
Patent Citations
Tool
JP2020001122A
Proposal apparatus for amendment of patent classification, proposal method for amendment of patent classification, and program
JP2022048781A
Optical connector system
JP2023012345A
Text processing method and device, computer equipment and storage medium
CN115859176A
Automatic tag impartment device, automatic tag impartment method, automatic tag impartment program and recording medium recording the program
JP2008310626A
Cited By
Method for generating a request to a database
US20260087251A1