Document processing device, document processing method, and program
The document processing device and method address the inefficiency in document creation by assigning category information and importance to divided sentences, enabling efficient text extraction and reducing the time required for document creation.
Patent Information
- Application Number
- PCT/JP2023/041356
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-22
AI Technical Summary
Existing document processing methods struggle to efficiently extract appropriate text for document creation, leading to time-consuming document creation processes, not only in the medical field but also in other domains.
A document processing device and method that involves acquiring divided sentences from documents, assigning category information and importance based on similarity with existing documents, and storing this information for efficient text extraction and document creation.
This approach reduces the effort required for document creation by providing document creators with relevant text segments based on category information and importance, thereby streamlining the document creation process.
Smart Images

Figure JP2023041356_22052025_PF_FP_ABST
Abstract
Description
Document processing device, document processing method, and program
[0001] The present disclosure relates to a document processing device, a document processing method, and a program.
[0002] In medical care, patient information such as a patient's medical history is recorded in a medical record, and various documents can be created by document creators such as doctors based on the information in the medical record. Examples of documents that can be created include a "discharge summary" when a patient is discharged from the hospital, a "letter of referral" when a patient is transferred to another hospital, and an "academic document" when a research report is submitted. However, when creating various documents from patient information, document creators must extract sentences from the patient information in the medical record, but extracting appropriate sentences for the document to be created can be difficult, resulting in a problem of time-consuming document creation.
[0003] Here, Patent Document 1 describes a method of determining the importance of each element included in a medical record document, estimating which elements to adopt from the importance, and generating a summary using the adopted elements. Specifically, Patent Document 1 determines the importance of elements based on whether elements of a past medical record document have been adopted in the summary, and determines the importance of elements of a new medical record document based on the importance of the elements.
[0004] Japanese Patent Application Laid-Open No. 2020-38602
[0005] However, in the technology described in Patent Document 1, importance is determined based on past sentences and used for new sentences, but it is unclear whether such importance is appropriate for the new sentence. Therefore, it remains difficult to extract sentences that are more appropriate for the document to be created, and document creation is time-consuming. Furthermore, this problem can occur not only when creating documents in the medical field, but also when creating any document.
[0006] Therefore, an object of the present disclosure is to solve the above-mentioned problem of document creation being time-consuming.
[0007] A document processing device according to one embodiment of the present disclosure includes: an acquisition unit that acquires divided sentences obtained by dividing text in a document into predetermined units; an assignment unit that assigns predetermined category information to each of the divided sentences based on a plurality of the divided sentences; and an association unit that associates and stores the category information assigned to the divided sentences. A document processing method according to one embodiment of the present disclosure includes: acquiring divided sentences obtained by dividing text in a document into predetermined units; assigning predetermined category information to each of the divided sentences based on the plurality of divided sentences; and storing the divided sentences in association with the category information assigned to the divided sentences. A program according to one embodiment of the present disclosure includes: causing a computer to execute the following processes: acquiring divided sentences obtained by dividing text in a document into predetermined units; assigning predetermined category information to each of the divided sentences based on the plurality of divided sentences; and storing the divided sentences in association with the category information assigned to the divided sentences. a document processing device according to an embodiment of the present disclosure, comprising: an acquisition unit that acquires first divided sentences obtained by dividing text in a created document created from a predetermined document into predetermined units, and acquires second divided sentences obtained by dividing text in another document into predetermined units, a determination unit that calculates a similarity of the second divided sentences to each of the first divided sentences according to a preset criterion and determines an importance of the second divided sentences based on the similarity, and an association unit that associates the determined importance with the second divided sentences and stores the same.A document processing method according to an embodiment of the present disclosure, comprising: an acquisition unit that acquires first divided sentences obtained by dividing text in a created document created from a predetermined document into predetermined units, and acquires second divided sentences obtained by dividing text in the other document into predetermined units, calculates a similarity of the second divided sentences to each of the first divided sentences according to a preset criterion, and determines an importance of the second divided sentences based on the similarity, and stores the determined importance with the second divided sentences.Furthermore, a program that is one form of the present disclosure is configured to cause a computer to execute the following processes: acquire first divided sentences obtained by dividing text in a created document created from a specified document into predetermined units; acquire second divided sentences obtained by dividing text in another document into predetermined units; calculate the similarity of the second divided sentences to each of the plurality of first divided sentences based on a predetermined criterion; determine the importance of the second divided sentences based on the similarity; and associate and store the determined importance with the second divided sentences.
[0008] With the above-described configuration, the present disclosure can reduce the effort required for document creation.
[0009] FIG. 1 is a block diagram showing a configuration of a document processing device according to the present disclosure. FIG. 1 is a diagram showing a state of processing by a document processing device according to the present disclosure. FIG. 1 is a diagram showing a state of processing by a document processing device according to the present disclosure. FIG. 2 is a diagram showing a state of processing by a document processing device according to the present disclosure. FIG. 2 is a diagram showing a state of processing by a document processing device according to the present disclosure. FIG. 3 is a diagram showing a state of processing by a document processing device according to the present disclosure. FIG. 3 is a flowchart showing the processing operation of a document processing device according to the present disclosure. FIG. 4 is a flowchart showing the processing operation of a document processing device according to the present disclosure. FIG. 4 is a block diagram showing a hardware configuration of a document processing device according to the present disclosure. FIG. 5 is a block diagram showing a configuration of a document processing device according to the present disclosure. FIG. 5 is a block diagram showing a configuration of a document processing device according to the present disclosure.
[0010] First Embodiment A first embodiment of the present disclosure will be described with reference to the drawings. Note that the drawings may be relevant to any embodiment.
[0011] [Configuration] The document processing device 10 in this embodiment is used to support document creators, such as doctors, in creating medical documents. In the medical field, patient information, such as a patient's medical history, is recorded in an electronic medical record. Various documents, such as a "discharge summary" when a patient is discharged from the hospital, a "letter of referral" when a patient is transferred to another hospital, and an "academic document" when a research report is submitted, can be created by referencing the text in the electronic medical record. In this embodiment, the document processing device 10 associates various pieces of information with the text recorded in the electronic medical record and stores them. This allows the document creator to use the text in creating the document based on the associated information, thereby supporting the document creator in creating the document.
[0012] In this embodiment, the document to be processed by the document processing device 10 (predetermined document) is described as patient information such as medical history recorded in an electronic medical record, but the document to be processed is not limited to documents in the medical field and may be documents in any field. For example, the document to be processed may be a document recording work records at a construction site, and by associating and storing various information with each sentence in such a document, it may be used to support the creation of new documents such as reports and summaries.
[0013] The document processing device 10 is composed of one or more information processing devices each including a calculation device and a storage device. As shown in FIG. 1 , the document processing device 10 includes a division unit 11, an assignment unit 12, an association unit 13, and an extraction unit 14. The functions of the division unit 11, the assignment unit 12, the association unit 13, and the extraction unit 14 can be realized by the calculation device executing a program for realizing each function stored in the storage device. The document processing device 10 also includes an electronic medical record storage unit 16 and a text information storage unit 17. The electronic medical record storage unit 16 and the text information storage unit 17 are each composed of a storage device. The document processing device 10 is also connected to a user terminal 20, which is an information processing terminal operated by a document creator such as a doctor. Each component is described in detail below.
[0014] As will be described later, the document processing device 10 is configured with a function for assigning category information and importance to sentences (segments) in an electronic medical record. Therefore, the following description will clearly state the function related to assigning category information and the function related to assigning importance, and each component will be explained. However, the document processing device 10 is not limited to having both the function related to assigning category information and the function related to assigning importance, and may be configured with only one of these functions.
[0015] The electronic medical record storage unit 16 stores electronic medical records (or other documents) that contain patient information, such as the patient's medical history, entered by a doctor or the like. It is assumed that each electronic medical record is configured as a separate document for each patient. The patient information is composed of multiple sentences and is recorded as text data. For example, as shown in FIG. 2 , the text data that constitutes patient information is recorded by date and includes sentences that describe the medical treatment details and the patient's condition on that day. Here, the text data that constitutes patient information recorded in the electronic medical record shown in FIG. 2 is divided into sentences, which are predetermined units of sentences, by processing by the document processing device 10, as described below, and becomes a registration target document to which category information and importance are assigned.
[0016] The electronic medical record storage unit 16 also stores electronic medical records (predetermined documents) containing patient information such as past patient medical history. As shown in FIG. 3 , the past electronic medical records first contain patient information composed of text data, as described above, and also contain summary information (created documents) composed of text data summarizing the patient information. The summary information is, for example, created by a doctor using the patient information recorded in the electronic medical record. Therefore, the summary information may include some of the text of the patient information in the electronic medical record. The summary information may also be, for example, a "discharge summary," "letter of referral," or "academic document" created by a document creator, such as a doctor, for a past patient. The text data of the patient information and summary information recorded in the electronic medical record shown in FIG. 3 are used in the importance assignment process by the document processing device 10, as described below, and serve as training documents. The summary information does not necessarily have to be recorded in the same electronic medical record containing the patient information from which it was generated; it may exist as a document separate from the electronic medical record. In this case, the summary information may be associated with the patient information from which it was generated.
[0017] As will be described later, the text information storage unit 17 stores each sentence, which is a divided sentence obtained by dividing the patient information included in the electronic medical record. Also, each sentence is stored in association with category information and importance, which will be assigned to each sentence, as will be described later. In other words, the text information storage unit 17 stores each sentence, which is a divided sentence (second divided sentence) divided from the patient information of the electronic medical record, which is the document to be registered, as shown in FIG. 2, with category information and importance assigned to each sentence.
[0018] The division unit 11 (acquisition unit) divides text data, which is patient information included in the electronic medical record, into divided sentences, which are sentences in predetermined units. For example, the division unit 11 divides the text data, which is patient information, at each line break, period, or comma, and divides it into divided sentences in predetermined units, such as line break units, period units, or comma units. Note that the units by which the division unit 11 divides the text data, which is patient information, are not limited to the units described above, and may be divided in any units.
[0019] Here, the dividing unit 11 divides the text data, which is the patient information in the electronic medical record, which is the document to be registered as shown in FIG. 2, into divided sentences T n On the other hand, the dividing unit 11 divides the text data, which is the patient information in the electronic medical record, which is the training document as shown in FIG. 3, into training divided sentences t n , t an (n=1, 2, 3, ...). In FIG. 5, each training sentence is assigned a code t n , t an and the training segmentation information t without summary correspondence, which will be described later. n and the training segmented sentences t an At this point, each training segment is divided into two parts, each code t n , t an The segmentation unit 11 segments the training document into training segmented sentences t n , t an For the training segmentation sentence t n , t an is included in the summary information of the same training document, and annotation information indicating that the training segmented sentences included in the summary information are corresponding to the summary is added. an In this way, when the dividing unit 11 divides the patient information of the electronic medical record, which is the training document, into training divided information t n (other segmented sentences) and the training segmented sentence t an (first divided sentence). Note that when all or part of a training divided sentence is included in the summary information, the dividing unit 11 may use the training divided sentence as a training divided sentence corresponding to the summary. Furthermore, the dividing unit 11 may check whether the training divided sentence is included in summary information in another document associated with the training document, rather than in the same electronic medical record as the training document, and generate a training divided sentence corresponding to the summary.
[0020] The division unit 11 divides the above-mentioned divided sentence Tn (Second divided sentence), training divided sentence t without summary correspondence n (Other segmented sentences), training segmented sentence t an The division unit 11 is not limited to dividing and generating the first divided sentence from the patient information in the electronic medical record as described above. For example, the division unit 11 may divide the first divided sentence T n , training segmented sentence t without summary correspondence n , the training segmented sentences t an In this case, the training segmented sentences t an The training sentences may be manually annotated with annotation information corresponding to the summaries.
[0021] In the above, the segmentation unit 11 extracts training segmented sentences t n and the training segmented sentences t an In the above example, the summary information already generated is divided into predetermined units of divided sentences, and the divided sentences are used as training divided sentences t corresponding to the summary. an In this case, the training segmented sentences t an As described above, there is no need to add annotation information to the training segmented sentences t an will be used.
[0022] The assignment unit 12 assigns each of the above-mentioned segmented sentences T n In this embodiment, the assigning unit 12 assigns preset category information to the segmented sentence T (second segmented sentence). n Category information is assigned to
[0023] The classification model will now be described. First, the training data used to train the classification model includes multiple divided sentences, each associated with corresponding category information. The multiple divided sentences are divided from patient information in the same electronic medical record, and the order in which they are written in the electronic medical record is set. That is, the order in which the multiple divided sentences are written is set to the same order as the order in which they are written in the electronic medical record. For example, the order in which the multiple divided sentences are written is set to the same order as the order in which they are written in the electronic medical record. Furthermore, the dates written in the electronic medical record are set for the multiple divided sentences. For example, different dates are set for each divided sentence, and the order in which they are written is set in chronological order. Category information in the training data is associated with each of the multiple divided sentences. Multiple types of category information, such as radiology, pathology, family, medical history, and others, are set in advance. Of these, category information corresponding to the type of summary information in which the divided sentence is actually used, is associated with the divided sentence. For example, if the summary information in which the divided sentence is used is a "discharge summary," category information such as "family" and "pathology" set to correspond to the "discharge summary" is associated with the divided sentence.
[0024] Then, a classification model is trained using the above-described training data. Specifically, the classification model receives input of multiple segmented sentences as training data, and machine-learns the relationships between the multiple segmented sentences so as to output category information associated with each segmented sentence in the training data. That is, the classification model receives input of multiple segmented sentences and performs supervised learning using category information corresponding to each segmented sentence as training data. As a result, the classification model is configured to receive input of multiple segmented sentences whose corresponding category information is unknown, and output category information that can correspond to each segmented sentence. In particular, since the multiple segmented sentences used as training data are written in a predetermined order in the electronic medical record from which they are divided, the classification model is constructed to output category information for each segmented sentence taking into account the order in which the multiple segmented sentences are written. However, the method for training the classification model is not limited to the above-described machine learning method, and any method may be used for training.
[0025] Then, the assignment unit 12 assigns to the classification model learned as described above a plurality of registration target divided sentences T n By inputting n At this time, the assigning unit 12 outputs category information corresponding to each of the plurality of segmented sentences T n may be input to the classification model. In other words, the order of the entries of the multiple divided sentences to be registered is set to be the same as the order of entries in the electronic medical record, and for example, the same consecutive order of entries as in the electronic medical record may be set. Furthermore, the multiple divided sentences to be registered may be set to have the date entered in the electronic medical record, or may be set to be in date order. In this way, multiple divided sentences to be registered T that are set to have the same order of entries as in the electronic medical record that is the source of division may be input. n By inputting the above into the classification model, multiple segmented sentences T n The order of the entries is also appropriately adjusted to each segmented sentence T nIn other words, as described above, since the classification model is trained taking into consideration the order in which the multiple segmented sentences are written, the assignment unit 12 can assign category information to each of the multiple segmented sentences to be registered based on the order in which the multiple segmented sentences to be registered are written. For example, as shown in FIG. 6, the assignment unit 12 assigns category information to each of the multiple segmented sentences to be registered, 1 , T 2 , T 3 , . . . are input to the classification model, and each segmented sentence T 1 , T 2 , T 3 , . . . can be assigned category information (other, medical history, family, . . .).
[0026] The above-described method of assigning category information is merely an example, and the assigning unit 12 may assign category information using other methods. For example, the assigning unit 12 may assign category information to each of the multiple registration-target segmented sentences based on the relationship between the contents of the multiple registration-target segmented sentences, without using the classification model described above. Specifically, the assigning unit 12 may assign category information corresponding to a specific registration-target segmented sentence based on words included in a specific registration-target segmented sentence and words included in other registration-target segmented sentences, which are included in the multiple registration-target segmented sentences. In this case, the category information may also be assigned based on the order in which the words included in the specific registration-target segmented sentence and the words included in the other registration-target segmented sentences are written. For example, the category information corresponding to a specific registration-target segmented sentence may be assigned based on the number of occurrences or the order in which specific words appear in the multiple registration-target segmented sentences.
[0027] Furthermore, the assignment unit 12 (determination unit) determines whether each of the above-mentioned segmented sentences T n In this embodiment, the assigning unit 12 performs a process of assigning importance to each of the training segment sentences t n , t an Each segmented sentence T to be registered n Based on the similarity of each segmented sentence T n In particular, the assignment unit 12 calculates the importance of each segment sentence T n and the training segmented sentences t anBased on the similarity between n At this time, the importance of the training segmented sentence t an may be generated by dividing the patient information and adding annotation information, or may be generated by dividing the summary information and adding no annotation information. n and the training segmented sentences t an We will explain the case where the importance of the segmented sentence T to be registered is calculated using the above, but as will be described later, the importance of the segmented sentence T to be registered may also be calculated using training segmented sentences corresponding to summaries divided from summary information.
[0028] Specifically, the attachment unit 12 first calculates the segmented sentences T n and each of the training segmented sentences t n , t an In order to calculate the similarity between the segmented sentences T and T, the segmented sentences are converted into fixed-length vectors having the same vector length. At this time, the assignment unit 12 converts each segmented sentence T to be registered into a fixed-length vector having the same vector length using a machine-learned conversion model. n and each training segment t n , t an and are converted into fixed-length vectors. The converted vectors are assumed to be machine-learned so that, for example, a predetermined sentence is input, the sentence is converted into fixed-length vector data, and output. As a result, the annotation unit 12 converts each training segmented sentence t without corresponding summary, as shown in FIG. n is input to the conversion model to obtain the training segmented sentence vector t vn and the training segmented sentences t an is input to the conversion model to obtain the training segmented sentence vectors t van and each segmented sentence T n is input to the conversion model to create the segmented sentence vector T vn The attachment unit 12 is not limited to using the machine-learned conversion model described above, and may convert each divided sentence into a fixed-length vector by any method.
[0029] Then, the annotation unit 12 uses each vector converted to a fixed length as described above to generate each training segmented sentence vector t vn and the training segmented sentence vector t van and each divided sentence vector T vn Specifically, as shown in FIG. 8, the assignment unit 12 calculates the similarity between the target segmented sentence vector T v From each training segmented sentence vector t v and the training segmented sentence vector t va The distance between each of the fixed-length vectors is calculated, and a value based on this distance is calculated as the similarity. Here, the closer the distance between the fixed-length vectors, the higher the similarity. Note that the distance between the fixed-length vectors can be calculated using, for example, Euclidean distance or Manhattan distance. However, this distance may be calculated using any method.
[0030] Next, the assignment unit 12 calculates the target segmented sentence vector T v k training segmented sentence vectors t v and the training segmented sentence vector t va and among them, the training divided sentence vector t va Based on the number of divided sentences, the divided sentence vector T v As an example, FIG. 8 illustrates the case where k=5, and as shown in the dotted circle, the importance of the segmented sentence vector T to be registered is calculated. v The five training segmented sentence vectors t v , t va , and each of the five training segmented sentence vectors t v , t va Among them, two are training segmented sentence vectors t va Therefore, the importance is calculated as "2". v The k nearest training segmentation vectors t v , t vaAmong them, the training divided sentence vector t va The importance is calculated based on the number of training sentence vectors t va The greater the number, the higher the calculated importance.
[0031] Then, the assignment unit 12 assigns each divided sentence vector T v For each training segmented sentence vector t v , t va 9, the assignment unit 12 calculates the distance to the target segment sentence T 1 , T 2 , T 3 , . . . can be assigned a level of importance.
[0032] The above-described method for calculating the importance is an example, and the assigning unit 12 may calculate the importance by other methods. For example, the assigning unit 12 may calculate the importance by using the divided sentence vector T v The distance from the training segmented sentence vector t is within a predetermined threshold. v , t va Among them, the training divided sentence vector t va In this way, the assignment unit 12 may calculate the importance based on the number of divided sentence vectors T v The distance from the training segmented sentence vector t, i.e., the similarity, satisfies a preset condition. v , t va Among them, the training divided sentence vector t va The assignment unit 12 may calculate the importance based on the number of divided sentence vectors T v The training segmented sentence vector t corresponding to the summary whose similarity based on the distance between the va In this way, the assignment unit 12 may calculate the importance based on the number of divided sentence vectors T v The similarity of the training sentence vector t va The importance may be calculated based on the number of
[0033] In addition, the assignment unit 12 further assigns each registration target divided sentence vector Tv For example, the assigning unit 12 may calculate the importance of each of the segmented sentence vectors T v k training divided sentence vectors t v , t va Among them, the training divided sentence vector t v and the average distance to each summary-corresponding training segmented sentence vector t va The importance may be calculated based on the ratio of the average distance to the target segmented sentence vector T v t, the nearest neighbors in distance from the training segmented sentence vector t va In this way, the assigning unit 12 may calculate the importance based on the distance value to the registration target divided sentence vector T v and the training segmented sentence vector t va The importance may be calculated based on the distance, that is, the similarity, between the two.
[0034] In the above, each segmented sentence vector T v For the training segmented sentence vector t without summary correspondence, v and the training segmented sentence vector t va The example shows the case where the distance between each of the training divided sentence vectors t va For example, as described above, the distance to the training segmented sentence t corresponding to the summary segmented from the summary information may be calculated. a is prepared in advance, each segmented sentence vector T v For the training segmented sentence vector t va In this case, for example, the distance between the training segmented sentence vector t and the summary vector t is within the threshold. va and the number of training segmented sentence vectors t va The importance may be calculated based on the average value of the distances.
[0035] Here, the assigning unit 12 calculates the distance between fixed-length vectors as the similarity, but the method is not necessarily limited to calculating the distance between fixed-length vectors as the similarity. For example, the similarity may be calculated by determining the degree of correspondence or relevance of characters included in each divided sentence. Furthermore, the assigning unit 12 may determine the importance based on the similarity in any manner, such as by determining that the higher the similarity, the higher the importance.
[0036] As described above, the assigning unit 12 may perform only one of the processes of assigning category information to each of the divided sentences T to be registered and the process of assigning importance to each of the divided sentences T to be registered.
[0037] The associating unit 13 associates the category information and importance assigned as described above with each of the divided sentences T to be registered, and stores them in the text information storage unit 17. For example, as shown in FIG. 6, category information is associated with each of the divided sentences T to be registered, and as shown in FIG. 9, importance is associated with each of the divided sentences T to be registered and stored. At this time, the associating unit 13 may associate both category information and importance with each of the divided sentences T to be registered, as shown in FIG. 9, or may associate importance with each of the divided sentences T to be registered without associating it with category information. For example, as described above, if the assigning unit 12 does not perform the process of assigning category information and only performs the process of assigning importance, only importance is associated with each of the divided sentences T to be registered and stored.
[0038] In addition, the associating unit 13 also associates information specifying the electronic medical record containing the patient information from which the divided sentence T was divided with the divided sentence T to be registered, and stores the information in the text information storage unit 17.
[0039] The extraction unit 14 extracts the segmented sentences T to be registered stored in the sentence information storage unit 17 based on the associated category information and importance. In particular, when the extraction unit 14 receives a sentence extraction request for the corresponding electronic medical record from a user terminal 20, which is an information processing terminal of a document creator such as a doctor who creates documents based on patient information in the electronic medical record, the extraction unit 14 extracts the segmented sentences T to be registered in response to the sentence extraction request. At this time, the sentence extraction request includes information identifying the electronic medical record and the type of document to be newly created. For example, the type of document to be created may be "discharge summary," "letter of referral," "academic document," etc. However, the sentence extraction request is not limited to the type of document to be created, and may also include information about the document, such as the name of the document. Furthermore, the sentence extraction request does not necessarily have to include information about the document to be created.
[0040] The extraction unit 14 then extracts the registration-target segmented sentences T from the registration-target segmented sentences T associated with the electronic medical record information specified by the sentence extraction request, based on category information and importance. For example, the extraction unit 14 extracts the registration-target segmented sentences T associated with category information corresponding to the type of prepared document included in the sentence extraction request. It is assumed that category information corresponding to each type of prepared document is set in advance, and that such information is stored in the extraction unit 14. As an example, the prepared document type "discharge summary" is associated with category information such as "family" and "pathology." It is also possible to associate category information with information about the prepared document, such as the name of the prepared document. In this way, the extraction unit 14 can extract the registration-target segmented sentences T associated with the corresponding category information, even when the sentence extraction request includes information about the prepared document, such as the name of the prepared document.
[0041] The extraction unit 14 then outputs the segmented sentence T to be registered that has been extracted based on the category information to the document creator who made the request. For example, the extraction unit 14 outputs the extracted segmented sentence T to be registered to the user terminal 20 operated by the document creator so that it is displayed. At this time, the extraction unit 14 may output the segmented sentence T to be registered together with the associated category information so that it is displayed.
[0042] Furthermore, the extraction unit 14 may further extract the registration-target segmented sentences T from the registration-target segmented sentences T extracted based on the category information based on the associated importance. For example, the extraction unit 14 may select and extract the registration-target segmented sentences T whose associated importance is equal to or greater than a threshold from the registration-target segmented sentences T extracted based on the category information. Then, the extraction unit 14 outputs only the selected registration-target segmented sentences T extracted based on the category information whose importance is higher than the others and equal to or greater than the threshold to be displayed on the user terminal 20 operated by the document creator. At this time, the extraction unit 14 may output the registration-target segmented sentences T so that the associated category information and importance are displayed together with the registration-target segmented sentences T. Note that the extraction unit 14 may extract a predetermined number of registration-target segmented sentences T whose associated importance is high, or may extract the registration-target segmented sentences T whose importance meets a preset standard.
[0043] Furthermore, the extraction unit 14 may extract the segmented sentences T to be registered based on the associated importance, without taking into account the associated category information. For example, in response to a sentence extraction request, the extraction unit 14 may extract the segmented sentences T to be registered that satisfy a preset criterion, such as the associated importance being equal to or greater than a threshold or being in the top predetermined number. As an example, the extraction unit 14 extracts the segmented sentences T to be registered based on the associated importance, regardless of whether the sentence extraction request includes information such as the type of document to be created, and outputs the segmented sentences T to be registered for display on the user terminal 20 operated by the document creator. At this time, the extraction unit 14 may output the segmented sentences T to be registered so that the associated importance is displayed together with the segmented sentences T to be registered.
[0044] [Operation] Next, the operation of the document processing device 10 described above will be explained. The following will be explained separately for a process of assigning category information to a divided sentence, a process of assigning importance to a divided sentence, and a process of extracting a registered divided sentence. It is assumed that the document processing device 10 stores information on electronic medical records, that is, electronic medical records that are registration target documents to which the above-mentioned category information and importance are assigned, and electronic medical records that are training documents that are electronic medical records of past patients and include already created summary information.
[0045] First, the process of assigning category information to divided sentences will be described. The document processing device 10 acquires an electronic medical record, which is a document to be registered as shown in Fig. 2, and divides text data, which is patient information in the electronic medical record, into divided sentences, which are sentences in predetermined units (step S1 in Fig. 10). For example, the document processing device 10 divides the text data, which is patient information, into divided sentences in predetermined units, such as line break units, periods, and commas, and assigns the divided sentences T to be registered as shown in Fig. 4. n Let's say.
[0046] Next, the document processing apparatus 10 performs the following steps to register each segmented sentence T n Specifically, the document processing apparatus 10 assigns preset category information to a plurality of segmented sentences T to be registered to a classification model that has been trained in advance. n By inputting n The classification model is trained by machine learning to receive multiple segmented sentences T divided from patient information in the same electronic medical record as input, and to output category information, which is training data associated with each segmented sentence. For this reason, the document processing device 10 inputs multiple segmented sentences T to be registered divided from patient information in the same electronic medical record, which is the same document to be registered, to the classification model. n By inputting n The category information for each segmented sentence T n As described above, the classification model is trained taking into consideration the order in which the multiple segmented sentences are written in the electronic medical record. Therefore, the document processing device 10 can assign category information to the multiple segmented sentences T to be registered, in which the order in which they are written in the electronic medical record is set, to the classification model. n By entering the above, category information can be assigned according to the order of entries in the electronic medical record.
[0047] Thereafter, the document processing device 10 associates the assigned category information with each of the segmented sentences T to be registered and stores them (step S3 in FIG. 10). For example, as shown in FIG. 6, each of the segmented sentences T to be registered is associated with category information and stored.
[0048] Next, a process of assigning importance to divided sentences will be described. The document processing device 10 acquires an electronic medical record, which is a document to be registered as shown in Fig. 2, and divides text data, which is patient information in the electronic medical record, into divided sentences, which are sentences in predetermined units (step S11 in Fig. 11). For example, the document processing device 10 divides the text data, which is patient information, into divided sentences in predetermined units, such as line break units, periods, and commas, and assigns the divided sentences T to be registered as shown in Fig. 4. n In addition, the document processing device 10 acquires an electronic medical record, which is a training document as shown in Fig. 3, and divides the text data, which is patient information in the electronic medical record, into divided sentences, which are sentences of a predetermined unit, in the same manner as described above (step S11 in Fig. 11), and generates training divided sentences t n , t an Let's say.
[0049] Next, the document processing apparatus 10 generates training segmented sentences t n , t an , and annotation information indicating that the summary information in the electronic medical record corresponds to the summary is added to the training segmented sentences t an In this way, the document processing apparatus 10 extracts training segmentation information t without corresponding summary from the electronic medical record, which is a training document, as shown in FIG. n and the training segmented sentences t an and generate.
[0050] Next, the document processing apparatus 10 performs the following steps to register each segmented sentence T n and each training segment t n , t an and into fixed-length vectors having the same vector length (step S13 in FIG. 11). At this time, the document processing apparatus 10 converts each of the segmented sentences T nand each training segment t n , t an By inputting the above, each segmented sentence vector T vn , each training segmented sentence vector t vn , the training segmented sentence vector t van , you get.
[0051] Next, the document processing apparatus 10 uses each vector converted to a fixed length to generate a segmented sentence T n (Step S14 in FIG. 11). Specifically, the document processing apparatus 10 assigns importance to each training segmented sentence vector t vn and the training segmented sentence vector t van and each divided sentence vector T vn The document processing apparatus 10 then calculates the distance between each training segmented sentence vector t vn and the training segmented sentence vector t van and each divided sentence vector T vn Based on the distance, the segmented sentence vector T vn For example, as shown in FIG. 8, the importance of the segmented sentence vector T v k training segmented sentence vectors t v and the training segmented sentence vector t va and among them, the training divided sentence vector t va Based on the number of divided sentences, the divided sentence vector T v At this time, the document processing apparatus 10 calculates the importance of the segmented sentence T to be divided corresponding to the segmented sentence vector T v For the training segmented sentence vector t va However, the document processing apparatus 10 calculates a higher importance value for the segmented sentence vector T v and each training segment t n , t an The importance may be calculated by any method based on the similarity between the two.
[0052] Thereafter, the document processing apparatus 10 associates the calculated importance with each of the segmented sentences T to be registered and stores the importance (step S15 in FIG. 11). For example, as shown in FIG. 9, the document processing apparatus 10 associates the importance with each of the segmented sentences T to be registered and stores the importance.
[0053] The document processing device 10 may perform both the process of assigning category information and the process of assigning importance to each of the segmented sentences T to be registered, or may perform only one of these processes. The document processing device 10 may associate and store either category information or importance with each of the segmented sentences T to be registered, or may store both in association with each of the segmented sentences T to be registered.
[0054] Furthermore, when calculating the importance, the document processing apparatus 10 extracts training segmentation information t n and the training segmented sentences t an It is not necessary to perform the process of generating each training segment t n , t an Furthermore, when calculating the importance, the document processing device 10 does not necessarily need to convert each divided sentence into a fixed-length vector, but may calculate the similarity of each divided sentence using another method and calculate the importance from the similarity.
[0055] Next, the process of extracting registered divided sentences will be described. The document processing device 10 receives a sentence extraction request from the user terminal 20 of a document creator such as a doctor. At this time, the document processing device 10 receives the sentence extraction request including information specifying the electronic medical record and the type of document to be created (step S21 in FIG. 12).
[0056] Next, the document processing device 10 extracts the segmented sentences T to be registered that are associated with the electronic medical record information specified by the sentence extraction request. At this time, the document processing device 10 specifies category information corresponding to the type of document to be created that is included in the sentence extraction request (step S22 in FIG. 12), and extracts the segmented sentences T to be registered that are associated with the specified category information (step S23 in FIG. 12).
[0057] Next, the document processing apparatus 10 selects and extracts the segmented sentences T to be registered from the segmented sentences T extracted based on the category information, further based on the associated importance (step S24 in FIG. 12). For example, the document processing apparatus 10 selects and extracts the segmented sentences T to be registered whose associated importance is equal to or greater than a threshold from the segmented sentences T to be registered extracted based on the category information.
[0058] Thereafter, the document processing device 10 outputs the segmented sentences T to be registered, which have been extracted based on the category information and further selected and extracted based on importance, to the user terminal 20 of the requesting document creator for display (step S25 in FIG. 12). At this time, the document processing device 10 may output the segmented sentences T to be registered so that the associated category information and importance are displayed together with the segmented sentences T to be registered.
[0059] The document processing device 10 may output the segmented sentences T to be registered extracted based on the category information without selecting them based on importance. The document processing device 10 may also extract the segmented sentences T to be registered based only on importance without extracting the segmented sentences T to be registered based on the category information, and output the extracted segmented sentences T to be registered.
[0060] As described above, the document processing device 10 of this embodiment assigns category information to each of the divided sentences, based on the multiple divided sentences obtained by dividing text in a document such as an electronic medical record, and stores the associated information. Therefore, the divided sentences can be extracted based on the associated category information and used to create a document corresponding to the category information, thereby reducing the effort required for document creation. In other words, the divided sentences can be presented to the document creator based on the category information, which can assist the document creator in making decisions during document creation and reduce the effort required for document creation. Furthermore, by assigning category information based on the order in which the multiple divided sentences are written in the document from which they are divided, more appropriate category information can be assigned.
[0061] Furthermore, according to the document processing device 10 of this embodiment, divided sentences obtained by dividing text in a document such as an electronic medical record are assigned importance based on their similarity to divided sentences in a created document such as a summary, and the divided sentences are stored in association with each other. Therefore, divided sentences can be extracted based on the associated importance and used in document creation, thereby reducing the effort required for document creation. In other words, divided sentences can be presented to the document creator based on their importance, which can assist the document creator in making decisions during document creation and reduce the effort required for document creation.
[0062] Second Embodiment Next, a second embodiment of the present disclosure will be described with reference to the drawings. This embodiment shows an outline of the configuration of the document processing device described in the above embodiment. Note that Figures 13 to 15 are diagrams for explaining the configuration, and these diagrams may be relevant to any of the embodiments.
[0063] First, the hardware configuration of the document processing device 100 will be described with reference to Fig. 13. The document processing device 100 is configured as a general information processing device, and is equipped with the following hardware configuration, for example: CPU (Central Processing Unit) 101 (arithmetic unit); ROM (Read Only Memory) 102 (storage device); RAM (Random Access Memory) 103 (storage device); programs 104 loaded into RAM 103; storage device 105 storing programs 104; drive device 106 for reading and writing data from and to a storage medium 110 external to the information processing device; communication interface 107 for connecting to a communication network 111 external to the information processing device; input / output interface 108 for inputting and outputting data; and bus 109 for connecting the various components.
[0064] 13 shows an example of the hardware configuration of an information processing device that is the document processing device 100, and the hardware configuration of the information processing device is not limited to the above-described case. For example, the information processing device may be configured with only a part of the above-described configuration, such as excluding the drive device 106. Furthermore, the information processing device may use a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an MPU (Micro Processing Unit), an FPU (Floating Point Number Processing Unit), a PPU (Physics Processing Unit), a TPU (Tensor Processing Unit), a quantum processor, a microcontroller, or a combination thereof, instead of the above-described CPU.
[0065] The document processing device 100 can be equipped with an acquisition unit 121, an assignment unit 122, and an association unit 123 shown in FIG. 14 by having the CPU 101 acquire and execute the program group 104. The program group 104 is stored in advance in the storage device 105 or the ROM 102, for example, and is loaded into the RAM 103 and executed by the CPU 101 as needed. The program group 104 may be supplied to the CPU 101 via the communication network 111, or may be stored in advance in the storage medium 110, and the drive device 106 may read and supply the program to the CPU 101. However, the acquisition unit 121, assignment unit 122, and association unit 123 described above may be constructed using dedicated electronic circuits for realizing such means.
[0066] The acquiring unit 121 acquires divided sentences obtained by dividing text in a document into predetermined units. The assigning unit 122 assigns preset category information to each of the divided sentences based on the plurality of divided sentences. The associating unit 123 associates the category information assigned to the divided sentences with the divided sentences and stores them.
[0067] With the above configuration, the present disclosure assigns category information to each of the divided sentences obtained by dividing the text in a document, associates the divided sentences, and stores them. As a result, the divided sentences can be extracted based on the associated category information and used to create documents corresponding to the category information, thereby reducing the effort required for document creation.
[0068] 15 can be constructed and equipped with the acquisition unit 131, determination unit 132, and association unit 133 shown in FIG. 15 by having the CPU 101 acquire and execute the program group 104. The program group 104 may be stored in advance in the storage device 105 or ROM 102, for example, and loaded into the RAM 103 and executed by the CPU 101 as needed. The program group 104 may be supplied to the CPU 101 via the communication network 111, or may be stored in advance in the storage medium 110, with the drive device 106 reading out the program and supplying it to the CPU 101. However, the acquisition unit 131, determination unit 132, and association unit 133 described above may be constructed using dedicated electronic circuits for realizing such means.
[0069] The acquiring unit 131 acquires first divided sentences obtained by dividing text in a created document created from a predetermined document into predetermined units, and acquires second divided sentences obtained by dividing text in another document into predetermined units. The determining unit 132 calculates the similarity of the second divided sentences to each of the multiple first divided sentences according to a predetermined criterion, and determines the importance of the second divided sentences based on the similarity. The associating unit 133 associates the determined importance with the second divided sentences and stores them.
[0070] With the above-described configuration, the present disclosure assigns importance to second divided sentences obtained by dividing text in a document based on their similarity to first divided sentences in the created document, associates them, and stores them. As a result, the second divided sentences can be extracted based on the associated importance and used in document creation, thereby reducing the effort required for document creation.
[0071] At least one or more of the functions of the above-described acquiring unit 121, assigning unit 122, and associating unit 123 may be executed by an information processing device installed and connected anywhere on the network, that is, may be executed by so-called cloud computing. Also, at least one or more of the functions of the above-described acquiring unit 131, determining unit 132, and associating unit 133 may be executed by an information processing device installed and connected anywhere on the network, that is, may be executed by so-called cloud computing.
[0072] The above-described program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-RWs, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program can also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can be supplied to a computer via wired communication paths such as electric wires and optical fibers, or via wireless communication paths.
[0073] Although the present disclosure has been described above with reference to the above-described embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each of the above-described embodiments can be combined with other embodiments as appropriate.
[0074] <Addendum> Some or all of the above embodiments can also be described as the following addendum. Below, an outline of the configurations of a document processing device, a document processing method, and a program according to the present disclosure will be described. However, the present disclosure is not limited to the following configurations. (Addendum A1) A document processing device comprising: an acquisition unit that acquires divided sentences obtained by dividing text in a document into predetermined units; an assignment unit that assigns predetermined category information to each of the divided sentences based on a plurality of the divided sentences; and an association unit that associates the category information assigned to the divided sentences with the divided sentences and stores them. (Addendum A2) The document processing device according to Addendum A1, wherein the assignment unit assigns the category information to each of the divided sentences based on a plurality of the divided sentences in the same document. (Addendum A3) The document processing device according to Addendum A2, wherein the assignment unit assigns the category information to each of the divided sentences based on the order in which the plurality of the divided sentences are written in the same document. (Appendix A4) The document processing device according to Appendix A3, wherein the assignment unit assigns the category information to each of the divided sentences based on the divided sentences that are written in succession within the same document. (Appendix A5) The document processing device according to Appendix A1, wherein dates are assigned to the text in the document, and the assignment unit assigns the category information to each of the divided sentences based on the divided sentences that are separated from the text that are written in the same document and that are given different dates. (Appendix A6) The document processing device according to Appendix A5, wherein the assignment unit assigns the category information to each of the divided sentences based on the divided sentences that are written in succession within the same document. (Appendix A7) The document processing device according to Appendix A1, further comprising an extraction unit that extracts the divided sentences based on the category information associated with the divided sentences. (Appendix A8) The document processing device according to appendix A7, wherein the extraction unit receives information related to a created document, and extracts the divided sentences associated with the category information corresponding to the information related to the created document.(Appendix A9) The document processing device according to Appendix A7, wherein the extraction unit accepts a predetermined type of created document and extracts the segmented sentences associated with the category information corresponding to the type of created document. (Appendix A10) The document processing device according to Appendix A1, wherein the assignment unit assigns the category information to each of the segmented sentences by inputting the segmented sentences to a machine learning model configured to input the segmented sentences and output the category information corresponding to each of the segmented sentences. (Appendix A11) A document processing method comprising: acquiring segmented sentences obtained by dividing text in a document into predetermined units; assigning predetermined category information to each of the segmented sentences based on the plurality of segmented sentences; and storing the category information assigned to the segmented sentences in association with the segmented sentences. (Appendix A12) The document processing method according to Appendix A11, wherein the category information is assigned to each of the segmented sentences based on the plurality of segmented sentences in the same document. (Appendix A13) The document processing method according to Appendix A12, wherein the category information is assigned to each of the divided sentences based on the order in which the divided sentences are written in the same document. (Appendix A14) The document processing method according to Appendix A11, wherein the divided sentences are extracted based on the category information associated with the divided sentences. (Appendix A15) A computer-readable storage medium storing a program that causes a computer to execute processes of obtaining divided sentences obtained by dividing text in a document into predetermined units, assigning predetermined category information to each of the divided sentences based on the plurality of divided sentences, and storing the category information assigned to the divided sentences in association with them.(Appendix B1) A document processing device comprising: an acquisition unit that acquires first divided sentences obtained by dividing text in a created document created from a predetermined document into predetermined units, and acquires second divided sentences obtained by dividing text in another document into predetermined units, a determination unit that calculates a similarity of the second divided sentence to each of the plurality of first divided sentences according to a preset criterion and determines an importance of the second divided sentence based on the similarity, and an association unit that associates the determined importance with the second divided sentence and stores it. (Appendix B2) A document processing device according to Appendix B1, wherein the determination unit determines the importance based on the number of first divided sentences for which the similarity of the second divided sentence to the first divided sentence satisfies a preset condition. (Appendix B3) The document processing device according to Appendix B1, wherein the acquisition unit acquires the first divided sentence and other divided sentences other than the first divided sentence, which are divided sentences obtained by dividing text in the specified document into predetermined units, included in the created document, and the determination unit calculates the similarity of the second divided sentence to each of the first divided sentence and the other divided sentences, and determines the importance of the second divided sentence based on the similarity of the second divided sentence to a plurality of the first divided sentences and the other divided sentences. (Appendix B4) The document processing device according to Appendix B3, wherein the determination unit determines the importance based on the number of first divided sentences among the first divided sentences and the other divided sentences for which the similarity of the second divided sentence to the first divided sentence and the other divided sentences satisfies a preset condition. (Appendix B5) The document processing device according to Appendix B4, wherein the determination unit determines the importance based on the number of first divided sentences among a predetermined number of the first divided sentences and the other divided sentences in descending order of the similarity of the second divided sentence to the first divided sentence and the other divided sentence. (Appendix B6) The document processing device according to Appendix B3, wherein the determination unit converts the first divided sentence, the other divided sentence, and the second divided sentence into fixed-length vectors, and calculates the similarity of the second divided sentence to each of the first divided sentence and the other divided sentence using the converted fixed-length vectors.(Appendix B7) The document processing device according to Appendix B6, wherein the determination unit inputs the first segmented sentence, the other segmented sentence, and the second segmented sentence into a machine learning model configured to convert a predetermined sentence into the fixed-length vector, and converts them into the fixed-length vector. (Appendix B8) The document processing device according to Appendix B1, further comprising: an extraction unit that extracts the second segmented sentence based on the importance associated with the second segmented sentence. (Appendix B9) The document processing device according to Appendix B1, wherein the determination unit assigns predetermined category information to each of the second segmented sentences based on a plurality of the second segmented sentences, and the associating unit associates the category information assigned to the second segmented sentence with and stores the second segmented sentence. (Appendix B10) The document processing device according to Appendix B9, further comprising: an extraction unit that extracts the second segmented sentence based on the category information and the importance associated with the second segmented sentence. (Appendix B11) The document processing device according to Appendix B10, wherein the extraction unit accepts information related to a created document, and extracts the second divided sentences from the second divided sentences associated with the category information corresponding to the information about the created document based on the importance associated with the second divided sentences. (Appendix B12) The document processing device according to Appendix B10, wherein the extraction unit accepts a predetermined type of created document, and extracts the second divided sentences from the second divided sentences associated with the category information corresponding to the type of created document based on the importance associated with the second divided sentences. (Appendix B13) A document processing method comprising: acquiring first divided sentences obtained by dividing text in a created document created from a specified document into predetermined units, and acquiring second divided sentences obtained by dividing text in another document into predetermined units; calculating a similarity of the second divided sentences to each of a plurality of the first divided sentences according to a predetermined criterion, determining an importance of the second divided sentence based on the similarity; and storing the determined importance with respect to the second divided sentences.(Appendix B14) The document processing method according to Appendix B13, wherein a second divided sentence is extracted based on the importance associated with the second divided sentence. (Appendix B15) The document processing method according to Appendix B13, wherein, based on a plurality of the second divided sentences, predetermined category information is assigned to each of the second divided sentences, and the category information assigned to the second divided sentence is associated with the second divided sentence and stored. (Appendix B16) The document processing method according to Appendix B15, wherein a second divided sentence is extracted based on the category information and the importance associated with the second divided sentence. (Appendix B17) A computer-readable storage medium storing a program that causes a computer to execute the following processes: acquiring first divided sentences obtained by dividing text in a created document created from a specified document into specified units, and acquiring second divided sentences obtained by dividing text in another document into specified units; calculating the similarity of the second divided sentences to each of the multiple first divided sentences based on a predetermined criterion; determining the importance of the second divided sentences based on the similarity; and associating and storing the determined importance with the second divided sentences.
[0075] REFERENCE SIGNS LIST 10 Document processing device 11 Dividing unit 12 Assigning unit 13 Associating unit 14 Extracting unit 16 Electronic medical record storage unit 17 Text information storage unit 20 User terminal 100 Document processing device 101 CPU 102 ROM 103 RAM 104 Program group 105 Storage device 106 Drive device 107 Communication interface 108 Input / output interface 109 Bus 110 Storage medium 111 Communication network 121 Acquisition unit 122 Assigning unit 123 Associating unit 131 Acquisition unit 132 Determination unit 133 Associating unit
Claims
1. A document processing device comprising: an acquisition unit that acquires first divided sentences obtained by dividing text in a created document created from a specified document into specified units, and acquires second divided sentences obtained by dividing text in another document into specified units; a determination unit that calculates a similarity of the second divided sentences to each of a plurality of the first divided sentences based on a preset criterion, and determines an importance of the second divided sentences based on the similarity; and an association unit that associates and stores the determined importance with the second divided sentences.
2. A document processing device according to claim 1, wherein the determination unit determines the importance based on the number of first divided sentences for which the similarity of the second divided sentence to the first divided sentence satisfies a preset condition.
3. A document processing device as described in claim 1, wherein the acquisition unit acquires a first divided sentence and other divided sentences other than the first divided sentence contained in the created document, the first divided sentence being a divided sentence obtained by dividing the text in the specified document into predetermined units, and the determination unit calculates the similarity of the second divided sentence to each of the first divided sentence and the other divided sentences, and determines the importance of the second divided sentence based on the similarity of the second divided sentence to a plurality of the first divided sentences and the other divided sentences.
4. A document processing device as described in claim 3, wherein the determination unit determines the importance based on the number of first divided sentences among the first divided sentence and the other divided sentences for which the similarity of the second divided sentence to the first divided sentence and the other divided sentences satisfies a preset condition.
5. A document processing device as described in claim 4, wherein the determination unit determines the importance based on the number of first divided sentences among a predetermined number of the first divided sentences and the other divided sentences in order of the highest similarity of the second divided sentence to the first divided sentence and the other divided sentences.
6. A document processing device as described in claim 3, wherein the determination unit converts the first divided sentence, the other divided sentence, and the second divided sentence into fixed-length vectors, and calculates the similarity of the second divided sentence to each of the first divided sentence and the other divided sentences using the converted fixed-length vectors.
7. A document processing device as described in claim 6, wherein the determination unit converts a specified sentence into the fixed-length vector by inputting the first divided sentence, the other divided sentence, and the second divided sentence into a machine learning model configured to convert a specified sentence into the fixed-length vector.
8. A document processing device according to claim 1, further comprising: an extraction unit that extracts the second divided sentence based on the importance associated with the second divided sentence.
9. A document processing device as described in claim 1, wherein the determination unit assigns predetermined category information to each of the second divided sentences based on a plurality of the second divided sentences, and the association unit associates the category information assigned to the second divided sentence with the second divided sentence and stores it.
10. A document processing device according to claim 9, further comprising an extraction unit that extracts the second divided sentence based on the category information and the importance associated with the second divided sentence.
11. A document processing device as described in claim 10, wherein the extraction unit accepts information related to a created document, and extracts the second divided sentence from among the second divided sentences to which the category information corresponding to the information about the created document is associated, based on the importance associated with the second divided sentence.
12. A document processing device as described in claim 10, wherein the extraction unit accepts a predetermined type of created document, and extracts the second divided sentences from among the second divided sentences associated with the category information corresponding to the type of created document, based on the importance associated with the second divided sentences.
13. A document processing method comprising: obtaining first divided sentences obtained by dividing text in a created document created from a specified document into predetermined units; obtaining second divided sentences obtained by dividing text in another document into predetermined units; calculating a similarity of the second divided sentences to each of a plurality of the first divided sentences based on a preset criterion; determining an importance of the second divided sentences based on the similarity; and storing the determined importance in association with the second divided sentences.
14. A document processing method according to claim 13, further comprising the step of extracting the second divided sentence based on the importance associated with the second divided sentence.
15. A document processing method as claimed in claim 13, comprising the steps of: assigning pre-set category information to each of the second divided sentences based on a plurality of the second divided sentences; and storing the category information assigned to each of the second divided sentences in association with the second divided sentences.
16. A document processing method according to claim 15, further comprising the step of extracting the second divided sentence based on the category information and the importance associated with the second divided sentence.
17. A computer-readable storage medium having stored therein a program for causing a computer to execute a process of obtaining first divided sentences obtained by dividing text in a created document created from a specified document into predetermined units, and obtaining second divided sentences obtained by dividing text in another document into predetermined units, calculating a similarity of the second divided sentences to each of the first divided sentences based on a preset criterion, determining an importance of the second divided sentences based on the similarity, and storing the determined importance in association with the second divided sentences.
Citation Information
Patent Citations
Document abstract processing method and device, equipment and medium
CN114461795A
Device and method for document processing and storage medium storing document processing program
JP1999053396A
Medical record summary information generating device, medical record summary information generation method, and program
JP2020038602A