Method for summarizing tasks, task summarization system, and computer program
Patent Information
- Application Number
- JP2023125439
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-05-10
- Filing Date
- 2023-08-01
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2043-08-01
AI Technical Summary
Existing methods for organizing tasks from discussions are inefficient and time-consuming due to complex and detailed task summaries, wasting significant time in meetings.
A to-do organizing system utilizing a front-end and back-end language model, implemented through machine learning, to summarize tasks by token division and generate concise summary results based on the number of characters in the text data.
The system effectively expands the applicability of language models and provides a versatile function to summarize tasks, especially from lengthy discussions, by generating information-dense summaries with reduced repetitive content.
Smart Images

Figure 00000010_0000 
Figure 00000011_0000
Abstract
Description
[Technical field]
[0001] The present invention relates to a method for organizing tasks, in particular to a method for organizing tasks applied to text data.The present invention further relates to a system and a computer program for organizing tasks applied to text data. [Background technology]
[0002] In many teams, in order to achieve the team's common goal, members hold meetings to discuss and set up multiple step-by-step tasks. However, in reality, meetings often drag on and the tasks that are discussed end up being complicated and detailed, so it often takes a lot of time to put together the tasks. Therefore, the problem that the present invention aims to solve is to help put together tasks using modern technology. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] China Patent Application Publication No. 111277589 Summary of the Invention [Problem to be solved by the invention]
[0004] Therefore, an object of the present invention is to provide a method, a system, and a computer program for summarizing to-do items, which are capable of summarizing to-do items from discussion data. [Means for solving the problem]
[0005] The method for synthesizing tasks is executed by a system for synthesizing tasks. The system for synthesizing tasks stores a front-end language model and a back-end language model realized by machine learning technology.
[0006] The method for summarizing tasks includes the steps of: A) performing token splitting on target text data, and determining whether the number of split tokens is equal to or greater than a predetermined threshold; B) if it is determined that the number of split tokens is equal to or greater than the predetermined threshold, using a front-end language model to generate pre-processed text data based on the target text data, the pre-processed text data being expressed in a natural language and having a number of characters less than the number of characters in the target text data, and then using a back-end language model to identify N meanings indicated by the pre-processed text data from the pre-processed text data, based on the N meanings, and generate and output a first summary result including N task messages that correspond to the N meanings, are expressed in a natural language, and each indicate a task, where N is an integer equal to or greater than 1; and C) if it is determined that the number of split tokens is not equal to or greater than the predetermined threshold, using the back-end language model to identify M meanings indicated by the target text data from the target text data, based on the M meanings, and generate and output a second summary result including M task messages that correspond to the M meanings, are expressed in a natural language, and each indicate a task, where M is an integer equal to or greater than 1.
[0007] The task orchestration system includes a processing unit and a storage unit electrically connected to the processing unit.
[0008] The storage unit stores a front-end language model and a back-end language model that are realized by machine learning techniques.
[0009] The processing unit is configured to perform token division on the target text data, determine whether the number of divided tokens is equal to or greater than a predetermined threshold, and if it is determined that the number of divided tokens is equal to or greater than the predetermined threshold, use a front-end language model to generate pre-processed text data based on the target text data, the pre-processed text data being expressed in a natural language and having a number of characters less than the number of characters in the target text data, and then use a back-end language model to identify N meanings indicated by the pre-processed text data from the pre-processed text data, based on the N meanings, generate and output a first summary result including N task messages each corresponding to the N meanings, expressed in a natural language, and indicating to-do items, where N is an integer equal to or greater than 1; if it is determined that the number of divided tokens is not equal to or greater than the predetermined threshold, use a back-end language model to identify M meanings indicated by the target text data from the target text data, based on the M meanings, generate and output a second summary result including M task messages each corresponding to the M meanings, expressed in a natural language, and indicating to-do items, where M is an integer equal to or greater than 1.
[0010] The computer program, when executed by a computer system, causes the computer system to perform the aforementioned aggregation method using a front-end language model and a back-end language model implemented by machine learning. Effect of the Invention
[0011] By executing the to-do summarizing method according to the present invention, when the number of divided tokens of the target text data is determined to be equal to or greater than a predetermined threshold (i.e., the number of characters in the target text data is relatively large), the to-do summarizing system inputs the target text data into a front-end language model to obtain pre-processed text data, and then inputs the pre-processed text data into a back-end language model to obtain a summarizing result. In this way, when the back-end language model has a limit on the number of input characters, the present invention can expand the application range of the back-end language model and provide a more versatile to-do summarizing function. The present invention uses two language models to realize a more versatile to-do summarizing system, which can summarize to-dos from a recording file of a discussion by multiple people, such as a meeting, or a written record thereof.
[0012] Other features and advantages of the present invention will become apparent in the following detailed description of the embodiments which proceeds with reference to the accompanying drawings. [Brief description of the drawings]
[0013] [Figure 1] 1 is a block diagram showing an example of a system for summarizing tasks according to an embodiment of the present invention and a user device to which the system is applied; [Diagram 2] 1 is a flowchart illustrating a method for summarizing tasks according to the embodiment; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] Before describing the present invention in more detail, unless otherwise specified, the term "electrically connect" in this specification is used to describe a coupling relationship between computer hardware (e.g., electronic systems, equipment, devices, units, parts, etc.), and refers to a "wired electrical connection" in which multiple computer hardware pieces are physically connected via conductor or semiconductor materials, or a "wireless electrical connection" in which wireless data transmission is realized using wireless communication technology (e.g., wireless network, Bluetooth (registered trademark), electrical induction, etc.). Meanwhile, unless otherwise specified, the term "electrically connect" in this specification further refers to a "direct electrical connection" in which multiple computer hardware pieces are directly connected to each other, or an "indirect electrical connection" in which multiple computer hardware pieces are connected to each other via other computer hardware.
[0015] 1, a to-do list system 1 of the present invention is configured to be electrically connected to a user-side device 5. The user-side device 5 is, for example, a smartphone, a tablet computer, a notebook computer, or a desktop computer used by a user.
[0016] The system 1 for organizing to-dos includes a processing unit 11 electrically connected to the user device 5, and a storage unit 12 electrically connected to the processing unit 11. More specifically, in this embodiment, the system 1 for organizing to-dos is, for example, a computer device, but may also be a server device. The processing unit 11 is a processor realized by an integrated circuit and has the functions of command transmission and reception and data calculation, and the storage unit 12 is a data storage device (e.g., a hard disk, a hard disk array, or other types of computer-readable storage media) for storing digital data. In a similar embodiment, the processing unit 11 may be a processing circuit having a processor, and the storage unit 12 may be a collection of multiple storage devices of the same or different types. Furthermore, in another embodiment, the system 1 for organizing to-dos may be multiple computers or server devices electrically connected to each other, in which case the processing unit 11 is a collection of processors or processing circuits that each of the multiple computers or server devices has, and the storage unit 12 is a collection of data storage devices that each of the multiple computers or server devices has. Therefore, the computer hardware realization of the system 1 for organizing to-dos is not limited to this embodiment.
[0017] The storage unit 12 stores a speech processing model M0, a front-end language model LM1, and a back-end language model LM2.
[0018] The voice processing model M0 is realized by machine learning, for example, using voice data, which is a recording file representing the voices of multiple speakers, as training data. As a result, the voice processing model M0 can recognize speakers using speaker separation based on voiceprint recognition for voice data representing the voices of multiple speakers, separate the voices of the voice data for each speaker, and distinguish the speech content of each speaker. The voice processing model M0 can also use speech-to-text to generate corresponding text data based on the voice represented by the voice data. Note that training of the voice processing model M0 can be realized by conventional technology and is not the point of the present invention, so will not be described in detail.
[0019] The front-end language model LM1 and the back-end language model LM2 are pre-trained language models realized by machine learning, using text data, which is a character record showing the contents of a conversation between multiple people, as training data. This allows natural language processing to be performed on the input text data by a generative method. More specifically, in this embodiment, the front-end language model LM1 is preferably BLOOMZ, but may be a pre-trained language model capable of generating text expressed in a natural language, such as BLOOM, MT0, GPT-2, or T5. Meanwhile, in this embodiment, the back-end language model LM2 is preferably GPT-3, but may be a pre-trained language model capable of generating text expressed in a natural language, such as GPT-4, GPT-3.5, or GPT-2.
[0020] In this specification, "generative" (also called "abstractive") means that a language model generates output text data based on input text data using natural language generation technology. As a person having ordinary knowledge in the technical field to which the present invention belongs knows, "generative" means that a language model understands input text data and then generates output text data of a new document using natural language generation technology, so that the output text data may contain expressions that are not included in the input text data. For example, the output text data may contain words or sentences that are not included in the input text data, or may express the contents of the input text data more concisely, or may summarize the contents of the input text data in bullet points or in a table. As a result, the "generative" method in this specification is different from the "extractive" method in which words and sentences included in the input text data are extracted and combined to produce output text data.
[0021] Referring to FIG. 2, a method for consolidating tasks executed by the system 1 for consolidating tasks of the present embodiment is shown.
[0022] In step S1, the processing unit 11 of the to-do system 1 obtains speech data representing the speech of multiple speakers, and then generates corresponding text data based on the speech data using the speech processing model M0.
[0023] More specifically, the processing unit 11 uses the speech processing model M0 to generate corresponding text data through speaker separation and speech-to-text technology, so that the corresponding text data includes multiple speech parts each corresponding to one of multiple speakers, and the speech parts can indicate the order of speech of each speaker and the content of each speech.
[0024] In addition, in this embodiment, the voice data is transmitted from the user-side device 5 to the processing unit 11 by, for example, manual operation of the user. In this embodiment, the voice data is, for example, a recording file of a face-to-face meeting or an online meeting. In other embodiments, the processing unit 11 may read voice data stored in an external storage device (for example, a USB flash drive) or receive voice data from a voice input device (for example, a microphone), and the means for obtaining the voice data of the processing unit 11 is not limited to this embodiment.
[0025] In step S2, the processing unit 11 performs token division on the text data generated in step 1 (hereinafter referred to as target text data) to obtain a plurality of divided tokens. In this embodiment, each divided token is one character or a combination of multiple characters divided from the target text data, that is, a token. In addition, the processing unit 11 performs token division based on a token table pre-stored in the storage unit 12. For example, in this embodiment, the processing unit 11 divides the text data of "natural language" into two tokens, "natural" and "language", based on the token table. In other embodiments, the processing unit 11 may divide the text data of "natural language" into four tokens, "self", "natural", "language", and "language", and is not limited to this embodiment.
[0026] In step S3, the processing unit 11 determines whether the number of divided tokens is equal to or greater than a predetermined threshold. The predetermined threshold may be set to, for example, 2000, but can be freely set and adjusted according to actual situations and needs, and is not limited to a fixed value. If it is determined that the number of divided tokens is equal to or greater than the predetermined value, the flow proceeds to step S4, and if it is determined that the number of divided tokens is smaller than the predetermined value, the flow proceeds to step S7.
[0027] When it is determined that the number of divided tokens is equal to or greater than the predetermined threshold, in step S4, since the number of divided tokens being equal to or greater than the predetermined threshold indicates that the number of characters in the target text data is relatively large, the processing unit 11 inputs the target text data to the front-end language model LM1, and uses the front-end language model LM1 to generate pre-processed text data based on the target text data, the pre-processed text data being expressed in a natural language and having a number of characters less than that of the target text data, in a generative manner. The pre-processed text data is a summary of the target text data.
[0028] More specifically, in this embodiment, the processing unit 11 inputs the target text data and soft prompts corresponding to the "multiple speakers" text data into the front-end language model LM1. The soft prompts (also called "continuous prompts") are predicted in advance by a language model (which may be, but is not limited to, the front-end language model LM1) using a prompt learning technique of prompt engineering. The prompt learning may be, but is not limited to, prefix tuning, tuning initialized with discrete prompts, or hard-soft prompt hybrid tuning. The soft prompts are expressed by vectors or by numerical values in other non-natural languages. The front-end language model LM1 uses the soft prompts to understand the overall meaning of the "multiple speakers" input text data (i.e., the target text data) and summarize one or more discussion themes, and then uses its own attention (attention mechanism) to generate output text data for parts of the input text data that are highly relevant to the discussion theme (i.e., ignoring parts that are less relevant or irrelevant). Attention is implemented by the front-end language model LM1 during the training phase using the gradient descent method. Note that attention is a conventional technique and will not be described in detail in this specification.
[0029] In addition, unlike hard prompts (also called "discrete prompts"), soft prompts can effectively avoid a situation in which the output of a language model differs greatly due to a slight difference in the input. In other words, soft prompts can improve the stability and reliability of a language model. Therefore, when the front-end language model LM1 is made to use a specific summary generation strategy according to the characteristics of text data showing multiple speakers, using soft prompts is more effective than using hard prompts. Making the front-end language model LM1 use a specific summary generation strategy according to the characteristics of text data showing multiple speakers means, for example, generating a summary only for a part corresponding to a response to a "question and answer" for text data including a "question and answer", or generating a summary only for a part corresponding to a result of a discussion that has been agreed upon (i.e., a statement that has not been rejected or refuted) for text data including a "semantically contradictory dialogue". In addition, the front-end language model LM1 can understand text data having the above-mentioned characteristics by machine learning (natural language understanding) and can generate a summary using a specific summary generation strategy (natural language generation). In addition, the training and specific operation of the front-end language model LM1 are not the point of this specification, so they will not be described in detail.
[0030] As a result, the front-end language model LM1 can ignore parts of the target text data that are less relevant to the context and generate pre-processed text data for parts that are highly relevant, so that when the number of characters in the target text data is relatively large, the target text data can be simplified. Also, since the front-end language model LM1 generates pre-processed text data by a generative method, when the target text data contains a lot of repetitive content, the present invention can better summarize the content of the target text data and generate pre-processed text data with high information density compared to an extractive method.
[0031] In step S5, the processing unit 11 inputs the pre-processed text data into the back-end language model LM2, and uses the back-end language model LM2 to identify N meanings represented by the pre-processed text data from the pre-processed text data, and generates a first summary result by a generative method based on the N meanings. N is an integer equal to or greater than 1. The meanings represented by the pre-processed text data are expressed, for example, as sentences representing plans or imperative sentences such as "I plan to do XX" or "I must do XX", and are identified by the back-end language model LM2. Meanwhile, the first summary result is realized, for example, as a text file, and further includes N task messages corresponding to the N meanings. Each task message is expressed in natural language and indicates the task represented by the corresponding meaning, for example, but is not limited to, collecting specific materials and submitting them by a specific date, visiting a specific client during a specific time period, and periodically reporting the progress of a specific task.
[0032] In step S6, the processing unit 11 outputs the first summary result. More specifically, in this embodiment, the processing unit 11 transmits the first summary result to the user-side device 5, and causes the user-side device 5 to display the first summary result to the user. In other embodiments, the processing unit 11 may transmit the first summary result to a display device (not shown) electrically connected to the processing unit 11 for display, or may transmit the first summary result to one or more pre-set e-mail addresses, which is not limited to this embodiment.
[0033] When it is determined that the number of divided tokens is smaller than the predetermined threshold, in step S7, the fact that the number of divided tokens is smaller than the predetermined threshold indicates that the number of characters in the target text data is relatively small, so the processing unit 11 inputs the target text data into the back-end language model LM2, and uses the back-end language model LM2 to identify M meanings indicated by the target text data from the target text data, and generates a second summary result by a generative method based on the M meanings. M is an integer equal to or greater than 1. The second summary result is realized, for example, as a text file, similar to the first summary result, and further includes M mission messages corresponding to the M meanings. Each mission message is expressed in a natural language and indicates the task indicated by the corresponding meaning, and may be, but is not limited to, for example, material collection, client visit, progress report, etc.
[0034] In step S8, the processing unit 11 outputs the second summary result. More specifically, in this embodiment, similar to step S6, the processing unit 11 transmits the second summary result to the user-side device 5, and causes the user-side device 5 to display the second summary result to the user. In other embodiments, similar to step S6, the processing unit 11 may transmit the second summary result to a display device electrically connected to the processing unit 11 to display it, or may transmit the second summary result to one or more pre-set email addresses, and is not limited to this embodiment.
[0035] It should be understood that Fig. 2 and steps S1 to S8 are merely illustrative of the method of organizing tasks of the present invention. Even if steps S1 to S8 are combined, divided, or the order is changed, if the same effect can be obtained in a substantially similar manner to this embodiment, it corresponds to an embodiment of the method of organizing tasks of the present invention and should be included in the scope of the present invention. Therefore, Fig. 2 and the above steps S1 to S8 do not limit the present invention.
[0036] In this embodiment, the computer program includes a front-end language model LM1, a back-end language model LM2, and a speech processing model M0. When the computer program of this embodiment is executed by a computer system (e.g., one computer device or a server device, or a combination of multiple computer devices or server devices), the computer system becomes the above-mentioned task consolidation system 1, and causes the computer system to execute the task consolidation method using the front-end language model LM1, the back-end language model LM2, and the speech processing model M0. In another embodiment, the front-end language model LM1, the back-end language model LM2, and the speech processing model M0 may be stored in a remote server. When the computer program is executed by a computer system, the computer system accesses the front-end language model LM1, the back-end language model LM2, and the speech processing model M0 via a network.
[0037] In the present invention, by executing the method for summarizing to-dos, the system 1 for summarizing to-dos, when it is determined that the number of divided tokens of the target text data is equal to or greater than a predetermined threshold (i.e., the number of characters of the target text data is relatively large), obtains pre-processed text data with a smaller number of characters based on the target text data using the front-end language model LM1, and then generates a first summary result from the pre-processed text data using the back-end language model LM2. In this way, when the back-end language model LM2 has a limit on the number of input characters, the present invention can expand the application range of the back-end language model LM2 and provide a more versatile function for summarizing to-dos. In addition, since the front-end language model LM1 generates pre-processed text data using a generative method, when the target text data contains a lot of repetitive content, it is possible to summarize the content of the target text data using an extractive method, and to generate pre-processed text data with a high information density to be input to the back-end language model LM2. Thus, the present invention uses a language model that utilizes two generative methods to realize a more versatile system 1 for summarizing to-dos, and can summarize to-dos from, for example, a recording file of a discussion such as a meeting or a written record thereof, thereby reliably achieving the object of the present invention.
[0038] In the above description, for the purpose of explanation, numerous specific details are set forth to provide a thorough understanding of the embodiments. However, it will be apparent to one skilled in the art that one or more other embodiments may be implemented without the specific details. In addition, in the description of "one embodiment" or "one embodiment" in this specification, it should be understood that all descriptions accompanied by ordinal numbers or other indications may be included in the specific implementation of the present invention having specific aspects, structures, and features. Furthermore, in this specification, multiple variations are sometimes incorporated into one embodiment, drawing, or description thereof, but this is for the purpose of streamlining the specification and for the purpose of understanding the multiple aspects of the present invention, and one or more features or specific examples in one embodiment may be implemented together with one or more features or specific examples in other embodiments, where appropriate, in the implementation of the present invention.
[0039] Although the embodiments and modified examples of the present invention have been described above, the present invention is not limited to these, and is intended to encompass all modifications and equivalent configurations as various configurations falling within the spirit and scope of the broadest interpretation. [Explanation of symbols]
[0040] 1. A system for organizing your to-dos 11 Processing Unit 12 Storage Unit M0 Audio Processing Model LM1 Front-end Language Model LM2 Backend Language Models 5 User side equipment S1-S8 Step
Claims
1. 1. A method of to-do scheduling implemented by a to-do scheduling system, comprising: The system for consolidating tasks stores a front-end language model and a back-end language model realized by machine learning technology, and the method for consolidating tasks includes: A) performing token division on target text data and determining whether the number of divided tokens is equal to or greater than a predetermined threshold; B) when it is determined that the number of the divided tokens is equal to or greater than the predetermined threshold, using the front-end language model to generate pre-processed text data based on the target text data, the pre-processed text data being expressed in a natural language and having a number of characters less than the number of characters in the target text data, and then using the back-end language model to identify N meanings indicated by the pre-processed text data from the pre-processed text data, and based on the N meanings, generate and output a first summary result including N task messages respectively corresponding to the N meanings, expressed in a natural language, and indicating tasks, where N is an integer equal to or greater than 1; C) when it is determined that the number of the divided tokens is not equal to or greater than the predetermined threshold, using the back-end language model to identify M meanings indicated by the target text data from the target text data, and based on the M meanings, generate and output a second summary result including M mission messages each corresponding to the M meanings, each expressed in a natural language, and each indicating an action to be taken, where M is an integer equal to or greater than 1. How to organize your to-dos.
2. In step B), the pre-processed text data is generated in a generative manner using the front-end language model, and then the first summary result is generated in a generative manner using the back-end language model; 2. The method of claim 1, wherein in step C), the second summary result is generated by a generative method using the back-end language model.
3. 3. The method of claim 2, wherein in step B), the pre-processed text data is generated using the front-end language model by inputting the target text data and soft prompts predicted by a language model into the front-end language model.
4. 2. The method of claim 1, further comprising, prior to step A), a step of: D) receiving voice data and generating text data of the target including a plurality of speech portions each corresponding to one of the plurality of speakers based on the voice data representing the voices of a plurality of speakers.
5. A processing unit; a storage unit electrically connected to the processing unit; The storage unit stores a front-end language model and a back-end language model realized by machine learning technology; The processing unit includes: Tokenizing the target text data and determining whether the number of tokens obtained by the tokenizing is equal to or greater than a predetermined threshold; if it is determined that the number of the divided tokens is equal to or greater than the predetermined threshold, generate pre-processed text data based on the target text data using the front-end language model, the pre-processed text data being expressed in a natural language and having a number of characters less than the number of characters in the target text data, and then identify N meanings indicated by the pre-processed text data from the pre-processed text data using the back-end language model, and generate and output a first summary result based on the N meanings, the first summary result including N task messages respectively corresponding to the N meanings, expressed in a natural language, and indicating tasks, where N is an integer equal to or greater than 1; When it is determined that the number of the divided tokens is not equal to or greater than the predetermined threshold, the back-end language model is used to identify M meanings indicated by the target text data from the target text data, and based on the M meanings, a second summary result is generated and output, the second summary result including M mission messages each corresponding to the M meanings, each expressed in a natural language, and each indicating an action to be taken, where M is an integer equal to or greater than 1. A system for organizing things.
6. The processing unit includes: if it is determined that the number of the divided tokens is equal to or greater than the predetermined threshold, generating the pre-processed text data by a generative technique using the front-end language model, and then generating the first summary result by a generative technique using the back-end language model; The system for consolidating tasks according to claim 5, configured to generate the second consolidation result by a generative method using the back-end language model when it is determined that the number of divided tokens is not greater than or equal to the predetermined threshold.
7. 7. The to-do summary system of claim 6, wherein the processing unit is configured to generate the pre-processed text data using the front-end language model by inputting the target text data and soft prompts predicted by a language model into the front-end language model.
8. 6. The to-do list system of claim 5, wherein the processing unit is further configured to receive audio data and generate, based on the audio data representing voices of a plurality of speakers, text data of the target including a plurality of speech portions each corresponding to one of the plurality of speakers.
9. 5. A computer program product which, when executed by a computer system, causes the computer system to perform the method of any one of claims 1 to 4 using a front-end language model and a back-end language model implemented by machine learning.