Minutes processing method, device, equipment and media

By entering the conference text into the -do recognition model and tension judgment model, and determining the conference task statement, the problem of insufficient efficiency and accuracy of meeting record text conversion and task intention extraction in the prior art is solved, and a more efficient and accurate task extraction effect is achieved.

JP7676561B2Active Publication Date: 2025-05-14BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023544227
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-27
Filing Date
2022-01-05
Publication Date
2025-05-14
Estimated Expiration
2042-01-05

AI Technical Summary

Technical Problem

The prior art has efficiency and accuracy problems in the text conversion and task intention extraction process of conference records, resulting in the decision-making process being inefficient and accurate enough.

Method used

By entering the meeting text into the -do recognition model, the initial task statement is determined; then input the initial task statement into the tension judgment model to determine the tension result; finally, the conference task statement is determined based on the tension result, thereby improving the accuracy and efficiency of task extraction.

Benefits of technology

Through this method, the completed sentences are avoided from being misidentified as meeting task statements, which significantly improves the accuracy of task statements, and improves the user's work efficiency in processing task statements, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676561000001
    Figure 0007676561000001
  • Figure 0007676561000002
    Figure 0007676561000002
  • Figure 0007676561000003
    Figure 0007676561000003
Patent Text Reader

Abstract

A method, device, equipment, and medium for processing minutes of a meeting. The method includes a step (101) of acquiring a meeting text of a meeting audio / video, a step (102) of inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence, a step (103) of inputting the initial ToDo sentence into a tense judgment model to determine a tense result of the initial ToDo sentence, and a step (104) of determining a meeting ToDo sentence in the initial ToDo sentence based on the tense result. According to the above method, by recognizing the meeting text of the meeting audio / video and then adding a tense judgment, it is possible to increase the accuracy of determining the meeting ToDo sentence, and further increase the user's work efficiency by using the meeting ToDo sentence, thereby improving the user's experience.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims the benefit of priority to a Chinese patent application bearing application number 202110113700.1 and entitled "Method, Apparatus, Instrument and Medium for Processing Minutes," filed with the State Intellectual Property Office of the People's Republic of China on January 27, 2021, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to the technical field of meeting recognition, and in particular to a method, device, equipment and medium for processing minutes of a meeting. [Background technology]

[0003] With the continuous development of intelligent devices and multimedia technology, online conferencing via intelligent devices has become increasingly used in daily life and office life due to its remarkable expression in terms of communication efficiency and information storage.

[0004] After the meeting, the audio and video are converted into text by recognition processing, and the ToDo sentences containing the task intent can be determined from the text. However, determining the ToDo sentences has problems of low efficiency and low accuracy. Summary of the Invention [Means for solving the problem]

[0005] In order to solve or at least partially solve the above technical problems, the present disclosure provides a method, device, apparatus, and medium for processing minutes.

[0006] An embodiment of the present disclosure comprises: obtaining a conference text of the conference audio-video; inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence; inputting the initial ToDo sentence into a tense judgment model to determine a tense outcome of the initial ToDo sentence; determining a meeting ToDo sentence in the initial ToDo sentence based on the tense result; Provide a method for handling minutes, including:

[0007] An embodiment of the present disclosure comprises: receiving a user's display trigger operation for a target transcript in a minutes display interface, in which the minutes display interface displays a conference audio-video, a conference text of the conference audio-video, and the target transcript; displaying the target sentence and related sentences of the target sentence; The present invention further provides a method for processing minutes, including:

[0008] An embodiment of the present disclosure comprises: A text acquisition module for acquiring the conference text of the conference audio-video; an initial ToDo module for inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence; a tense judgment module for inputting the initial ToDo sentence into a tense judgment model to determine a tense outcome of the initial ToDo sentence; a meeting ToDo module for determining a meeting ToDo sentence in the initial ToDo sentence based on the tense result; The present invention further provides a minutes processing device, including:

[0009] An embodiment of the present disclosure comprises: A display trigger module for receiving a user's display trigger operation for a target transcript in a minutes display interface, the minutes display interface displaying a conference audio / video, a conference text of the conference audio / video, and the target transcript; a display module for displaying the target sentence and related sentences of the target sentence; The present invention further provides a minutes processing device, including:

[0010] An embodiment of the present disclosure further provides an electronic device including a processor and a memory for storing instructions executable by the processor, the processor being used to realize a minutes processing method according to an embodiment of the present disclosure by reading and executing the executable instructions from the memory.

[0011] The embodiment of the present disclosure further provides a computer-readable storage medium having stored thereon a computer program for executing the minutes processing method according to the embodiment of the present disclosure. Effect of the Invention

[0012] The technical solution according to the embodiment of the present disclosure has the following advantages over the conventional technology. In the method for processing minutes according to the embodiment of the present disclosure, the method includes the steps of obtaining a meeting text from a meeting audio / video, inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence, inputting the initial ToDo sentence into a tense judgment model to determine a tense result of the initial ToDo sentence, and determining a meeting ToDo sentence in the initial ToDo sentence according to the tense result. According to the above technical solution, by recognizing the meeting text from the meeting audio / video and then adding tense judgment, it is possible to avoid recognizing an already completed sentence as a meeting ToDo sentence, and greatly improve the accuracy of determining the meeting ToDo sentence, and further improve the user's work efficiency by the meeting ToDo sentence, thereby improving the user's experience effect. [Brief description of the drawings]

[0013] The foregoing and other features, advantages and aspects of each embodiment of the present disclosure will become more apparent with reference to the following specific embodiments in conjunction with the accompanying drawings. The same or similar drawing reference numerals refer to the same or similar elements throughout the drawings. It should be understood that the drawings are schematic and that parts and elements are not necessarily drawn to scale.

[0014] [Figure 1] 1 is a flowchart of a method for processing minutes according to an embodiment of the present disclosure; [Diagram 2] 1 is a flowchart of a method for processing minutes according to another embodiment of the present disclosure; [Diagram 3] 1 is a schematic diagram of a minutes display interface according to an embodiment of the present disclosure; [Figure 4] 1 is a schematic diagram illustrating a configuration of a meeting minutes processing device according to an embodiment of the present disclosure; [Diagram 5] 1 is a schematic diagram illustrating a configuration of a meeting minutes processing device according to an embodiment of the present disclosure; [Figure 6] 1 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, the embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be realized in various forms and should not be construed as being limited to the embodiments described herein, but rather that these embodiments are provided for a deeper and more complete understanding of the present disclosure. It should also be understood that the drawings and embodiments of the present disclosure are only given for illustrative purposes and do not limit the scope of protection of the present disclosure.

[0016] It should be understood that the steps described in the method embodiments of the present disclosure may be performed in a different order and / or in parallel. Additionally, method embodiments may include additional steps and / or omit the performance of steps that are illustrated. The scope of the present disclosure is not limited in this respect.

[0017] As used herein, the term "comprising" and variations thereof mean an open-ended inclusion, i.e., "including but not limited to." The term "based on" means "based at least in part on." The term "in one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one other embodiment," and the term "some embodiments" means "at least some embodiments." Additionally, definitions relating to other terms are provided below.

[0018] It should be noted that the concepts of "first", "second", etc. referred to in this disclosure are used only to distinguish different devices, modules or units, but do not limit the order or interdependence of functions performed by these devices, modules or units.

[0019] It will be appreciated by those of skill in the art that modifications such as "a" and "a plurality" referred to in the present disclosure are intended to be illustrative rather than limiting and should be construed as "one or more" unless the context clearly dictates otherwise.

[0020] The names of messages or information exchanged between devices in the embodiments of the present disclosure are used for illustrative purposes only and are not used to limit the scope of these messages or information.

[0021] After the meeting, the meeting audio and video can be converted into text by recognition processing. However, since the content of the meeting text is usually large, how to quickly and accurately extract sentences containing task intentions is particularly important. The content of the meeting is a record of discussing one or more topics, and often leads to a certain conclusion or association with many other topics in the end. In addition, many tasks that need to be completed during the meeting are often allocated, and since the meeting text of the meeting contains a large number of characters, if it is possible to select tasks that contain intentions (todo) that need to be completed, the effort required for organizing the minutes can be greatly reduced. Among them, ToDo sentences can be one type of intention. However, currently, there are problems with low efficiency and low accuracy in determining ToDo sentences. In order to solve the above problems, an embodiment of the present disclosure provides a method for processing minutes. Hereinafter, the method will be described with reference to a specific embodiment.

[0022] 1 is a flowchart of a method for processing minutes according to an embodiment of the present disclosure. The method can be executed by a minutes processing device. Here, the device can be realized by software and / or hardware and generally integrated into electronic equipment. As shown in FIG. 1, the method can include the following steps:

[0023] Step S101: obtaining a conference text of a conference audio-video by a processing device.

[0024] Meeting Audio-Video means the audio and / or video used to record the meeting process, and Meeting Text means the textual content obtained by processing the meeting Audio-Video through speech recognition.

[0025] In an embodiment of the present disclosure, the processing device can obtain a conference text obtained by audio-video processing, and the processing device can also obtain the conference text by obtaining the conference audio-video and processing the conference audio-video.

[0026] Step S102: The processing device inputs the meeting text into a ToDo recognition model to determine an initial ToDo sentence.

[0027] The ToDo recognition model is a pre-trained deep learning model for recognizing ToDo intent sentences from meeting text, and the specific deep learning model used is not limited.

[0028] In an embodiment of the present disclosure, before step S102 is performed, the processing device can also generate a ToDo recognition model. The ToDo recognition model is generated by the following method. That is, the ToDo recognition model is obtained by training an initial single classification model based on the positive samples of the ToDo sentence. Considering the boundaryless nature of the negative samples, the embodiment of the present disclosure takes the ToDo recognition model as a single classification model as an example for description. The single classification model is a special classification task model, and the training samples used in this model only have the tag of a positive class, and other samples are classified into another class. It may be understood that the boundary of the positive samples is determined, and the data outside the boundary is classified into another class.

[0029] The positive sample of the ToDo sentence may be a sample with a positive tag, i.e., a sample determined as a meeting ToDo sentence. The number of positive samples of the ToDo sentence is not limited and can be set according to the actual situation. Specifically, the processing device inputs the positive samples of the ToDo sentence into the initial single classification model to perform model training, and obtains a trained single classification model, i.e., a ToDo recognition model.

[0030] In an embodiment of the present disclosure, the step of the processing device inputting the meeting text into the ToDo recognition model to determine the initial ToDo sentence may include the step of the processing device converting the text sentence in the meeting text into a sentence vector and inputting the sentence vector into the ToDo recognition model to determine the initial ToDo sentence. The text sentence is obtained by sentence segmenting or dividing the meeting text, and the number of the text sentences may be multiple.

[0031] The processing device converts each text sentence included in the meeting text into a sentence vector by an embedding layer, inputs each sentence vector into a pre-trained ToDo recognition model to predict the classification result of the ToDo sentence, and determines the sentence having a return value as the initial ToDo sentence. Since the ToDo recognition model is a single classification model, it may be understood that the classification is performed by calculating the radius and center of a sphere, where the sphere is the boundary of the positive samples, and the space inside the sphere represents the distribution space of the positive samples of the ToDo sentence.

[0032] In the above solution, the processing device uses a single classification model to recognize ToDo sentences from the meeting text, thereby reducing the amount of data required to train the deep learning model, improving the model training efficiency, and improving the recognition accuracy.

[0033] Step S103: The processing device inputs the initial ToDo sentence into the tense judgment model to determine the tense result.

[0034] The tense judgment model is a pre-trained model, similar to the above ToDo recognition model, and is used to further make tense judgment on the initial ToDo sentence recognized in the previous step, and the specific deep learning model used is not limited. Tense is a form that characterizes actions, behaviors, and states under various time conditions. The tense results may include past tense, present tense, future tense, etc. The past tense is used to represent the past time, the present tense is used to represent the current time, and the future tense is used to represent the future time.

[0035] Specifically, the processing device can recognize the meeting text through the ToDo recognition model to determine the initial ToDo sentence, and then input the initial ToDo sentence into a pre-trained tense judgment model to further perform tense judgment and determine the tense result. The tense judgment model can be a three-classification model.

[0036] Step S104: The processing device determines a meeting ToDo sentence in the initial ToDo sentence based on the tense result.

[0037] Meeting ToDo statements are different from initial ToDo statements and refer to statements that contain the final ToDo intentions.

[0038] Specifically, the step of determining the meeting ToDo sentence in the initial ToDo sentence based on the tense result may include the step of determining the initial ToDo sentence whose tense result is the future tense as the meeting ToDo sentence. After determining the tense result of each of the above initial ToDo sentences, the processing device sets the initial ToDo sentence whose tense result is the future tense as the meeting ToDo sentence, and deletes the initial ToDo sentence whose tense result is the past tense and the present tense, so as to finally obtain the meeting ToDo sentence.

[0039] In an embodiment of the present disclosure, the processing device recognizes the ToDo intent of the meeting text through a deep learning model, thereby assisting in sorting the meeting ToDo sentences in the minutes, and improving the user's work efficiency. Compared with the conventional machine learning method, the ToDo recognition model uses a single classification model, so that the judgment accuracy of the negative sample can be greatly improved, and the negative sample of the ToDo intent sentence has no boundary, so the judgment accuracy of the model is high, and the user experience can be greatly improved.

[0040] In the processing method of minutes according to the embodiment of the present disclosure, the processing device obtains the meeting text of the meeting audio-video; inputs the meeting text into a ToDo recognition model to determine an initial ToDo sentence; inputs the initial ToDo sentence into a tense judgment model to determine a tense result of the initial ToDo sentence; and determines the meeting ToDo sentence in the initial ToDo sentence according to the tense result. According to the above technical solution, by recognizing the meeting text of the meeting audio-video and then adding tense judgment, it is possible to avoid recognizing an already completed sentence as a meeting ToDo sentence, and greatly improve the accuracy of determining the meeting ToDo sentence, and further improve the user's work efficiency by the meeting ToDo sentence, and improve the user's experience effect.

[0041] In some embodiments, after obtaining the conference text of the conference audio-video, the method may further include: dividing the conference text into sentences to obtain a plurality of text sentences; and filtering the text sentences by preprocessing the text sentences based on a predetermined rule. Optionally, the preprocessing of the text sentences based on the predetermined rule includes removing the text sentences that are missing an intent word, and / or removing the text sentences whose string length is less than a length threshold, and / or removing the text sentences that are missing a noun.

[0042] The text sentences are obtained by sentence segmenting or dividing the meeting text, specifically, dividing the meeting text according to punctuation marks to convert the meeting text into a plurality of text sentences. The predetermined rule may be a rule for processing a plurality of text sentences, but is not specifically limited thereto, for example, the predetermined rule may be to remove dead words and / or to remove duplicated words.

[0043] In an embodiment of the present disclosure, a meeting text can be divided into sentences to obtain multiple text sentences, and then a word division process is performed on each text sentence to obtain a result of the word division process, and the text sentences can be preprocessed based on a predetermined rule and the result of the word division process to filter the text sentences, and the preprocessed text sentences are more likely to be ToDo sentences. The step of preprocessing the text sentences can include a step of searching the result of the word division process of each text sentence, determining whether it contains an intention word and / or a noun, and deleting the text sentences that are missing the intention word and / or a noun. An intention word refers to a pre-organized word that may contain a ToDo intention. For example, if a text sentence contains the word "needs to be completed", it may have a ToDo intention, and "needs to be completed" is an intention word. In an embodiment of the present disclosure, a thesaurus can be set to store multiple intention words and / or nouns for preprocessing.

[0044] And / or, the step of preprocessing the text sentences may include a step of determining the length of each text sentence, comparing the length of each text sentence with a length threshold, and deleting text sentences whose length is less than the length threshold. The length threshold refers to a preset numerical value of the length of a sentence, and if a text sentence is too short, it may not be a sentence, so by setting the length threshold, it is possible to delete text sentences that are too short.

[0045] Optionally, the step of preprocessing the text sentence based on the predetermined rule may include a step of performing sentence pattern matching on the text sentence based on a predetermined sentence pattern, and deleting the text sentence that does not satisfy the predetermined sentence pattern. The predetermined sentence pattern may be understood as a sentence pattern that is likely to include a ToDo intention. The predetermined sentence pattern may include various sentence patterns, for example, the predetermined sentence pattern may be subject + preposition + time word + verb + object, and for the corresponding sentence, take "Mr. Wang, please finish your homework tomorrow" as an example, and this sentence is a ToDo sentence. Each text sentence is matched with the predetermined sentence pattern, and the text sentence that does not satisfy the predetermined sentence pattern is deleted.

[0046] In an embodiment of the present disclosure, after obtaining a meeting text, the text sentences contained in the meeting text can be pre-processed according to a number of predetermined rules. Because the predetermined rules are related to ToDo intent, the pre-processed text sentences are more likely to be ToDo sentences, and further improve the efficiency and accuracy of determining subsequent ToDo sentences.

[0047] 2 is a flowchart of a minutes processing method according to another embodiment of the present disclosure. The method can be executed by a minutes processing device. Here, the device can be realized by software and / or hardware and generally integrated into electronic equipment. As shown in FIG. 2, the method can include the following steps:

[0048] Step S201: The processing device accepts a user's display trigger operation for a target transcript in the minutes display interface, and the minutes display interface displays the meeting audio-video, the meeting text of the meeting audio-video, and the target transcript.

[0049] The minutes display interface refers to an interface for displaying minutes generated in advance. The meeting audio / video and the meeting text are separately displayed in different areas of the minutes display interface. The minutes display interface may be provided with areas such as an audio / video area, a subtitle area, and a minutes display area for respectively displaying contents related to the meeting, such as the meeting audio / video, the meeting text of the meeting audio / video, and the minutes. The display trigger operation refers to an operation that triggers the display of the meeting ToDo sentence in the minutes, and the specific method is not limited. For example, the display trigger operation may be a click operation and / or a hovering operation on the meeting ToDo sentence.

[0050] The recorded sentence refers to a sentence in the minutes, and is displayed in the minutes display area. The recorded sentence includes a meeting ToDo sentence, and the meeting ToDo sentence is a recorded sentence corresponding to the recording type and is a ToDo sentence determined in the above embodiment. The minutes refers to the main content of the meeting generated by processing the meeting audio-video. The minutes may be of various types, and in the embodiment of the present disclosure, the minutes may include at least one of the following types: agenda, agenda, discussion, conclusion, and ToDo, and the meeting ToDo sentence is a sentence belonging to the type of ToDo.

[0051] In an embodiment of the present disclosure, when a user browses content in the minutes display interface, the client terminal can accept the user's display trigger operation for one target record sentence in the minutes.

[0052] Illustratively, FIG. 3 is a schematic diagram of a minutes display interface according to an embodiment of the present disclosure. As shown in FIG. 3, the first area 11 in the minutes display interface 10 displays the minutes, the top of the first area 11 displays the conference video, the second area 12 displays the conference text, and the bottom of the minutes display interface 10 displays the conference audio, which may specifically include the timeline of the conference audio. FIG. 3 shows five types of minutes, including the agenda, agenda, discussion, conclusion, and ToDo, and the ToDo list includes three meeting ToDo sentences. The arrow in FIG. 3 may indicate a display trigger operation for the first meeting ToDo sentence.

[0053] The meeting text in FIG. 3 can be divided into subtitle segments based on the various users participating in the meeting, and the subtitle segments of three users, user 1, user 2, and user 3, are illustrated. In FIG. 3, the top of the minutes display interface 10 further displays the theme of the meeting, "Team Review Meeting," and related contents of the meeting. In the figure, "2019.12.20 10:00 AM" indicates the start time of the meeting, "1h30m30s" indicates that the duration of the meeting is 1 hour 30 minutes 20 seconds, and "16" indicates the number of participants. It should be understood that the minutes display interface 10 in FIG. 3 is only an example, and the position of the content contained therein is also an example, and the specific position and display method can be set according to the actual situation.

[0054] Step S202: The processing device displays the target record sentence and the related sentences of the target record sentence.

[0055] The related sentence is a subtitle sentence that is included in the meeting text and is positionally associated with the target record sentence. The number of related sentences can be set according to the actual situation, for example, the related sentences can be two subtitle sentences located before and after the target record sentence in the meeting text. The number can be two. The subtitle sentence can be one component of the meeting text and is obtained by dividing the meeting text. The meeting text includes multiple subtitle sentences, but the specific number is not limited.

[0056] In an embodiment of the present disclosure, the step of displaying the target transcript and the relevant sentences of the target transcript may include the step of displaying the target transcript and the relevant sentences of the target transcript in a floating window of a minutes display interface. The floating window is displayed in an area of ​​the minutes display interface, and the specific position of the floating window can be set according to actual circumstances, for example, the position of the floating window can be any position that does not block the current target transcript.

[0057] After receiving the display trigger operation for the target record sentence, the processing device can display one floating window to the user, and display the target record sentence and the relevant sentences of the target record sentence in the floating window. In the embodiment of the present disclosure, by displaying the target record sentence and a plurality of sentences before and after the target record sentence, it is possible to avoid the user's difficulty in understanding when the target record sentence is displayed alone, and to make the user's content easier to understand, and to improve the display effect of the record sentence.

[0058] 3, for example, the first underlined meeting ToDo sentence in the ToDo list of the minutes displayed in the first area 11 is the target meeting ToDo sentence. When a display trigger is performed on the target ToDo sentence, the target meeting ToDo sentence and related sentences of the target ToDo sentence are displayed in the floating window 13. The related sentences displayed in the floating window 13 in the figure are one sentence before and one sentence after the target meeting ToDo sentence.

[0059] In some embodiments, the method for processing minutes may further include playing the meeting audio-video based on a relevant duration of the target transcript, and highlighting the relevant subtitle of the target transcript in the meeting text. The relevant subtitle of the target transcript refers to the subtitle corresponding to the target transcript in the subtitle text, and the relevant duration of the target transcript refers to the duration of the original meeting audio corresponding to the relevant subtitle in the meeting audio-video. The relevant duration may include a start time and an end time.

[0060] After receiving the user's display trigger operation for the target transcript, the processing device can play the conference audio-video at a start time in a relevant period of the target transcript, stop playing the conference audio-video at an end time, jump the conference text to a position of the relevant subtitle of the target transcript, and highlight the relevant subtitle of the target transcript in a predetermined manner. Optionally, the predetermined manner can be any feasible display manner that can be distinguished from other parts of the conference text, for example, including but not limited to at least one of highlighting, bolding, and underlining.

[0061] In the above solution, the user can realize the associated interaction of the related content in the meeting audio / video and the meeting text through the interactive trigger of the transcript in the meeting minutes display interface, improving the user's interactive experience. In addition, the three-way associated interaction between the transcript, the meeting audio / video and the meeting text allows the user to intuitively understand the relationship between the three, which is more helpful for the user to accurately understand the content of the meeting.

[0062] In addition, unless inconsistent, each step and feature in the embodiments of the present disclosure can be mutually superimposed and combined with other embodiments of the present disclosure (including, but not limited to, the embodiment shown in FIG. 1 and specific implementation methods of the specific embodiment).

[0063] In the minutes processing solution according to the embodiment of the present disclosure, the processing device accepts a user's display trigger operation for the target transcript in the minutes display interface on which the meeting audio-video, the meeting text of the meeting audio-video, and the target transcript are displayed, and displays the target transcript and the relevant sentences of the target transcript. According to the above technical solution, after determining a more accurate transcript, the processing device can present the transcript and a number of sentences before and after it after accepting a user trigger for one of the transcripts, thereby avoiding the user's difficulty in understanding when the target transcript is displayed alone, making it easier for the user to understand the content, improving the display effect of the transcript, and further improving the user's experience effect.

[0064] 4 is a schematic diagram of a processing device for minutes according to an embodiment of the present disclosure. The device can be realized by software and / or hardware and generally integrated into electronic equipment. As shown in FIG. 4, the device includes: a text acquisition module 401 for acquiring a conference text of the conference audio-video; an initial ToDo module 402 for inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence; a tense decision module 403 for inputting the initial ToDo sentence into a tense decision model to determine a tense outcome of the initial ToDo sentence; and a meeting ToDo module 404 for determining a meeting ToDo sentence in the initial ToDo sentence based on the tense result.

[0065] Optionally, the initial ToDo module 402 may specifically: The text sentences in the meeting text are converted into sentence vectors, and the sentence vectors are input into the ToDo recognition model to determine initial ToDo sentences. The ToDo recognition model is a single classification model.

[0066] Optionally, the apparatus further comprises a model training module, specifically comprising: Based on the positive samples of ToDo sentences, an initial single classification model is used to obtain the ToDo recognition model.

[0067] Optionally, the meeting ToDo module 404 may specifically: The tense result is used to determine an initial ToDo sentence whose tense result is in the future tense as a meeting ToDo sentence.

[0068] Optionally, the apparatus further includes a pre-processing module, which, after obtaining the conference text of the conference audio-video, comprises: Segment the meeting text to obtain a plurality of text sentences; It is used to filter the text sentences by pre-processing them based on predefined rules.

[0069] Optionally, the pre-processing module specifically comprises: Removing text sentences that are missing intent words, and / or Removing text sentences whose string length is less than a length threshold; and / or It is used to remove text sentences that are missing nouns.

[0070] Optionally, the pre-processing module specifically comprises: It is used to perform sentence pattern matching on the text sentences based on a predetermined sentence pattern, and to delete the text sentences that do not satisfy the predetermined sentence pattern.

[0071] The meeting minutes processing device according to the embodiment of the present disclosure obtains the meeting text of the meeting audio / video through the collaboration between the modules, inputs the meeting text into a ToDo recognition model to determine an initial ToDo sentence, inputs the initial ToDo sentence into a tense judgment model to determine the tense result of the initial ToDo sentence, and determines the meeting ToDo sentence in the initial ToDo sentence according to the tense result. According to the above technical solution, by recognizing the meeting text of the meeting audio / video and then adding the tense judgment, it is possible to avoid recognizing an already completed sentence as a meeting ToDo sentence, greatly improving the accuracy of determining the meeting ToDo sentence, and further improving the user's work efficiency by the meeting ToDo sentence, and improving the user's experience effect.

[0072] 5 is a schematic diagram of a meeting minutes processing device according to an embodiment of the present disclosure. The device can be realized by software and / or hardware and generally integrated into electronic equipment. As shown in FIG. 5, the device includes: A display trigger module 501 for receiving a user's display trigger operation for a target transcript in a minutes display interface, the minutes display interface displaying a conference audio / video, a conference text of the conference audio / video, and the target transcript; a display module 502 for displaying the target transcript and related sentences of the target transcript.

[0073] Optionally, the related sentence comprises a subtitle sentence positionally associated with the target transcript sentence in the meeting text, the meeting text comprising a plurality of the subtitle sentences, and the target transcript sentence comprising a target meeting ToDo sentence.

[0074] Optionally, the display module 502 may specifically: The target transcript and related sentences of the target transcript are displayed in a floating window of a minutes display interface.

[0075] Optionally, the apparatus further comprises: The present invention further includes an association interaction module for playing the meeting audio-video based on a relevant period of the target transcript and highlighting a relevant subtitle of the target transcript in the meeting text.

[0076] The minutes processing device according to the embodiment of the present disclosure accepts a user's display trigger operation for a target transcript in a minutes display interface through the collaboration between the modules, and the minutes display interface displays the meeting audio-video, the meeting text of the meeting audio-video, and the target transcript, and displays the target transcript and the relevant sentences of the target transcript. According to the above technical solution, after determining a more accurate transcript, the transcript and a plurality of sentences before and after the target transcript are presented after the user's trigger for one of the transcripts is accepted, which avoids the user's difficulty in understanding when the target transcript is presented alone, and makes the user's understanding of the content easier, improves the display effect of the transcript, and further improves the user's experience effect.

[0077] FIG. 6 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. Referring to FIG. 6 below, a structural schematic diagram of an electronic device 600 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 600 in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals, etc.), and fixed terminals such as digital televisions and desktop computers. The electronic device shown in FIG. 6 is merely an example and should not impose any limitations on the functions and scope of use of the embodiment of the present disclosure.

[0078] 6, the electronic device 600 may include a processing unit (e.g., CPU, graphic processor, etc.) 601 that may perform various appropriate operations and processes according to programs stored in a read only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. The RAM 603 also stores various programs and data necessary for operating the electronic device 600. The processing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0079] Typically, input devices 606 including touch screens, touch pads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc., output devices 607 including liquid crystal displays (LCDs), speakers, vibrating computers, etc., storage devices 608 including magnetic tapes, hard disks, etc., and communication devices 609 may be connected to the I / O interface 605. The communication devices 609 allow the electronic device 600 to communicate and exchange data with other devices, wirelessly or via wires. Although FIG. 6 illustrates the electronic device 600 having various devices, it should be understood that it is not required to implement or include all of the devices illustrated. Alternatively, more or fewer devices may be implemented or included.

[0080] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowcharts may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 609, or may be installed from the storage device 608 or ROM 602. When the computer program is executed by the processing device 710, the above-mentioned functions limited to the processing method for minutes according to the embodiment of the present disclosure are performed.

[0081] It should be noted that the computer readable medium referred to in this disclosure may be a computer readable signal medium or a computer readable storage medium, or any combination of the two. The computer readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, apparatus or device, or any combination of the above. More specific examples of computer readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, a computer readable storage medium may be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, apparatus or device. In this disclosure, a computer readable signal medium may include a data signal propagated in baseband or as part of a carrier wave in which computer readable program code is stored. Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may transmit, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium may be transmitted using any suitable medium, including, but not limited to, electrical wires, fiber optic cables, RF (radio frequency), or any suitable combination of the above.

[0082] In some embodiments, the client terminals and servers may communicate utilizing any network protocol now known or later developed, such as HTTP (HyperText Transfer Protocol), and may interconnect with any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as networks now known or later developed.

[0083] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated in the electronic device.

[0084] The computer-readable medium stores one or more programs, which, when executed by the electronic device, cause the electronic device to execute the steps of: acquiring a conference text of a conference audio / video, inputting the conference text into a ToDo recognition model to determine an initial ToDo sentence, inputting the initial ToDo sentence into a tense judgment model to determine a tense result of the initial ToDo sentence, and determining a conference ToDo sentence in the initial ToDo sentence based on the tense result.

[0085] Alternatively, the computer-readable medium stores one or more programs, which, when executed by the electronic device, cause the electronic device to execute a step of receiving a user's display trigger operation for a target transcript in a minutes display interface, in which the minutes display interface displays a conference audio-video, a conference text of the conference audio-video, and the target transcript, and a step of displaying the target transcript and a related sentence of the target transcript.

[0086] Also, the computer program code for carrying out the operations of the present disclosure can be written in one or more programming languages ​​or combinations thereof. Such programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and the like, as well as traditional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. When a remote computer is involved, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., connected via an Internet connection by an Internet service provider).

[0087] The flowcharts and block diagrams in the drawings illustrate possible system architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or part of code that includes one or more executable instructions for implementing a certain logical function. It should be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order from the order shown. For example, two blocks shown in succession may actually be executed substantially in parallel or may be executed according to the reverse order, depending on the related functions. It should be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams may be implemented by a dedicated hardware-based system for executing a certain function or operation, or by a combination of dedicated hardware and computer instructions.

[0088] The units mentioned in the embodiments of the present disclosure may be realized in software or hardware, and the names of the units, if any, are not intended to be limitations on the units themselves.

[0089] The functions described herein may be performed, at least in part, by one or more hardware logic components. For example, but not limited to, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.

[0090] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in combination with an instruction execution system, apparatus or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection by one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a compact disk read-only memory (CD-ROM), optical storage, magnetic storage, or any suitable combination of the above.

[0091] According to one or more embodiments of the present disclosure, the present disclosure provides a method for producing a cellular membrane comprising: obtaining a conference text of the conference audio-video; inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence; inputting the initial ToDo sentence into a tense judgment model to determine a tense outcome of the initial ToDo sentence; determining a meeting ToDo sentence in the initial ToDo sentence based on the tense result.

[0092] According to one or more embodiments of the present disclosure, in the method for processing minutes of the present disclosure, the step of inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence includes: The method includes converting text sentences in the meeting text into sentence vectors and inputting the sentence vectors into the ToDo recognition model to determine initial ToDo sentences, where the ToDo recognition model is a single classification model.

[0093] According to one or more embodiments of the present disclosure, in the method for processing minutes of the present disclosure, the ToDo recognition model is generated in the following manner: The ToDo recognition model is obtained by training an initial single classification model based on the positive samples of ToDo sentences.

[0094] According to one or more embodiments of the present disclosure, in the method for processing minutes of the present disclosure, the step of determining a meeting ToDo sentence in the initial ToDo sentence based on the result of the tense includes: The method includes determining an initial ToDo sentence whose tense result is in the future tense as a meeting ToDo sentence.

[0095] According to one or more embodiments of the present disclosure, in a method for processing minutes of a meeting according to the present disclosure, after a step of obtaining a meeting text of a meeting audio-video, Sentence-segmenting the meeting text to obtain a plurality of text sentences; and filtering the text sentences by pre-processing the text sentences based on predefined rules.

[0096] According to one or more embodiments of the present disclosure, in the method for processing minutes of a meeting according to the present disclosure, the step of preprocessing the text sentences based on a predetermined rule includes: Removing text sentences that are missing intent words; and / or Removing text sentences whose string length is less than a length threshold; and / or This includes removing text sentences that are missing nouns.

[0097] According to one or more embodiments of the present disclosure, in the method for processing minutes of a meeting according to the present disclosure, the step of preprocessing the text sentences based on a predetermined rule includes: The method includes a step of performing sentence pattern matching on the text sentences based on a predetermined sentence pattern, and deleting text sentences that do not satisfy the predetermined sentence pattern.

[0098] According to one or more embodiments of the present disclosure, the present disclosure provides a method for producing a cellular membrane comprising: receiving a user's display trigger operation for a target transcript in a minutes display interface, in which the minutes display interface displays a conference audio-video, a conference text of the conference audio-video, and the target transcript; displaying the target transcript and related sentences of the target transcript.

[0099] According to one or more embodiments of the present disclosure, in the method for processing minutes of the present disclosure, the related sentences include subtitle sentences that are positionally associated with the target record sentence in the meeting text, the meeting text includes a plurality of the subtitle sentences, and the target record sentence includes a target meeting ToDo sentence.

[0100] According to one or more embodiments of the present disclosure, in the method for processing minutes of the present disclosure, the step of displaying the target transcript and the related sentences of the target transcript includes: The method includes displaying the target transcript and related sentences of the target transcript in a floating window of a minutes display interface.

[0101] According to one or more embodiments of the present disclosure, in a method for processing minutes according to the present disclosure, The method further includes playing the meeting audio-video based on a relevant period of the target transcript, and highlighting a relevant subtitle of the target transcript in the meeting text.

[0102] According to one or more embodiments of the present disclosure, the present disclosure provides a method for producing a cellular membrane comprising: A text acquisition module for acquiring the conference text of the conference audio-video; an initial ToDo module for inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence; a tense judgment module for inputting the initial ToDo sentence into a tense judgment model to determine a tense outcome of the initial ToDo sentence; A meeting ToDo module for determining a meeting ToDo sentence in the initial ToDo sentence based on the tense result is provided.

[0103] According to one or more embodiments of the present disclosure, in the processing device for minutes of a meeting according to the present disclosure, the initial ToDo module specifically: The text sentences in the meeting text are converted into sentence vectors, and the sentence vectors are input into the ToDo recognition model and used to determine initial ToDo sentences, where the ToDo recognition model is a single classification model.

[0104] According to one or more embodiments of the present disclosure, in the processing device for minutes of a meeting according to the present disclosure, the device further comprises: A model training module is included for obtaining the ToDo recognition model by training an initial single classification model based on positive samples of ToDo sentences.

[0105] According to one or more embodiments of the present disclosure, in the meeting minutes processing device according to the present disclosure, the meeting ToDo module specifically: The tense result is used to determine an initial ToDo sentence whose tense result is in the future tense as a meeting ToDo sentence.

[0106] According to one or more embodiments of the present disclosure, in the processing device of the minutes of a meeting according to the present disclosure, the device further includes a pre-processing module, the pre-processing module comprising: After getting the conference audio-video conference text, Segment the meeting text to obtain a plurality of text sentences; It is used to filter the text sentences by pre-processing them based on predefined rules.

[0107] According to one or more embodiments of the present disclosure, in the processing device for minutes of a meeting according to the present disclosure, the pre-processing module specifically includes: Removing text sentences that are missing intent words, and / or Removing text sentences whose string length is less than a length threshold; and / or It is used to remove text sentences that are missing nouns.

[0108] According to one or more embodiments of the present disclosure, in the processing device for minutes of a meeting according to the present disclosure, the pre-processing module specifically includes: It is used to perform sentence pattern matching on the text sentences based on a predetermined sentence pattern, and to delete the text sentences that do not satisfy the predetermined sentence pattern.

[0109] According to one or more embodiments of the present disclosure, the present disclosure provides a method for producing a cellular membrane comprising: A display trigger module for receiving a user's display trigger operation for a target transcript in a minutes display interface, the minutes display interface displaying a conference audio / video, a conference text of the conference audio / video, and the target transcript; a display module for displaying the target transcript and related sentences of the target transcript.

[0110] According to one or more embodiments of the present disclosure, in a minutes processing device according to the present disclosure, the related sentences include subtitle sentences that are positionally associated with the target record sentence in the meeting text, the meeting text includes a plurality of the subtitle sentences, and the target record sentence includes a target meeting ToDo sentence.

[0111] According to one or more embodiments of the present disclosure, in the processing device for minutes of a meeting according to the present disclosure, the display module specifically includes: The target transcript and related sentences of the target transcript are displayed in a floating window of a minutes display interface.

[0112] According to one or more embodiments of the present disclosure, in the processing device for minutes of a meeting according to the present disclosure, the device further comprises: An associated interaction module is included for playing the meeting audio-video based on a relevant time period of the target transcript and for highlighting a relevant subtitle of the target transcript in the meeting text.

[0113] According to one or more embodiments of the present disclosure, the present disclosure provides a method for producing a cellular membrane comprising: A processor; a memory for storing instructions executable by said processor; The processor provides an electronic device used to realize any one of the minutes processing methods disclosed herein by reading and executing the executable instructions from the memory.

[0114] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having stored thereon a computer program for executing any one of the minutes processing methods according to the present disclosure.

[0115] The above description merely describes the preferred embodiments of the present disclosure and the technical principles utilized. It should be understood by those skilled in the art that the scope of the disclosure of the present disclosure is not limited to the technical solution formed by the specific combination of the above technical features, but also includes other technical solutions formed by any combination of the above technical features or equivalent features without departing from the concept of the above disclosure, such as replacing the above features with technical features having similar functions (but not limited to) disclosed in the present disclosure.

[0116] Also, although operations are described in a particular order, this should not be construed as requiring that these operations be performed in the particular order or sequence shown. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although the above description includes some specific implementation details, this should not be construed as limiting the scope of the disclosure. Certain features described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments alone or in any suitable subcombination.

[0117] Although the present subject matter has been described in language specific to structural features and / or logical operations of a method, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features and operations described above, but rather, the specific features and operations described above are merely example forms for implementing the claims.

Claims

1. A method for processing minutes of a meeting executed by an electronic device, comprising the steps of: obtaining a conference text of the conference audio / video; inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence, the ToDo recognition model being used to recognize a ToDo intent sentence from the meeting text; inputting the initial ToDo sentence into a tense judgment model to determine a tense outcome of the initial ToDo sentence; determining a meeting ToDo sentence in the initial ToDo sentence based on the tense result, the step including determining an initial ToDo sentence whose tense result is a future tense as a meeting ToDo sentence; The method according to claim 1, further comprising:

2. The step of inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence includes: converting a text sentence in the meeting text into a sentence vector, and inputting the sentence vector into the ToDo recognition model to determine an initial ToDo sentence, wherein the ToDo recognition model is a single classification model; 2. The method of claim 1 .

3. The ToDo recognition model is generated in a manner to obtain the ToDo recognition model by training an initial single classification model based on positive samples of ToDo sentences; 2. The method of claim 1 .

4. After the step of obtaining a conference text of the conference audio-video, Sentence-segmenting the meeting text to obtain a plurality of text sentences; - filtering the text sentences by pre-processing the text sentences based on predefined rules; 2. The method of claim 1, further comprising:

5. The step of preprocessing the text sentences based on predetermined rules comprises: Removing text sentences that are missing intent words; and / or Removing text sentences whose string length is less than a length threshold; and / or removing text sentences that are missing nouns; 5. The method of claim 4.

6. The step of preprocessing the text sentences based on predetermined rules comprises: performing a sentence pattern matching on the text sentence based on a predetermined sentence pattern, and deleting the text sentence that does not satisfy the predetermined sentence pattern; 5. The method of claim 4.

7. A method for processing minutes of a meeting executed by an electronic device, comprising the steps of: receiving a user's display trigger operation for a target transcript in a minutes display interface, in which a conference audio / video, a conference text of the conference audio / video, and the target transcript are displayed in the minutes display interface; displaying the target sentence and related sentences of the target sentence; Including, The method of claim 1 , wherein the target record includes the meeting ToDo statement determined based on the processing method of claim 1 .

8. The related sentence includes a subtitle sentence positionally associated with the target record sentence in the meeting text, the meeting text includes a plurality of the subtitle sentences, and the target record sentence includes a target meeting ToDo sentence.

8. The method of claim 7.

9. The step of displaying the target sentence and related sentences of the target sentence includes: displaying the target transcript and related sentences of the target transcript in a floating window of a minutes display interface; 8. The method of claim 7.

10. The method further includes playing the conference audio-video based on a relevant period of the target transcript, and highlighting a relevant subtitle of the target transcript in the conference text.

8. The method of claim 7.

11. A minutes processing device, comprising: a text acquisition module for acquiring a conference text of the conference audio / video; an initial ToDo module for inputting the meeting text into a ToDo recognition model to determine an initial ToDo sentence, the ToDo recognition model being used to recognize a ToDo intent sentence from the meeting text; a tense decision module for inputting the initial ToDo sentence into a tense decision model to determine a tense outcome of the initial ToDo sentence; a meeting ToDo module for determining a meeting ToDo sentence in the initial ToDo sentence based on the tense result, the meeting ToDo module determining an initial ToDo sentence whose tense result is a future tense as a meeting ToDo sentence; An apparatus comprising:

12. A minutes processing device, comprising: a display trigger module for receiving a user's display trigger operation for a target transcript in a minutes display interface, the minutes display interface displaying a conference audio / video, a conference text of the conference audio / video, and the target transcript; a display module for displaying the target sentence and related sentences of the target sentence; Including, The apparatus according to claim 1 , wherein the target record sentence comprises the meeting ToDo sentence determined based on the processing method of claim 1 .

13. A processor; a memory for storing instructions executable by said processor; The processor reads the executable instructions from the memory and executes them to realize the method for processing minutes of a meeting according to any one of claims 1 to 10. An electronic device comprising:

14. 1. A computer-readable storage medium, comprising: A computer program is stored in the computer, the computer program being used to execute the method for processing minutes of a meeting according to any one of claims 1 to 10. A computer-readable storage medium comprising:

Citation Information

Patent Citations

  • Conference summary processing method and device, server and readable storage medium

    CN110533382A

  • Sentence tense recognition method and equipment based on dependency syntax and readable storage medium

    CN112069800A

  • Conference information management system, and program and storage medium for attaining function of the system

    JP2008242798A

  • To-do automatic extraction system

    JP2013250598A

  • Management of commitments and requests extracted from communications and content

    JP2018522325A