Meeting minutes document generation method and device, electronic equipment and storage medium
By splitting and analyzing video conference recordings, key information is automatically extracted and meeting minutes are generated, solving the problem of complex manual recording and achieving efficient automated minute generation.
Patent Information
- Application Number
- CN202210936369.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-03
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-08-03
AI Technical Summary
Current methods for generating video conference minutes rely on manual recording and post-processing, resulting in data redundancy, complex processing, and wasted effort.
By splitting the meeting recording file, extracting sub-video keywords and meeting topics, determining the playback speed, and simultaneously displaying keywords and related words, the system can receive user input to generate meeting minutes files.
It improves the efficiency of meeting minutes generation, reduces manpower burden, simplifies processing complexity, and automates the extraction of key information and document generation.
Smart Images

Figure CN115329129B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a conference summary file generation method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the popularization of Internet technology and cloud technology, video conference systems are widely used in remote work communication, improving work communication efficiency. At present, the conference summary of the video conference is usually recorded by the participants through manual photographing, screen recording, audio recording and the like, and then manually arranged in the later period. However, the data obtained by the above-mentioned method is relatively large and the corresponding relationship is not clear, which is troublesome and time-consuming in the later period. SUMMARY
[0003] In order to solve the above-mentioned problems in the prior art, the present application provides a conference summary file generation method, device, electronic equipment and storage medium, which can automatically analyze the conference screen recording file, extract the key information and display it to the user, and then assist the user to output the corresponding conference summary file, thereby reducing the labor burden and improving the efficiency.
[0004] In a first aspect, the embodiments of the present application provide a conference summary file generation method, which comprises:
[0005] Splitting the conference screen recording file to obtain at least one sub-video;
[0006] Determining the keywords and conference fields of each sub-video in the at least one sub-video, and determining the association words of each sub-video according to the keywords and conference fields of each sub-video;
[0007] Determining the playback speed of each sub-video according to the keywords and conference fields of each sub-video;
[0008] Playing each sub-video to the user in sequence according to the playback rule of each sub-video, and synchronously displaying the keywords and association words of each sub-video to the user;
[0009] Receiving the input information of the user, and generating the conference summary file according to the input information.
[0010] In a second aspect, the embodiments of the present application provide a conference summary file generation device, which comprises:
[0011] The analysis module is configured to split the conference screen recording file to obtain at least one sub-video, determine the keywords and conference fields of each sub-video in the at least one sub-video, and determine the association words of each sub-video according to the keywords and conference fields of each sub-video, and determine the playback speed of each sub-video according to the keywords and conference fields of each sub-video;
[0012] The display module is configured to sequentially play each sub-video to the user according to the playing rule of each sub-video, and synchronously display the keywords and the associated words of each sub-video to the user.
[0013] The generation module is configured to receive input information of the user, and generate a conference minutes file according to the input information.
[0014] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory is configured to store a computer program. The processor is configured to execute the computer program stored in the memory, so that the electronic device executes the method of the first aspect.
[0015] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program. The computer program causes a computer to execute the method of the first aspect.
[0016] In a fifth aspect, a computer program product is provided, which includes a non-transitory computer readable storage medium storing a computer program. The computer is operable to cause a computer to execute the method of the first aspect.
[0017] The implementation of the embodiments of the present application has the following beneficial effects:
[0018] In the embodiments of the present application, first, the conference recording file is split to divide the longer conference video into multiple small sub-videos for analysis, which improves the efficiency and reduces the complexity of processing. Then, the keywords and the conference field of each sub-video are determined, and the extracted keywords are associated and derived according to the conference field to obtain a series of associated words. At the same time, the playing speed of each sub-video is determined according to the keywords and the conference field of each sub-video. Finally, each sub-video is sequentially played to the user according to the playing rule of each sub-video, and the keywords and the associated words of each sub-video are synchronously displayed to the user, and then the input information of the user is received, and a conference minutes file is generated according to the input information. Thus, the conference recording file can be automatically analyzed, and the key information is extracted and displayed to the user, and then the user can output the corresponding conference minutes file, which reduces the labor burden and improves the efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0020] Figure 1A schematic diagram of a hardware structure of a conference minutes file generation device provided by an embodiment of the present application;
[0021] Figure 2 A schematic diagram of a conference minutes file generation method provided by an embodiment of the present application;
[0022] Figure 3 A schematic diagram of a method for determining keywords and conference fields of each sub-video in at least one sub-video provided by an embodiment of the present application;
[0023] Figure 4 A schematic diagram of a topological relationship graph established according to the keywords of each sub-video provided by an embodiment of the present application;
[0024] Figure 5 A schematic diagram of a topological relationship graph provided by an embodiment of the present application;
[0025] Figure 6 A schematic diagram of a method for determining a playing speed of each sub-video according to the keywords and conference fields of each sub-video provided by an embodiment of the present application;
[0026] Figure 7 A schematic diagram of a display interface provided by an embodiment of the present application;
[0027] Figure 8 A functional module composition block diagram of a conference minutes file generation device provided by an embodiment of the present application;
[0028] Figure 9 A schematic diagram of a structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative work fall within the scope of protection of the present application.
[0030] The terms "first", "second", "third", and "fourth" and the like in the description and in the claims of the present application are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances such that the descriptive terms used in the specification are, under some circumstances, interchangeable with the precursor or successor terms used in the claims to describe an element. Furthermore, the use of the terms in the detailed description is merely to facilitate reading of the specification and they are not to be construed to limit the scope of the application. The terms "comprises", "comprising", "includes", "including", "has", "having" and their variants are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of steps or elements is not necessarily limited to the listed steps or elements, but can include additional or other steps or elements not expressly listed or inherent to such process, method, article, or apparatus.
[0031] Reference herein to "an embodiment" means that a particular feature, structure, result, or characteristic described in connection with an embodiment can be included in at least one embodiment of the application. The appearances of the phrase "in an embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is expressly understood that any of the embodiments described herein can be combined with any of the other embodiments unless specifically noted otherwise.
[0032] First, refer to Figure 1 , Figure 1 A hardware structure schematic diagram of a conference minutes file generation device provided by an embodiment of the application. The conference minutes file generation device 100 includes at least one processor 101, a communication line 102, a memory 103, and at least one communication interface 104.
[0033] In the embodiment, the processor 101 can be a general central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of programs of the application.
[0034] The communication line 102 can include a path for transmitting information between the above-mentioned components.
[0035] The communication interface 104 can be any kind of device (such as an antenna, etc.) such as a transceiver for communicating with other devices or communication networks, for example, Ethernet, RAN, wireless local area networks (WLAN), etc.
[0036] The memory 103 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0037] In this embodiment, the memory 103 can exist independently and be connected to the processor 101 via the communication line 102. Alternatively, the memory 103 can be integrated with the processor 101. The memory 103 provided in this embodiment is typically non-volatile. The memory 103 stores computer execution instructions for implementing the scheme of this application, and its execution is controlled by the processor 101. The processor 101 executes the computer execution instructions stored in the memory 103 to implement the method provided in the following embodiments of this application.
[0038] In an optional implementation, the computer execution instructions may also be referred to as application code, and this application does not specifically limit this terminology.
[0039] In an optional implementation, processor 101 may include one or more CPUs, for example... Figure 1 CPU0 and CPU1 in the CPU.
[0040] In an optional implementation, the meeting minutes generation device 100 may include multiple processors, such as... Figure 1 Processors 101 and 107 are shown in the diagram. Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0041] In an optional embodiment, if the conference summary document generation apparatus 100 is a server, for example, it can be a standalone server, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and big data and artificial intelligence platform. The conference summary document generation apparatus 100 can further include an output device 105 and an input device 106. The output device 105 communicates with the processor 101 and can display information in various ways. For example, the output device 105 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 106 communicates with the processor 101 and can receive user input in various ways. For example, the input device 106 can be a mouse, a keyboard, a touch screen device, or a sensor device, etc.
[0042] The conference summary document generation apparatus 100 described above can be a general-purpose device or a special-purpose device. The embodiments of the present application do not limit the type of the conference summary document generation apparatus 100.
[0043] In the following, the conference summary document generation method disclosed by the present application will be described:
[0044] Referring to Figure 2 , Figure 2 A flowchart of a conference summary document generation method provided by an embodiment of the present application is shown. The conference summary document generation method includes the following steps:
[0045] 201: Split the conference recording file to obtain at least one sub-video.
[0046] In the present embodiment, the conference recording file can be split according to different speakers, or can be split according to different content themes, or other splitting methods for video files in the art can also be applicable to the present embodiment, which is not limited herein.
[0047] 202: Determine the keywords and conference fields of each sub-video in the at least one sub-video, and determine the association words of each sub-video according to the keywords and conference fields of each sub-video.
[0048] In the present embodiment, a method for determining the keywords and conference fields of each sub-video in the at least one sub-video is provided, specifically, as shown inFigure 3 As shown in Figure 3 , the method includes:
[0049] 301: Determine the content display area in each sub-video, and perform text recognition on the content in the content display area to obtain the first text information.
[0050] In this embodiment, the OCR (Optical Character Recognition) method is adopted to perform text recognition on the content in the content display area of each sub-video to obtain the first text information.
[0051] 302: Perform audio recognition on each sub-video to obtain the second text information.
[0052] In this embodiment, any technology that can achieve speech-to-text conversion in the art can be applied to this embodiment, and no limitation is made here.
[0053] 303: Determine at least one first keyword in the first text information and at least one second keyword in the second text information.
[0054] Specifically, taking the first text information as an example, it can be segmented to obtain at least one segment. Exemplarily, the N-gram segmentation method with an order of 2 or 3 can be used to segment the first sentence. Specifically, the N-gram segmentation method is a method of splitting a sentence into a sequence of segments each consisting of N characters, and each segment is called an N-gram. Exemplarily, when N = 2, the N-gram segmentation method can be called bi-gram (binary gram). When using bi-gram for segmentation, two adjacent characters of the text to be segmented are output in sequence. For example, for the text "Semi-automatic generation of meeting minutes documents", if bi-gram is used for segmentation, the segments "meeting", "discussion minutes", "minutes", "document", "semi-automatic", "automatic generation" can be obtained.
[0055] Meanwhile, in this embodiment, a keyword refers to a word that can reflect the key information in the text, and these words often appear in the form of nouns in the text. Based on this, after obtaining at least one segment, the grammar of the first text information can be analyzed to obtain the grammar features, and the词性 information of each segment in the at least one segment can be determined according to the grammar features. Then, the segments with the词性 information of nouns are screened out as candidate words.
[0056] Furthermore, by obtaining the Term Frequency–Inverse Document Frequency (TF-IDF) score of each candidate word, the importance of the candidate word is determined. Then, candidate words with scores lower than a preset threshold, i.e., common words, are eliminated, and the remaining candidate words are used as the first keywords corresponding to the first text information.
[0057] In this embodiment, the method for determining at least one second keyword in the second text information is similar to the method for determining at least one first keyword in the first text information, and will not be described again here.
[0058] 304: Deduplicat at least one primary keyword and at least one secondary keyword to obtain the keywords for each sub-video.
[0059] In this embodiment, since at least one first keyword and at least one second keyword are extracted from first text information and second text information, respectively, and the first text information and second text information are respectively derived from video footage and audio in the same sub-video, it is inevitable that there will be duplicate words among the extracted keywords. Therefore, before processing, it is necessary to remove these duplicate keywords to ensure that each keyword is different after set.
[0060] 305: Build a topological relationship diagram based on the keywords of each sub-video.
[0061] In this embodiment, such as Figure 4 As shown, firstly, the keywords for each sub-video can be randomly selected c! times, resulting in c! keyword groups. Specifically, c+1 represents the number of keywords in each sub-video. That is, if a sub-video has n keywords, then the keywords for that sub-video will be randomly selected (n-1)! times. For example, if a sub-video has 6 keywords, then the keywords for that sub-video will be randomly selected 5! times, or 15 times, resulting in 15 keyword groups.
[0062] In this embodiment, each random selection involves choosing any two different keywords from the keywords of each sub-video, and the keywords selected in any two random selections are not exactly the same. In short, each random selection will choose two different keywords, hereinafter referred to as the third keyword and the fourth keyword, and the final c! keyword groups selected will not be completely identical.
[0063] Then, the relevance coefficients between the third and fourth keywords in each keyword group can be determined, resulting in c! relevance coefficients. Specifically, the relevance coefficients between the third and fourth keywords can be determined by calculating the semantic similarity between them.
[0064] Finally, c+1 keywords of each sub-video can be taken as c+1 nodes respectively, and each correlation coefficient of the c! correlation coefficients can be taken as an edge between two nodes corresponding to two keywords in the keyword group corresponding to each correlation coefficient, to obtain the topological relationship graph.
[0065] For example, if the keywords A, B, C and D are present, 3! times of random extraction are required to obtain the keyword groups [A, B], [A, C], [A, D], [B, C], [B, D] and [C, D], and the correlation coefficients corresponding to each keyword group are calculated as e[A, B] = f, e[A, C] = g, e[A, D] = h, e[B, C] = i, e[B, D] = j and e[C, D] = k. Based on this, the keywords A, B, C and D are taken as nodes respectively, f is taken as an edge between node A and node B, g is taken as an edge between node A and node C, h is taken as an edge between node A and node D, i is taken as an edge between node B and node C, j is taken as an edge between node B and node D, and k is taken as an edge between node C and node D, to obtain the topological relationship graph as shown in Figure 5 .
[0066] 306: Matching in the preset knowledge network according to the topological relationship graph to determine the conference field of each sub-video.
[0067] In the embodiment, the preset knowledge network is formed based on the knowledge in the field related to the enterprise holding the conference after being sorted out, and the essence is a network composed of various field nodes, various keyword nodes and the relationship between the nodes. Based on this, by matching the topological relationship graph formed above in the knowledge network, the most similar part to the topological relationship graph is found out, and then the field node associated with the part is traced, to obtain the conference field of each sub-video. In addition, in the possible embodiment, there can be multiple corresponding field nodes, at this time, the field information corresponding to the field node with the highest correlation coefficient can be taken as the conference field of each sub-video.
[0068] 203: Determining the playback speed of each sub-video according to the keyword and the conference field of each sub-video.
[0069] In the embodiment, a method for determining the playback speed of each sub-video according to the keyword and the conference field of each sub-video is provided, as shown in Figure 6 , the method comprises:
[0070] 601: Determining the first importance of each sub-video according to the keyword and the conference field of each sub-video.
[0071] In the embodiment, the relevance between each keyword and the conference field can be calculated, and then the average of all the relevance is taken as the first importance of each sub-video.
[0072] 602: Determine the main speaker of each sub-video, and determine the second importance of each sub-video according to the main speaker.
[0073] In the embodiment, the second importance of each sub-video can be determined according to the position of the main speaker of each sub-video and the relevance between the main speaker and the conference field of each sub-video. Specifically, an importance score can be set for each position of the enterprise, and then the corresponding importance score is obtained after the position of the main speaker is determined, and the importance score and the relevance between the main speaker and the conference field of each sub-video are normalized to determine the weight of the importance score and the relevance. Then, the importance score and the relevance are weighted and summed according to the weight to obtain the second importance of each sub-video.
[0074] 603: Determine the number of repetitions of each sub-video, and determine the third importance of each sub-video according to the number of repetitions and the position of each sub-video in the repeated video.
[0075] In the embodiment, for the video with repeated content, the sub-video that first appears in it can be marked as important, and based on this, the third importance can be represented by formula ①:
[0076]
[0077] Wherein, x3 represents the third importance, a represents the number of repetitions, and b represents the position sequence number of each sub-video in the repeated video.
[0078] 604: Determine the fourth importance of each sub-video according to the first importance, the second importance and the third importance.
[0079] In the embodiment, the fourth importance can be represented by formula ②:
[0080]
[0081] Wherein, x1 represents the first importance, x2 represents the second importance, x3 represents the third importance, and x4 represents the fourth importance.
[0082] 605: Determine the playback speed of each sub-video according to the fourth importance.
[0083] In the embodiment, the mapping relationship between the importance interval and the playback speed can be set in advance, and then after the fourth importance of each sub-video is obtained, the playback speed corresponding to the importance interval in which the fourth importance is located can be determined as the playback speed of the sub-video according to the fourth importance.
[0084] 204: Play each sub-video sequentially to the user according to the playback rules of each sub-video, and simultaneously display the keywords and related words of each sub-video to the user.
[0085] In this embodiment, the interface displayed to the user can be divided into three areas: one for displaying the sub-video, another for displaying the keywords and related terms corresponding to the sub-video, and the third for displaying the user's input information. For example, as shown... Figure 7 The image shows a display interface. Area 1 displays sub-videos, area 2 displays keywords and related terms for the sub-videos, and area 3 displays user input information.
[0086] 205: Receive user input information and generate meeting minutes based on the input information.
[0087] In this embodiment, before step 205, a corresponding minutes template can be obtained based on the meeting domain of each sub-video, and the minutes template can be displayed to the user in the first area. Specifically, this first area is the area for displaying user input content, and it follows the previous method. Figure 7 In the example, the first region is region 3.
[0088] Therefore, in this embodiment, only the first information input by the user in the first area can be received, and then a meeting minutes file can be generated based on the first information and the minutes template. Specifically, the user can see the matched meeting minutes template in the first area, and can manually select the appropriate template if they feel it is incorrect. After selecting the template, the user can fill in information in the corresponding position in the first area according to the displayed template, combined with sub-videos, keywords, and related words. Thus, the system can fill in the user's input information into the corresponding position in the template based on the currently displayed template and the position of the information filled in by the user, thereby generating the meeting minutes file.
[0089] In summary, the conference summary file generation method provided by the application first splits the conference recording file to divide the longer conference video into multiple small sub-videos for analysis, thereby improving efficiency and reducing processing complexity. Then, the keywords and conference fields of each sub-video are determined, and the extracted keywords are associated and derived according to the conference fields to obtain a series of associated words. Meanwhile, the playback speed of each sub-video is determined according to the keywords and conference fields of each sub-video. Finally, each sub-video is played to the user in sequence according to the playback rule of each sub-video, and the keywords and associated words of each sub-video are synchronously displayed to the user, and then the input information of the user is received to generate a conference summary file according to the input information. In this way, the conference recording file can be automatically analyzed, and the key information is extracted and displayed to the user, thereby assisting the user to output the corresponding conference summary file, reducing the labor burden and improving the efficiency.
[0090] Reference Figure 8 , Figure 8 A functional module composition block diagram of a conference summary file generation device provided by an embodiment of the application is provided. As shown in Figure 8 , the conference summary file generation device 800 comprises:
[0091] An analysis module 801 is configured to split the conference recording file to obtain at least one sub-video, determine the keywords and conference fields of each sub-video in the at least one sub-video, determine the associated words of each sub-video according to the keywords and conference fields of each sub-video, and determine the playback speed of each sub-video according to the keywords and conference fields of each sub-video.
[0092] A display module 802 is configured to play each sub-video to the user in sequence according to the playback rule of each sub-video, and synchronously display the keywords and associated words of each sub-video to the user.
[0093] A generation module 803 is configured to receive the input information of the user, and generate a conference summary file according to the input information.
[0094] In the embodiment of the application, in terms of determining the playback speed of each sub-video according to the keywords and conference fields of each sub-video, the analysis module 801 is specifically configured to:
[0095] determine the first importance of each sub-video according to the keywords and conference fields of each sub-video;
[0096] determine the main speaker of each sub-video, and determine the second importance of each sub-video according to the main speaker;
[0097] determine the number of repetitions of each sub-video, and determine the third importance of each sub-video according to the number of repetitions and the position of each sub-video in the repeated video;
[0098] According to the first importance, the second importance and the third importance, a fourth importance of each sub-video is determined;
[0099] According to the fourth importance, a playback speed of each sub-video is determined.
[0100] In an embodiment of the present application, the third importance can be represented by formula (3):
[0101]
[0102] wherein x3 represents the third importance, a represents the number of repetitions, and b represents the position sequence number of each sub-video in the repeated video.
[0103] In an embodiment of the present application, the fourth importance can be represented by formula (4):
[0104]
[0105] wherein x1 represents the first importance, x2 represents the second importance, x3 represents the third importance, and x4 represents the fourth importance.
[0106] In an embodiment of the present application, in determining the keywords and the conference field of each sub-video in the at least one sub-video, the analysis module 801 is specifically configured to:
[0107] determine a content display area in each sub-video, and perform text recognition on the content in the content display area to obtain first text information;
[0108] perform audio recognition on each sub-video to obtain second text information;
[0109] determine at least one first keyword in the first text information and at least one second keyword in the second text information;
[0110] perform a de-duplication set on the at least one first keyword and the at least one second keyword to obtain the keywords of each sub-video;
[0111] establish a topological relationship graph according to the keywords of each sub-video;
[0112] match in a preset knowledge network according to the topological relationship graph to determine the conference field of each sub-video.
[0113] In an embodiment of the present application, in establishing the topological relationship graph according to the keywords of each sub-video, the analysis module 801 is specifically configured to:
[0114] For each sub-video, keywords are randomly selected c! times to obtain c! keyword groups, where c+1 is the number of keywords in each sub-video. Each time, any two different keywords are randomly selected from the keywords of each sub-video, and the keywords selected in any two random selections are not exactly the same. Each keyword group in c! keyword groups includes a third keyword and a fourth keyword.
[0115] Determine the relevance coefficient between the third and fourth keywords in each keyword group to obtain c! relevance coefficients;
[0116] Each sub-video's c+1 keywords are treated as c+1 nodes, and each of the c! relevance coefficients is used as an edge between the two nodes corresponding to the two keywords in the keyword group corresponding to each relevance coefficient, thus obtaining a topological graph.
[0117] In an embodiment of the present invention, before receiving user input information and generating meeting minutes based on the input information, the display module 802 is further configured to:
[0118] Obtain the corresponding minutes template based on the meeting area of each sub-video, and display the minutes template to the user in the first area;
[0119] Based on this, in receiving user input information and generating meeting minutes files based on the input information, the generation module 803 is specifically used for:
[0120] Receive the first information entered by the user in the first area, and generate a meeting minutes file based on the first information and the minutes template.
[0121] See Figure 9 , Figure 9 This is a schematic diagram of the structure of an electronic device provided for an embodiment of this application. For example... Figure 9 As shown, the electronic device 900 includes a transceiver 901, a processor 902, and a memory 903. These are connected via a bus 904. The memory 903 stores computer programs and data, and can transfer data stored in the memory 903 to the processor 902.
[0122] Processor 902 is used to read the computer program in memory 903 and perform the following operations:
[0123] The meeting recording file is split into at least one sub-video;
[0124] Identify the keywords and conference domains for each sub-video in at least one sub-video, and determine the associated words for each sub-video based on the keywords and conference domains of each sub-video;
[0125] determine a play speed of each sub-video according to the keywords and the conference field of each sub-video;
[0126] play each sub-video to the user in sequence according to the play rules of each sub-video, and synchronously display the keywords and the associated words of each sub-video to the user;
[0127] receive input information of the user, and generate a conference minutes file according to the input information.
[0128] In an embodiment of the present application, in the aspect of determining the play speed of each sub-video according to the keywords and the conference field of each sub-video, the processor 902 is specifically configured to perform the following operations:
[0129] determine a first importance of each sub-video according to the keywords and the conference field of each sub-video;
[0130] determine a main speaker of each sub-video, and determine a second importance of each sub-video according to the main speaker;
[0131] determine a repetition number of each sub-video, and determine a third importance of each sub-video according to the repetition number and a position of each sub-video in a repeated video;
[0132] determine a fourth importance of each sub-video according to the first importance, the second importance and the third importance;
[0133] determine the play speed of each sub-video according to the fourth importance.
[0134] In an embodiment of the present application, the third importance can be represented by formula ⑤:
[0135]
[0136] wherein x3 represents the third importance, a represents the repetition number, and b represents a position serial number of each sub-video in the repeated video.
[0137] In an embodiment of the present application, the fourth importance can be represented by formula ⑥:
[0138]
[0139] wherein x1 represents the first importance, x2 represents the second importance, x3 represents the third importance, and x4 represents the fourth importance.
[0140] In an embodiment of the present application, in the aspect of determining the keywords and the conference field of each sub-video in the at least one sub-video, the processor 902 is specifically configured to perform the following operations:
[0141] Determine a content display area in each sub-video, and perform text recognition on the content in the content display area to obtain first text information;
[0142] Perform audio recognition on each sub-video to obtain second text information;
[0143] Determine at least one first keyword in the first text information and at least one second keyword in the second text information;
[0144] Perform a deduplication set on the at least one first keyword and the at least one second keyword to obtain a keyword of each sub-video;
[0145] Establish a topological relationship graph according to the keyword of each sub-video;
[0146] Match in a preset knowledge network according to the topological relationship graph to determine a conference field of each sub-video.
[0147] In an embodiment of the present application, in terms of establishing a topological relationship graph according to the keyword of each sub-video, the processor 902 is specifically configured to perform the following operations:
[0148] Randomly select the keyword of each sub-video c+1 times to obtain c+1 keyword groups, wherein c+1 is the number of keywords of each sub-video, each random selection selects any two different keywords in the keywords of each sub-video, and the keywords selected in any two random selections are not completely the same, and each keyword group in the c+1 keyword groups includes a third keyword and a fourth keyword;
[0149] Determine the correlation coefficient between the third keyword and the fourth keyword in each keyword group respectively to obtain c+1 correlation coefficients;
[0150] Take the c+1 keywords of each sub-video as c+1 nodes respectively, and take each correlation coefficient in the c+1 correlation coefficients as an edge between two nodes corresponding to two keywords in the keyword group corresponding to the correlation coefficient, to obtain a topological relationship graph.
[0151] In an embodiment of the present application, before receiving the input information of the user and generating the conference minutes file according to the input information, the processor 902 is further configured to perform the following operations:
[0152] Obtain a corresponding minutes template according to the conference field of each sub-video, and display the minutes template to the user in a first area;
[0153] Based on this, in terms of receiving the input information of the user and generating the conference minutes file according to the input information, the processor 902 is specifically configured to perform the following operations:
[0154] The first information input by the user in the first area is received, and a conference minutes file is generated according to the first information and a minutes template.
[0155] It should be understood that the conference minutes file generation apparatus in the present application can include a smart phone (such as an Android phone, an iOS phone, a Windows Phone phone, etc.), a tablet computer, a palm computer, a notebook computer, a mobile Internet device (MID), a robot, a wearable device, etc. The conference minutes file generation apparatuses described above are only examples and are not exhaustive, and include but are not limited to the conference minutes file generation apparatuses described above. In actual applications, the conference minutes file generation apparatuses described above can also include a smart vehicle terminal, a computer device, etc.
[0156] From the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software in combination with a hardware platform. Based on such an understanding, all or part of the technical solutions of the present application that contribute to the background art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments of the present application.
[0157] Therefore, the embodiments of the present application also provide a computer readable storage medium storing a computer program, which is executed by a processor to implement some or all steps of any conference minutes file generation method as described in the above method embodiments. For example, the storage medium can include a hard disk, a floppy disk, an optical disk, a magnetic tape, a magnetic disk, a USB flash disk, a flash memory, etc.
[0158] The embodiments of the present application also provide a computer program product including a non-transitory computer readable storage medium storing a computer program, which is operable to cause a computer to execute some or all steps of any conference minutes file generation method as described in the above method embodiments.
[0159] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.
[0160] In the above-described embodiments, the description of each embodiment focuses on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0161] In several embodiments provided in the present application, it should be understood that the disclosed apparatus can be implemented by other means. For example, the apparatus embodiments described above are only illustrative, and for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical or other forms.
[0162] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., can be located in one place or can be distributed to a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0163] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software program module.
[0164] The integrated unit, if realized in the form of a software program module and sold or used as an independent product, can be stored in a computer readable memory. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes or the whole or part of the technical solutions can be embodied in the form of a software product, which is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0165] Those skilled in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by instructing the relevant hardware with programs, and the programs can be stored in a computer readable memory, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0166] The above has carried out the detailed introduction to the embodiments of the application, the principle and the implementation of the application are described in the text by applying the specific examples, the above implementation of the method is only for helping understanding the method and its core idea of the application; at the same time, for the general technical personnel in the art, according to the idea of the application, the specific implementation and the application range will have the change, the above is described, the content of the specification should not be understood as the limitation of the application.
Claims
1. A method of generating a meeting minutes document, characterized by, The method comprises: splitting the conference recording file to obtain at least one sub-video; determining keywords and conference fields of each sub-video in the at least one sub-video, and determining association words of each sub-video according to the keywords and the conference fields of each sub-video; determining a playback speed of each sub-video according to the keywords and the conference fields of each sub-video; playing each sub-video to a user in sequence according to the playback rules of each sub-video, and synchronously displaying the keywords and the association words of each sub-video to the user; receiving input information of the user, and generating the conference minutes file according to the input information; wherein the determining the playback speed of each sub-video according to the keywords and the conference fields of each sub-video comprises: determining a first importance degree of each sub-video according to the keywords and the conference fields of each sub-video; determining a main speaker of each sub-video, and determining a second importance degree of each sub-video according to the main speaker; determining a repetition number of each sub-video, and determining a third importance degree of each sub-video according to the repetition number and a position of each sub-video in a repeated video; determining a fourth importance degree of each sub-video according to the first importance degree, the second importance degree and the third importance degree; determining the playback speed of each sub-video according to the fourth importance degree.
2. The method of claim 1, wherein, The third importance degree satisfies the following formula: wherein x3 represents the third importance degree, a represents the repetition number, and b represents a position serial number of each sub-video in a repeated video.
3. The method according to claim 1 or 2, characterized in that, The fourth importance degree satisfies the following formula: wherein x1 represents the first importance degree, x2 represents the second importance degree, x3 represents the third importance degree, and x4 represents the fourth importance degree.
4. The method of claim 1, wherein, The determining the keywords and the conference fields of each sub-video in the at least one sub-video comprises: determining a content display area in each sub-video, and performing text recognition on content in the content display area to obtain first text information; performing audio recognition on each sub-video to obtain second text information; determining at least one first keyword in the first text information, and determining at least one second keyword in the second text information; performing a de-duplication set on the at least one first keyword and the at least one second keyword to obtain the keywords of each sub-video; establishing a topological relationship graph according to the keywords of each sub-video; matching in a preset knowledge network according to the topological relationship graph to determine the conference fields of each sub-video.
5. The method of claim 4, wherein, The establishing a topological relationship graph according to the keywords of each sub-video comprises: randomly selecting the keywords of each sub-video c times to obtain c keyword groups, wherein c+1 is the number of the keywords of each sub-video, each time of random selection selects any two different keywords in the keywords of each sub-video, and the keywords selected in any two times of random selection are not completely the same, and each keyword group in the c keyword groups comprises a third keyword and a fourth keyword; determine a correlation coefficient between the third keyword and the fourth keyword in each keyword group respectively, to obtain c correlation coefficients; take the c+1 keywords of each sub-video as c+1 nodes respectively, and take each correlation coefficient in the c correlation coefficients as an edge between two nodes corresponding to two keywords in the keyword group corresponding to the correlation coefficient, to obtain the topological relationship graph.
6. The method of claim 1, wherein, Before the receiving the input information of the user and the generating the conference minutes file according to the input information, the method further comprises: obtaining a corresponding minutes template according to the conference field of each sub-video, and displaying the minutes template to the user in a first area; the receiving the input information of the user and the generating the conference minutes file according to the input information, comprising: receiving first information input by the user in the first area, and generating the conference minutes file according to the first information and the minutes template.
7. A meeting summary document generating apparatus characterized by comprising: The device comprises: an analysis module, configured to split a conference recording file to obtain at least one sub-video, determine keywords and a conference field of each sub-video in the at least one sub-video, determine association words of each sub-video according to the keywords and the conference field of each sub-video, and determine a playback speed of each sub-video according to the keywords and the conference field of each sub-video; a display module, configured to sequentially play each sub-video to a user according to a playback rule of each sub-video, and synchronously display the keywords and the association words of each sub-video to the user; a generation module, configured to receive input information of the user, and generate the conference minutes file according to the input information; In terms of determining the playback speed of each sub-video according to the keywords and the conference field of each sub-video, the analysis module is specifically configured to: determine a first importance of each sub-video according to the keywords and the conference field of each sub-video; determine a main speaker of each sub-video, and determine a second importance of each sub-video according to the main speaker; determine a repetition number of each sub-video, and determine a third importance of each sub-video according to the repetition number and a position of each sub-video in a repeated video; determine a fourth importance of each sub-video according to the first importance, the second importance and the third importance; determine the playback speed of each sub-video according to the fourth importance.
8. An electronic device, comprising: A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method for changing fluid-medium file broadcasting speed
CN101075949A
Conference summary generation method and device, electronic equipment and storage medium
CN111666746A