Conference summary information extraction method, system and device and storage medium

By extracting semantic information in meeting records and filtering exception statements, using pre-trained models to extract key information and converting it into structured data, the problem of difficult to retrieve and manage handwritten meeting records is solved, and more efficient information understanding and management is achieved.

CN120045609APending Publication Date: 2025-05-27SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Patent Information

Application Number
CN202411917615.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Handwritten meeting minutes relying on physical media, it is difficult to effectively retrieve and manage, and the handwriting and format vary from person to person, which increases the difficulty of later search.

Method used

By extracting conference data from conference record documents, using convolutional neural networks and long and short-term memory networks to extract font contour features and semantic vectors, combining pre-established knowledge vector databases for exception statement filtering, writing propt prompt words to extract key information using pre-trained models and convert them into structured minutes information.

Benefits of technology

It realizes the conversion of unstructured meeting records into structured data, thereby better understanding of the meeting content, simplifying the post-retrieval and management process, and reducing difficulties caused by word marks and format differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045609A_ABST
    Figure CN120045609A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and particularly provides a conference summary information extraction method, system and device and a storage medium, and the method comprises the steps: extracting conference data from a conference record document; performing abnormal statement filtering processing on the conference data by utilizing a pre-established knowledge vector library; compiling a prompt prompt word, and extracting key information from the filtered conference data based on the prompt prompt word by using a pre-training model; and converting the key information into structured summary information. According to the method and the device, the document characters are converted into the structured data from the unstructured characters, so that the conference content can be better understood.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a method, system, device and storage medium for extracting conference minutes information. Background Art

[0002] As a traditional way of recording, handwritten meeting minutes are still widely used in some occasions. However, this method has a significant drawback, that is, it cannot be effectively retrieved and viewed later.

[0003] Handwritten records often rely on physical media such as paper and pen, which makes it inconvenient to store and carry information. Over time, a large number of handwritten records may pile up, not only taking up space but also difficult to manage. When you need to find the minutes of a specific meeting, you may have to spend a lot of time and energy flipping through these paper documents, which is not only inefficient but also prone to errors.

[0004] In addition, the handwriting and format of handwritten records may vary from person to person, further increasing the difficulty of later retrieval. Some people's handwriting may be difficult to read, while some records may lack uniform formatting and annotations, making it more difficult to extract and understand the information. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method, system, device and storage medium for extracting meeting minutes information to solve the above-mentioned technical problems.

[0006] In a first aspect, the present invention provides a method for extracting meeting minutes information, comprising: Extract meeting data from meeting minutes documents; Using a pre-established knowledge vector library to filter abnormal statements from the conference data; Writing prompt words, and using the pre-trained model to extract key information from the filtered conference data based on the prompt words; The key information is converted into structured minutes information.

[0007] In an optional implementation, extracting conference data from a conference record document includes: Extracting font outline features from the conference record document using a convolutional neural network to obtain a feature vector; The feature vector is input into a pre-trained long short-term memory network integrating an attention mechanism to obtain a semantic vector.

[0008] In an optional implementation, using a pre-established knowledge vector library to filter abnormal statements from the conference data includes: Use the cosine similarity method to match and compare the semantic vectors of the meeting with the vectors in the knowledge vector library; According to the matching results, calculate the similarity or distance between each semantic vector and the knowledge vector library; If the similarity between a certain semantic vector and the knowledge vector library is lower than a pre-set threshold, then regard it as an abnormal statement; Filter out all abnormal statements and only retain the normal semantic vectors with a higher matching degree to the knowledge vector library; Convert the remaining semantic vectors into meeting data in text format.

[0009] In an optional implementation, write prompt words, and use a pre-trained model to extract key information from the filtered meeting data based on the prompt words, including: Write prompt words and integrate the prompt words into a structured text template. The prompt words include meeting time, location, host, theme, participants, and meeting content; Use the pre-trained BERT model to extract key information from the filtered meeting data based on the prompt words.

[0010] In an optional implementation, convert the key information into structured minutes information, including: Construct a minutes information template; Map the extracted key information to the template and fill it into the corresponding positions; Sort the structured minutes information according to the meeting time and display and output it in sequence according to the sorting.

[0011] In a second aspect, the present invention provides a meeting minutes information extraction system, including: A data input module for extracting meeting data from a meeting record document; A first processing module for filtering abnormal statements from the meeting data using a pre-established knowledge vector library; A second processing module for writing prompt words and using a pre-trained model to extract key information from the filtered meeting data based on the prompt words; A data output module for converting the key information into structured minutes information.

[0012] In an optional implementation, the data input module includes: A feature extraction unit for extracting font contour features from the meeting record document using a convolutional neural network to obtain feature vectors; A feature recognition unit, configured to input the feature vector into a pre-trained long short-term memory network with a fusion attention mechanism to obtain a semantic vector.

[0013] In an optional embodiment, the first processing module includes: A first calculation unit, configured to use the cosine similarity method to match and compare the semantic vector of the meeting with the vectors in the knowledge vector library; A second calculation unit, configured to calculate the similarity or distance between each semantic vector and the knowledge vector library according to the matching result; A determination unit, configured to regard a semantic vector as an abnormal statement if its similarity with the knowledge vector library is lower than a preset threshold; A filtering unit, configured to filter out all abnormal statements and only retain the normal semantic vectors with a high matching degree with the knowledge vector library; A conversion unit, configured to convert the remaining semantic vectors into meeting data in text format.

[0014] In a third aspect, a device is provided, including: A memory, configured to store a meeting minutes information extraction program; A processor, configured to implement the steps of the meeting minutes information extraction method provided in the first aspect when executing the meeting minutes information extraction program.

[0015] In a fourth aspect, a computer-readable storage medium is provided, on which a meeting minutes information extraction program is stored. When the meeting minutes information extraction program is executed by a processor, the steps of the meeting minutes information extraction method provided in the first aspect are implemented.

[0016] The beneficial effects of the present invention are that the meeting minutes information extraction method, system, device and storage medium provided by the present invention can better understand the meeting content by converting the document text from unstructured text into structured data.

[0017] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.

[0020] Figure 2It is a schematic block diagram of a system according to an embodiment of the present invention.

[0021] Figure 3 It is a schematic structural diagram of a device provided by an embodiment of the present invention. Detailed implementation manners

[0022] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0024] The meeting minutes information extraction method provided by the embodiment of the present invention is executed by a computer device. Correspondingly, the meeting minutes information extraction system runs in the computer device.

[0025] Figure 1 It is a schematic flowchart of a method according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be a meeting minutes information extraction system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0026] Such as Figure 1 As shown, the method includes: S1. Extract meeting data from the meeting record document; S2. Use the pre-established knowledge vector library to perform abnormal statement filtering processing on the meeting data; S3. Compile a prompt, and use a pre-trained model to extract key information from the filtered meeting data based on the prompt; S4. Convert the key information into structured minutes information.

[0027] In an embodiment of the present invention, based on step S1, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation manner.

[0028] Suppose there is a handwritten meeting record document that contains discussions about a certain project. The goal is to use CNN and LSTM to identify the font features in the document and parse out the semantic information therein for subsequent tasks such as information retrieval or content classification.

[0029] Processing procedure: S101. Image preprocessing First of all, it is necessary to perform image preprocessing on the meeting record document. This includes converting the document image into a format suitable for CNN processing (such as grayscale image or binary image), and performing necessary image enhancement operations (such as denoising, smoothing, etc.) to improve the image quality.

[0030] S102. Extract font features using CNN Next, the preprocessed image is input into the CNN. The CNN will extract the features in the image layer by layer in depth, especially the font contour features. These features may include the thickness, direction, curvature of the strokes, and the overall style of the font (such as handwriting, printed font, etc.).

[0031] At the last layer of the CNN, these features will be encoded into a series of feature vectors. These vectors precisely represent the font features in the document numerically, providing a basis for subsequent processing.

[0032] S103. Input the feature vectors into LSTM Then, these feature vectors rich in font information are input into a pre-trained LSTM integrated with an attention mechanism. The LSTM will process these feature vectors one by one and capture the long-term dependencies between them.

[0033] In this process, the attention mechanism will play an important role. It will dynamically adjust the attention to different feature vectors in order to more accurately understand the semantic content in the document. For example, when a key term or phrase appears in the document, the attention mechanism may increase the attention to the feature vectors of that part in order to better capture its semantic information.

[0034] S104. Semantic parsing and output Finally, the LSTM will output a series of semantic vectors containing rich semantic information. These vectors can be further used for tasks such as information retrieval, content classification, and sentiment analysis.

[0035] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0036] After being processed by the LSTM, a series of semantic vectors containing rich semantic information are obtained. Next, the following steps will be carried out: S201. Perform matching using the cosine similarity method Cosine similarity is a metric for measuring the similarity of the directions of two vectors, with a value range of [-1, 1]. The closer the value is to 1, the closer the directions of the two vectors are, that is, the higher the similarity. Calculate the cosine similarity between the semantic vector of the meeting and each vector in the knowledge vector library to evaluate the similarity between them.

[0037] S202. Calculate similarity or distance Based on the calculation results of cosine similarity, the similarity between each semantic vector and the knowledge vector library can be obtained. To represent this similarity more intuitively, it can also be converted into a distance metric (such as 1 minus the cosine similarity) for subsequent analysis and comparison.

[0038] S203. Identify abnormal statements Set a predefined threshold for determining whether a semantic vector matches the knowledge vector library. If the similarity of a semantic vector to the knowledge vector library is lower than this threshold, it is regarded as an abnormal statement. These abnormal statements may contain new, unknown information or have a large difference from the information in the knowledge vector library.

[0039] S204. Filter abnormal statements Filter out all the identified abnormal statements and only retain the normal semantic vectors with a high degree of matching to the knowledge vector library. This can ensure that the information processed subsequently is accurate and reliable, while reducing noise and interference.

[0040] S205. Convert to meeting data in text format Finally, convert the remaining semantic vectors into meeting data in text format. This can be achieved by mapping the semantic vectors to a predefined vocabulary or phrase library. The converted text data can be more conveniently used for subsequent information retrieval, content classification, sentiment analysis and other tasks.

[0041] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0042] S301. Write Prompt prompts First, we need to write a series of prompt prompts according to the characteristics of the meeting record data. These prompt prompts will be used to guide the BERT model to extract key information from the data. These prompt prompts include but are not limited to: Meeting time: Used to extract the specific time information of the meeting, such as date, start time and end time.

[0043] Meeting Location: Used to extract the specific location of the meeting or the platform information for online meetings.

[0044] Host: Used to extract the name or identity of the meeting host.

[0045] Meeting Theme: Used to extract the core discussion topic or issue of the meeting.

[0046] Attendees: Used to extract the list or roles of the attendees.

[0047] Meeting Content: Used to extract the main discussion content, decision results, or follow-up action plans of the meeting.

[0048] S302. Integrate Structured Text Template Next, we integrate these prompt words into a structured text template. This template will clearly show the information positions corresponding to each prompt word and facilitate information extraction by the BERT model.

[0049] S303. Extract Key Information Using Pre-trained BERT Model Now, we have the structured text template and prompt words. Next, we will use the pre-trained BERT model to extract key information from the filtered meeting data.

[0050] The BERT model is a pre-trained language model based on the Transformer architecture, which can extract key information by understanding the context of the text. We input the meeting data into the BERT model and use the prompt words as a guide to let the model locate and extract the information corresponding to each prompt word in the data.

[0051] For example, when the model processes the prompt word "Meeting Time", it will try to extract the specific information related to the meeting time from the data and fill it into the corresponding position in the template. Similarly, the model will process other prompt words and extract the corresponding information.

[0052] S304. Output Structured Information Finally, the BERT model will output a text template containing structured information, where each prompt word is replaced with the specific information extracted from the meeting data. This structured information can be conveniently used for subsequent information retrieval, content classification, data analysis, and other tasks.

[0053] Among them, for the topic classification of the meeting content, the meeting minutes can be manually classified according to the content of the meeting minutes, and the classification results can be used as a dataset. We can divide the dataset into a training set, a test set, and a validation set according to a certain proportional relationship; each dataset sample is the meeting minutes content and its classification label.

[0054] Set appropriate downstream classification task parameters according to the number of classification labels; Initialize the hyperparameters of the model; Fine-tune and train the above-mentioned training set, test set, and test set as required. During the fine-tuning process, evaluate according to a certain evaluation criterion, and optimize the parameters according to the evaluation results to achieve the best effect; Use the third-party library of Python to deploy the fine-tuned model; Users can provide feedback on the classification labels, correct the meeting minutes content with incorrect classification labels, and the feedback results can be included in the new dataset, which can be used for iterative optimization of the classification model.

[0055] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.

[0056] S401. Construct a minutes information template Before extracting the key information, we need to first construct a minutes information template. This template will be constructed based on the prompt words defined before and contain information fields corresponding to each prompt word.

[0057] S402. Map and fill the extracted key information with the template After extracting the key information using the BERT model, we need to map this information with the minutes information template and fill them into the corresponding fields. This can be achieved by writing a mapping function that will receive the structured information output by the BERT model and match and fill it with the fields in the template.

[0058] For example, if the BERT model outputs a structured information containing the meeting time, we can fill this information into the "<time field>" position in the template. Similarly, we can also fill other key information into the corresponding fields.

[0059] S403. Sort and display the structured minutes information according to the meeting time Finally, we need to sort the structured minutes information according to the meeting time and display the output in sequence according to the sorting result. This can be achieved by writing a sorting function that will receive a list of structured information containing multiple meeting minutes and sort them according to the meeting time field.

[0060] After sorting, we can display the sorted list of meeting minutes for the user to view and analyze. This can be achieved by presenting a table or list containing the sorted meeting minutes on the user interface.

[0061] The display page shows the meeting minutes information in chronological order of the meeting time. Users can query by customizing filtering conditions, such as time, location, host, meeting name, etc., to view the information of a certain meeting.

[0062] Users can provide feedback on the extracted meeting information, correct errors in the meeting minutes information, and the feedback results can be incorporated into a new dataset, which can be used for iterative optimization of the extraction effect of the large model.

[0063] In some embodiments, the meeting minutes information extraction system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the meeting minutes information extraction system can be stored in the memory of the computer device and executed by at least one processor to perform the functions of meeting minutes information extraction (see Figure 1 description).

[0064] In this embodiment, according to the functions it performs, the meeting minutes information extraction system can be divided into multiple functional modules, such as Figure 2 shown. The functional modules of system 200 may include: a data input module 210, a first processing module 220, a second processing module 230, and a data output module 240. The modules referred to in the present invention refer to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0065] The data input module is used to extract meeting data from the meeting record document; The first processing module is used to perform abnormal statement filtering processing on the meeting data by using a pre-established knowledge vector library; The second processing module is used to write prompt words and use a pre-trained model to extract key information from the filtered meeting data based on the prompt words; The data output module is used to convert the key information into structured minutes information.

[0066] Optionally, as an embodiment of the present invention, the data input module includes: A feature extraction unit, configured to extract font contour features from the meeting record document by using a convolutional neural network to obtain a feature vector; A feature recognition unit, configured to input the feature vector into a long short-term memory network integrating an attention mechanism that has been pre-trained to obtain a semantic vector.

[0067] Optionally, as an embodiment of the present invention, the first processing module includes: A first calculation unit, configured to use the cosine similarity method to match and compare the semantic vector of the meeting with the vectors in the knowledge vector library; A second calculation unit, configured to calculate the similarity or distance between each semantic vector and the knowledge vector library according to the matching result; A determination unit, configured to regard a semantic vector as an abnormal statement if its similarity with the knowledge vector library is lower than a preset threshold; A filtering unit, configured to filter out all abnormal statements and only retain normal semantic vectors with a high matching degree with the knowledge vector library; A conversion unit, configured to convert the remaining semantic vectors into meeting data in text format.

[0068] Figure 3 The method for extracting meeting minutes information provided in the embodiments of the present application can be applied to a device. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown in the figure, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.

[0069] Among them, the device 300 may include: a processor 310, a memory 320, and a communication unit 330. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It may be a bus structure, a star structure, or may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0070] Among them, the memory 320 can be used to store the execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 can execute some or all of the steps in the above method embodiments.

[0071] The processor 310 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 320, and by calling the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor can be composed of an integrated circuit (IC). For example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 310 can include only a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single arithmetic core or can include multiple arithmetic cores.

[0072] The communication unit 330 is used to establish a communication channel so that the storage device can communicate with other devices. It receives user data sent by other devices or sends user data to other devices.

[0073] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the various embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), etc.

[0074] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disc, etc., various media that can store program codes, including several instructions to enable a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0075] For the same and similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0076] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems or modules can be in electrical, mechanical or other forms.

[0077] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0078] In addition, in each embodiment of the present invention, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0079] Although the present invention has been described in detail by referring to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention / Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of changes or substitutions, and all should be covered within the protection scope of the present invention.

Claims

1. A method for extracting meeting minutes information, characterized in that: include: Extract meeting data from meeting minutes documents; Using a pre-established knowledge vector library to filter abnormal statements from the conference data; Writing prompt words, and using the pre-trained model to extract key information from the filtered conference data based on the prompt words; The key information is converted into structured minutes information.

2. The method according to claim 1, characterized in that Extract meeting data from meeting record documents, including: Extracting font outline features from the conference record document using a convolutional neural network to obtain a feature vector; The feature vector is input into a pre-trained long short-term memory network integrating an attention mechanism to obtain a semantic vector.

3. The method according to claim 2, characterized in that Using a pre-established knowledge vector library to filter abnormal statements from the conference data includes: Use the cosine similarity method to match and compare the semantic vector of the conference with the vectors in the knowledge vector library; According to the matching results, the similarity or distance between each semantic vector and the knowledge vector library is calculated; If the similarity between a semantic vector and the knowledge vector library is lower than a preset threshold, it will be regarded as an abnormal sentence; Filter out all abnormal sentences and only keep normal semantic vectors that have a high degree of match with the knowledge vector library; Convert the remaining semantic vectors into conference data in text format.

4. The method according to claim 3, characterized in that Write prompt words, and use the pre-trained model to extract key information from the filtered meeting data based on the prompt words, including: Write prompt words and integrate the prompt words into a structured text template, wherein the prompt words include the meeting time, location, host, topic, participants, and meeting content; The pre-trained BERT model is used to extract key information from the filtered meeting data based on the prompt word.

5. The method according to claim 1, characterized in that Convert the key information into structured minutes information, including: Construct a minutes information template; Map the extracted key information with the template and fill it into the corresponding position; Sort the structured minutes information by meeting time and display them in order.

6. A meeting minutes information extraction system, characterized in that: include: A data input module, used to extract meeting data from meeting record documents; A first processing module, configured to filter abnormal statements from the conference data using a pre-established knowledge vector library; The second processing module is used to write prompt words, and use the pre-trained model to extract key information from the filtered conference data based on the prompt words; The data output module is used to convert the key information into structured minutes information.

7. The system according to claim 6, characterized in that The data input module comprises: A feature extraction unit, used to extract font outline features from the conference record document using a convolutional neural network to obtain a feature vector; The feature recognition unit is used to input the feature vector into a pre-trained long short-term memory network fused with an attention mechanism to obtain a semantic vector.

8. The system according to claim 7, characterized in that The first processing module comprises: A first computing unit is used to match and compare the semantic vector of the conference with the vectors in the knowledge vector library using a cosine similarity method; A second calculation unit, used to calculate the similarity or distance between each semantic vector and the knowledge vector library according to the matching result; A judgment unit, used to regard a certain semantic vector as an abnormal sentence if the similarity between the semantic vector and the knowledge vector library is lower than a preset threshold; A filtering unit is used to filter out all abnormal sentences and retain only normal semantic vectors that have a high degree of matching with the knowledge vector library; The conversion unit is used to convert the remaining semantic vectors into conference data in text format.

9. A device, characterized in that: include: A memory for storing a meeting minutes information extraction program; A processor is used to implement the steps of the meeting minutes information extraction method as described in any one of claims 1 to 5 when executing the meeting minutes information extraction program.

10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a meeting minutes information extraction program, and when the meeting minutes information extraction program is executed by a processor, the steps of the meeting minutes information extraction method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method and device for generating conference summary and electronic equipment

    CN117316161A

  • Intelligent generation method and system of conference summary

    CN118709672A

  • Conference summary generation method and device, terminal and computer readable storage medium

    CN119150814A

Cited By

  • Conference information processing system and terminal

    CN120786018A