Intelligent conference recording system
Through the multimodal data fusion and analysis of the intelligent conference record system, the problems of insufficient speech recognition accuracy and insufficient multimodal data processing capabilities in the prior art are solved, and efficient conference record and decision-making support are achieved.
Patent Information
- Application Number
- CN202510568188.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-01
AI Technical Summary
The existing conference recording method faces the problems of insufficient speech recognition accuracy, difficulty in identifying handwritten notes, lack of intelligent search functions and insufficient multimodal data processing capabilities in complex conference scenarios, resulting in low information utilization and difficulty in supporting efficient decision-making.
It provides an intelligent conference record system, including a conference information collection module, a data pre-cleaning module, a conference information transcription module and an intelligent conference content analysis module. Through speech recognition, image analysis, multimodal data fusion and cross-conference association analysis, structured conference record reports are generated.
It significantly improves the recognition accuracy of voice and handwritten notes, supports deep fusion and intelligent analysis of multimodal data, improves information utilization, and enhances decision-making support capabilities.
Smart Images

Figure CN120415931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of conference recording, and in particular to an intelligent conference recording system. Background Art
[0002] With the development and widespread application of technologies like artificial intelligence, big data analytics, and cloud computing, efficient information processing and knowledge management have become key factors in optimizing individual capabilities, enhancing corporate competitiveness, and strengthening government governance. In this context, meeting minutes, as an indispensable component of decision-making, communication, collaboration, and knowledge sharing, have a direct impact on the development of individuals and organizations, as well as the success of strategic implementation.
[0003] However, existing meeting recording methods face the following technical problems in complex meeting scenarios:
[0004] (1) Insufficient speech recognition accuracy: When multiple people are speaking or there is a lot of background noise, the accuracy of existing speech recognition technology is low, resulting in a high recognition error rate of the transcribed text.
[0005] (2) Difficulty in recognizing handwritten notes: The recognition of handwritten notes is easily affected by factors such as writing quality and font style, resulting in overall recognition errors of the text.
[0006] (3) Lack of intelligent search function: Existing conference software tools are unable to perform intelligent search of conference content and / or cross-conference correlation analysis, resulting in low information utilization and difficulty in supporting efficient decision-making.
[0007] (4) Insufficient multimodal data processing capabilities: Existing conference software tools are unable to effectively integrate multimodal data such as voice, images, and text, resulting in insufficient comprehensiveness and structuring of conference content. Summary of the Invention
[0008] (1) Technical problems solved
[0009] In view of the deficiencies of the prior art, the present invention provides an intelligent conference recording system that can solve at least one of the above technical problems.
[0010] (2) Technical solution
[0011] To solve the above technical problems, the present invention provides the following technical solutions: an intelligent conference recording system, comprising:
[0012] Meeting information collection module, used to collect meeting information;
[0013] The conference information transcription module is used to transcribe the conference information to obtain conference data;
[0014] The intelligent analysis module for meeting content is used to perform intelligent analysis and processing on meeting data to generate a meeting record report.
[0015] Preferably, the intelligent meeting recording system further includes a data pre-cleaning module for pre-cleaning the meeting information collected by the meeting information acquisition module; the meeting information transcription module is further used to transcribe the pre-cleaned meeting information.
[0016] Preferably, the meeting information includes voice, image, video and / or text.
[0017] Preferably, the transcription process specifically includes performing speech recognition on the voice to obtain meeting data.
[0018] Preferably, the transcription process specifically includes parsing and extracting the image content of the image to obtain meeting data.
[0019] Preferably, the intelligent analysis and processing specifically includes deeply fusing the meeting data from meeting information in different formats.
[0020] Preferably, the intelligent analysis module for meeting content is also used to perform cross-meeting correlation analysis.
[0021] Preferably, the intelligent meeting recording system further includes a data editing module for editing the meeting data obtained by the meeting information transcription module.
[0022] Preferably, the intelligent meeting recording system further includes a meeting management module for classifying and marking meetings.
[0023] Preferably, the intelligent meeting recording system further includes a display module for displaying the meeting record report generated by the intelligent analysis module for meeting content.
[0024] (III) Beneficial Effects
[0025] Compared with the prior art, the present invention provides an intelligent meeting recording system with the following beneficial effects: The intelligent meeting recording system of the present invention includes functional modules such as a meeting information acquisition module, a data pre-cleaning module, a meeting information transcription module, and an intelligent analysis module for meeting content. Through the intelligent analysis module for meeting content, the present invention can deeply fuse multi-modal data such as voice, image, and text to provide users with structured data and intelligent analysis; and the present invention can support cross-meeting correlation analysis to better support users in making decisions; in addition, the present invention can effectively improve the speech recognition accuracy and the recognition accuracy of handwritten notes. In the above manner, the present invention proposes an intelligent meeting recording system that integrates efficient meeting recording and intelligent information analysis, which can improve the overall working efficiency of meeting recording. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is the principle block diagram of an intelligent conference recording system according to the present invention;
[0027] Figure 2 This is the step flow chart of an intelligent conference recording system according to the present invention. Specific embodiments
[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0029] The present invention provides an intelligent conference recording system, which includes functional modules such as a conference information collection module, a conference information transcription module, and a conference content intelligent analysis module. Specifically, the conference information collection module is used to collect conference information; further, the conference information transcription module is used to transcribe the conference information to obtain conference data; further, the conference content intelligent analysis module is used to perform intelligent analysis processing on the conference data to generate a conference recording report.
[0030] Preferably, the intelligent conference recording system of the present invention further includes a data pre-cleaning module, which is used to pre-clean the conference information collected by the conference information collection module to remove data such as conference interruptions, non-conference content, and unclear judgments, ensuring the accuracy of the analysis results. The conference information transcription module is further used to transcribe the pre-cleaned conference information.
[0031] Specifically, the above-mentioned conference information collection module can specifically collect conference information through user terminal devices such as mobile phones, tablets, and computers. The conference information collected by the conference information collection module can include conference information in different formats such as voice, image, video, and / or text; the conference information can specifically include relevant information content of the conference such as the conference theme, conference time, speaker, conference content, and conference scene. The above-mentioned conference information transcription module and conference content intelligent analysis module are preferably deployed on a cloud server, which can reduce the dependence on local resources of the client device, improve the cross-platform compatibility of the system of the present invention, and optimize the user experience at the same time.
[0032] Further, the above data pre - cleaning module is preferably used to process the speech collected by the meeting information collection module as follows: (1) Meeting session boundary detection: Use pydub or Librosa to extract the audio energy curve of the speech, detect long - term low - energy segments (such as continuous silence > 30 seconds), and mark them as non - meeting sessions. (2) Voice activity detection (VAD) and silence filtering: Dynamically segment valid speech segments, adopt a VAD model based on energy entropy or deep learning (such as Webrtc VAD, YAMNet) to identify the boundaries between speech and silence, remove short - term silences and background noises (< 1 second, such as page - turning sounds, keyboard tapping sounds); set the minimum length of the speech segment (such as > 1.5 seconds) to filter out extremely short meaningless sounds (such as coughing, sighing). (3) Meeting interruption recognition: Conduct abnormal event detection, detect sudden noises (such as applause, equipment noise) or long - term silence (> 2 minutes), and combine with the meeting agenda logic (such as tea breaks, recesses) to mark them as meeting interruption segments. When processing later, you can choose to delete or separately annotate them; multi - modal assistance, combine the meeting information in video format, and synchronously analyze the changes in the video screen (such as black screen, camera switching) to assist in judging the interruption points.
[0033] The transcription process of the above-mentioned meeting information transcription module may specifically include performing speech recognition processing on the speech to obtain meeting data, that is, performing speech recognition processing on the above-mentioned meeting information in speech format to obtain meeting data. Preferably, the speech recognition processing can adopt deep learning and voiceprint recognition technologies, supporting multi-language and speaker recognition, noise suppression, and real-time transcription: (1) For the front-end audio preprocessing and noise suppression of the speech: First, adopt a deep learning noise reduction model: preferably adopt a time-domain modeling method (such as Conv-TasNet, DPRNN, LSTM-Transformer) or a frequency-domain method (such as a noise reduction autoencoder DAE based on complex CNN), input the noisy speech into the deep learning noise reduction model, and thus output clean speech; Second, adopt multi-microphone array fusion: combine beamforming hardware technology to enhance the target speaker signal and suppress environmental noise (such as keyboard sounds, air conditioner sounds); Third, real-time optimization: adopt a streaming processing framework (such as PyTorch Streaming, TensorFlow Lite), supporting low-latency online inference (single-frame processing time < 30 ms). (2) Speaker recognition and speaker segmentation: First, voiceprint feature extraction: use a pre-trained speaker embedding model (such as ECAPA-TDNN, ResNet-TDNN, etc.) to generate d-vector or x-vector, and extract speaker identity features (frame-level embedding); Second, adopt a real-time clustering algorithm: online clustering based on a sliding window (such as the Viterbi algorithm, dynamic time warping DTW), and real-time segment audio segments of different speakers (timestamp accuracy ≤ 200 ms); Third, multi-speaker tracking: combine Kalman Filter or Hidden Markov Model (HMM) to handle speaker alternation scenarios and reduce the segmentation error rate (ER < 5%); Fourth, register speakers (import the voiceprint library before the meeting) and dynamically create files for unknown speakers, and update the speaker list in real time through cosine similarity matching.(3) Multilingual Speech Recognition: First, unified model architecture: Use end-to-end models (such as Transformer-based wav2vec2.0, Whisper, HuBERT), and switch the output space through language tags to support multilingual mixed recognition of Chinese, English, Japanese, etc.; Second, transfer learning and adaptation: For low-resource languages, use pre-trained models (such as mBART, XLSR) to fine-tune on a small amount of data in the target language to solve the cross-language acoustic difference problem; Third, language model fusion: Train independent N-gram or Transformer language models (LMs) for each language, and switch through weighted finite state transducers (WFSTs) or dynamic decoding strategies to improve the recognition accuracy in complex scenarios (word error rate WER < 8%); (4) Real-time Transcription and Result Integration: First, streaming processing framework: Build a real-time pipeline based on Kafka and Flask Streaming, input audio stream → noise reduction → speaker diarization → ASR → result alignment (timestamp accuracy ≤ 100 ms); Second, format output: Structured output according to speaker ID (such as Speaker A / B), timestamp (HH:MM:SS), and text content, support export in formats such as Markdown and Word, and automatically break lines and paragraphs; Third, post-processing optimization: Correct colloquial errors (such as filtering "um" and "ah") and proper noun corrections through a rule engine. Through the above methods, the conference information transcription module can integrate multiple functions such as noise reduction, multilingual recognition, speaker separation, and real-time transcription, significantly improving the efficiency and accuracy of digital processing of conference content.
[0034] In addition, the transcription process of the above-mentioned meeting information transcription module may specifically include parsing and extracting image content from images to obtain meeting data, that is, parsing and extracting image content from meeting information in image format. The meeting information transcription module can specifically implement image content parsing and extraction through OCR (Optical Character Recognition) and image enhancement algorithms, supporting handwritten font conversion and image optimization; preferably, the meeting information transcription module can specifically parse and extract image content from images through Tesseract OCR and Google Vision API to obtain meeting data in text form. In addition, for images including handwritten fonts, the meeting information transcription module preferably adopts the following image processing operations: (1) Preprocessing such as image enhancement, skew correction, and region segmentation to optimize image quality; (2) Feature extraction: Input the preprocessed image into a deep learning model such as a convolutional neural network CNN for feature extraction; (3) Handwritten character recognition: The extracted features can be input into a convolutional recurrent neural network CRNN, which is suitable for single-line handwritten text, and processes variable-length sequences through CTC (Connectionist Temporal Classification), supporting the recognition of cursive characters; the features can also be input into a Transformer-based model: such as Vision Transformer (ViT) or a model combined with an attention mechanism (such as Attention-OCR), using global context information to improve the recognition accuracy of complex layouts (multi-line and multi-column handwritten); the features can also be input into few-shot / zero-shot learning: for special symbols and abbreviations that appear rarely in the meeting, quickly adapt to new handwriting styles through meta-learning, reducing the data annotation cost. (4) Inference and post-processing to improve text accuracy: Language model error correction: Combining NLP techniques (such as N-gram, BERT), correct the recognition results contextually. For example, correct "meeting record rate" to "meeting minutes"; Format restoration: Retain the paragraph structure, punctuation marks, and table framework in the handwritten text (such as locating the table boundary by detecting horizontal / vertical lines) to ensure the layout consistency between the electronic text and the original notes; Multi-modal fusion: If the meeting record contains speech-to-text transcription at the same time, align the handwritten text and speech content through timestamp or location information to cross-verify the recognition results (such as the handwritten "Q3 target" corresponding to "third-quarter target" in the speech).
[0035] Preferably, the intelligent analysis process of the above-mentioned meeting content intelligent analysis module may specifically include deeply fusing the meeting data of meeting information from different formats into structured meeting data, thereby generating a meeting record report. That is to say, the meeting content intelligent analysis module deeply fuses multi-modal meeting data such as speech, images, and text, so as to provide users with structured data and intelligent analysis.
[0036] Furthermore, preferably, the intelligent analysis module for meeting content realizes the deep integration of multi-modal data such as speech, images, and text in the meeting record. Specifically, it can include four processing processes: (1) data alignment, (2) feature fusion, (3) structured modeling, and (4) intelligent analysis, so as to convert unstructured data into retrievable and analyzable structured information and generate an intelligent meeting record report, which is described in detail as follows: (1) Multi-modal data preprocessing (it can be understood that data preprocessing is achieved through the transcription processing of the above-mentioned meeting information transcription module) and data alignment: First, standardize each modal data. For speech data (i.e., meeting information in speech format), the meeting information transcription module preferably generates a text sequence with timestamps (including speaker tags, such as "[00:05] Zhang San: The goal this time is a 15% increase in Q3 revenue") through real-time speech recognition (such as Whisper, Vosk); for image data, meeting images (whiteboard, handwritten notes, PPT photos, etc.) are converted into text blocks with position coordinates (such as "[upper left corner of the whiteboard] Meeting agenda: 1. Quarterly summary") through OCR technology (such as Tesseract + handwritten model), and at the same time, metadata of non-text elements such as charts and formulas in the image are extracted (such as "bar chart - KPI completion rate of each department"); for text data, preprocess prior texts such as meeting agendas and historical documents, and label key entities (time, place, person name, task words) through named entity recognition (NER) to build a domain knowledge base (such as a common meeting term library). Second, the above data alignment includes cross-modal time / space alignment: For time alignment, based on the timestamps of speech transcription, match the image capture time in the same time period (such as the EXIF time of a mobile phone photo) or the PPT page turning time of screen sharing to establish a time mapping relationship of "speech - image - text" (such as "00:10 - 00:15 corresponds to the handwritten content on the whiteboard"); for space alignment: For handwritten notes / whiteboard images in the meeting record, associate the keywords in the speech transcription text through coordinate positioning (such as the character Bounding Box output by OCR) (such as when the speech mentions "next week's focus", it corresponds to the handwritten position of this phrase in the image) to achieve a two-way index of "content - position".(2) Multi-modal Feature Fusion: First, low-level feature fusion (early fusion): Vector embedding layer: For speech, generate acoustic feature vectors (such as 100-dimensional / frame) through a pre-trained speech model (such as HuBERT); for text, tokenize the speech-to-text transcription and OCR text (such as Chinese jieba tokenization), and generate context word vectors (768-dimensional / word) through BERT / Word2Vec; for images, extract visual features from the OCR region images (such as 2048-dimensional image vectors output by ResNet), or generate image patch embeddings using a multi-modal model (such as ViT); Fusion method: Concatenate (such as speech vector + text vector + image vector) or weighted sum the multi-modal vectors, and input them into subsequent neural networks (such as LSTM / Transformer) for joint modeling to capture cross-modal semantic associations (such as the phrase "please look at the whiteboard" in speech corresponding to a specific region in the image). Second, high-level semantic fusion (late fusion): Includes decision-level fusion. The speech-to-text transcription extracts meeting elements (conclusions, to-dos, risks) through NLP, the image OCR text extracts handwritten highlights (such as keywords marked with "★"), and the prior text extracts the agenda structure; Integrate the structured results of different modalities through a rule engine or a graph model (such as a knowledge graph). Third, spatio-temporal attention fusion (intermediate fusion): Transformer-based multi-modal attention mechanism: Input: Time series features of speech (T×D1), position features of images (L×D2), word sequences of text (N×D3); Further model the dependencies between modalities through cross-modal attention, for example: A certain word in the speech ("goal") focuses on the visual features of the corresponding handwritten word in the image to enhance semantic understanding. (3) Structured Meeting Data Generation: First, Meeting Element Extraction and Modeling: First, perform entity extraction: Use a joint extraction model (such as BERT+CRF) to extract six types of core entities from the fused data: Time ("Q3" → "the third quarter of 2025"), Location (online / offline), Person (speaker, responsible person), Task (to-do items), Data ("a 15% increase"), Sentiment Words ("key", "attention required"); Second, perform relationship modeling: Construct a meeting element relationship graph (knowledge graph); Finally, perform structured decoding: Decode the fused features into a preset structured format (JSON / table) through template matching (such as "to-do item = task + responsible person + time") or a sequence-to-sequence model (such as T5).Second, complementary filling of multi-modal content, including missing information completion and format unification: Missing information completion: If words are missed in speech-to-text transcription (e.g., "Q3 target" is not recognized), the handwritten text "Q3 target" in image OCR is used for supplementation; if OCR errors occur due to blurred images ("Zhang San" is recognized as "Zhang Shan"), it is corrected through the speaker label in the speech (the speaking period of Zhang San); Format unification: The texts in different modalities (spoken expressions in speech, handwritten abbreviations, formal terms in PPT) are processed through text normalization (e.g., "Q3" is unified as "the third quarter") to ensure the consistency of structured data. (4) Intelligent analysis and report generation: First, conduct in-depth analysis, including meeting summary generation, keyword and trend analysis, intelligent Q&A and retrieval: For meeting summary generation: Based on multi-modal input, use summary models (such as BART, PEGASUS) to generate multi-dimensional summaries: Decision summary: Extract "meeting conclusions" and "points of contention" (identify high-conflict dialogue segments through sentiment analysis); Action summary: Classify to-do items by responsible person / time node and associate with the handwritten key markings in the image (such as asterisks, underlines); For keyword and trend analysis: Count high-frequency words (combining speech word frequency, handwritten annotation frequency) to generate a keyword cloud (such as "revenue growth", "channel expansion"); Analyze the evolution of topics based on the timeline (such as the topic shift from "current situation analysis" to "solution discussion" and then to "risk assessment"); For intelligent Q&A and retrieval: Build a meeting knowledge base to support natural language queries (such as "What is the Q3 target mentioned by Zhang San?"), and quickly locate the original text segment (speech audio segment + corresponding handwritten note image) through multi-modal retrieval (timestamp + speaker + keyword). Second, generate visual reports, including dynamic time axis, responsibility matrix diagram, sentiment / keyword heat map: For the dynamic time axis, synchronously display the speech waveform, OCR text position, and PPT page-turning record, and click on the time point to view the corresponding image and text; For the responsibility matrix diagram, generate a Gantt chart according to the three dimensions of "responsible person - task - time", and mark the task source with multi-modal evidence (speech segment icon + handwritten image thumbnail); For the sentiment / keyword heat map, overlay the keyword heat (the depth of color represents the mention frequency) on the meeting video screenshot or whiteboard image, and mark the speaker's sentiment tendency (positive / neutral / negative).
[0037] In addition, preferably, the intelligent analysis module for meeting content can also be used for cross-meeting correlation analysis. Conduct cross-meeting correlation analysis in meeting records, aiming to mine information such as logical relationships, theme evolution, task tracking, and decision-making closed-loop among meetings by integrating data from multiple meetings, providing a basis for project management, trend analysis, and decision-making support. Preferably, the intelligent analysis module for meeting content conducts cross-meeting correlation analysis through the following processing methods: (1) Standardize and structure cross-meeting data: First, unify data formats and extract metadata. Specifically, for basic metadata: extract meeting ID, time, duration, participants (name / role), meeting type (regular meeting / special meeting / review meeting, etc.), affiliated project / department, etc.; for content structuring: among them, for text data, extract keywords, named entities (person names, project names, task names, etc.), resolution items (task description, responsible person, deadline, status), problem list, risk points through NLP technology, and for multimodal data (if any): speech-to-text transcription (combining historical meeting speech recognition results), image OCR text (handwritten notes, whiteboard content), attachments (document links, tabular data). Second, align time and space dimensions. Specifically, for timestamp standardization: unify the meeting time into ISO format, and mark the quarter / month / project phase (such as "the 3rd meeting in the requirements phase"); for project / business line association: mark the affiliated project, product line, or business segment for each meeting to ensure data can be grouped and aggregated during cross-project analysis. (2) Core correlation analysis dimensions and methods: First, theme and content association to identify theme evolution and repeated problems among meetings: Specifically, for theme modeling (LDA / BTM): conduct theme modeling on the text of multiple meetings to generate the theme distribution of each meeting (such as "technical solution", "progress risk", "resource coordination"), calculate the theme similarity (cosine similarity) among meetings, and cluster relevant meetings. Example: It is found that the "2025Q1 project kick-off meeting" and the "2025Q2 requirements review meeting" share the "technical architecture" theme, with a similarity of 0.85. For keyword co-occurrence network: construct a cross-meeting keyword co-occurrence graph, with nodes being high-frequency keywords (such as "API interface", "performance optimization") and edges being the frequency of keywords appearing simultaneously in different meetings, visualizing the theme association path. For problem tracking: through named entity recognition and text matching, track the discussion frequency, solution evolution, and change in responsibility attribution of the same problem (such as "server latency") in multiple meetings. Second, conduct task and resolution association for closed-loop management and progress tracking: Specifically, for resolution item mapping: generate a unique identifier for each resolution item (such as "TASK-20250425-01"), and match the update records of the same task across meetings (such as "in the 20250505 meeting, the status of TASK-20250425-01 was updated to 'completed'").For dependency analysis: Identify the dependencies between resolution items through a rules engine or a graph database (such as Neo4j). For example, the resolution of "purchase servers" in Meeting A is a prerequisite for the resolution of "deploy new services" in Meeting B. If the former is delayed, the risk of the latter is automatically marked. In addition, for progress quantification: Calculate the average time taken for a task from creation to closure and the cross-meeting update frequency to evaluate the execution efficiency of the team (such as "the average number of cross-meeting closed loops for requirement change tasks is 3"). Third, perform the association of personnel and roles, and the collaboration network and knowledge flow: Specifically, for the participant association graph: Construct the co-occurrence network of participants across meetings, analyze the core members (frequent attendees), information hubs (bridge figures connecting different meetings), and collaboration blind spots (departments that have never participated in meetings together), and visualize the collaboration network through Gephi. The size of the nodes represents the number of meetings attended, and the edge weight represents the co-occurrence times. For speech pattern analysis: Combine historical speech recognition data to count the speaking duration and keyword preferences of the same person in different meetings (such as "Li Si frequently mentions 'cost control' in cross-departmental meetings"), and identify their roles (decision-makers / executors / technical experts). Fourth, perform the association of time and trends, long-term evolution, and periodic analysis: Specifically, for time series analysis: Sort the meetings by time and draw trend charts of key indicators (such as "the number of risk issues increases with the project phase" and "the proportion of discussions related to customer complaints in the end-of-quarter meetings increases"), and use the Prophet or ARIMA model to predict the high-frequency topics that may appear in future meetings. For periodic pattern recognition: Discover recurring topics (such as "resource allocation is always discussed in the monthly regular meeting") or seasonal issues (such as "Q4 meetings focus on discussing the annual budget") to assist in optimizing the meeting agenda.
[0038] In addition, the intelligent meeting recording system of the present invention may further include a data editing module for editing the meeting data obtained by the meeting information transcription module, so that the user can edit and correct the transcribed text meeting data.
[0039] Preferably, the intelligent meeting recording system of the present invention may further include a meeting management module for classifying and marking meetings, so that the user can classify, mark, and manage the meetings.
[0040] Specifically, the intelligent meeting recording system of the present invention further includes a display module for displaying the meeting record report generated by the meeting content intelligent analysis module.
[0041] In addition, the intelligent meeting recording system of the present invention further includes a database that integrates different meeting environment data, including voice records, visual materials, text summaries, key points, and participant details, etc., to better enhance the intelligent search and cross-meeting association parsing functions of the system.
[0042] Compared with the prior art, the present invention provides an intelligent meeting recording system, which has the following beneficial effects: The intelligent meeting recording system of the present invention includes functional modules such as a meeting information collection module, a data pre-cleaning module, a meeting information transcription module, and a meeting content intelligent analysis module. Through the meeting content intelligent analysis module, the present invention can deeply integrate multi-modal data such as speech, images, and texts, providing users with structured data and intelligent analysis; and the present invention can support cross-meeting correlation analysis to better support users in decision-making; in addition, the present invention can effectively improve the speech recognition accuracy and the handwritten note recognition accuracy. In the above manner, the present invention proposes an intelligent meeting recording system that integrates efficient meeting recording, intelligent information analysis, and meeting the personalized needs of users, which can improve the overall working efficiency of meeting recording.
[0043] It should be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0044] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent meeting recording system, characterized in that, Including: A meeting information collection module, configured to collect meeting information; A meeting information transcription module, configured to perform transcription processing on the meeting information to obtain meeting data; A meeting content intelligent analysis module, configured to perform intelligent analysis processing on the meeting data to generate a meeting record report.
2. The intelligent meeting recording system according to claim 1, characterized in that: It further includes a data pre-cleaning module, configured to perform pre-cleaning processing on the meeting information collected by the meeting information collection module; the meeting information transcription module is further configured to perform transcription processing on the pre-cleaned meeting information.
3. The intelligent meeting recording system according to claim 1, wherein: The meeting information includes voice, image, video, and / or text.
4. The intelligent meeting recording system according to claim 3, wherein: The transcription processing specifically includes performing speech recognition processing on the voice to obtain the meeting data.
5. The intelligent meeting recording system according to claim 3, wherein: The transcription processing specifically includes performing image content parsing and extraction on the image to obtain the meeting data.
6. The intelligent meeting recording system according to claim 3, wherein: The intelligent analysis processing specifically includes performing deep fusion on the meeting data from the meeting information in different formats.
7. The intelligent meeting recording system according to claim 1, characterized in that: The meeting content intelligent analysis module is further configured to perform cross-meeting correlation analysis.
8. The intelligent conference recording system according to claim 1, wherein: It further includes a data editing module, configured to edit the meeting data obtained by the meeting information transcription module.
9. The intelligent meeting recording system according to claim 1, wherein: It further includes a meeting management module, configured to classify and label meetings.
10. The intelligent meeting recording system according to claim 1, characterized in that: It further includes a display module, configured to display the meeting record report generated by the meeting content intelligent analysis module.
Citation Information
Cited By
Intelligent teaching evaluation and diagnosis system and method based on multi-mode audio and video analysis
CN120997010A
Event processing method and system based on large language model
CN121260185A
LLM enhancement-based multi-speaker voice recognition and voiceprint matching system
CN121306145A