Conference summary automatic generation method and system based on OCR technology
Multimodal data is collected through multi-view cameras and microphone arrays, combined with PaddleOCR and CNN-LSTM networks, real-time capture and structured output of conference minutes is achieved, solving the problems of low efficiency in generating traditional conference minutes and missing information, and achieving efficient and accurate multimodal data fusion and automated management.
Patent Information
- Application Number
- CN202510820674.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Traditional conference minutes generation relies on manual recording efficiency and is prone to missing key information. Existing automatic generation technology cannot fully capture visual information and non-verbal signals, lacks real-time fusion and structured output capabilities of multimodal data, and cannot meet the needs of efficient conference management.
By deploying a multi-view camera and microphone array to collect multi-modal data, using PaddleOCR, voice recognition and CNN-LSTM network detection gestures, combining multi-modal time synchronization and dynamic weight allocation models, real-time capture and structured output of conference content is achieved.
It significantly improves the accuracy and generation efficiency of meeting minutes, realizes closed-loop automated management of meeting information, has real-time error correction and intelligent distribution capabilities, and is adapted to various meeting scenarios.
Smart Images

Figure CN120337862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and artificial intelligence, and particularly to a method and system for automatically generating meeting minutes based on OCR technology. Background Art
[0002] Traditional meeting minutes generation relies on manual recording, which is inefficient and prone to missing key information. Existing automatic generation technologies are mostly based on a single modality, such as speech recognition, and cannot comprehensively capture visual information in meetings, such as whiteboard content and projection documents, and non-verbal signals, such as speakers' gestures. For example, pure speech recognition may result in transcription errors due to accents and background noise, while static text analysis is difficult to identify dynamic emphasis content. In addition, existing technologies lack the ability to fuse multi-modal data in real time and output structured data, and cannot meet the needs of efficient meeting management. Summary of the Invention
[0003] The object of the present invention is to overcome one or more deficiencies of the prior art, and provide a method and system for automatically generating meeting minutes based on OCR technology. By integrating PaddleOCR, speech recognition, and gesture action detection, real-time capture of meeting content, multi-modal data fusion, and structured output are realized, significantly improving the accuracy and generation efficiency of meeting minutes.
[0004] The object of the present invention is achieved by the following technical solutions:
[0005] A method for automatically generating meeting minutes based on OCR technology, comprising the following steps:
[0006] Synchronously collect a visual image stream including whiteboard content, projection images, and speakers' gestures, and a surrounding audio stream through a multi-view camera and a microphone array deployed in the venue;
[0007] Perform PaddleOCR text detection, recognition, and layout analysis on the visual image stream in sequence to obtain structured text data; perform end-to-end speech recognition and entity relationship extraction on the surrounding audio stream to generate timestamped speech text; gesture action detection step: perform real-time detection of speakers' gestures through a CNN-LSTM network model, and mark the gesture action types and time nodes corresponding to key content;
[0008] Multi-modal data processing: Align visual text, speech text, and gesture marked data based on a multi-modal time synchronization algorithm, and use an attention mechanism to construct a dynamic weight allocation model, denoted as Attention weight, is the text embedding vector recognized by OCR, is the standard speech text embedding vector; is the embedding dimension, and the corresponding formula is: , where is the weight matrix, is the bias term, is the sigmoid function, and then automatically adjusts the fusion priority of the OCR text and the speech text in the corresponding period according to the gesture action type. Let be the fused text representation, and the calculation formula is: ;
[0009] According to the predefined structured semantic template, automatically fill the fused multi-modal data into the structured fields of the topic classification, decision-making matters, and action item list to generate a meeting minutes document with key annotations.
[0010] Furthermore, in the gesture action detection step, the CNN-LSTM network model is trained in the following manner:
[0011] Use the Jester public gesture dataset for pre-training to extract gesture spatio-temporal features;
[0012] Use the custom gesture data in the meeting scenario for transfer learning to optimize the parameters of the action classifier and update the parameters , where is the first-order matrix estimate of the gradient at the th iteration, is the second-order estimate of the gradient, is the learning rate, is a small constant;
[0013] The custom gesture data includes: pointing, circle selection, erasing, and zooming gestures;
[0014] The output result includes the confidence score of the gesture action and the corresponding image region coordinates.
[0015] Furthermore, the multi-modal time synchronization algorithm is specifically as follows:
[0016] Through the hardware clock calibration module of the camera and the microphone, control the timestamp error between the visual frame and the audio frame within the set range;
[0017] Based on the dynamic time warping algorithm, perform non-linear alignment on the OCR text generation time, speech recognition time, and gesture detection time.
[0018] Furthermore, the construction method of the dynamic weight allocation model includes:
[0019] When a "circle selection" or "pointing" gesture is detected, increase the weight coefficient of the OCR text in the corresponding area to the set multiple of the speech text;
[0020] When an "erasing" gesture is detected, trigger the filtering mechanism for the data in the corresponding period;
[0021] The weight coefficient is dynamically adjusted by a reinforcement learning model trained with meeting historical data.
[0022] Furthermore, the structured semantic template can be customized by users, including:
[0023] Configurable field mapping rules: Automatically map the entity types in multimodal data to the template fields;
[0024] Adjustable format output rules: Generate a rich text format document with a table of contents index, key points highlighted in red, and attachment associations.
[0025] Furthermore, it also includes a real-time error correction mechanism:
[0026] Perform context semantic verification on the OCR recognition results, detect text coherence through the BERT model. Let the sentence , take the CLS token vector output by BERT as the sentence vector, and the cosine similarity calculation formula: , trigger secondary recognition for recognition results with a confidence level less than the set value, and the calculation formula: , where is a hyperparameter for adjusting the slope;
[0027] Among them, is the OCR recognition sentence to be verified, is the context reference sentence, is the sentence 's semantic vector (the CLS token sentence vector output by the BERT model), is the sentence 's semantic vector (the sentence vector of the context reference sentence with the same dimension as );
[0028] Perform speaker separation processing on the speech recognition results, and correct cross-speaker transcription errors in combination with the meeting room seat layout information.
[0029] Furthermore, the multimodal data processing process is deployed on edge computing devices, specifically:
[0030] Use NVIDIA Jetson AGX Orin as an edge node to achieve local real-time inference and end-to-end processing latency;
[0031] Encrypt and transmit the original data to the cloud through the network for model update and historical data archiving.
[0032] Furthermore, the PaddleOCR processing module contains a trainable domain adaptation model:
[0033] For professional fields, fine-tune the text detection branch and recognition branch of PP-OCRv4 with a small amount of labeled data;
[0034] Perform end-to-end recognition in a mixed handwritten and printed text scenario.
[0035] Furthermore, after the meeting minutes document is generated, it also includes an intelligent distribution step:
[0036] Extract key participant information from the document through NLP technology, automatically generate to-do reminder and push it to the enterprise collaboration platform;
[0037] Generate an audio summary file with a timestamp index and associate it with the corresponding paragraph of the meeting minutes.
[0038] A meeting minutes automatic generation system based on OCR technology, the system includes:
[0039] Multimodal acquisition terminal: including a multi-view camera with 4K resolution, an array microphone and a hardware clock synchronization module;
[0040] Edge computing processing unit: integrated with PaddleOCR engine, PaddleSpeech speech recognition engine and custom gesture detection engine;
[0041] Intelligent output module: including a semantic template engine, a rich text generator and a cross-platform API interface (supporting integration with Feishu / DingTalk / Outlook);
[0042] Each module of the system is connected through a time synchronization bus to achieve nanosecond-level time alignment of multimodal data.
[0043] The beneficial effects of the present invention are:
[0044] (1) Realize the intelligent generation of meeting minutes through multimodal data fusion technology: use hardware clock calibration and dynamic time warping algorithm to ensure nanosecond-level synchronization of multi-source data, combine CNN-LSTM network to detect the gestures of speakers in real time and dynamically adjust the weights of OCR and speech texts, and strengthen the priority of key information;
[0045] (2) Automatically generate structured documents through configurable semantic templates and intelligent field mapping, significantly improving the standardization; use BERT model to verify the coherence of OCR, correct speech errors in combination with seat layout, and achieve low-latency processing based on edge computing devices to ensure reliability;
[0046] (3)Fine-tune the OCR model for professional scenarios, support mixed recognition of handwritten or printed text, and optimize the generalization ability of gesture detection through transfer learning; The whole process covers data collection, processing, error correction, minutes generation, and intelligent distribution, realizing the closed-loop automation of meeting information from capture to management, with high accuracy, high efficiency, and strong adaptability. Description of the Drawings
[0047] Figure 1 The flowchart of the steps of a method for automatically generating meeting minutes based on OCR technology provided in Embodiment 1;
[0048] Figure 2 The structural diagram of a system for automatically generating meeting minutes based on OCR technology provided in Embodiment 1;
[0049] Figure 3 The flowchart of the execution steps of a method and system for automatically generating meeting minutes based on OCR technology provided in Embodiment 2;
[0050] Figure 4 The flowchart of the execution steps of a method and system for automatically generating meeting minutes based on OCR technology provided in Embodiment 2. Detailed Embodiments
[0051] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0052] Embodiment 1
[0053] Refer to Figure 1 , a method for automatically generating meeting minutes based on OCR technology, includes the following steps:
[0054] Synchronously collect the visual image stream including whiteboard content, projection screen, and speaker gestures and the surrounding audio stream through multi-view cameras and microphone arrays deployed in the venue;
[0055] Perform PaddleOCR text detection, recognition, and layout analysis on the visual image stream in sequence to obtain structured text data; perform end-to-end speech recognition and entity relationship extraction on the surrounding audio stream to generate timestamped speech text; Gesture action detection step: Real-time detect the speaker's gestures through the CNN-LSTM network model, and mark the gesture action types and time nodes corresponding to the key content;
[0056] Multimodal Data Processing: Align visual text, speech text, and gesture marker data based on a multimodal time synchronization algorithm, construct a dynamic weight allocation model using an attention mechanism, and automatically adjust the fusion priority of OCR text and speech text in the corresponding time period according to the gesture action type;
[0057] According to a predefined structured semantic template, automatically fill the fused multimodal data into the structured fields of the topic classification, decision-making items, and action item list, and generate a meeting minutes document with key annotations.
[0058] In the gesture action detection step, the CNN-LSTM network model is trained in the following way:
[0059] Use the Jester public gesture dataset for pre-training to extract gesture spatio-temporal features;
[0060] Use custom gesture data in the meeting scenario for transfer learning to optimize the parameters of the action classifier;
[0061] The custom gesture data includes: pointing, circle selection, erasing, and zooming gestures;
[0062] The output result includes the confidence score of the gesture action and the corresponding image region coordinates.
[0063] The multimodal time synchronization algorithm is specifically as follows:
[0064] Through the hardware clock calibration module of the camera and microphone, control the timestamp error between the visual frame and the audio frame within the set range;
[0065] Based on the dynamic time warping algorithm, perform non-linear alignment on the OCR text generation time, speech recognition time, and gesture detection time.
[0066] The construction method of the dynamic weight allocation model includes:
[0067] When a "circle selection" or "pointing" gesture is detected, increase the weight coefficient of the OCR text in the corresponding area to the set multiple of the speech text;
[0068] When an "erasing" gesture is detected, trigger the filtering mechanism for the data in the corresponding time period;
[0069] The weight coefficient is dynamically adjusted by a reinforcement learning model trained with meeting historical data.
[0070] The structured semantic template supports user-defined configuration, including:
[0071] Configurable field mapping rules: Automatically map the entity types in the multimodal data to the template fields;
[0072] Adjustable format output rules: Support generating rich text format documents with table of contents indexing, key points highlighted in red, and attachment associations.
[0073] During the execution of this method, a real-time error correction mechanism is also set up:
[0074] Perform context semantic verification on the OCR recognition results, detect text coherence through the BERT model, and set sentences , and take the CLS token vector output by BERT as the sentence vector, and the cosine similarity calculation formula: , trigger secondary recognition for recognition results with a confidence level less than the set value, and the calculation formula: , where is the hyperparameter for adjusting the slope;
[0075] Among them, is the OCR recognition sentence to be verified, is the context reference sentence, is the sentence 's semantic vector (the CLS token sentence vector output by the BERT model), is the sentence 's semantic vector (the sentence vector of the context reference sentence with the same dimension as );
[0076] Perform speaker separation processing on the speech recognition results, and correct cross-speaker transcription errors by combining the meeting room seat layout information.
[0077] The multi-modal data processing process is deployed on edge computing devices, specifically:
[0078] Use NVIDIA Jetson AGX Orin as the edge node to achieve local real-time inference and end-to-end processing latency;
[0079] Encrypt and transmit the original data to the cloud through the network for model update and historical data archiving.
[0080] The PaddleOCR processing module contains a trainable domain adaptation model:
[0081] For professional fields, fine-tune the text detection branch and recognition branch of PP-OCRv4 with a small amount of labeled data;
[0082] Perform end-to-end recognition in the mixed scenario of handwritten and printed texts.
[0083] After the meeting minutes document is generated, it also includes intelligent distribution steps:
[0084] Extract key participant information from documents through NLP technology, automatically generate to-do reminders and push them to the enterprise collaboration platform;
[0085] Support the generation of audio summary files with timestamp indexes and associate them with the corresponding paragraphs of the meeting minutes.
[0086] See Figure 2 , provided a meeting minutes automatic generation system based on OCR technology, the system includes:
[0087] Multimodal acquisition terminal: includes a multi-view camera with 4K resolution, an array microphone and a hardware clock synchronization module;
[0088] Edge computing processing unit: integrated with PaddleOCR engine, PaddleSpeech speech recognition engine and custom gesture detection engine;
[0089] Intelligent output module: includes a semantic template engine, a rich text generator and a cross-platform API interface (supporting integration with Feishu / DingTalk / Outlook);
[0090] Each module of the system is connected through a time synchronization bus to achieve nanosecond-level time alignment of multimodal data.
[0091] Embodiment 2
[0092] In the meeting room, 3 high-definition cameras each perform their own duties, one is aimed at the main podium projection screen, and the other two are aimed at the whiteboards on the left and right sides respectively; an 8-channel ring array microphone is evenly embedded in the edge of the meeting table; the edge computing device is placed in the corner of the meeting room and connected to other devices through network cables; the system server is deployed in the company's computer room to ensure stable operation. This scenario setting lays a foundation for comprehensively and accurately collecting meeting data and provides environmental support for realizing efficient automatic generation of meeting minutes.
[0093] See Figures 3 - 4 , the specific implementation steps are as follows:
[0094] The multimodal data acquisition process includes:
[0095] S1. Hardware deployment and parameter configuration:
[0096] S1.1 Camera Deployment: When installing the camera for shooting the projection screen, the staff will repeatedly view the live video through the management interface and continuously fine-tune the angle and focal length until the projected content completely covers the screen without distortion and the text is clear and legible. At the same time, the settings of high-definition resolution and frame rate can ensure that even if the projected images are switched quickly, there will be no problems such as blurred images or trailing shadows, thus completely capturing the text, charts and other information on each page of the projection. For the camera shooting the whiteboard, during installation, not only should the entire whiteboard be within the camera's field of view, but also the pitch angle and distance of the camera should be precisely adjusted according to the size and height of the whiteboard to ensure that the writing content at any position on the whiteboard can be clearly captured, avoiding the situation where some text cannot be recognized due to the viewing angle problem, and significantly improving the comprehensiveness and accuracy of the meeting visual information collection.
[0097] S1.2 Microphone Deployment: After installing the circular array microphone at the edge of the meeting table, on-site audio collection tests will be carried out. The staff will speak at different positions and volumes in the meeting room, and observe the volume and sound quality of each channel through the monitoring interface of the audio processor, and finely adjust the automatic gain control and noise suppression parameters. For example, when it is found that the sound in a certain corner is weak, the gain of that channel will be appropriately increased; if the environmental noise is not filtered thoroughly, the noise suppression threshold will be further optimized. In this way, the collected audio is clear and stable, effectively reducing the interference of background noise on speech recognition, greatly improving the accuracy of speech content collection, and providing high-quality audio data for accurately transcribing the meeting speech subsequently.
[0098] S1.3 Clock Synchronization: After starting the clock calibration program, the system will continuously monitor the deviation between the time of the camera and the microphone and the standard time source. If a slight deviation in the time of a certain device is found, a synchronization signal will be automatically sent again for calibration to ensure that before the meeting starts, the time error of all devices is controlled within a very small range. This multiple calibration and real-time monitoring mechanism ensures a high degree of consistency in time for visual and audio data, provides a solid guarantee for the accurate fusion of subsequent multi-modal data, avoids information confusion caused by time asynchronization, and greatly improves the accuracy and integrity of the meeting information collection.
[0099] S2. Data Acquisition Process:
[0100] Before the S2.1 meeting starts, the operator will conduct a comprehensive and detailed inspection in the system management interface. In addition to checking the camera images and the microphone's sound collection status, the operator will also check whether the network connection is stable and whether the device storage space is sufficient. If it is found that the camera images are stuck, the operator will immediately troubleshoot network or device problems, such as restarting the camera or checking the network cable connection. If there is noise in the microphone, the operator will readjust the audio processor settings or check whether the microphone is interfered. Through these preventive inspections and treatments, potential faults in the data collection process are eliminated at the source, effectively ensuring the reliability and comprehensiveness of the meeting data collection.
[0101] After the "Start Recording" button is clicked, the data collection device will respond quickly. While starting to collect data, the camera and the microphone will automatically generate their respective data collection logs, recording information such as the start time and the device status. During the data transmission process, a redundant transmission mechanism is adopted, that is, the same data is transmitted multiple times through different network paths. If a transmission path fails, the system can automatically switch to the backup path to ensure that the data is not lost or interrupted. This reliable data transmission method provides double protection for the real-time and complete collection of meeting data, enabling every image and sound during the meeting to be transmitted to the edge computing device in a timely and accurate manner, providing a solid data foundation for the real-time generation of meeting minutes.
[0102] After receiving the data, the edge computing device will conduct a preliminary verification of the temporary files. For example, it will check whether the number of frames in the video file is continuous and whether the duration of the audio file matches the timestamp. If it is found that the file is damaged or data is missing, a retransmission request will be immediately sent to the collection device to ensure that the stored temporary files are complete and usable. In addition, the temporary files will be classified and named according to the file type and the collection time, facilitating subsequent quick retrieval and processing. This strict data verification and management method further improves the accuracy and integrity of the data collection, providing high-quality data resources for the subsequent data processing link.
[0103] The multi-modal data processing process includes:
[0104] S3. OCR text recognition and processing:
[0105] S3.1 Text Detection: When the edge computing device reads the camera image stream frame by frame, to improve the detection efficiency, a block detection strategy is adopted. That is, each frame of the image is divided into multiple small blocks, and text region detection is performed on each small block in parallel. During the detection process, for images with complex backgrounds, such as text regions with pattern backgrounds in projection screens, the OCR recognition model will use special image preprocessing algorithms to enhance the contrast between the text and the background, making the text region easier to detect. At the same time, when screening invalid text boxes, in addition to considering the area and aspect ratio, the texture features within the text box will also be analyzed. If it is judged as a non-text pattern region, it will be excluded. Through these optimization measures, the accuracy and efficiency of text region detection have been greatly improved, ensuring that all the text content in the conference visual information is accurately recognized without omission, and significantly enhancing the comprehensiveness of conference text information collection.
[0106] S3.2 Character Recognition: When preprocessing the text region image, in addition to conventional size adjustment and noise reduction, adaptive processing will also be performed for different fonts, font sizes, and writing styles. For example, for handwritten fonts, a special stroke enhancement algorithm will be used to highlight the stroke features of the text, facilitating the recognition network to better extract text features; for small-font text, image magnification and super-resolution processing will be performed to improve text clarity. During the recognition process, when encountering characters with low confidence, the system will not only mark them for verification but also combine the context information where the characters are located to obtain parameters , and then perform semantic speculation and candidate character matching. For example, if "com[low-confidence character]er" is recognized, based on the context, it can be speculated that the character may be "put", and then the character "put" will be recognized and verified again. This multi-strategy character recognition and verification mechanism effectively improves the accuracy of character recognition. Even in complex conference scenarios, various text contents can be accurately recognized, further enhancing the accuracy of conference information collection.
[0107] S3.3 Layout Analysis: When the layout analysis tool is working, it will comprehensively use a variety of analysis methods. In addition to classification based on rules and machine learning, visual features such as the layout and spacing of text will also be considered. For example, for list items, features such as the distance between the bullet point and the text and the indentation between list items will be analyzed to accurately judge; for table content, the table structure will be recognized by detecting horizontal and vertical lines and the distribution law of text. When annotating the hierarchical relationship of text and paragraph attribution, the conference agenda information and semantic understanding results will be combined. For example, if a certain paragraph starts with "Next, discuss the project progress" and corresponds to the "Project Progress Discussion" session in the agenda, it will be attributed to this agenda item. Through this detailed layout analysis and content classification, the generated structured text data is well-organized, providing a good foundation for generating high-quality meeting minutes later, and greatly enhancing the structured degree and readability of the meeting minutes.
[0108] S4. Speech Recognition and Semantic Analysis:
[0109] S4.1 Audio Transcription: When denoising the audio collected by the microphone, a combination of multiple denoising algorithms is used. First, a denoising algorithm based on a statistical model is used to remove common stationary background noise; then, for sudden short-term noises such as door opening / closing sounds and mobile phone ringtones, a denoising method based on time-domain analysis is used for processing. In a speech recognition system, to improve the adaptability to speeches with different speaking speeds and accents, a large amount of speech data with different styles is used to train and optimize the model. When generating speech text, grammar checking and correction are performed on the text in real time, such as automatically adding punctuation marks and correcting word order errors. These measures make the audio transcription more accurate and fluent. Even in a complex conference speaking environment, speech can be quickly and accurately converted into text, providing efficient and reliable speech content support for the real-time generation of meeting minutes, and significantly improving the efficiency of meeting information processing.
[0110] S4.2 Entity Relationship Extraction: When the semantic analysis model extracts key entities and analyzes entity relationships, a large amount of knowledge graphs in the conference domain are combined. The knowledge graph contains various conference-related concepts, relationships, and attributes, such as the relationships between "project" and "person in charge" and "time node". When processing speech text, the model matches and infers the information in the text with the knowledge graph. For example, when it is recognized that "Zhang San is responsible for developing new functions", through the knowledge graph, it can be known that "Zhang San" is the "person in charge" and "developing new functions" is the "task", and their relationship is automatically established. At the same time, the model continuously learns new conference data and updates the knowledge graph to meet the needs of different types of conferences. This entity relationship extraction method based on the knowledge graph can deeply understand the conference speech content, accurately extract key information, make the meeting minutes content more refined and accurate, highlight the key points of the meeting, and greatly improve the quality and practicality of the meeting minutes.
[0111] S5. Gesture Action Detection:
[0112] S5.1 Model Training and Deployment:
[0113] S5.1.1 Pre-training: When pre-training using a public gesture dataset, in order to enable the model to better learn the spatio-temporal features of gestures, data augmentation techniques are employed. For example, operations such as randomly rotating, scaling, and adding noise to video data are carried out to expand the diversity of training data. During the training process, the training status of the model is monitored in real time, such as changes in the loss function value and the improvement of accuracy. If it is found that the model shows overfitting, that is, the model performs well on the training set but its performance deteriorates on the validation set, the training parameters will be adjusted in a timely manner, such as reducing the learning rate and increasing the regularization term. Through these optimization measures, the pre-trained model can fully learn the common features of various gestures, laying a solid foundation for accurate recognition in the meeting scenario and effectively improving the generalization ability and accuracy of the gesture detection model.
[0114] S5.1.2 Fine-tuning: When collecting gesture video data from company internal meetings, various gesture actions in different speaker and different meeting scenarios are covered. When annotating the data, not only the gesture types are annotated, but also information such as the start and end times and the amplitude of the gesture is recorded in detail. During the fine-tuning process, in order to accelerate the training speed and improve the training effect, an optimization strategy of transfer learning is adopted. That is, some network layer parameters related to general gesture feature extraction in the pre-trained model are frozen, and only the network layers related to the recognition of specific gestures in the meeting scenario are trained. At the same time, according to the characteristics of the training data, the hyperparameters of the model are adjusted, such as batchsize and the number of training epochs. After fine-tuning, the model can accurately recognize common gesture actions in the meeting, providing a reliable guarantee for accurately marking the key content of the meeting and further enhancing the comprehensiveness and accuracy of meeting information collection.
[0115] S5.2 Real-time detection process: After the camera image is input into the gesture recognition module, in order to improve the detection speed, a cascade detection architecture is adopted. First, a lightweight network model is used to quickly screen the image to determine whether there is a possible gesture area; if a suspicious area is detected, a more complex and accurate model is used for detailed detection and recognition. When extracting the spatio-temporal features of the gesture, a time series analysis method is combined. Not only the gesture image features of the current frame are analyzed, but also the gesture changes in the previous and next few frames are considered to more accurately judge the type of gesture action. For example, for the "circling" gesture, the movement trajectory and shape changes of the gesture from the starting position to the ending position are observed. This real-time, efficient and accurate gesture detection process can capture every key gesture of the speaker in a timely manner and cooperate seamlessly with other modal data, enabling the meeting minutes to completely record the non-verbal key emphasized content in the meeting and significantly enhancing the comprehensiveness and effectiveness of meeting information collection.
[0116] The multi-modal data fusion process includes:
[0117] S6. Time Synchronization: When performing time synchronization, to address time deviations caused by factors such as network transmission delays, a dynamic time compensation algorithm is adopted. This algorithm predicts the time delay during data transmission based on historical data and the current network conditions, and adjusts the timestamp in advance. Meanwhile, during the data fusion process, continuous time consistency verification is carried out on the synchronized data. For example, at regular intervals, it is checked whether the content of different modality data at the same time point is logically consistent. If inconsistencies are found, time alignment and adjustment are performed again. Through this dynamic adjustment and continuous verification mechanism, a high degree of consistency of multi-modal data in the time dimension is ensured, enabling the fused data to truly and accurately reflect the actual situation of the meeting, providing strong support for generating accurate meeting minutes, and effectively improving the accuracy and reliability of meeting information processing.
[0118] S7. Dynamic Weight Allocation:
[0119] S7.1 When the "selection" gesture is detected, the dynamic weight allocation mechanism further analyzes the selection range and duration. If the selection range is large and the duration is long, it indicates that this part of the content may be very important, and the weight of the corresponding region OCR text will be increased accordingly. At the same time, to make the weight allocation more reasonable, the frequency and importance of the text in this region in the speech text are also considered for comprehensive judgment. For example, if the text in the selected region is emphasized multiple times in the speech, the weight will be increased again. This refined weight allocation method can highlight the key information in the meeting more effectively, making the meeting minutes more focused on the key content, greatly improving the readability and practicality of the meeting minutes, and effectively enhancing the accuracy and effectiveness of meeting information processing.
[0120] S7.2 After the "erasure" gesture is detected, in addition to marking and removing the corresponding period data, the system also performs correlation analysis on the relevant data before and after. If it is found that the erased content has supplementary explanations or corrections in the subsequent voice or visual information, the relevant information will be retained and integrated. For example, if a certain piece of text on the whiteboard is erased and the speaker mentions in the speech "the part that was just erased should be modified like this", the system will correlate the subsequent modification content with the previously erased content to ensure the integrity and accuracy of the meeting minutes content. In addition, the generated filtering log will detail the relevant information of the erasure gesture, such as the time, location, and summary of the erased content, facilitating subsequent review and auditing, and further enhancing the reliability and traceability of meeting information processing.
[0121] When the S7.3 system analyzes historical meeting data to adjust the parameters of the dynamic weight allocation model, it will adopt a variety of data analysis methods. In addition to counting the association frequencies between different gestures and important information, it will also use clustering analysis methods to classify similar meeting scenarios and gesture usage patterns, and formulate personalized weight allocation strategies for different types of meetings. For example, for technical solution discussion meetings and project progress report meetings, different weight adjustment rules will be set according to their respective characteristics and common gesture usage habits. At the same time, meeting organizers and those who often use meeting minutes will be invited to participate in the evaluation, and the parameters will be further optimized according to their feedback. Through this multi-dimensional and multi-agent participation optimization method, the dynamic weight allocation model can continuously adapt to various meeting scenarios, improve the quality and efficiency of meeting information processing, and make the meeting minutes better meet the user's needs.
[0122] The structured output process includes:
[0123] S8. Semantic template configuration:
[0124] The commonly used meeting minutes templates preset by the S8.1 system for technology companies fully consider the characteristics and requirements of technology industry meetings during design. Each section in the template is carefully planned. For example, the "Basic Meeting Information" section contains necessary fields such as meeting name, time, location, host, and participants, facilitating a quick understanding of the meeting overview; the "Agenda Discussion" section sets sub-fields such as topics, discussion content, and relevant personnel's speeches in the order of the meeting agenda, enabling a detailed record of the meeting discussion process. This preset template provides a scientific and reasonable basic structural framework for the meeting minutes, making the generated meeting minutes have a unified format and specification, greatly improving the standardization and readability of the meeting minutes, and facilitating users to quickly read and manage the meeting minutes.
[0125] When users make modifications in the template editing interface, the S8.2 system provides an intuitive and convenient operation method. For example, fields can be added or deleted through simple drag-and-drop operations, and the name and format settings can be modified by double-clicking on the field. When adding custom fields, the system will automatically prompt relevant setting options, such as field type (text, date, number, etc.), whether it is required, and display order. After the modification is completed, the system will preview the template effect in real time, facilitating users to make adjustments. This flexible and easy-to-use template customization function meets the personalized needs of different users and different meeting types for the minutes format. Whether it is a small internal department meeting or a large cross-departmental project meeting, meeting minutes that meet the requirements can be generated through customizing the template, significantly enhancing the adaptability and flexibility of the system.
[0126] When setting field mapping rules in S8.3, the system adopts a combination of intelligent matching and manual adjustment. First, the system automatically performs preliminary field mapping based on the types and semantic features of entities in the multimodal data. For example, recognized time entities are automatically mapped to time-related fields in the template. For some complex or ambiguous cases, users can manually adjust the mapping relationship. At the same time, the system also supports batch setting of mapping rules. For entities of the same type, they can be set to multiple corresponding fields at one time. In addition, the system saves the mapping rules set by users. When encountering similar meeting data next time, it will automatically apply the set rules, improving the efficiency and accuracy of meeting minutes generation and making the content of the meeting minutes more structured and organized.
[0127] The process of obtaining the document in S9 includes:
[0128] When the system fills the fused multimodal data in the template format in S9.1, it will perform strict data verification and format conversion. For data obtained from different modalities, it will check its integrity and accuracy, such as checking whether the text recognized by OCR is complete and whether there are missing paragraphs in the text recognized by speech. During the filling process, the data will be converted accordingly according to the format requirements of the template fields. For example, time data will be converted to the specified date and time format, and list item data will be typeset according to the list format of the template. At the same time, to ensure the accuracy of data filling, the filled content will be checked again to ensure that each field is filled with the corresponding data accurately. This rigorous data processing and filling method makes the generated meeting minutes complete, accurate, well-structured, improving the quality and readability of the meeting minutes and facilitating users to comprehensively understand the meeting situation.
[0129] When automatically highlighting key content in S9.2, the system not only relies on gesture markings and voice emphasis but also combines the results of semantic analysis. For example, for key decisions, important goals, etc. mentioned in the meeting, even without obvious gesture or voice emphasis, they will be highlighted according to semantic importance. When automatically associating attachments, the system will identify the relevance between the attachments and the meeting content. For example, if a certain page in a PPT file is relevant to a certain topic discussed in the meeting, it will accurately associate that page of the PPT at the corresponding position in the meeting minutes. When adding a timestamp-indexed audio summary, it will automatically generate reasonable index nodes according to the chapter division and key content distribution of the meeting content. For example, indexes are set at the beginning and end of each topic discussion to facilitate users to quickly locate the audio clips they are interested in. The comprehensive application of these functions further highlights the key points of the meeting, enriches the form and content of the meeting minutes, and greatly improves the convenience of using the meeting minutes.
[0130] The method described in this embodiment realizes the intelligent generation of meeting minutes through multi-modal data fusion technology (integrating visual, voice, and gesture signals): using hardware clock calibration and dynamic time warping algorithms to ensure nanosecond-level synchronization of multi-source data, combining CNN-LSTM networks to detect speaker gestures (such as selection and erasure) in real time to dynamically adjust the weights of OCR and speech texts, and strengthening the priority of key information; automatically generating structured documents (including table of contents index, key points highlighted in red, and attachment association) through configurable semantic templates and intelligent field mapping, significantly improving the standardization; using the BERT model to verify the coherence of OCR, correcting speech errors in combination with the seat layout, and achieving low-latency processing based on edge computing devices to ensure reliability; fine-tuning the OCR model for professional scenarios, supporting mixed recognition of handwritten / printed characters, and optimizing the generalization ability of gesture detection through transfer learning; covering the entire process from data collection, processing, and error correction to minutes generation and intelligent distribution (such as to-do push, audio summary), realizing the closed-loop automation of meeting information from capture to management, with high precision, high efficiency, and strong adaptability.
[0131] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and modifications made by those skilled in the art that do not depart from the spirit and scope of the present invention shall all be within the protection scope of the appended claims of the present invention.
Claims
1. A method for automatically generating meeting minutes based on OCR technology, characterized in that, It includes the following steps: Synchronously collect the visual image stream containing whiteboard content, projection screen, and speaker gestures, as well as the surrounding audio stream, through multi-view cameras and microphone arrays deployed at the venue; Perform PaddleOCR text detection, recognition, and layout analysis on the visual image stream in sequence to obtain structured text data; perform end-to-end speech recognition and entity relationship extraction on the surrounding audio stream to generate timestamped speech text; Gesture action detection step: Real-time detect the speaker's gestures through the CNN-LSTM network model, and mark the gesture action types and time nodes corresponding to key content; Multimodal data processing: Align visual text, speech text, and gesture marker data based on a multimodal time synchronization algorithm. Use the attention mechanism to construct a dynamic weight allocation model, denoted as the attention weight, is the text embedding vector recognized by OCR, is the standard speech text embedding vector; is the embedding dimension, and the corresponding formula is: , where is the weight matrix, is the bias term, is the sigmoid function, and then automatically adjust the fusion priority of the OCR text and speech text in the corresponding period according to the gesture action type. Let be the fused text representation, and the calculation formula is: ; According to the predefined structured semantic template, automatically fill the fused multi-modal data into the structured fields of topic classification, decision-making matters, and action item lists, and generate a meeting minutes document with key annotations.
2. The automatic generation method of meeting minutes based on OCR technology according to claim 1, characterized in that, In the gesture action detection step, the CNN-LSTM network model is trained in the following way: Use the Jester public gesture dataset for pre-training to extract gesture spatio-temporal features; Use the custom gesture data in the meeting scenario for transfer learning, optimize the parameters of the action classifier, and update the parameters , where is the first-order matrix estimate of the gradient at the th iteration, is the second-order estimate of the gradient, is the learning rate, is a small constant; The custom gesture data includes: pointing, circle selection, erasing, and zooming gestures; The output result includes the confidence score of the gesture action and the corresponding image region coordinates.
3. The automatic generation method of meeting minutes based on OCR technology according to claim 1, characterized in that, The multi-modal time synchronization algorithm is specifically as follows: Through the hardware clock calibration module of the camera and microphone, control the timestamp error between the visual frame and the audio frame within the set range; Based on the dynamic time warping algorithm, perform non-linear alignment on the OCR text generation time, speech recognition time, and gesture detection time.
4. A method for automatically generating meeting minutes based on OCR technology according to claim 1, characterized in that, The construction method of the dynamic weight allocation model includes: When the "circle selection" or "pointing" gesture is detected, increase the weight coefficient of the OCR text in the corresponding area to the set multiple of the speech text; When the "erasing" gesture is detected, trigger the filtering mechanism for the data in the corresponding period; The weight coefficient is dynamically adjusted by the reinforcement learning model trained with historical meeting data.
5. A method for automatically generating meeting minutes based on OCR technology according to claim 1, characterized in that, The structured semantic template can be customized by the user and includes: Configurable field mapping rules: Automatically map the entity types in the multi-modal data to the template fields; Adjustable format output rules: Generate a rich text format document with a table of contents index, key highlighting, and attachment association.
6. The automatic generation method of meeting minutes based on OCR technology according to claim 1, characterized in that, There is also a real-time error correction mechanism: Perform context semantic verification on the OCR recognition results, detect text coherence through the BERT model, and set the sentence , and take the CLS token vector output by BERT as the sentence vector. The cosine similarity calculation formula is: . Trigger secondary recognition for recognition results with a confidence level less than the set value. The calculation formula is: , where is a hyperparameter for adjusting the slope; Among them, is the OCR recognition sentence to be verified, is the context reference sentence, is the sentence 's semantic vector, is the sentence 's semantic vector; Perform speaker separation processing on the speech recognition results, and correct the cross-speaker transcription errors in combination with the meeting room seating layout information.
7. A method for automatically generating meeting minutes based on OCR technology according to claim 1, characterized in that In the process of multi-modal data processing, it is deployed on edge computing devices, specifically: Use NVIDIA Jetson AGX Orin as the edge node to achieve local real-time inference and end-to-end processing delay; Encrypt and transmit the original data to the cloud through the network for model update and historical data archiving.
8. The automatic generation method of meeting minutes based on OCR technology according to claim 1, characterized in that The PaddleOCR processing module contains a trainable domain adaptation model: For professional fields, fine-tune the text detection branch and recognition branch of PP-OCRv4 through labeled data; Perform end-to-end recognition in the mixed scenario of handwritten and printed texts.
9. A method for automatically generating meeting minutes based on OCR technology according to claim 1, characterized in that, After the meeting minutes document is generated, there is also an intelligent distribution step: Extract the key participant information in the document through NLP technology, automatically generate to-do reminder and push it to the enterprise collaboration platform; Generate an audio summary file with a timestamp index and associate it with the corresponding paragraphs of the meeting minutes.
10. A meeting minutes automatic generation system based on OCR technology, which uses a meeting minutes automatic generation method according to any one of claims 1-9, is characterized in that, The system includes: Multimodal acquisition terminal: including multi-view cameras, array microphones, and a hardware clock synchronization module; Edge computing processing unit: integrating the PaddleOCR engine, PaddleSpeech speech recognition engine, and a custom gesture detection engine; Intelligent output module: including a semantic template engine, a rich text generator, and a cross-platform API interface; Each module of the system is connected through a time synchronization bus to achieve nanosecond-level time alignment of multimodal data.
Citation Information
Patent Citations
Conference summary generation method and device, computer equipment and storage medium
CN112466306A
Engineering surveying and mapping pavement flatness detection device and method
CN119642751A
Conference summary generation method and device, electronic equipment and storage medium
CN119720980A
An innovative system and method for automatically generating meeting minutes and intelligently refining them
CN119783644A
Conference memory enhancement method and equipment for conference tablet and medium
CN119988591A