Meeting summary automatic generation method based on multi-source heterogeneous information fusion

By combining multi-source heterogeneous information fusion technology and natural language processing, accurate speaker recognition and meeting minutes generation in multi-speaker scenarios are achieved, the problem of poor identification results in the existing technology is solved, and the accuracy and comprehensiveness of meeting minutes are improved.

CN120045701APending Publication Date: 2025-05-27ZHEJIANG UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510100951.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the speaker's identity in multispeaker scenarios and generate complete and accurate meeting minutes, especially in noise, overlapping voice or dynamic scenarios.

Method used

Using a method based on multi-source heterogeneous information fusion, combining video, audio signals, face recognition, sound source positioning, voiceprint recognition technology and natural language processing, the accurate identification of speaker identity and automatic generation of conference content is achieved by collecting multi-source information in real time.

Benefits of technology

It improves the accuracy of speaker recognition and the accuracy of speech content, can efficiently and accurately identify and calibrate each speaker and his speech content, and improves the comprehensiveness and accuracy of meeting minutes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045701A_ABST
    Figure CN120045701A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic conference summary generation method based on multi-source heterogeneous information fusion. The method comprises the following steps: extracting facial features of a speaker by using a face recognition technology and recognizing a face identity of the speaker; acquiring an audio signal through a microphone array, and recognizing a voiceprint identity of a speaker by using a voiceprint recognition technology; in combination with video and audio information, through a multi-source heterogeneous information fusion technology, alignment of video and audio data is carried out in time; through a sound source positioning technology, a voiceprint identity and a face identity are matched and aligned in space, and the identity of a speaker is accurately positioned and recognized; the spokesman and the speaking content are calibrated and separated, and it is ensured that the identity of the spokesman is accurately matched with the speaking content; and generating a spokesman abstract and a conference summary according to the calibrated speech content by using natural language processing and a deep learning model. The method is suitable for various conference scenes, the speaking content of different spokesmen can be recognized, the spokesman abstract is generated, and the efficiency and accuracy of conference summary generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and involves the comprehensive application of video and audio signal acquisition, face recognition, sound source localization, voiceprint recognition, speech-to-text conversion, and deep learning models. Specifically, it relates to a method for automatically generating meeting minutes based on the fusion of multi-source heterogeneous information. Background Art

[0002] With the advent of the information age, especially in the fields of enterprises, education, government, and scientific research institutions, meetings have become an important form of daily communication and decision-making. In a modern office environment, meeting minutes, as a record of the meeting results, play an important role. The existing methods for generating meeting minutes mainly rely on manual recording or a single information source (such as only audio or only video), and there are problems such as low efficiency, poor accuracy, and incomplete content. In a complex meeting scenario with multiple participants, it is difficult to accurately identify the speaker's identity and automatically record their speech content. In addition, the accuracy of single-modal methods drops significantly when facing noise, overlapping speech, or dynamic scenarios. Therefore, there is an urgent need for a multi-modal fusion automation method to improve the efficiency and accuracy of meeting minute generation.

[0003] Currently, some research and technologies have attempted to achieve automatic meeting minute generation through technologies such as speech recognition, video analysis, and natural language processing. For example, the current method for automatically generating meeting minutes based on speech recognition only relies on a single audio data source and cannot accurately distinguish the speech content of multiple speakers. Especially in an environment where multiple people speak simultaneously or there is a lot of background noise, the recognition effect is poor. In addition, although video data can provide certain clues about the speaker, existing video analysis methods usually cannot effectively fuse audio information, resulting in low face recognition accuracy. In complex scenarios, factors such as uneven lighting and frequent movement of participants make it difficult for existing video analysis methods to work stably.

[0004] Natural language processing technology has achieved success in many text generation tasks. For example, pre-trained language models such as BERT and GPT have demonstrated powerful capabilities in text generation and understanding. However, existing meeting minutes generation methods often lack context relevance and coherence in practical applications. The generated content often omits key information, resulting in insufficient practicality and accuracy of the meeting minutes (Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). "BERT: Pre-training of deep bidirectional transformers for language understanding." Proceedings of NAACL-HLT 2019, 4171-4186. https: / / doi.org / 10.18653 / v1 / N19-1423.).

[0005] Therefore, the existing technologies have not been able to fully utilize video, audio, speech-to-text, and natural language processing technologies to fuse multi-source heterogeneous information for efficient and accurate meeting minutes generation. Especially in the multi-speaker scenario, how to accurately identify the speaker's identity and extract the corresponding speech content remains an important challenge in the current technology.

[0006] In view of the deficiencies of the existing technologies, the present invention proposes an intelligent meeting minutes automatic generation method based on multi-source heterogeneous information fusion, aiming to effectively solve the problems of multi-speaker identification, information fusion, and the accuracy and efficiency in meeting minutes generation. Summary of the Invention

[0007] The present invention proposes a meeting minutes automatic generation method based on multi-source heterogeneous information fusion, which combines multiple technologies such as video, audio signals, face recognition, sound source localization, voiceprint recognition technology, and natural language processing. By collecting multi-source information in real time, it realizes the accurate identification of the speaker's identity and the automatic generation of meeting content. By introducing dynamic database technology to manage face identity and voiceprint identity information, the efficiency and accuracy of speaker identification are improved.

[0008] The object of the present invention is achieved by the following technical solutions: A meeting minutes automatic generation method based on multi-source heterogeneous information fusion, comprising the following steps:

[0009] (1) Real-time capture the video signal of the meeting scene through a wide-angle camera, extract the facial features of the speaker using face detection technology, and obtain the speaker's face identity through face recognition technology

[0010] (2) Obtain the audio signal of the meeting scene through the microphone array, extract the speaker's voiceprint features using voiceprint recognition technology, and identify the speaker's voiceprint identity

[0011] (3) Combine the video signal and the audio signal, and adopt multi-source heterogeneous information fusion technology, specifically including:

[0012] Time alignment: Use dynamic time warping technology to perform time alignment on the video V t and the audio signal A t to ensure the consistency of the speech content and the speaker's identity on the time axis;

[0013] Spatial matching: Use sound source localization technology to calculate the sound source position L t (x, y, z) and the face position P video (x, y), calculate the matching relationship through the Euclidean distance, and match and align the voiceprint identity and the face identity to obtain the final speaker identity

[0014] Speech content calibration: Combine the results of time alignment and spatial matching to calibrate the speaker identity and the speech content, and generate the final speaker identity and its corresponding speech content;

[0015] (4) Through natural language processing technology, generate a speaker summary from the calibrated speech content and further generate a complete meeting minutes.

[0016] Furthermore, the face recognition technology adopts deep convolutional neural network algorithms, including MTCNN and YOLOv5-face algorithms, for detecting and recognizing the speaker's facial features.

[0017] Furthermore, the voiceprint recognition technology is based on a deep learning model, including the x-vector or i-vector model, and uses the spectral features of the audio signal to extract the speaker's voiceprint features.

[0018] Furthermore, the sound source localization technology calculates the sound source position L t (x, y, z) through the microphone array combined with the generalized cross-correlation time delay estimation algorithm.

[0019] Furthermore, the dynamic time warping technology is used to perform time alignment on the time series V of the video signal t and the time series A of the audio signal t and calculate the synchronous mapping relationship between the time series using the optimal path algorithm.

[0020] Further, the identity of the final speaker in the multi-source heterogeneous information fusion is comprehensively calibrated through the time alignment and space matching results.

[0021] Further, a natural language processing model is used to perform semantic analysis and extractive summary generation on the calibrated speech content.

[0022] The beneficial effects of the present invention are as follows: Through video analysis, speech recognition, deep learning, and multi-modal data fusion technologies, and by combining wide-angle cameras, sound source localization, and voiceprint recognition technologies, the accuracy of speaker recognition and the accuracy of speech content are improved. Each speaker and their speech content can be efficiently and accurately identified and calibrated, and complex meeting scenarios can be effectively handled, improving the comprehensiveness and accuracy of speaker summaries and meeting minutes. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of the method for automatically generating meeting minutes based on multi-source heterogeneous information fusion of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0024] Based on multi-source heterogeneous information fusion and deep learning technologies, the present invention combines video, audio, speech-to-text, and natural language processing algorithms to accurately identify the identity of speakers and extract corresponding speech content, and finally generate structured meeting minutes. The following describes the detailed implementation manners of the present invention with reference to the drawings.

[0025] See Figure 1 , a method for automatically generating meeting minutes based on multi-source heterogeneous information fusion provided by the present invention includes the following steps:

[0026] Step 1: Video acquisition and face recognition.

[0027] 1.1 Video acquisition.

[0028] Two 180-degree cameras are arranged in the meeting scenario to achieve 360-degree video acquisition coverage. The video stream is processed frame by frame, and each frame image I t is stored as a two-dimensional pixel matrix. The video stream is decoded through tools such as OpenCV and Dlib to obtain a continuous frame sequence, and the video signal V t can be expressed as:

[0029] V t ={I 1 , I 2 , ……, I n}

[0030] where n represents the total number of video frames, and I t is the image of the t-th frame. Each frame image I tConvert to grayscale or normalized RGB images to reduce computational complexity.

[0031] 1.2 Storage and marking of face recognition feature vectors.

[0032] Use the YOLOv5-face algorithm for face detection and extract the face region R in the video frame t . YOLOv5-face is an object detection algorithm based on YOLO, which directly outputs the bounding box coordinates of the face.

[0033] Formula:

[0034] R t = FaceDetection(I t )

[0035] where R t is all detected face regions (bounding boxes). FaceDetection is an object detection algorithm based on YOLO. The detected face region R t is represented by a bounding box:

[0036] R t = {(x 1 , y 1 , x 2 , y 2 ) 1 , (x 1 , y 1 , x 2 , y 2 ) 2 , ……, (x 1 , y 1 , x 2 , y 2 ) n}

[0037] where: (x 1 , y 1 ) represents the coordinates of the upper left corner of the face region, and (x 2 , y 2 ) represents the coordinates of the lower right corner of the face region. n is the number of faces detected in the frame.

[0038] Use face recognition technology (deep convolutional neural network CNN, such as algorithms like MTCNN, YOLOv5-face, etc.) to extract the face feature F of each R t . Classic models such as FaceNet or ResNet-50 take a normalized face image as input and output a feature vector F face ; Represent the extracted feature as: face ;

[0039] Fface = FaceFeatureExtractor(R t )

[0040] Input different feature vectors F face into the identity database FaceDB, recording identity data bits (facial identity).

[0041] Feature similarity calculation:

[0042]

[0043] Among them, is the j-th facial feature vector in FaceDB, and Similarity represents that the feature similarity calculation method uses the cosine similarity method.

[0044]

[0045] Take the maximum value among all similarity scores as the confidence level of face recognition:

[0046]

[0047] Confidence level S video The closer it is to 1, the more reliable the face recognition result is.

[0048] If the similarity S j is less than the set threshold, it is considered a new speaker, generating a new facial identity and storing F face into FaceDB:

[0049]

[0050] FaceDB is an identity database used to record and mark the features of recognized speakers and use them for identity matching during subsequent audio recognition.

[0051] Dynamic database initialization:

[0052] Initialize a dynamic database or data structure (such as a hash table or index table) to store the facial feature vector F face .

[0053] Hash table: Key: Unique identifier Value: Corresponding facial feature vector F face , marking status, timestamp and other information. For example:

[0054] Database = {ID_1: {F_face: [...], Status: "Unconfirmed", Timestamp: T1}, ID_2: {F_face: [...], Status: "Confirmed", Timestamp: T2},...}

[0055] Each entry in the database records the following information: a unique identifier generated (such as ). The corresponding face feature vector F face . The marked value of the initial state (such as "Unconfirmed").

[0056] Determine whether it is a new person: For each detected face region R t , after extracting the feature vector F face , calculate its similarity with the feature vectors already stored in the database:

[0057]

[0058] where is the feature vector of the j-th record in the database. If all similarities S j are less than the set threshold (such as S < 0.7), then determine that this face is a newly emerged participant.

[0059] Assign an identifier to the new person:

[0060] If a new face feature F face is detected, add a new entry to the database and assign a new unique identifier

[0061]

[0062] where is the newly generated identifier. F face is the feature vector corresponding to this face. The status is marked as "Unconfirmed" and awaits subsequent verification.

[0063] Repeat person matching:

[0064] First, perform similarity comparison. For the face feature vector F extracted for each frame face , calculate its similarity with all the stored features in the database:

[0065] If there is a similarity S j above the threshold (such as S ≥ 0.7), then determine that the current face is the same person as the person corresponding to a certain record in the database.

[0066] Second, update the identifier. If a matching record is found, directly use the corresponding Speaker identifier for the current frame: S j ≥0.7

[0067] Calculate the spatial coordinates P of the speaker video (x, y).

[0068] For each detected face region R t , according to the bounding box coordinates (x 1 , y 1 , x 2 , y 2 ), calculate the center point coordinates of the face:

[0069] Associate the coordinates P video (x, y) of each speaker with their identity to obtain the complete speaker calibration information:

[0070]

[0071] Step 2: Audio acquisition and sound source localization.

[0072] 2.1 Audio acquisition and preprocessing.

[0073] Arrange 12 microphone arrays in the meeting room to collect audio signals in real time. The audio signal can be expressed as:

[0074] A t ={a 1 (t), a 2 (t), ……, a 12 (t)}

[0075] where a i (t) represents the audio signal collected by the i-th microphone at time t. To improve the positioning accuracy, the following preprocessing is performed on the collected audio signal: remove background noise (such as using a band-pass filter to extract the human voice frequency band), and normalize the signal strength for subsequent calculations.

[0076] 2.2 Time difference estimation.

[0077] For sound source localization, use the microphone array to calculate the sound source position L t (x, y, z) through the generalized cross-correlation function (GCC-PHAT) and time difference of arrival (TDOA) technology:

[0078] (1) Calculate the cross-correlation function:

[0079] Select two microphones (such as microphone i and j) and calculate the cross-correlation function of the signals they receive:

[0080] Rij R(τ) = ∫ a i (t) · a j (t + τ) dτ

[0081] Where: R ij (τ) represents the cross - correlation function between the i - th and j - th microphones. τ represents the time delay.

[0082] In actual calculations, cross - correlation is usually implemented through the Fast Fourier Transform (FFT) and the Inverse Fourier Transform (IFFT):

[0083] R ij (τ) = IFFT(FFT(a i (t)) · FFT(a j (t) * ))

[0084] (2) Determine the time difference:

[0085] The time delay Δt corresponding to the peak of the cross - correlation function R ij (τ): ij :

[0086]

[0087] (3) Calculate the distance difference: According to the distance d ij between microphones and the speed of sound c (usually c = 343 m / s), calculate the distance difference from the sound source to the two microphones:

[0088] Δd ij = c · Δt ij

[0089] Where: Δt ij is the time difference between the signals received by the i - th and j - th microphones, d ij is the distance difference from the sound source to the microphones, and c is the speed of sound.

[0090] 2.3 Calculation of the sound source position.

[0091] (1) Using the time - difference estimation results of multiple microphones, calculate the three - dimensional spatial position L t (x, y, z) of the sound source through geometric relationships. Assume that the position of each microphone in the microphone array is known and is expressed as:

[0092] MIC i = (x i , y i , z i ), i ∈ [1, 12]

[0093] Where, MIC i is the position coordinate of the i - th microphone.

[0094] (2) Establish the distance equation group:

[0095] Assume the sound source is at (x, y, z). The distance difference Δd between the sound source and the i-th and j-th microphones ij Satisfies the following equation:

[0096]

[0097] The above equations represent the geometric relationship between the sound source location and the microphones.

[0098] (3) Solving simultaneous equations:

[0099] Choose N microphones, you can construct up to For example, if a 12-microphone array is used, a maximum of 12*11 / 2=66 distance difference equations can be constructed. In actual calculations, at least 4 microphones are usually required (3 equations can solve 3 unknowns), and the redundant equations of 12 microphones can improve the robustness and accuracy of positioning.

[0100] All distance difference equations are combined to form a nonlinear system of equations:

[0101]

[0102] (4) Using numerical solution:

[0103] Because the system of equations is nonlinear, numerical methods (such as the least squares method) are usually used to solve the sound source position (x, y, z). The objective function is:

[0104]

[0105] Use gradient descent or other optimization algorithms to iteratively solve the sound source position L t (x,y,z).

[0106] Extract voiceprint features F using Mel-frequency cepstral coefficients (MFCC) and deep learning models (such as LSTM or CNN, including x-vector or i-vector models) audio , and generate voiceprint feature vectors, and dynamically store them in the voiceprint identity database (VoiceDB). The operation process is as follows: Feature similarity calculation:

[0107]

[0108] in, is the k-th voiceprint feature vector in VoiceDB.

[0109] Feature similarity calculation:

[0110]

[0111] Similarity indicates that the feature similarity calculation method adopts the cosine similarity method:

[0112]

[0113] Take the maximum value among all similarity scores as the credibility of face recognition:

[0114]

[0115] S audio The closer it is to 1, the more reliable the voiceprint recognition result is.

[0116] If S j is less than the set threshold, it is considered a new speaker, and a new voiceprint identity is generated and F audio is stored in VoiceDB:

[0117]

[0118] Step 3: Multi-source heterogeneous information fusion.

[0119] The face information P video (x, y) in video acquisition and the sound source localization and voiceprint recognition information L t (x, y, z), F audio are temporally and spatially aligned to obtain the final speaker identity

[0120] Use dynamic time warping (DTW) or maximum mutual information technology to synchronize the video and audio signals in time to ensure the consistency of the speech content and the speaker identity on the time axis.

[0121] Through the timestamp and synchronize the video and audio signals to ensure the temporal consistency of the multi-source signals. The alignment method formula is:

[0122]

[0123] where is the fused timestamp, and are the timestamps of the video and audio respectively, and TimeAlignment is the time alignment process.

[0124] Specifically, time alignment (DTW): The dynamic time warping (DTW) method is used to synchronize the video and audio signals in time:

[0125]

[0126] Among them, V t [i] and A t [j] are the timestamps of the video frame and the audio frame.

[0127] Spatial matching: Calculate the sound source position L t (x, y, z) and the Euclidean distance between the face position P video (x, y) to complete the speaker matching:

[0128]

[0129] Select the speaker corresponding to the smallest distance d min as the matching object.

[0130] Based on the comprehensive matching results, calibrate and fuse the identities: The output of the identity fusion function f is a unique identity label

[0131]

[0132] The specific operations are as follows:

[0133] Calibrate the logic and set the distance threshold d threshold :

[0134] If d min > d threshold , it is considered that the sound source and the face do not belong to the same speaker, and the unmatched result is output:

[0135]

[0136] If d min < d threshold : Proceed to the next step of credibility fusion calculation.

[0137] Similarity weighted fusion, combining the identity credibility S video and S audio , and allocate weights ω min and ω video and ω audio :

[0138]

[0139] α: Weight adjustment coefficient, controlling the influence of distance on the weight.

[0140] ωvideo +ω audio = 1

[0141] Calculate the total credibility S of the fused identity based on the weights fusion :

[0142] S fusion = ω video ·S video + ω audio ·S audio

[0143] If S fusion ≥ Threshold (set the credibility threshold, e.g., 0.8), then it is determined that and belong to the same speaker, and output the fused identity:

[0144]

[0145] If S fusion < Threshold, it is determined as unmatched, and output the undetermined identity:

[0146]

[0147] Output the calibration result, associate the fused identity label with the time-aligned data, and output the complete speech record:

[0148]

[0149] Among them, the dynamic database format is as follows:

[0150] FaceDB: Store the face features and identity identifiers of the speakers, in the following format:

[0151] [{"ID":"ID_1","FaceFeature":[...],"Status":"Active"},

[0152] {"ID":"ID_2","FaceFeature":[...],"Status":"Active"}]

[0153] VoiceDB: Store the voiceprint features and identity identifiers of the speakers, in the following format:

[0154] [{"ID":"ID_1","VoiceFeature":[...],"Status":"Active"},

[0155] {"ID":"ID_2","VoiceFeature":[...],"Status":"Active"}]。

[0156] Step 4: Speaker Summary and Meeting Minutes Generation.

[0157] Text Processing: Using natural language processing (NLP) technology, extract the speaker's speech content and generate structured text. Specifically include: Automatic Speech Recognition (ASR):

[0158] C i = ASR(A t )

[0159] where C i is the calibrated sequence of each speaker and their speech content.

[0160] Semantic Analysis and Summary Generation: Use deep learning models (such as BERT or GPT) to perform semantic analysis and extractive summarization on the text, generating a summary S i for each speaker:

[0161] S i = Summarization(C i )

[0162] Meeting Minutes Generation: According to the speaker's identity and the speech text C i , generate a complete meeting text record T:

[0163] T = {(ID 1 , T 1 ), (ID 2 , T 2 ), ……(ID i , T i )}

[0164] where ID i is the speaker's identity and T i is the corresponding speech content.

[0165] Use deep learning models (such as BERT or GPT) to perform semantic analysis and extractive summarization on the text, generating the meeting minutes M:

[0166] M = Summarization(T)

[0167] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

[0168] The above embodiments are only used to illustrate the design concept and characteristics of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A method for automatically generating meeting minutes based on multi-source heterogeneous information fusion, characterized in that: The following steps are involved: (1) Use a wide-angle camera to capture the video signal of the conference scene in real time, use face detection technology to extract the speaker's facial features, and use face recognition technology to obtain the speaker's facial identity (2) The audio signal of the conference scene is obtained through the microphone array, and the voiceprint recognition technology is used to extract the voiceprint features of the speaker and identify the speaker's voiceprint identity. (3) Combining video signals and audio signals, using multi-source heterogeneous information fusion technology, specifically including: Time alignment: Use dynamic time warping technology to align the video V t and audio signal A t Perform time alignment to ensure the consistency of speech content and speaker identity on the timeline; Spatial matching: Use sound source localization technology to calculate the sound source position L t (x,y,z) and face position P video (x, y), the matching relationship is calculated by Euclidean distance, and the voiceprint identity and face identity Perform matching and alignment to obtain the final speaker identity Speech content calibration: Combine the results of time alignment and spatial matching to calibrate the speaker identity and speech content to generate the final speaker identity and the corresponding speech content; (4) Through natural language processing technology, the calibrated speech content is used to generate a speaker summary, and further generate a complete meeting minutes.

2. The method for automatically generating meeting minutes based on multi-source heterogeneous information fusion according to claim 1 is characterized in that: The face recognition technology uses deep convolutional neural network algorithms, including MTCNN and YOLOv5-face algorithms, to detect and identify the facial features of the speaker.

3. The method for automatically generating meeting minutes based on multi-source heterogeneous information fusion according to claim 1 is characterized in that: The voiceprint recognition technology is based on a deep learning model, including an x-vector or i-vector model, which uses the spectral features of the audio signal to extract the voiceprint features of the speaker.

4. The method for automatically generating meeting minutes based on multi-source heterogeneous information fusion according to claim 1 is characterized in that: The sound source localization technology calculates the sound source position L by combining a microphone array with a generalized cross-correlation delay estimation algorithm. t (x,y,z).

5. The method for automatically generating meeting minutes based on multi-source heterogeneous information fusion according to claim 1 is characterized in that: The dynamic time warping technique is used to transform the time series V of the video signal t and the time series A of the audio signal t Perform time alignment and use the optimal path algorithm to calculate the synchronization mapping relationship between time series.

6. The method for automatically generating meeting minutes based on multi-source heterogeneous information fusion according to claim 1 is characterized in that: The final speaker identity of the multi-source heterogeneous information fusion It is calibrated comprehensively through time alignment and space matching results.

7. The method for automatically generating meeting minutes based on multi-source heterogeneous information fusion according to claim 1 is characterized in that: The natural language processing model is used to perform semantic analysis and extractive summary generation on the calibrated speech content.

Citation Information

Cited By

  • Conference summary automatic generation method and system based on OCR technology

    CN120337862A

  • A meeting minutes automatic generation method and system based on OCR technology

    CN120337862B

  • Multi-speaker identification method and device, equipment and storage medium

    CN120656451A

  • Conference information output method and device, terminal equipment and storage medium

    CN121365142A

  • Intelligent voice conference behavior analysis method and system based on multi-source perception

    CN121462328A