Conference summary generation method and device based on multi-modal fusion, equipment and storage medium

By using a multimodal fusion method to identify speakers by combining audio and video data, the problem of errors in meeting minutes in existing technologies has been solved, and more accurate meeting minutes generation has been achieved.

CN121842346APending Publication Date: 2026-04-10SHENZHEN DINSTAR TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, intelligent meeting systems rely on voiceprints or single visual features for identification, which makes the generated meeting minutes prone to errors.

Method used

A multimodal fusion approach is used to collect audio and video data from the meeting scene. By combining voiceprint recognition and visual identity recognition, the speaker's identity is determined and meeting minutes are generated.

Benefits of technology

Accurate identification of meeting speakers leads to more accurate meeting minutes, improving the accuracy and reliability of the identification process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121842346A_ABST
    Figure CN121842346A_ABST
Patent Text Reader

Abstract

The invention discloses a conference summary generation method and device based on multi-modal fusion, equipment and a storage medium, and relates to the technical field of information processing, and the conference summary generation method based on multi-modal fusion comprises the steps: collecting audio data and video data of a conference scene; processing the audio data to obtain a voiceprint recognition result; processing the video data to obtain a visual identity recognition result; determining the identity of a speaker based on the voiceprint recognition result and the visual identity recognition result; and generating a conference summary according to the speaker identity, the audio data and the video data. According to the invention, the identity of the spokesman of the conference is accurately identified, so that a more accurate conference summary is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, and in particular to a method, apparatus, device, and storage medium for generating meeting minutes based on multimodal fusion. Background Technology

[0002] With the development of intelligent meeting systems, voiceprint or single visual features are currently widely used for identification, which makes the generated meeting minutes prone to errors. Therefore, how to accurately identify the speaker's identity and generate more accurate meeting minutes remains a problem that needs to be solved.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide a method, apparatus, device, and storage medium for generating meeting minutes based on multimodal fusion, aiming to solve the technical problem of how to accurately identify the speaker's identity in a meeting, thereby generating more accurate meeting minutes.

[0005] To achieve the above objectives, this application proposes a method for generating meeting minutes based on multimodal fusion, the method comprising: Collect audio and video data from the meeting scenario; The audio data is processed to obtain the voiceprint recognition result; The video data is processed to obtain visual identity recognition results; The speaker's identity is determined based on the voiceprint recognition result and the visual identity recognition result; Meeting minutes are generated based on the speaker's identity, the audio data, and the video data.

[0006] In one embodiment, the step of collecting audio and video data from the conference scene includes: Audio data from the meeting scene is collected using a uniformly distributed microphone array; Based on the audio data, the sound source is located to obtain the speaker's location information; The speaker's location information is used to control the adjustment of the shooting angle of the edge camera; Video data of the meeting scene is captured by the central camera and the adjusted edge cameras.

[0007] In one embodiment, the step of processing the audio data to obtain the voiceprint recognition result includes: The audio data is subjected to noise reduction and speech activity detection to segment out effective audio segments; Extract the voiceprint features of the effective audio segments; The voiceprint features are matched with a preset voiceprint feature library to obtain the voiceprint recognition result.

[0008] In one embodiment, the step of processing the video data to obtain a visual identity recognition result includes: Detect and extract facial images, identification images, and lip movement sequence frames from the video data; The face image is matched with a preset facial feature database to obtain the face recognition result; The identity matching result is obtained based on the identity image; The synchronization between the lip movement sequence frames and the audio data is analyzed to obtain a lip synchronization score. The facial recognition result, the identity matching result, and the lip-synchronization scoring result are determined as the visual identity recognition result.

[0009] In one embodiment, before the step of determining the speaker's identity based on the voiceprint recognition result and the visual identity recognition result, the method further includes: Obtain the voiceprint confidence score of the voiceprint recognition result and the visual confidence score of the visual identity recognition result; When the voiceprint confidence level is less than the preset voiceprint threshold, audio data is re-acquired and the latest voiceprint recognition result is identified. When the visual confidence level is less than the preset voiceprint threshold, video data is re-acquired and the latest visual identity recognition result is identified.

[0010] In one embodiment, the step of determining the speaker's identity based on the voiceprint recognition result and the visual identity recognition result includes: Obtain the voiceprint confidence score of the voiceprint recognition result, and obtain the facial recognition confidence score, identity identifier matching result, and lip synchronization scoring result from the visual identity recognition result; The speaker's identity is determined based on the voiceprint confidence score, the facial recognition confidence score, the identity identifier matching result, and the lip synchronization score.

[0011] In one embodiment, the step of generating meeting minutes based on the speaker's identity, the audio data, and the video data includes: The content of the spoken text is determined based on the audio data; Extract the target video segment from the video data during the speech; The speaker's identity, the content of the speech text, and the target video clip are used to generate meeting minutes.

[0012] Furthermore, to achieve the above objectives, this application also proposes a meeting minutes generation device based on multimodal fusion, the meeting minutes generation device based on multimodal fusion comprising: The acquisition module is used to acquire audio and video data in the meeting scenario; An audio module is used to process the audio data to obtain voiceprint recognition results; The video module is used to process the video data to obtain visual identity recognition results; The determination module is used to determine the speaker's identity based on the voiceprint recognition result and the visual identity recognition result; The generation module is used to generate meeting minutes based on the speaker's identity, the audio data, and the video data.

[0013] Furthermore, to achieve the above objectives, this application also proposes a meeting minutes generation device based on multimodal fusion, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the meeting minutes generation method based on multimodal fusion as described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the meeting minutes generation method based on multimodal fusion as described above.

[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the meeting minutes generation method based on multimodal fusion as described above.

[0016] This application provides a method for generating meeting minutes based on multimodal fusion. The method involves collecting audio and video data from a meeting scenario; processing the audio data to obtain a voiceprint recognition result; processing the video data to obtain a visual identity recognition result; determining the speaker's identity based on the voiceprint and visual identity recognition results; and generating meeting minutes based on the speaker's identity, the audio data, and the video data. This application simultaneously collects and separately processes the meeting's audio and video data to obtain voiceprint and visual identity recognition results, and then fuses and compares the two types of data to determine the speaker's identity. This method can accurately identify the speaker's identity and generate more accurate meeting minutes. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the meeting minutes generation method based on multimodal fusion provided in this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the meeting minutes generation method based on multimodal fusion provided in this application; Figure 3 A simplified flowchart illustrating the meeting minutes generation method based on multimodal fusion provided in Embodiment 1 of this application; Figure 4 This is a schematic diagram of the module structure of the meeting minutes generation device based on multimodal fusion according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the meeting minutes generation method based on multimodal fusion in the embodiments of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] This application collects audio and video data from a meeting scenario; processes the audio data to obtain a voiceprint recognition result; processes the video data to obtain a visual identity recognition result; determines the speaker's identity based on the voiceprint recognition result and the visual identity recognition result; and generates meeting minutes based on the speaker's identity, the audio data, and the video data.

[0024] With the development of intelligent meeting systems, voiceprint or single visual features are currently widely used for identification, which makes the generated meeting minutes prone to errors. Therefore, how to accurately identify the speaker's identity and generate more accurate meeting minutes remains a problem that needs to be solved.

[0025] This application obtains voiceprint recognition results and visual identity recognition results by simultaneously collecting and processing audio and video data of the meeting, and then fuses and compares the two types of data to determine the speaker's identity. This can accurately identify the speaker's identity and generate more accurate meeting minutes.

[0026] Based on this, embodiments of this application provide a method for generating meeting minutes based on multimodal fusion, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the meeting minutes generation method based on multimodal fusion of this application.

[0027] In this embodiment, the meeting minutes generation method based on multimodal fusion includes steps S10 to S40: Step S10: Collect audio and video data from the meeting scenario; It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as a meeting minutes generation device based on multimodal fusion. The following description uses a meeting minutes generation device based on multimodal fusion as an example to illustrate this embodiment and the subsequent embodiments.

[0028] It should be noted that before the meeting, multi-dimensional identity data of attendees needs to be entered and a related database established. When attendees enter, they input their names through an interactive display module. The system records a 30-second self-introduction by each attendee and extracts a 128-dimensional voiceprint feature vector. The system's video capture module needs to capture images of attendees' frontal faces, extract 256-dimensional facial feature vectors, and capture images of attendees' name tags or table signs, extracting features such as name tag or table sign numbers and company logos. Finally, the "name, voiceprint features, facial features, and name tag / table sign features" are linked and stored.

[0029] In one feasible approach, the steps of collecting audio and video data of the meeting scene include: collecting audio data of the meeting scene through a uniformly distributed microphone array; locating the sound source based on the audio data to obtain the speaker's directional information; controlling the adjustment of the shooting angle of the edge cameras based on the speaker's directional information; and collecting video data of the meeting scene through the center camera and the adjusted edge cameras.

[0030] It should be noted that an 8-microphone array can be used, evenly distributed along the edge of the conference table at a spacing of 1.2m, to collect voice signals from the entire area. Beamforming technology is used to focus the speaker's voice and filter out ambient noise to obtain the audio data for the meeting scene. One central camera has a 120° horizontal field of view, and four edge directional cameras have a 60° horizontal field of view, supporting ±30° motorized rotation. The central camera captures panoramic video of the core area, while the edge cameras capture video of the edge areas. Furthermore, the edge cameras can adjust their angles based on the sound source localization results, accurately capturing images of the speaker's face, lips, and name tag / table label.

[0031] Step S20: Process the audio data to obtain the voiceprint recognition result; It should be noted that the original audio data is first preprocessed (e.g., noise reduction), and then the voiceprint recognition model is run to extract the voiceprint feature vector representing the speaker's biometric characteristics from the processed audio data. This feature vector is then compared with the pre-registered speaker voiceprint feature database to finally output a voiceprint recognition result.

[0032] In one feasible approach, the step of processing the audio data to obtain a voiceprint recognition result includes: performing noise reduction and speech activity detection on the audio data to segment out effective audio segments; extracting voiceprint features from the effective audio segments; and matching the voiceprint features with a preset voiceprint feature library to obtain a voiceprint recognition result.

[0033] It should be noted that filtering algorithms, such as Normalized Least Mean Squares (NLMS), can be applied to adaptively reduce noise in the original audio data, suppressing environmental noise and reverberation. Subsequently, active segments containing human voices are distinguished from silent or noise-only inactive segments in the audio data, and the continuous audio stream is segmented into independent, effective audio segments containing only human voices. Next, a voiceprint feature extraction model, such as a Convolutional Neural Network-Long Short-Term Memory Hybrid Model (CNN-LSTM), is used to generate a fixed-dimensional mathematical vector from the effective audio segments that uniquely represents the speaker's voice characteristics—the voiceprint feature vector. Finally, the cosine similarity of this voiceprint feature vector with the voiceprint feature vectors of all registered speakers pre-stored in the database is calculated to obtain the voiceprint recognition result. For example, a voiceprint recognition result could be "Li Si, confidence level 91%".

[0034] Step S30: Process the video data to obtain visual identity recognition results; It should be noted that the visual identity recognition results include facial recognition results, lip movement timing features, and identity identifier matching results, such as employee badge or table tag matching results.

[0035] In one feasible approach, the step of processing the video data to obtain a visual identity recognition result includes: detecting and extracting a face image, an identity identifier image, and lip movement sequence frames from the video data; matching the face image with a preset facial feature library to obtain a face recognition result; obtaining an identity identifier matching result based on the identity identifier image; analyzing the synchronization between the lip movement sequence frames and the audio data to obtain a lip synchronization scoring result; and determining the face recognition result, the identity identifier matching result, and the lip synchronization scoring result as the visual identity recognition result.

[0036] It should be noted that a ResNet-50 model can be used to extract facial features and match them with facial features in a repository to output facial recognition results (including confidence scores); a Three-Dimensional Convolutional Neural Network (3D CNN) model can be used to extract lip movement temporal features and calculate a consistency score (0-100 points) with the current speech rhythm to obtain lip synchronization scores; and a Convolutional Recurrent Neural Network (CRNN) can be used to recognize employee badge / table label numbers and identifiers, and output employee badge / table label matching results (including successful / failed matching indicators).

[0037] Step S40: Determine the speaker's identity based on the voiceprint recognition result and the visual identity recognition result; It should be noted that, based on the preset fusion strategy, cross-validation and comprehensive judgment are performed on the voiceprint recognition results and visual identity recognition results, which can yield a more reliable and accurate final speaker identity determination result than a single modality.

[0038] In one feasible approach, before the step of determining the speaker's identity based on the voiceprint recognition result and the visual identity recognition result, the method further includes: obtaining the voiceprint confidence score of the voiceprint recognition result and the visual confidence score of the visual identity recognition result; when the voiceprint confidence score is less than a preset voiceprint threshold, re-collecting audio data and identifying the latest voiceprint recognition result; when the visual confidence score is less than the preset voiceprint threshold, re-collecting video data and identifying the latest visual identity recognition result.

[0039] It should be noted that before making a fusion decision, it's possible to determine whether the voiceprint recognition and visual identity recognition results are usable. Specifically, this can be done by extracting a confidence value from the voiceprint recognition result and simultaneously extracting a comprehensive or key dimension (such as facial recognition confidence) visual confidence value from the visual identity recognition result to assess usability. For example, if the voiceprint confidence is less than 90%, the facial confidence is less than 85%, or the employee badge / table sign matching fails, the edge camera can be triggered to adjust its angle and re-capture the employee badge / table sign or facial image for secondary verification. This improves the accuracy of the final result.

[0040] Step S50: Generate meeting minutes based on the speaker's identity, the audio data, and the video data.

[0041] It should be noted that after obtaining the identity of the final speaker, the identity information, the corresponding speaking time period, the text content converted from the audio data, and key images of the relevant speaking moment extracted from the video data (such as speaker close-ups) are automatically associated and integrated. Finally, this information is organized according to a preset structured format (such as items including timestamps, speakers, speaking content, and attached images) to output a complete meeting minutes document.

[0042] In one feasible approach, the step of generating meeting minutes based on the speaker's identity, the audio data, and the video data includes: determining the speech text content based on the audio data; extracting the target video segment from the video data; and generating meeting minutes using the speaker's identity, the speech text content, and the target video segment.

[0043] It should be noted that the basic data "Speaker Identity (Name + ID Card / Table Label Number) - Speaking Time - Speech-to-Text Result" is generated by associating the speaker's identity, the content of their speech, and the target video segment. Then, keyframes (containing clear faces and ID cards / table labels) of the speaking time are extracted from the video stream and associated with the corresponding speaking segments in the basic data. Finally, the minutes are organized according to the "speaking time order" to generate a meeting summary with a table of contents (categorized by speaker) and search tags (ID card / table label number, keywords), which is stored in the storage module and displayed in the interactive display module.

[0044] This embodiment collects audio and video data from a meeting scenario; processes the audio data to obtain a voiceprint recognition result; processes the video data to obtain a visual identity recognition result; determines the speaker's identity based on the voiceprint recognition result and the visual identity recognition result; and generates meeting minutes based on the speaker's identity, the audio data, and the video data. This embodiment simultaneously collects and separately processes the meeting's audio and video data to obtain voiceprint recognition and visual identity recognition results, and then fuses and compares the two types of data to determine the speaker's identity. This allows for accurate identification of the meeting's speakers, thereby generating more accurate meeting minutes.

[0045] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S40 also includes steps S401 to S402: Step S401: Obtain the voiceprint confidence score of the voiceprint recognition result, and obtain the facial recognition confidence score, identity matching result, and lip synchronization score of the visual identity recognition result; It should be noted that, in order to improve the accuracy of speaker identification results, voiceprint confidence, facial recognition confidence, identity matching results, and lip-sync scoring results can be used together for comprehensive judgment from multiple dimensions to reduce the false judgment rate.

[0046] Step S402: Determine the speaker's identity based on the voiceprint confidence score, the facial recognition confidence score, the identity matching result, and the lip synchronization score.

[0047] Understandably, if the voiceprint confidence score is ≥90%, the facial confidence score is ≥85%, the lip consistency score is ≥80, and the employee badge / table sign matches successfully, the identity can be directly confirmed. If the voiceprint confidence score is ≥90%, but the facial confidence score is <85% or the employee badge / table sign match fails, the edge camera is triggered to adjust its angle and re-capture the employee badge / table sign / face for secondary verification. If the secondary verification of the employee badge / table sign match is successful, combined with the lip consistency score ≥80, the identity is confirmed. If the voiceprint confidence score is <90%, the employee badge / table sign recognition result is the core. If the employee badge / table sign matches successfully, combined with the facial confidence score ≥80 and the lip consistency score ≥75, the identity is confirmed. If the employee badge / table sign is not recognized, sound source localization is initiated, the corresponding area camera is controlled to focus and capture, and facial and employee badge / table sign information is re-collected for verification.

[0048] This embodiment obtains the voiceprint confidence score of the voiceprint recognition result, and the facial recognition confidence score, identity identifier matching result, and lip synchronization score from the visual identity recognition result; the speaker's identity is determined based on the voiceprint confidence score, the facial recognition confidence score, the identity identifier matching result, and the lip synchronization score. This embodiment comprehensively judges the speaker's identity through data from multiple dimensions, resulting in more accurate recognition results.

[0049] For example, to help understand the implementation flow of the meeting minutes generation method based on multimodal fusion obtained by combining this embodiment with the above embodiment one, please refer to... Figure 3 , Figure 3 This paper presents a simplified flowchart of a meeting minutes generation method based on multimodal fusion. Specifically: Before the meeting: Collect participants' voiceprints, facial features, and name tags / table labels, and establish an association database. At the start of the meeting: Simultaneously collect audio and video data. Audio preprocessing: Reduce noise and separate effective segments, and extract voiceprint features using an audio neural network. Video preprocessing: Detect faces / name tags / table labels, extract lip movements, and perform facial recognition, lip scoring, and name tag / table label recognition using a visual neural network. Then, cross-validate confidence levels to confirm identities, convert speech to text, and associate the data with "identity-name tag / table label-text-image" to generate structured minutes. After the meeting, output the complete meeting minutes.

[0050] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the meeting minutes generation method based on multimodal fusion of this application. Any simple transformations based on this technical concept are within the protection scope of this application.

[0051] This application also provides a meeting minutes generation device based on multimodal fusion, please refer to... Figure 4 The meeting minutes generation device based on multimodal fusion includes: Acquisition module 10 is used to acquire audio and video data in the meeting scenario; Audio module 20 is used to process the audio data to obtain voiceprint recognition results; Video module 30 is used to process the video data to obtain visual identity recognition results; The determination module 40 is used to determine the speaker's identity based on the voiceprint recognition result and the visual identity recognition result; The generation module 50 is used to generate meeting minutes based on the speaker's identity, the audio data, and the video data.

[0052] This embodiment collects audio and video data from a meeting scenario; processes the audio data to obtain a voiceprint recognition result; processes the video data to obtain a visual identity recognition result; determines the speaker's identity based on the voiceprint recognition result and the visual identity recognition result; and generates meeting minutes based on the speaker's identity, the audio data, and the video data. This embodiment simultaneously collects and separately processes the meeting's audio and video data to obtain voiceprint recognition and visual identity recognition results, and then fuses and compares the two types of data to determine the speaker's identity. This allows for accurate identification of the meeting's speakers, thereby generating more accurate meeting minutes.

[0053] In one embodiment, the acquisition module 10 is further configured to acquire audio data of the conference scene through a uniformly distributed microphone array; locate the sound source based on the audio data to obtain the speaker's location information; control the adjustment of the shooting angle of the edge cameras through the speaker's location information; and acquire video data of the conference scene through the center camera and the adjusted edge cameras.

[0054] In one embodiment, the audio module 20 is further configured to perform noise reduction and speech activity detection on the audio data, segment out effective audio segments, extract voiceprint features of the effective audio segments, and match the voiceprint features with a preset voiceprint feature library to obtain voiceprint recognition results.

[0055] In one embodiment, the video module 30 is further configured to detect and extract a face image, an identity identifier image, and lip movement sequence frames from the video data; match the face image with a preset facial feature library to obtain a face recognition result; obtain an identity identifier matching result based on the identity identifier image; analyze the synchronization between the lip movement sequence frames and the audio data to obtain a lip synchronization score result; and determine the face recognition result, the identity identifier matching result, and the lip synchronization score result as a visual identity recognition result.

[0056] In one embodiment, the determining module 40 is further configured to obtain the voiceprint confidence score of the voiceprint recognition result and the visual confidence score of the visual identity recognition result; when the voiceprint confidence score is less than a preset voiceprint threshold, audio data is re-acquired and the latest voiceprint recognition result is identified; when the visual confidence score is less than the preset voiceprint threshold, video data is re-acquired and the latest visual identity recognition result is identified.

[0057] In one embodiment, the determining module 40 is further configured to obtain the voiceprint confidence score of the voiceprint recognition result, and obtain the facial recognition confidence score, identity identifier matching result, and lip synchronization score of the visual identity recognition result; and determine the speaker's identity based on the voiceprint confidence score, the facial recognition confidence score, the identity identifier matching result, and the lip synchronization score.

[0058] In one embodiment, the generation module 50 is further configured to determine the content of the spoken text based on the audio data; extract the target video segment of the speech from the video data; and generate meeting minutes by combining the speaker's identity, the content of the spoken text, and the target video segment.

[0059] The meeting minutes generation device based on multimodal fusion provided in this application, employing the meeting minutes generation method based on multimodal fusion in the above embodiments, can solve the technical problem of how to accurately identify the speaker's identity in a meeting, thereby generating more accurate meeting minutes. Compared with the prior art, the beneficial effects of the meeting minutes generation device based on multimodal fusion provided in this application are the same as those of the meeting minutes generation method based on multimodal fusion provided in the above embodiments, and other technical features in the meeting minutes generation device based on multimodal fusion are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0060] This application provides a meeting minutes generation device based on multimodal fusion. The meeting minutes generation device based on multimodal fusion includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the meeting minutes generation method based on multimodal fusion in the above embodiment 1.

[0061] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a meeting minutes generation device based on multimodal fusion suitable for implementing embodiments of this application. The meeting minutes generation device based on multimodal fusion in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The meeting minutes generation device based on multimodal fusion shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0062] like Figure 5As shown, a multimodal fusion-based meeting minutes generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the multimodal fusion-based meeting minutes generation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the multimodal fusion-based meeting minutes generation device to wirelessly or wiredly communicate with other devices to exchange data. Although a multimodal fusion-based meeting minutes generation device with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0063] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0064] The meeting minutes generation device based on multimodal fusion provided in this application, employing the meeting minutes generation method based on multimodal fusion in the above embodiments, can solve the technical problem of how to accurately identify the speaker's identity in a meeting, thereby generating more accurate meeting minutes. Compared with the prior art, the beneficial effects of the meeting minutes generation device based on multimodal fusion provided in this application are the same as those of the meeting minutes generation method based on multimodal fusion provided in the above embodiments, and other technical features in this meeting minutes generation device based on multimodal fusion are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0065] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0066] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0067] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the meeting minutes generation method based on multimodal fusion in the above embodiments.

[0068] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0069] The aforementioned computer-readable storage medium may be included in a multimodal fusion-based meeting minutes generation device; or it may exist independently and not assembled into a multimodal fusion-based meeting minutes generation device.

[0070] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a multimodal fusion-based meeting minutes generation device, cause the multimodal fusion-based meeting minutes generation device to: collect audio and video data of a meeting scene; process the audio data to obtain a voiceprint recognition result; process the video data to obtain a visual identity recognition result; determine the speaker's identity based on the voiceprint recognition result and the visual identity recognition result; and generate meeting minutes based on the speaker's identity, the audio data, and the video data.

[0071] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0072] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0073] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0074] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described multimodal fusion-based meeting minutes generation method. This solves the technical problem of accurately identifying the speakers at a meeting, thereby generating more accurate meeting minutes. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the multimodal fusion-based meeting minutes generation method provided in the above embodiments, and will not be repeated here.

[0075] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the multimodal fusion-based meeting minutes generation method described above.

[0076] The computer program product provided in this application solves the technical problem of accurately identifying the speaker's identity in a meeting, thereby generating more accurate meeting minutes. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the meeting minutes generation method based on multimodal fusion provided in the above embodiments, and will not be repeated here.

[0077] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for generating meeting minutes based on multimodal fusion, characterized in that, The method includes: Collect audio and video data from the meeting scenario; The audio data is processed to obtain the voiceprint recognition result; The video data is processed to obtain visual identity recognition results; The speaker's identity is determined based on the voiceprint recognition result and the visual identity recognition result; Meeting minutes are generated based on the speaker's identity, the audio data, and the video data.

2. The method as described in claim 1, characterized in that, The steps for collecting audio and video data from the meeting scenario include: Audio data from the meeting scene is collected using a uniformly distributed microphone array; Based on the audio data, the sound source is located to obtain the speaker's location information; The speaker's location information is used to control the adjustment of the shooting angle of the edge camera; Video data of the meeting scene is captured by the central camera and the adjusted edge cameras.

3. The method as described in claim 1, characterized in that, The step of processing the audio data to obtain the voiceprint recognition result includes: The audio data is subjected to noise reduction and speech activity detection to segment out effective audio segments; Extract the voiceprint features of the effective audio segments; The voiceprint features are matched with a preset voiceprint feature library to obtain the voiceprint recognition result.

4. The method as described in claim 1, characterized in that, The steps for processing the video data to obtain visual identity recognition results include: Detect and extract facial images, identification images, and lip movement sequence frames from the video data; The face image is matched with a preset facial feature database to obtain the face recognition result; The identity matching result is obtained based on the identity image; The synchronization between the lip movement sequence frames and the audio data is analyzed to obtain a lip synchronization score. The facial recognition result, the identity matching result, and the lip-synchronization scoring result are determined as the visual identity recognition result.

5. The method as described in claim 1, characterized in that, Before the step of determining the speaker's identity based on the voiceprint recognition result and the visual identity recognition result, the method further includes: Obtain the voiceprint confidence score of the voiceprint recognition result and the visual confidence score of the visual identity recognition result; When the voiceprint confidence level is less than the preset voiceprint threshold, audio data is re-acquired and the latest voiceprint recognition result is identified. When the visual confidence level is less than the preset voiceprint threshold, video data is re-acquired and the latest visual identity recognition result is identified.

6. The method as described in claim 1, characterized in that, The step of determining the speaker's identity based on the voiceprint recognition result and the visual identity recognition result includes: Obtain the voiceprint confidence score of the voiceprint recognition result, and obtain the facial recognition confidence score, identity identifier matching result, and lip synchronization scoring result from the visual identity recognition result; The speaker's identity is determined based on the voiceprint confidence score, the facial recognition confidence score, the identity identifier matching result, and the lip synchronization score.

7. The method as described in claim 1, characterized in that, The step of generating meeting minutes based on the speaker's identity, the audio data, and the video data includes: The content of the spoken text is determined based on the audio data; Extract the target video segment from the video data during the speech; The speaker's identity, the content of the speech text, and the target video clip are used to generate meeting minutes.

8. A meeting minutes generation device based on multimodal fusion, characterized in that, The device includes: The acquisition module is used to acquire audio and video data in the meeting scenario; An audio module is used to process the audio data to obtain voiceprint recognition results; The video module is used to process the video data to obtain visual identity recognition results; The determination module is used to determine the speaker's identity based on the voiceprint recognition result and the visual identity recognition result; The generation module is used to generate meeting minutes based on the speaker's identity, the audio data, and the video data.

9. A meeting minutes generation device based on multimodal fusion, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the meeting minutes generation method based on multimodal fusion as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the meeting minutes generation method based on multimodal fusion as described in any one of claims 1 to 7.