Data processing method and device in conference scene and electronic equipment

By processing meeting data using a large-scale meeting language model and a multimodal visual language model, intelligent chapters, meeting minutes, and speaker summaries are automatically generated, solving the problems of difficult and time-consuming traditional manual processing and achieving efficient and accurate meeting data analysis.

CN120873155APending Publication Date: 2025-10-31GUANGZHOU SHIYUAN ELECTRONICS CO LTD +2

Patent Information

Application Number
CN202410449091.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In modern business and meeting environments, traditional manual processing and organization of meeting data is difficult and time-consuming, and it is hard to guarantee the accuracy and consistency of meeting minutes.

Method used

It uses a large language model for meetings and a large visual language model for multimodal meetings to process meeting data, automatically generating intelligent chapters, meeting minutes and speaker summaries, and providing efficient and accurate analysis results through unimodal and multimodal data processing respectively.

Benefits of technology

It reduces the workload of manual processing, improves the efficiency and accuracy of meeting data processing, and provides more comprehensive and richer analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873155A_ABST
    Figure CN120873155A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a data processing method and device in a conference scene and electronic equipment. The method comprises the following steps: acquiring conference data, including single-mode conference data and multi-mode conference data; preprocessing the conference data according to the type of the conference data to obtain preprocessed conference data; a large model corresponding to the preprocessed conference data is determined, the large model comprises a conference large language model and a multi-modal visual language large model, the conference large language model is used for processing single-modal conference data, and the multi-modal visual language large model is used for processing multi-modal conference data; and inputting the preprocessed conference data into a large model, and outputting a data result corresponding to a target task through the large model, the target task including at least one of an intelligent chapter, a conference summary, a spokesman summary and a conference to-do. According to the invention, comprehensive and rich analysis results of the conference data can be provided, the method is convenient and efficient, and the analysis results are more accurate and consistent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a data processing method, apparatus, and electronic device for a conference setting. Background Technology

[0002] In modern business and meeting environments, participating in meetings, recording content, and compiling meeting minutes are demanding yet crucial tasks. However, with the diversification of meeting data, including audio and video recordings, traditional manual processing and compilation have become increasingly difficult and time-consuming. Furthermore, the accuracy and consistency of meeting minutes present a challenge, as manual compilation is easily limited by environmental factors, individual memory, and comprehension. Therefore, providing an efficient, convenient, and accurate way to process and manage meeting data, enabling users to better understand meeting content, capture key information, and quickly handle pending tasks, is of great significance. Summary of the Invention

[0003] The main technical problem addressed by the embodiments of this application is how to obtain key meeting information efficiently, conveniently, and accurately.

[0004] To address the aforementioned technical problems, one technical solution adopted in this application is: providing a data processing method for a meeting scenario, comprising: acquiring meeting data, wherein the type of meeting data includes unimodal meeting data and multimodal meeting data; preprocessing the meeting data according to its type to obtain preprocessed meeting data; determining a large model corresponding to the preprocessed meeting data; the large model includes a meeting large language model and a multimodal visual language large model, wherein the meeting large language model is used to process the unimodal meeting data, and the multimodal visual language large model is used to process the multimodal meeting data; inputting the preprocessed meeting data into the large model, and outputting data results corresponding to a target task through the large model, wherein the target task includes at least one of intelligent chapters, meeting minutes, speaker summaries, and meeting tasks. The meeting data includes both unimodal and multimodal meeting data, and the two types of settings can adapt to different meeting needs to meet the analysis and processing requirements of different meeting scenarios. Processing unimodal meeting data using a large-scale meeting language model is advantageous because unimodal meeting data focuses on a single media format. This allows the large-scale meeting language model to offer higher processing speed and efficiency when handling large-scale unimodal meeting data, and it also has advantages such as in-depth mining of linguistic information and refined semantic understanding. These advantages enable the model to provide accurate and high-quality results in the analysis and processing of unimodal meeting data. Processing multimodal meeting data using a large-scale multimodal visual language model is also beneficial. Multimodal meeting data contains information from multiple media formats, and the large-scale multimodal visual language model can process this different media information simultaneously and perform cross-modal fusion analysis. For example, by associating image and text information, the model can understand the meeting content from multiple dimensions, providing more comprehensive and richer analytical results. Furthermore, both large-scale meeting language models and large-scale multimodal visual language models can automate various tasks when processing meeting data, such as generating intelligent chapters, meeting minutes, and speaker summaries, thereby reducing the workload of manual processing, making it convenient to use and improving efficiency. Moreover, large-scale models can apply rich context and semantic understanding when processing data, thus providing more accurate and consistent analytical results.

[0005] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a data processing device for a meeting scenario, comprising: a meeting data acquisition module for acquiring meeting data, wherein the type of meeting data includes unimodal meeting data and multimodal meeting data; a preprocessing module for preprocessing the meeting data according to its type to obtain preprocessed meeting data; a large model determination module for determining the large model corresponding to the preprocessed meeting data; wherein the large model includes a meeting large language model and a multimodal visual language large model, wherein the meeting large language model is used to process the unimodal meeting data, and the multimodal visual language large model is used to process the multimodal meeting data; and a target task generation module for inputting the preprocessed meeting data into the large model and outputting data results corresponding to the target task through the large model, wherein the target task includes at least one of intelligent chapters, meeting minutes, speaker summaries, and meeting tasks.

[0006] To solve the above-mentioned technical problems, another technical solution adopted in the embodiments of this application is: to provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0007] The data processing devices and electronic equipment used in the aforementioned meeting scenarios have the same beneficial effects as the data processing methods used in the aforementioned meeting scenarios. Attached Figure Description

[0008] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0009] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application;

[0010] Figure 2 This is an architecture diagram of a data processing system for a meeting scenario provided in an embodiment of this application;

[0011] Figure 3 This is a flowchart illustrating a data processing method in a meeting scenario provided in an embodiment of this application;

[0012] Figure 4 This is a flowchart of a method for preprocessing meeting data to obtain preprocessed meeting data, provided in an embodiment of this application.

[0013] Figure 5 This is a flowchart of a method for obtaining key images with time information provided in an embodiment of this application;

[0014] Figure 6 This is a flowchart of a method for outputting data results corresponding to the target task through the large model when the large model is the conference large language model, as provided in an embodiment of this application;

[0015] Figure 7 This is a flowchart of a method for outputting data results corresponding to a target task through the large model when the large model is the multimodal visual language large model, as provided in an embodiment of this application.

[0016] Figure 8 This is a flowchart of a data processing method in a meeting scenario provided by another embodiment of this application;

[0017] Figure 9 This is a flowchart of the training method for the conference large language model provided in the embodiments of this application;

[0018] Figure 10 This is a flowchart of the method for obtaining a vocabulary dataset provided in an embodiment of this application;

[0019] Figure 11 This is a flowchart of a method for obtaining a target instruction dataset provided in an embodiment of this application;

[0020] Figure 12 This is a flowchart of the training method for a multimodal visual language large model provided in the embodiments of this application;

[0021] Figure 13 This is a schematic diagram of the structure of a data processing device in a conference scenario provided in an embodiment of this application;

[0022] Figure 14 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0024] It should be noted that, unless otherwise specified, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device schematic diagram or the order in the flowchart.

[0025] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application.

[0026] As discussed in the background section, efficient, convenient, and accurate methods for processing and managing meeting data are of great significance in modern business and meeting environments. Traditional methods of manually processing and organizing meeting data have become difficult and time-consuming, and it is also difficult to guarantee the accuracy and consistency of meeting minutes.

[0027] Based on this, this application proposes a data processing method and apparatus for a meeting scenario. The method involves acquiring meeting data, including both unimodal and multimodal meeting data; preprocessing the meeting data according to its type to obtain preprocessed meeting data; determining a large model corresponding to the preprocessed meeting data; the large model including a meeting language model and a multimodal visual language model, whereby the meeting language model processes the unimodal meeting data and the multimodal visual language model processes the multimodal meeting data; inputting the preprocessed meeting data into the large model; and outputting data results corresponding to a target task, whereby the target task includes at least one of intelligent chapters, meeting minutes, speaker summaries, and meeting tasks. The inclusion of both unimodal and multimodal meeting data allows for adaptation to different meeting needs, satisfying the analysis and processing requirements of various meeting scenarios. Processing unimodal meeting data using a large-scale meeting language model is advantageous because unimodal meeting data focuses on a single media format. This allows the large-scale meeting language model to offer higher processing speed and efficiency when handling large-scale unimodal meeting data, and it also has advantages such as in-depth mining of linguistic information and refined semantic understanding. These advantages enable the model to provide accurate and high-quality results in the analysis and processing of unimodal meeting data. Processing multimodal meeting data using a large-scale multimodal visual language model is also beneficial. Multimodal meeting data contains information from multiple media formats, and the large-scale multimodal visual language model can process this different media information simultaneously and perform cross-modal fusion analysis. For example, by associating image and text information, the model can understand the meeting content from multiple dimensions, providing more comprehensive and richer analytical results. Furthermore, both large-scale meeting language models and large-scale multimodal visual language models can automate various tasks when processing meeting data, such as generating intelligent chapters, meeting minutes, and speaker summaries, thereby reducing the workload of manual processing, making it convenient to use and improving efficiency. Moreover, large-scale models can apply rich context and semantic understanding when processing data, thus providing more accurate and consistent analytical results.

[0028] First, let me explain the terms used in this application:

[0029] Monomodal meeting data refers to meeting data that contains only one data modality (such as text, audio, or video). Monomodal meeting data is limited to a single source of information. For example, monomodal meeting data could be a series of text meeting transcripts, a set of audio recordings, or a set of video recordings, where the meeting content contains only one of the following: text, audio, or video.

[0030] Multimodal meeting data refers to meeting data that includes multiple data modalities. This means that the meeting content comes from multiple different information sources, and these different data modalities can include text, audio, and video. Multimodal meeting data is richer and more diverse, providing more comprehensive meeting information. For example, multimodal meeting data can simultaneously include text chat logs between participants, audio recordings, and video recordings, as well as other possible data modalities, such as screen sharing or whiteboard content. By combining different data modalities, multimodal meeting data can provide more comprehensive and richer meeting information and context.

[0031] Large models are deep learning models with a massive number of parameters and computational power. With the development of deep learning and artificial intelligence, larger and more complex models have been built to handle more data and more complex tasks. Training large models typically requires substantial computational resources and data, and necessitates the use of distributed and parallel computing techniques to accelerate the training process. Large models have wide applications in various fields, including natural language processing, computer vision, and speech recognition. Common large models include GPT-3 (Generative Pre-trained Transformer 3), BERT (Bidirectional Encoder Representations from Transformers), and ResNet (Residual Network).

[0032] A large-scale language model for meetings refers to a language model specifically designed for processing meeting-related data. It can receive and process textual data related to meetings, such as meeting minutes, agendas, and notes. Pre-trained, this model can understand and generate natural language text and possesses contextual understanding capabilities for specific meeting content. It can be applied to various tasks, such as automatic meeting summary generation, automatic intelligent chapter generation, meeting minutes, speaker summaries, and meeting to-do lists, question-and-answer systems, and meeting topic extraction, providing more efficient and intelligent meeting support tools.

[0033] Multimodal Visual-Language Large Model (MVLM): This is a large-scale model that integrates visual and linguistic capabilities to process meeting data containing multiple data modalities (such as images, text, and audio). MVLM has the ability to process and understand multiple data modalities and can perform joint modeling and inference across different modalities. When processing multimodal meeting data, MVLM can simultaneously consider information such as images, text, and audio to achieve a more comprehensive and richer understanding and analysis of the meeting. For example, it can extract meeting content from PowerPoint slides, extract speech content from audio, analyze meeting content from text, and even recognize attendees' facial expressions from images.

[0034] This application embodiment, based on the aforementioned large-scale meeting language model and multimodal visual language model, analyzes and summarizes text, audio, and video generated during meetings. Through discourse normalization, the transcribed text is organized to generate highly readable minutes. Utilizing text summarization and multimodal image-text understanding capabilities, the system can summarize the main content of the meeting, the participants' speeches, and automatically generate a meeting to-do list. The system also has topic segmentation capabilities, generating intelligent chapters of the meeting based on chronological order. Furthermore, the system uses a keyword extraction algorithm to extract keywords mentioned in the meeting. Users only need to provide meeting data, and the system in this application embodiment can analyze and summarize the content of the meeting data to generate discourse-normalized minutes, meeting summaries, meeting conclusions, meeting to-do lists, speaker summaries, intelligent chapters of the meeting with chronological order, and keywords. In this way, users can easily obtain important information and summaries of the meeting, reducing the workload of manual organization and improving meeting efficiency and information readability.

[0035] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this application. The data processing method and apparatus for a conference scenario provided by this embodiment can be applied to this application scenario. Figure 1As shown, this application scenario includes a meeting setting, a server, and user terminals. The meeting setting includes participants, cameras, and projection equipment. In the meeting setting, participants can use various devices to engage in the meeting, such as mobile phones, computers, and tablets. Meeting data can be collected through microphones, such as those built into mobile phones or desktop microphones, to capture audio content during the meeting. Microphones can pick up sound information such as participants' speeches, discussions, and presentations. Cameras can also be used; participants can use camera devices, such as those built into computers, mobile phones, or dedicated conference cameras, to capture video content during the meeting. Cameras can capture participants' facial expressions, body language, and presentation content. Projection equipment can also be used; the meeting venue can be equipped with projectors to project presentation content or shared screens onto a large screen for participants to view. In addition to the above devices, the meeting setting can also include other sensors or devices, such as touchscreens, whiteboards, and styluses, for participants to interact and take notes.

[0036] These devices can collect meeting data (including audio, video, and text), which is then sent to a server for analysis and processing. Meeting data can be categorized into two types: unimodal meeting data and multimodal meeting data.

[0037] A server is used to process and store meeting data, perform preprocessing operations, and provide computing resources and services. A server can be a physical server, a dedicated hardware device for hosting applications and services, equipped with a high-performance processor, large-capacity storage, and high-speed network connectivity, capable of supporting large-scale data processing and storage needs. A server can also be a virtual server, where multiple virtual server instances are created on a single physical server using virtualization technology. Each virtual server instance has independent computing resources and an operating system environment, and can independently run and host applications. In this embodiment, the server receives meeting data collected from various devices in the meeting scenario (such as cameras and microphones), including audio, video, and other forms of data. The server performs preprocessing operations based on the type of meeting data, such as speech recognition, image processing, and text analysis, extracting and transforming useful information from the meeting data for subsequent processing and analysis. The server also provides computing resources and services to support large models in processing and analyzing preprocessed meeting data. Deep learning models such as a large meeting language model and a multimodal visual language model are deployed on the server. The large meeting language model is used to process unimodal meeting data, while the multimodal visual language model is used to process multimodal meeting data. By inputting the preprocessed meeting data into the corresponding large models, the system can generate data results for the target task, such as intelligent chapters, meeting minutes, speaker summaries, and meeting tasks.

[0038] A user terminal refers to a device or application used by a user to interact with a server. In the application scenarios of this application, the user terminal can be various devices or applications, such as computers, mobile phones, tablets, mini-programs, web interfaces, etc. The graphical interface of the user terminal provides functions for interacting with the server, including: viewing the meeting's table of contents, which includes options such as smart chapters, meeting minutes, speaker summaries, and meeting to-do lists; triggering task processing, where users can select and trigger specific target tasks on the graphical interface, such as generating smart chapters, meeting minutes, speaker summaries, or meeting to-do lists; and browsing meeting-related information, where users can browse the processing results of meeting data through the graphical interface, such as viewing smart chapters, meeting minutes, speaker summaries, or meeting to-do lists. The graphical interface of the user terminal can also provide the original text version of the meeting, i.e., the content without server optimization processing; the original audio or video can also be played or viewed by the user on the graphical interface. Optionally, the user terminal can also interact and communicate with the server to transmit user requests to the server and receive data and results returned by the server, so as to realize the interaction between the user and the conference data processing system. For example, if the user is not satisfied with the intelligent chapters, meeting minutes, speaker summaries or meeting to-do lists generated by the graphical interface, he / she can request the conference data processing system to re-summarize, such as by clicking the "Re-summarize" button on the graphical interface.

[0039] The above application scenario utilizes a server for processing and preprocessing meeting data, and employs a large meeting language model and a multimodal visual language model to process the data. The meeting agenda is displayed through a graphical interface on the user terminal, and corresponding tasks are triggered based on user needs to provide intelligent chapters, meeting minutes, speaker summaries, and meeting to-do lists. This approach offers advantages in automated processing, high efficiency, convenience, flexibility, customizability, and data storage and management, bringing convenience and benefits to meeting management and applications. It should be noted that... Figure 1 The application scenario shown is an example used to illustrate the data processing method in a meeting scenario according to the embodiments of this application. Other application scenarios may also exist.

[0040] Please see Figure 2 , Figure 2 This is an architecture diagram of a data processing system for a meeting scenario provided in an embodiment of this application.

[0041] The system accepts meeting data including text, audio, images, and video. Data entry methods include monomodal and multimodal recording, with the choice between monomodal and multimodal recording depending on the specific meeting type. Monomodal recording uses only one media format to record data, such as audio recording or text recording. Both the audio and text can be converted into a transcript containing time and speaker information. Multimodal recording uses multiple media formats simultaneously, such as audio, video, and text recording. This provides richer and more comprehensive data, including voice, images, and text. Monomodal recording is suitable for simple meeting scenarios where only text or audio recording is required. Multimodal recording is suitable for comprehensive analysis and understanding of meeting information, providing richer data sources and enabling the meeting data processing system to more comprehensively analyze and extract information from the meeting data.

[0042] Meeting data is obtained through the aforementioned single-modal or multi-modal recording methods, based on... Figure 2 The system architecture shown allows the acquired meeting data to be input into the aforementioned server, enabling the server to execute the data processing method provided in this embodiment for the meeting scenario. Please refer to... Figure 3 The method includes the following steps:

[0043] S11. Acquire meeting data, the types of which include single-modal meeting data and multimodal meeting data.

[0044] S12. Preprocess the meeting data according to the type of the meeting data to obtain preprocessed meeting data.

[0045] After obtaining the meeting data, a corresponding method is selected to preprocess the meeting data according to the recording mode corresponding to the meeting data. In this embodiment, when the meeting data is recorded in a single mode and is audio, the preprocessing of the audio may include: obtaining the speech content of the same speaker; filtering out interjections in the speech content to obtain preprocessed speech text, which is the preprocessed meeting data. Obtaining the speech content of the same speaker includes: converting the audio to text; identifying the speech activity parts in the audio to determine which time periods contain speech signals; for scenarios where multiple people speak simultaneously, using speaker separation technology (such as mixed speech separation) to separate the speech signals of different speakers in the audio; applying speaker recognition technology to the separated speech signals of each speaker to determine which speech segments belong to the same speaker; and, based on the speaker recognition results and combined with the converted text, merging the speech segments of the same speaker into a continuous speech content, wherein the speech segments can be sorted and merged according to their timestamps and speaker identities.

[0046] Optionally, when the unimodal recorded meeting data is text, preprocessing the text may include: text cleaning to remove noise and invalid characters (such as special symbols and punctuation); segmenting the cleaned text to obtain independent text fragments for each speaker or time period, which allows for better differentiation and analysis of different speakers' content; and filtering out interjections to obtain preprocessed meeting data. Optionally, when the unimodal recorded meeting data is image, preprocessing the image may include: image resizing to meet preset size requirements; image enhancement to improve image quality and visualization; and extracting and converting the image content into text, then performing the same preprocessing operations as described above to obtain preprocessed meeting data.

[0047] When the meeting data was obtained through multimodal recording, please refer to [link / reference]. Figure 4 The meeting data is preprocessed to obtain preprocessed meeting data, including:

[0048] S121. Filter the images in the multimodal conference data to obtain key images with time information.

[0049] Key images are pictures containing important information or representative content. These images may include important scenes related to the meeting's discussion topics, key points of the presentation, charts or images related to the presentation content, etc. The time information for key images can be the time the image was taken or a timestamp related to the meeting's timeline. Please refer to [link / reference]. Figure 5Step S121 specifically includes:

[0050] S1211. Obtain images from the multimodal conference data; extract images from the multimodal conference data, which may come from presentations, projector displays, screenshots, etc. in the conference.

[0051] S1212. Analyze the images, identify and extract images with preset user behaviors.

[0052] Among these methods, image processing and computer vision technologies can be used to analyze the extracted images, identify and extract images with preset user behaviors, such as page turning, screen casting, and screen touching.

[0053] S1213. Obtain images adjacent to the image with the preset user behavior, and calculate the pixel value difference between the adjacent images and the image with the preset user behavior.

[0054] After identifying the image with the preset user behavior, obtain its neighboring images. Calculate the pixel value difference between the neighboring images and the image with the preset user behavior. The pixel value difference can be quantified using image processing techniques, such as calculating pixel differences or structural similarity indices.

[0055] For a given image with a preset user behavior, there may be two adjacent images: the two images immediately before and after the image with the preset user behavior. In this case, the pixel value differences between these two adjacent images and the image with the preset user behavior are calculated separately.

[0056] When calculating the pixel value difference between two images, the number of changed pixels can be determined by calculating this difference, and the number of changed pixels can be used to determine whether the displayed content on the screen has changed. Specifically, two consecutive images are compared at the pixel level, specifically each pixel in the two images is compared, and the difference value is calculated. This can be done by calculating the Euclidean distance or absolute difference between the two pixels. Then, the number of pixels with a difference value greater than a preset threshold is counted to measure the number of changed pixels. Finally, the proportion of changed pixels to the total number of pixels is calculated, i.e., the number of changed pixels divided by the total number of pixels on the screen. If the proportion of changed pixels exceeds a preset threshold (e.g., 10%), it is determined that the displayed content of that image has changed.

[0057] S1214. When the difference in pixel values ​​is greater than a preset threshold, the time information corresponding to the adjacent images is obtained, and the adjacent images with the time information are identified as key images.

[0058] Key images can capture important information and key moments in a meeting. They can provide visual support for meeting minutes and reviews, and allow for the rapid acquisition of key information. Therefore, obtaining key images can provide richer, more intuitive, and more comprehensive meeting information, supporting meeting minutes, reviews, analyses, and decision-making.

[0059] S122. Identify the key image with time information to obtain the structured text content corresponding to the key image.

[0060] OCR (Optical Character Recognition) detection and recognition algorithms can be used to identify key images and convert the text content in the key images into structured text information. OCR technology can recognize text in images and convert it into editable text. Specifically: For each key image, necessary preprocessing operations, such as image enhancement and noise reduction, are first performed to improve the accuracy of subsequent OCR recognition; the preprocessed key image is then input into the OCR recognition algorithm for processing. The OCR algorithm analyzes and recognizes the text in the image and converts it into text format; after obtaining the OCR recognition result, the text content can be further processed and structured, which may include removing incorrectly recognized characters, formatting the text, splitting paragraphs, and extracting key information, in order to better understand and process the text data.

[0061] S123. Based on the time information of the key images, align the key images with the audio in the multimodal conference data to obtain the voice description text corresponding to each key image.

[0062] The process of extracting audio from multimodal conference data involves preprocessing the extracted audio, such as noise reduction and audio enhancement. The preprocessed audio is then input into a speech recognition algorithm, which analyzes the audio and converts it into a text-based speech description. Based on the time information of key images and the speech recognition results, image-speech alignment is performed. By comparing the timestamps of the key images and the speech recognition timestamps, the closest speech description text is found. Specifically, based on the timestamp matching relationship, the closest speech description is matched with the key images to form a correspondence.

[0063] S124. For each key image, compare the structured text content corresponding to the key image with the voice description text corresponding to the key image. When the similarity between the structured text content and the voice description text is greater than a preset threshold, the key image is determined to be the target key image.

[0064] The similarity calculation methods (such as cosine similarity, edit distance, etc.) can be used to compare the similarity between the structured text content and the spoken description text. This will generate a similarity score, representing the degree of similarity between the two. A preset threshold is used to determine whether the similarity between the structured text content and the spoken description text is high enough to determine whether the key image is the target key image. If the similarity is less than or equal to the preset threshold, the key image is not the target key image.

[0065] One method for calculating text similarity is to use pre-trained vector representation models (such as Word2Vec, BERT, GloVe, etc.) to convert text into vector representations. These models map the semantic information of the text into a high-dimensional vector space, so that the similarity between the texts can be measured by calculating the distance or similarity between the vectors. Specifically, this involves calculating the vector representations of the structured text content and the spoken description text, and then using relevant similarity measurement methods to calculate the similarity between them.

[0066] Based on the similarity calculation described above, the key images can be further filtered in step S121 to obtain more accurate key images, i.e., target key images.

[0067] S125. Obtain all the target key images and the corresponding voice description text of the target key images, wherein the target key images and their corresponding voice description text are the preprocessed meeting data.

[0068] Among them, the obtained target key images and their corresponding voice description text form key image-text pairs.

[0069] In this embodiment, the meeting data is preprocessed differently according to its type to obtain preprocessed meeting data. The preprocessed meeting data is then input into the corresponding large model based on its original type. If the preprocessed meeting data is unimodal meeting data, it is input into the large language model of the meeting; if the preprocessed meeting data is multimodal meeting data, it is input into the large visual language model of the multimodal meeting.

[0070] S13. Determine the large model corresponding to the preprocessed meeting data; the large model includes a meeting large language model and a multimodal visual language large model, the meeting large language model is used to process the single-modal meeting data, and the multimodal visual language large model is used to process the multimodal meeting data.

[0071] The large-scale language model for meetings is used to process unimodal meeting data, such as text, speech, or other single-modal data. The large-scale visual language model for multimodal meetings is used to process multimodal meeting data, such as data combining text, images, and audio. The large-scale language model for meetings and the large-scale visual language model for multimodal meetings can be used together to achieve comprehensive processing and analysis of different types of meeting data.

[0072] In this embodiment, the large language model for the conference can be a large language model containing a multi-layer transformer structure. This model predicts the probability of the next token tk+1 appearing based on the prefix t1, t2, t3, ..., tk of the input token sequence, i.e., p(tk+1|t1,t2,t3,...tk). This process is repeated in a word-chain manner until the generated token sequence reaches a certain length or other loop termination conditions are met, thus obtaining a complete string. The Transformer is a neural network architecture based on self-attention, widely used in natural language processing tasks. The Transformer network structure consists of an encoder and a decoder, each module composed of multiple identical stacked layers. The input sequence is represented as a series of tokens, where each token can be a word, character, or other discrete symbol. The model learns the correlation between each position and other positions based on the prefix of the input token sequence through a self-attention mechanism, and predicts the probability of the next token. In the conference language model, the input token sequence can be conference-related text data, such as meeting minutes or agenda items. By learning the contextual information in the text sequence, the model can predict the probability of the next token appearing, thus gradually generating a complete string. The training process of this conference language model is described in detail below.

[0073] The overall network architecture of the multimodal visual language large model comprises three components: a large language model, a visual encoder, and a visual language adapter. In this embodiment, the large language model can be the aforementioned conference large language model. The network architecture consisting of these three components enables comprehensive understanding and processing of images and text. The training process of this multimodal visual language large model is described in detail below.

[0074] S14. Input the preprocessed meeting data into the large model, and output the data results corresponding to the target task through the large model. The target task includes at least one of intelligent chapter, meeting minutes, speaker summary and meeting to-do list.

[0075] The target tasks are different types of data results generated from the processing of meeting data. The generation of target tasks is based on the analysis and processing of meeting data, aiming to provide a summary, refinement, and organization of the meeting content to help users better understand and utilize meeting information. In this embodiment, one or more target tasks can be selected for processing and the corresponding data results obtained according to user needs.

[0076] Intelligent chapters refer to the segmentation of text into chapters with independent themes or content. Intelligent chapter analysis aims to identify and extract key information from the text and divide it according to certain logical and semantic rules, making it easier for users to browse, navigate, and understand the text content. Similar to a book's table of contents, intelligent chapters provide an overview and navigation of the entire text, enabling users to quickly locate and understand content of interest. In a meeting setting, this can help participants better grasp the meeting's structure and main content.

[0077] Meeting minutes are documents or texts that record and summarize a meeting. They are a concise description of the meeting's content, discussions, and decisions, aiming to record the key points, outcomes, and actions taken so that participants can review, communicate, and follow up on the meeting's content. The accuracy and completeness of meeting minutes are crucial for effective meeting management and information communication.

[0078] A speaker's summary is a concise summary of each speaker's remarks, allowing attendees to quickly grasp the main points and key messages. Speaker summaries may include the speaker's identity, main viewpoints, key points, opinions, or suggestions. Through speaker summaries, attendees can better understand the meeting's theme and focus of discussion, enabling them to respond and make appropriate decisions.

[0079] Meeting to-do lists are lists of tasks generated during a meeting, allowing participants to know what needs to be done or what actions need to be taken afterward. Meetings typically produce various decisions, action plans, and tasks. The purpose of meeting to-do lists is to clearly document these items and assign them to relevant personnel or teams to ensure subsequent execution and follow-up. Meeting to-do lists can include task descriptions, responsible parties, priorities, etc. By recording and sharing meeting to-do lists, participants can clearly understand the tasks that need to be completed after the meeting, avoiding omissions and confusion, and ensuring that the implementation of decisions and meeting outcomes are translated into practical actions. This helps improve meeting efficiency, drive work progress, and ensure that meeting results are effectively implemented.

[0080] In this embodiment, the preprocessed meeting data is input into the large model, and the large model outputs the data results corresponding to the target task. The main method used is the use of prompts to guide the large model in generating the corresponding data results. By providing specific instructions or task descriptions in the prompts, the large model can generate corresponding content based on different prompts. A prompt refers to the input text or instructions provided to the large model when interacting with it. It is a piece of text provided in the input to clarify the task or guide the model to generate specific content. A prompt can include a problem description, task description, contextual information, or specific instructions to guide the large model in generating output results related to the target task. The prompt plays a crucial role in interacting with the large model and guiding output generation, helping the model understand task requirements and generate corresponding results.

[0081] Please refer to Figure 6 When the large model is the conference large language model, the step of inputting the preprocessed conference data into the large model and outputting the data results corresponding to the target task through the large model includes:

[0082] S1411. Obtain the prompt instruction corresponding to the target task;

[0083] S1412. Combine the spoken text and the prompting instructions corresponding to the target task to construct input data;

[0084] S1413. Input the constructed input data into the conference large language model to generate the content required for the target task.

[0085] The prompt instruction can be understood as the prompt mentioned above. It is used to tell the large language model of the conference the specific task to be completed or the expected output result. It can be a problem description, task description or specific instruction to guide the large model to generate content related to the target task.

[0086] The spoken text is preprocessed meeting data. The spoken text and the prompts for the target task are combined to construct the input data. This can be achieved by placing the spoken text before or after the prompts and separating them with appropriate delimiters. The constructed input data is a complete text string containing the spoken text and prompts relevant to the target task. This combination helps the large model understand the context and requirements of the task.

[0087] The constructed input data is fed into the conference language model. Based on the spoken text and prompts from the input data, the conference language model generates content relevant to the target task. This content can be generated text, suggestions, answers, summaries, etc., depending on the specific target task.

[0088] For example, if the target task is a meeting summary, the prompt input to the meeting's large language model can be in the following format. When generating content, the meeting's large language model extracts the text from the "instruction" field and generates the corresponding result according to its generation process. The meeting summary prompt design is as follows:

[0089] {

[0090] "instruction": "\n###{Speaker's Content}###\nYou are a meeting summary robot. Please summarize the above meeting discussion content using a contiguous sentence structure. The output should be fluent, concise, and in one paragraph (less than 250 characters):"

[0091] }

[0092] For example, if the target task is a meeting summary, the prompt input to the meeting's large language model can be in the following format. When generating content, the meeting's large language model extracts the text from the "instruction" field and generates the corresponding result according to its generation process. The meeting summary prompt design is as follows:

[0093] {

[0094] "Instruction": "\n###{Speaker's Content}###\n\nYou are a meeting summary robot, capable of summarizing the main points of the meeting discussion. Please summarize the main points of the above meeting discussion, requiring: 1. No unrelated content can be generated; 2. Output in bullet points."

[0095] }

[0096] For example, if the target task is a meeting schedule, the prompt input to the meeting's large language model can be in the following format. When generating content, the meeting's large language model extracts the text from the "instruction" field and generates the corresponding result based on its generation process. The meeting schedule prompt design is as follows:

[0097] {

[0098] "instruction": "\n###{Speaker's Content}###\nThe definition of meeting to-dos refers to matters discussed or decided during a meeting that require specific execution or processing after the meeting. The definition and processing methods for meeting to-dos can vary depending on the specific meeting purpose and agenda. Clearly defining to-dos during the meeting and following up and executing them afterward ensures the effective implementation and realization of the meeting's outcomes. You are an intelligent meeting assistant; please strictly adhere to the above definition of meeting to-dos and generate meeting to-dos for the above meeting content."

[0099] }

[0100] For example, if the target task is for the speaker to summarize, the prompt input to the conference language model can be in the following format. When generating content, the conference language model extracts the text from the "instruction" field and generates the corresponding result according to its generation process. The speaker summary prompt design is as follows:

[0101] {

[0102] "instruction": "\n###{Speaker's Content}###\nYou are an intelligent meeting assistant. The speakers for the above meeting content are distinguished by no.x. Please summarize the above meeting content according to the speaker distinction."

[0103] }

[0104] For example, if the target task is intelligent chapters, intelligent chapters require further incorporation of time information to construct the corresponding prompts. Specifically, this involves extracting the time information (which can be calculated in milliseconds) corresponding to the first character of each sentence based on the ending punctuation mark, and then appending this information to the beginning of the sentence. For instance, the text content recognized by ASR might be:

[1000] Hi everyone, welcome to this episode of Horoscope. I'm your host, Teacher Gu. Today we'll talk about the months for three of the twelve zodiac signs. The text after adding time stamp processing steps is as follows:

[1000] Hi everyone, welcome to this episode of Horoscope. I'm your host, Teacher Gu.

[6000] Today we'll talk about what the months are like for three of the 12 zodiac signs. All time information, along with the text containing that time information, is input into a large-scale language model for intelligent chapter analysis, including chapter segmentation and the generation of corresponding chapter summaries. Intelligent chapter analysis is performed in steps: first, the input is segmented into chapters to obtain the start time and topic for each chapter; then, summaries are generated for the text content within each chapter segment. The prompt input to the large-scale language model can be in the following format. When generating content, the large-scale language model extracts the text from the "instruction" field and generates the corresponding results according to its generation process. The prompt design has two forms: chapter segmentation and generating summaries for corresponding chapters. The chapter segmentation prompt design is as follows: { "instruction": "###{Speaker's content including time information}###. {All time information}. The content enclosed in ### is the transcript of a meeting, and the content enclosed in {} is all the start times corresponding to the content of this meeting. Please divide the meeting content into topics according to the discussion topics, and obtain the content of each topic and its corresponding start time. The output format is: Start Time: xxx, Topic: yyy." } The prompt for generating summaries for the corresponding chapters is designed as follows: { "instruction": "###{Content of the speaker corresponding to the chapter}###. Please summarize the above text. Output format: Summary: xx\n" } Please refer to Figure 7 When the large model is the multimodal visual language large model, the preprocessed meeting data includes multiple target key images, each target key image corresponding to a language description text. The process of inputting the preprocessed meeting data into the large model and outputting the data results corresponding to the target task through the large model includes: S1421. Obtain the prompting instruction corresponding to the target task; wherein, the prompting instruction is to tell the multimodal visual language large model the specific task to be completed or the expected output result. S1422. Combine each target key image, the voice description text corresponding to each target key image, and the prompt instructions corresponding to the target task to construct multiple input data; S1423. Input the constructed multiple input data into the multimodal visual language large model respectively to generate the result corresponding to each target key image; S1424. Input the results corresponding to each target key image into the multimodal visual language large model, and simultaneously input the prompt instructions corresponding to the target task into the multimodal visual language large model to generate the content required for the target task. In multimodal visual language tasks, each target key image is an important input element, providing rich information along with its corresponding language description text. For each target key image, it is combined with its corresponding language description text and the prompts for the target task to form a complete input data set. Since meeting scenarios typically involve multiple target key images, each target key image, along with its corresponding language description text and the prompts for the target task, constitutes one input data set, resulting in multiple input data sets. These multiple input data sets are fed into a multimodal visual language model to generate results corresponding to each target key image. These results are related to the target task; for example, if the target task is intelligent chapters, then the generated results for each target key image are the data results for the intelligent chapters corresponding to each target key image. This yields the intelligent chapters corresponding to all target key images. These intelligent chapters are then input into the multimodal visual language model, along with the prompts for the target task. These prompts are used to suggest the generation of intelligent chapters. The multimodal visual language model processes and analyzes the intelligent chapters corresponding to all target key images to obtain the data results for the intelligent chapters of the meeting data. For example, if the target task is a meeting summary, and the input is a multimodal video recording, then a multimodal visual-language model is requested. This model combines the target key images obtained from the multimodal recording with their corresponding audio descriptions. Multiple requests to the multimodal visual-language model yield multiple meeting summaries. These multiple summaries are then summarized and synthesized again by the multimodal visual-language model to create a unified meeting summary. The specific steps include: Step 1: The prompt input to the multimodal visual language model can be in the following format. When generating content, the multimodal visual language model extracts the text from the "instruction" field and generates the corresponding results according to the model's generation process. The meeting summary prompt design is as follows: { "instruction": "Picture 1: / absolute path / demo.jpeg\n###{Speaker's Content}###\nYou are a meeting summary robot. The content enclosed in ### is the spoken text content discussed during the meeting based on this image. Please combine the image information and the content enclosed in ### above, and use a complex sentence structure to summarize and analyze the image and meeting content. The output should be fluent, concise, and in one paragraph (less than 250 characters): } Multiple requests were made to the multimodal visual language large model, resulting in multiple conference summaries. Step 2: The multiple meeting summary results obtained in Step 1 are summarized and synthesized again using the model to create a unified meeting summary. The meeting summary prompt is designed as follows: { "instruction": "\n###{Results of Multiple Meeting Summaries}###\nYou are a meeting summary robot. Please summarize the results of the multiple meeting summaries above:" } For example, if the target task is to summarize a meeting, and the input is a multimodal video recording, then a multimodal visual-language model is requested. This model combines the target key images obtained from the multimodal recording with their corresponding audio descriptions. Multiple requests to the multimodal visual-language model yield multiple meeting summaries. These multiple summaries are then further summarized and synthesized using the multimodal visual-language model to create a unified meeting summary. The specific steps include: Step 1: The prompt input to the multimodal visual language model can be in the following format. When generating content, the multimodal visual language model extracts the text from the "instruction" field and generates the corresponding results according to the model's generation process. The prompt design summarized from the conference is as follows: { "instruction": "Picture 1: / absolute path / demo.jpeg\n###{Speaker's Content}###\nYou are a meeting summary robot. The content enclosed in ### is the spoken text content discussed during the meeting based on this image. Please summarize the main content of the meeting discussion, combining the image information and the content enclosed in ### above. Please summarize the main content of the meeting discussion above, requiring: 1. No unrelated content generated by yourself; 2. Output in bullet points. } Multiple requests were made to the multimodal visual language large model, and the results were summarized from multiple conferences. Step 2: The multiple meeting summary results obtained in Step 1 are further summarized and synthesized using a multimodal visual language model to form a unified meeting summary. The meeting summary prompt is designed as follows: { "instruction": "\n###{Results of Multiple Meeting Summaries}###\nYou are a meeting summary robot. Please summarize and categorize the results of the multiple meeting summaries above according to their format:" } For example, if the target task is a meeting schedule, and the input is a multimodal video recording, then a multimodal visual-language model is requested. This model combines the target key images obtained from the multimodal recording with their corresponding audio descriptions. Multiple requests to the multimodal visual-language model yield multiple meeting schedule results. These multiple results are then summarized and synthesized again using the multimodal visual-language model to create a unified meeting schedule. Specific steps include: Step 1: The prompt input to the multimodal visual language model can be in the following format. When generating content, the multimodal visual language model extracts the text from the "instruction" field and generates the corresponding results according to the model's generation process. The meeting to-do list prompt design is as follows: { "instruction": "Picture 1: / absolute path / demo.jpeg\n###{Speaker's Content}###\nThe definition of meeting to-dos refers to matters discussed or decided during a meeting that require specific execution or processing after the meeting. The definition and processing methods for meeting to-dos can vary depending on the specific meeting purpose and agenda. Clearly defining to-dos during the meeting and following up and executing them afterward ensures the effective implementation and realization of the meeting's outcomes. You are an intelligent meeting assistant. The content enclosed in ### is the spoken text content discussed based on this image during the meeting. Please strictly follow the definition of meeting to-dos above, combining the image information and the content enclosed in ### above, to generate meeting to-dos. } Multiple requests were made to the multimodal visual language large model, resulting in multiple pending conference results. Step 2: The multiple meeting to-do lists obtained in Step 1 are summarized and synthesized again using a multimodal visual language model to create a unified meeting to-do list. The meeting to-do list prompt is designed as follows: { "instruction": "\n###{Results of multiple pending meetings}###\nPlease summarize and generalize the results of the multiple pending meetings mentioned above:"} For example, if the target task is a speaker's summary, and the input is a multimodal video recording, then a multimodal visual language model is requested. The target key image obtained through multimodal recording is combined with its corresponding multimodal visual language model. Multiple requests to the multimodal visual language model yield multiple speaker summaries. These multiple summaries are then further summarized and synthesized using the multimodal visual language model to create a unified speaker summary. Specific steps include: Step 1: The prompt input to the multimodal visual language model can be in the following format. When generating content, the multimodal visual language model extracts the text from the "instruction" field and generates the corresponding results according to the model's generation process. The speaker's summary of the prompt design is as follows: { "instruction": "Picture 1: / absolute path / demo.jpeg\n###{Speaker's Content}###\nYou are a meeting summary robot. The content enclosed in ### is the conversational text content discussed during the meeting based on this image, differentiated by no.x. Please combine the image information and the content enclosed in ### above to summarize the content according to the speaker's distinction. } Multiple requests were made to the multimodal visual language model, and the results were summarized by multiple speakers. Step 2: The multiple speaker summaries obtained in Step 1 are further summarized and synthesized using a multimodal visual language model to form a unified speaker summary. The speaker summary prompt is designed as follows: { "instruction": "\n###{Results summarized by multiple speakers}###\nYou are a meeting summary robot. Please summarize the results summarized by multiple speakers above according to their format:" } For example, if the target task is intelligent chapters, and the input is multimodal recorded video, then a multimodal visual language model is requested. This model combines the target key images obtained from the multimodal recording with their corresponding time-added audio descriptions. Multiple requests are made to the multimodal visual language model to obtain the start time and topic of each chapter. Finally, a summary is generated from the text content of each chapter section using the multimodal visual language model. Specific steps include: Step 1: The prompt input to the multimodal visual language model can be in the following format. When generating content, the multimodal visual language model extracts the text from the "instruction" field and generates the corresponding results according to the generation process of the visual language model. The chapter-segmented prompt design is as follows: { "instruction": "Picture 1: / absolute path / demo.jpeg\n###{Speaker's content with added time information}###. {All time information}. You are a meeting summary robot. The content enclosed in ### is the spoken text content discussed based on this image during the meeting. The content enclosed in {} is all the start times corresponding to the content of this meeting. Please combine the image information and the content enclosed in ### and {} above to segment the meeting content into topics according to the discussion topics, and obtain the content of each topic and its corresponding start time. The output format is: Start time: xxx, Topic: yyy. \n" } Multiple requests were made to the multimodal visual language model to obtain the start time and corresponding topic for each chapter. Step 2: Based on the start time obtained in Step 1, the text content corresponding to each chapter interval can be obtained. The model then generates the corresponding title and summary. The corresponding prompt design is as follows: { "instruction": "###{Content of the speaker corresponding to the chapter}###. Please summarize the above text. Output format: Summary: xx\n" } The data processing method for conference scenarios provided in this application embodiment processes unimodal conference data using a large-scale conference language model. Since unimodal conference data focuses on a single media format, this allows the large-scale conference language model to provide higher processing speed and efficiency when processing large-scale unimodal conference data. It also offers advantages such as in-depth mining of linguistic information and refined semantic understanding. These advantages enable the model to provide accurate and high-quality results in the analysis and processing of unimodal conference data. Furthermore, the method processes multimodal conference data using a multimodal visual language model. Multimodal conference data contains information from multiple media formats. The multimodal visual language model can simultaneously process this different media information and perform cross-modal fusion analysis. For example, by associating image and text information, the model can understand the conference content from multiple dimensions, providing more comprehensive and richer analysis results. In addition, whether it is a large language model for meetings or a large visual language model for multimodal meetings, by using large models to process meeting data, various tasks can be performed automatically, such as generating intelligent chapters, meeting minutes and speaker summaries, thereby reducing the workload of manual processing, making it convenient to use and improving efficiency; moreover, large models can apply rich context and semantic understanding when processing data, thereby providing more accurate and consistent analysis results. Please see Figure 8 , Figure 8 This is a flowchart illustrating a data processing method in a meeting scenario, provided by another embodiment of this application. For example... Figure 8 As shown, the method includes the following steps: S21. Acquire meeting data, the types of which include single-modal meeting data and multimodal meeting data. S22. Preprocess the meeting data according to the type of the meeting data to obtain preprocessed meeting data. S23. Determine the large model corresponding to the preprocessed meeting data; the large model includes a meeting large language model and a multimodal visual language large model, the meeting large language model is used to process the single-modal meeting data, and the multimodal visual language large model is used to process the multimodal meeting data. S24. Display a table of contents on the graphical interface, including smart chapters, meeting minutes, speaker summaries, and meeting tasks. S25. When any event corresponding to the smart chapter, meeting minutes, speaker summary, and meeting to-do list in the directory is triggered, the preprocessed meeting data is input into the large model, and the data result corresponding to the target task is output through the large model, wherein the target task is the target task corresponding to the triggered event. Figure 8 The corresponding data processing methods in the meeting scenario and Figure 3The main difference is that, in this embodiment, the graphical interface displays a directory including smart chapters, meeting minutes, speaker summaries, and meeting tasks. When a user triggers any event in the directory, the pre-processed meeting data is input into the large model, which then outputs the target task data results corresponding to the triggered event. This embodiment provides an interactive user interface, allowing users to select specific target tasks for processing as needed. The interactive interface enables users to freely choose target tasks of interest. Different users may have different needs and priorities for different target tasks, thus allowing users to selectively process specific target tasks according to their own needs. This provides greater flexibility and customization, meeting users' personalized requirements. Furthermore, by allowing users to select specific target tasks, unnecessary tasks can be avoided, thereby improving processing efficiency. Finally, the interactive interface provides a personalized user experience, allowing users to customize their choices according to their needs and preferences. This enhances user engagement and satisfaction, providing a better user experience. In some embodiments, after obtaining the preprocessed meeting data, the method further includes: using the large model to perform discourse correction on the preprocessed meeting data to obtain discourse-corrected content; the discourse-corrected content is used to input the large model to generate the intelligent chapter, meeting minutes, speaker summary, and meeting to-do list through the large model. In this process, both unimodal and multimodal meeting data can undergo discourse shaping first. The corresponding large-scale model is used for discourse shaping, outputting the shaped content. This shaped content is then input into the large-scale model and combined with relevant prompts to output the data content corresponding to the target task. Discourse shaping refers to processing pre-processed meeting data to generate coherent and logical content. Through discourse shaping, various fragments, paragraphs, or elements in the meeting data can be integrated and sorted, presenting them in a logical order and matching the overall structure and theme of the meeting. This helps improve the quality and readability of generated intelligent chapters, meeting minutes, speaker summaries, and meeting to-do lists. Therefore, discourse shaping provides a better input foundation for subsequent large-scale model processing and target task generation. For example, for a transcript recorded in a single modal manner with speakers, the content of each speaker is input into a large-scale conference language model for discourse normalization, based on the speaker's distinction. The prompt input into the large-scale conference language model can be in the following format: When generating content, the large-scale conference language model extracts the text from the "instruction" field and generates the corresponding result according to the model's generation process. The discourse normalization prompt design is as follows: { "instruction": "Please standardize the text enclosed in ### below to make it more accurate, concise, and formal.\n\n-Accuracy: Correct typos.\n-Conciseness: Retain the key components of each sentence, remove unnecessary modifiers, conjunctions, meaningless expressions, and repetitive expressions.\n-Formality: Use written language, use verb-object expressions, and remove colloquial expressions such as 'um,' 'ah,' 'uh,' 'oh,' 'haha,' etc.\n\nText to be standardized:\n###{Speaker's content}###\" } A single meeting may involve multiple speakers. A large model can perform discourse normalization on each speaker's content, ultimately producing multiple normalized texts, which can then be used as the normalized content of each speaker. In some embodiments, after obtaining the preprocessed meeting data, the method further includes: extracting keywords based on the preprocessed meeting data, and displaying the extracted keywords on a graphical interface. Specifically, the most representative and information-rich words or phrases are extracted from the preprocessed meeting data for display on the graphical interface. These keywords can be extracted from audio-converted speech text and video-converted text-image pairs using a preset algorithm. The extracted keywords help provide functions such as summarizing and overviewing, navigation and guidance, identification of interests and concerns, and content retrieval and indexing, providing users with a better experience in browsing and searching meeting data. In some embodiments, please refer to Figure 9 This application provides a method for obtaining the aforementioned large language model of a conference, such as... Figure 9 As shown, the training methods for the conference large language model include: S31. Obtain the vocabulary dataset. A vocabulary dataset provides a vocabulary for text processing tasks, defining the input and output units of the model, as well as the encoding and decoding rules for the words. For example... Figure 10 As shown, the process of obtaining the vocabulary dataset includes: S311. Obtain a preset amount of raw text data; this can be large-scale text data obtained from sources such as the Internet, books, news articles, and social media. S312. Determine the segmentation method and vocabulary parameters; the segmentation method may be, for example, BPE (BytePair Encoding). Vocabulary parameters include vocabulary size, a list of supported languages, split_digits, and byte_fallback. S313. Using a preset word segmentation tool, the original text data is segmented based on the word segmentation method and the word list parameters to generate a first word list; the word segmentation tool can be Sentencepiece, etc. S314. Merge the first vocabulary with the original second vocabulary to obtain a merged expanded vocabulary, which constitutes the vocabulary dataset. The second vocabulary can be the vocabulary of the already obtained large model. In this embodiment, by merging the generated vocabulary with the already obtained vocabulary, the coverage of the vocabulary can be enriched, and the model's ability to process diverse text data can be improved. In this embodiment, a Chinese tokenizer model trained using the SentencePiece tool can be merged with the native tokenizer of the conference large language model. By merging the vocabulary, the efficiency of Chinese text processing and the accuracy of model generation can be improved. This method can effectively solve the problem of the native tokenizer when processing Chinese text. The native tokenizer's vocabulary may only contain a small number of Chinese characters, causing a single Chinese character to be segmented into multiple tokens (basic units that divide text into discrete units), thus reducing encoding and decoding efficiency. The Chinese tokenizer model trained using the SentencePiece tool can better process Chinese text, segmenting Chinese characters into more appropriate tokens. By merging the vocabularies of the two tokenizers, the Chinese word segmentation capabilities of SentencePiecetokenizer can be fully utilized while retaining other advantages of the original tokenizers. The merged tokenizer model can improve encoding and decoding efficiency in Chinese text processing tasks and better handle Chinese-specific language learning problems, thereby improving the accuracy of model generation. Optionally, the above method for merging vocabularies can be adjusted and optimized according to the specific task and data requirements to achieve the best Chinese text processing results. For English letters and numbers, no word segmentation is required because they are already separated by spaces or other specific symbols. Therefore, the native word segmenter can be used to meet the requirements for English letters and numbers. S32. Obtain the target instruction dataset. like Figure 11 As shown, the acquisition of the target instruction dataset includes: S321. Obtain the unlabeled dataset; S322. Generate an evaluation instruction for the unlabeled data in the unlabeled dataset, wherein the evaluation instruction includes an empty instruction and a non-empty instruction; S323. Delete the unlabeled data corresponding to the empty instruction from the unlabeled dataset to obtain a candidate instruction dataset, wherein the candidate instruction dataset consists of unlabeled data and non-empty instructions corresponding to the unlabeled data; S324. Perform quality scoring on the candidate instruction dataset to obtain unlabeled data whose quality score results are within a preset score range; S325. The unlabeled data of the quality scoring results within the preset score range, and the non-empty instructions corresponding to the unlabeled data are determined as the target instruction dataset. The above method generates instructions from unlabeled data by constructing appropriate prompts and filters them through quality scoring, thereby obtaining a target instruction dataset for training and evaluating the model. The methods for scoring the quality of the candidate instruction dataset may include manual evaluation, using automatic evaluation metrics to measure the quality of the candidate instruction dataset, or using pre-trained language models or other relevant models to score the candidate instruction dataset. S33. Use the vocabulary dataset and the target instruction dataset as training datasets. Before merging the vocabulary dataset and the target instruction dataset, it's advisable to analyze both to ensure their distribution across categories or samples is reasonable. If a category or sample is too small in the dataset, the model may struggle to learn its features during training, impacting performance. Conversely, an excessive number of a category or sample may cause the model to favor that category or sample while neglecting other categories or samples. S34. Obtain the large language model of the meeting to be trained. The process of obtaining the conference language model to be trained includes: adding noise to the embedding layer of the pre-acquired target language model to fine-tune the embedding layer of the target language model, thereby obtaining the conference language model to be trained. Wherein, the sampling range of the noise is Between these, ∂ is an adjustable parameter, L is the input length, and d is the dimension of the embedding layer. Specifically, the embedding layer is the embedding layer of the conference large language model. Adding noise to the embedding layer can improve the fine-tuning of instructions. Specifically, noise acts as a regularization mechanism, helping to mitigate overfitting; it introduces diversity into the embedding space, making similar inputs distinct in their embedding representations, thus increasing the model's sensitivity to inputs, enabling it to better distinguish different input samples and improve its ability to perceive subtle differences; moreover, the introduction of noise is equivalent to perturbing the input, which can be seen as a form of data augmentation. In summary, the introduction of noise helps increase the model's robustness, generalization ability, and ability to perceive subtle differences, while also serving as a regularization and data augmentation function. By adding noise to the embedding layer, the fine-tuning of instructions can be improved, making the large-scale language model to be trained more adaptable to the needs of specific tasks. S35. Pre-train the conference language model to be trained using the training dataset until a preset convergence condition is met. The pre-training of the conference language model is then considered complete, and the conference language model is obtained. The preset convergence condition refers to setting a condition to stop training during pre-training. When the model meets this condition, the pre-training process terminates, and pre-training is complete. For example, monitoring the change in the loss function during training; when the change in the loss function falls below a certain threshold or reaches a stable state, the pre-training process terminates. The training method for the large-scale language model for meetings provided in this embodiment can improve the model's performance in meeting-related tasks and scenarios, enhancing its language understanding, instruction processing, robustness, and generalization capabilities. By pre-training and merging different datasets, the model can better adapt to meeting-related language tasks and produce more accurate and fluent meeting-related outputs. In some embodiments, please refer to Figure 12 This application provides a method for obtaining the aforementioned multimodal visual language large model, such as... Figure 12 As shown, the training methods for a large multimodal visual language model include: S41. Obtain a multimodal visual language large model to be trained, wherein the multimodal visual language large model to be trained includes the conference large language model, the visual encoder, and the language adapter; S42. Obtain a training dataset, the training dataset including a first training dataset and a second training dataset, the first training dataset including target key images and language description text corresponding to the target key images, the second training dataset including text and target instruction data associated with the text; S43. While ensuring that the parameters of the conference large language model remain unchanged, the visual encoder and language adapter are pre-trained using the first training dataset until the preset first convergence condition is met, at which point the pre-training of the visual encoder and language adapter is determined to be complete; the visual encoder is used to process image input and generate image features, and the language adapter is used to combine image features with language features. S44. Using the first training dataset, perform multi-task pre-training on the conference large language model, visual encoder and language adapter until the preset second convergence condition is met, then determine that the multi-task pre-training of the conference large language model, visual encoder and language adapter is complete. S45. While ensuring that the parameters of the visual encoder remain unchanged after pre-training, the pre-trained conference language model and language adapter are fine-tuned using the second training dataset until the preset third convergence condition is met, at which point the fine-tuning of the conference language model and language adapter is determined to be complete. S46. The model consisting of a pre-trained and fine-tuned conference language model, a visual encoder, and a language adapter cascaded together is identified as a multimodal visual language model. The aforementioned conference language model, which is the same as the conference language model in the above embodiments, serves as a basic component and is capable of learning rich language representations and semantic relationships, providing powerful language processing capabilities for the entire multimodal visual language model. A visual encoder processes input images and generates a set of image features. It receives the input image and, through operations such as self-attention mechanisms and multilayer perceptrons, transforms it into a set of meaningful feature representations. Visual encoders are used to extract visual information from images, providing semantic understanding capabilities for large-scale multimodal visual language models. A language adapter combines image features with language features, reducing computational complexity while preserving semantic information. In the pre-training phase of the visual encoder and language adapter, a preset number of training epochs or iterations can be set as the first convergence condition. A preset number of training epochs or iterations can also be set as the second convergence condition; alternatively, the second convergence condition can be determined based on the model's performance metrics on a multi-task pre-training dataset, such as accuracy, loss function, or other task-specific metrics. In the instruction fine-tuning phase, the model reaching a preset performance metric or error range on the training dataset serves as the third convergence condition. The training method for multimodal visual language large models improves the model's comprehensive understanding, context perception, and feature fusion and interaction capabilities by comprehensively utilizing linguistic and visual information. This enables the model to more accurately understand and generate content in multimodal tasks and has better generalization ability. Please see Figure 13 This application provides a data processing device for a conference scenario, the device 50 including: The meeting data acquisition module 501 is used to acquire meeting data, including unimodal and multimodal meeting data. The preprocessing module 502 is used to preprocess the meeting data according to its type to obtain preprocessed meeting data. The large model determination module 503 is used to determine the large model corresponding to the preprocessed meeting data; the large model includes a meeting large language model and a multimodal visual language model, the meeting large language model being used to process the unimodal meeting data, and the multimodal visual language model being used to process the multimodal meeting data. The target task generation module 504 is used to input the preprocessed meeting data into the large model and output data results corresponding to the target task through the large model; the target task includes at least one of intelligent chapters, meeting minutes, speaker summaries, and meeting tasks. It should be noted that the data processing device in the above-described meeting scenario can execute the data processing method for the meeting scenario provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the method. For technical details not described in detail in the embodiments of the data processing device for the meeting scenario, please refer to the data processing method for the meeting scenario provided in the embodiments of this application. Figure 14 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. This electronic device can be used to execute the data processing method in the above-described conference scenario. Figure 14 As shown, the electronic device 60 includes: one or more processors 601 and a memory 602. Figure 14 Taking a processor 601 as an example, the processor 601 and the memory 602 can be connected via a bus or other means. Figure 14 Taking the example of a connection between China and Israel via a bus. The memory 602, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the data processing method in the conference scenario in the embodiments of this application (e.g., attached...). Figure 13 (The various modules shown). The processor 601 executes various functional applications and data processing of the electronic device by running non-volatile software programs, instructions, and modules stored in the memory 602, that is, it implements the data processing method in the meeting scenario of the above method embodiment. The memory 602 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the data processing device in the conference scenario. Furthermore, the memory 602 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 602 may optionally include memory remotely located relative to the processor 601, and these remote memories can be connected to the data processing device in the conference scenario via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The one or more modules are stored in the memory 602, and when executed by the one or more processors 601, they perform the data processing method for the conference scenario in any of the above method embodiments. The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application. The electronic devices in this application can exist in various forms, including but not limited to: ultra-mobile personal computer devices, servers, and other electronic devices with data interaction functions. This application provides a non-volatile computer-readable storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 14 One of the processors 601 enables the one or more processors to execute the data processing method in the conference scenario of any of the above method embodiments. This application provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, which, when executed by the electronic device, enable the electronic device to perform the data processing method in the conference scenario described in any of the above method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of this application as described above, which are not provided in detail for the sake of brevity; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A data processing method for a meeting scenario, characterized in that, include: Acquire meeting data, the types of which include single-modal meeting data and multimodal meeting data; The meeting data is preprocessed according to its type to obtain preprocessed meeting data; Determine the large model corresponding to the preprocessed meeting data; the large model includes a meeting large language model and a multimodal visual language large model, the meeting large language model is used to process the single-modal meeting data, and the multimodal visual language large model is used to process the multimodal meeting data; The preprocessed meeting data is input into the large model, and the large model outputs the data results corresponding to the target task. The target task includes at least one of intelligent chapters, meeting minutes, speaker summaries, and meeting tasks.

2. The method according to claim 1, characterized in that, Before performing the step of inputting the preprocessed meeting data into the large model, the method further includes: The graphical interface displays a table of contents including the smart chapters, meeting minutes, speaker summaries, and meeting tasks. When any event corresponding to the smart chapter, meeting minutes, speaker summary, and meeting to-do list in the directory is detected to be triggered, the step of inputting the preprocessed meeting data into the large model and outputting the data result corresponding to the target task through the large model is executed, wherein the target task is the target task corresponding to the triggered event.

3. The method according to claim 1, characterized in that, When the type of the meeting data is the single-modal meeting data, the step of preprocessing the meeting data according to the type of the meeting data to obtain preprocessed meeting data includes: The single-modal conference data is analyzed and processed to obtain the speech content of the same speaker; Filter out interjections from the speech content to obtain the speech text, which is the preprocessed meeting data.

4. The method according to claim 1, characterized in that, When the type of the meeting data is multimodal meeting data, the step of preprocessing the meeting data according to the type of the meeting data to obtain preprocessed meeting data includes: Images in the multimodal conference data are filtered to obtain key images with time information; The key images containing time information are identified to obtain the structured text content corresponding to the key images; Based on the time information of the key images, the key images are aligned with the audio in the multimodal conference data to obtain the voice description text corresponding to each key image; For each key image, the structured text content corresponding to the key image and the voice description text corresponding to the key image are compared. When the similarity between the structured text content and the voice description text is greater than a preset threshold, the key image is determined to be the target key image. Obtain all the target key images and the corresponding voice description text of the target key images. The target key images and their corresponding voice description text constitute the preprocessed meeting data.

5. The method according to claim 4, characterized in that, The step of filtering images from the multimodal conference data to obtain key images with time information includes: Acquire images from the multimodal conference data; Analyze the images to identify and extract images that exhibit preset user behaviors; Obtain images adjacent to the image with the preset user behavior, and calculate the pixel value difference between the adjacent images and the image with the preset user behavior; When the difference in pixel values ​​is greater than a preset threshold, the time information corresponding to the adjacent images is obtained, and the adjacent images with the time information are identified as key images.

6. The method according to claim 3, characterized in that, The large model is the conference large language model. The process involves inputting the preprocessed conference data into the large model and outputting the data results corresponding to the target task through the large model, including: Obtain the prompt instructions corresponding to the target task; The input data is constructed by combining the spoken text and the prompts corresponding to the target task. The constructed input data is input into the conference large language model to generate the content required for the target task.

7. The method according to claim 1, characterized in that, The training methods for the conference large language model include: Obtain the vocabulary dataset; Obtain the target instruction dataset; Use the vocabulary dataset and the target instruction dataset as training datasets; Obtain the large language model of the meeting to be trained; The conference language model to be trained is pre-trained using the training dataset until a preset convergence condition is met, at which point the pre-training of the conference language model to be trained is considered complete, and the conference language model is obtained.

8. The method according to claim 7, characterized in that, The acquisition of the vocabulary dataset includes: Retrieve a preset amount of raw text data; Determine the word segmentation method and vocabulary parameters; Using a preset word segmentation tool, the original text data is segmented based on the word segmentation method and the word list parameters to generate a first word list; The first vocabulary is merged with the original second vocabulary to obtain a merged expanded vocabulary, which constitutes the vocabulary dataset.

9. The method according to claim 7, characterized in that, The acquisition of the target instruction dataset includes: Obtain the unlabeled dataset; Generate evaluation instructions for the unlabeled data in the unlabeled dataset, the evaluation instructions including empty instructions and non-empty instructions; From the unlabeled dataset, delete the unlabeled data corresponding to the empty instruction to obtain a candidate instruction dataset, which consists of unlabeled data and non-empty instructions corresponding to the unlabeled data; The candidate instruction dataset is scored to obtain unlabeled data whose quality scores fall within a preset range; The unlabeled data whose quality score results are within a preset score range, and the non-empty instructions corresponding to the unlabeled data, are determined as the target instruction dataset.

10. The method according to claim 7, characterized in that, The process of obtaining the large language model of the conference to be trained includes: Noise is added to the embedding layer of the pre-acquired target large language model to fine-tune the embedding layer of the target large language model, thereby obtaining the conference large language model to be trained. Wherein, the sampling range of the noise is Between these, ∂ is an adjustable parameter, L is the input length, and d is the dimension of the embedding layer.

11. The method according to claim 7, characterized in that, In the process of pre-training the conference large language model to be trained using the training dataset, the method further includes: The context window of the conference large language model to be trained is expanded so that the size of the context window matches the length of the instructions in the target instruction dataset, where the instruction length is the number of characters in the instruction.

12. The method according to claim 4, characterized in that, The large model is the multimodal visual language large model. The preprocessed meeting data includes multiple target key images, each target key image corresponding to a language description text. The process involves inputting the preprocessed meeting data into the large model, and the large model outputting the data results corresponding to the target task, including: Obtain the prompt instructions corresponding to the target task; The key images of each target, the corresponding voice description text of each key image of the target, and the prompts and instructions corresponding to the target task are combined to construct multiple input data; The constructed input data are respectively input into the multimodal visual language large model to generate the result corresponding to each target key image; The results corresponding to each target key image are input into the multimodal visual language model, and the prompts corresponding to the target task are simultaneously input into the multimodal visual language model to generate the content required for the target task.

13. The method according to claim 1, characterized in that, The training methods for the multimodal visual language large model include: Obtain a large multimodal visual language model to be trained, wherein the large multimodal visual language model to be trained includes the conference large language model, a visual encoder, and a language adapter; Obtain a training dataset, which includes a first training dataset and a second training dataset. The first training dataset includes target key images and language description text corresponding to the target key images. The second training dataset includes text and target instruction data associated with the text. While keeping the parameters of the large language model of the conference unchanged, the visual encoder and language adapter are pre-trained using the first training dataset until a preset first convergence condition is met, at which point the pre-training of the visual encoder and language adapter is determined to be complete; the visual encoder is used to process image input and generate image features, and the language adapter is used to combine image features with language features. The first training dataset is used to perform multi-task pre-training on the conference large language model, visual encoder and language adapter until the preset second convergence condition is met, then the multi-task pre-training of the conference large language model, visual encoder and language adapter is determined to be complete. While ensuring that the parameters of the visual encoder remain unchanged after pre-training, the pre-trained conference language model and language adapter are fine-tuned using the second training dataset until the preset third convergence condition is met, at which point the fine-tuning of the conference language model and language adapter is determined to be complete. The model consisting of a pre-trained and fine-tuned conference language model, a visual encoder, and a language adapter cascaded together is identified as a multimodal visual language model.

14. The method according to any one of claims 1 to 13, characterized in that, After obtaining the preprocessed meeting data, the method further includes: The preprocessed meeting data is processed using the large model to obtain processed content. The processed content is then input into the large model, which generates the intelligent chapters, meeting minutes, speaker summaries, and meeting tasks.

15. The method according to any one of claims 1 to 13, characterized in that, After obtaining the preprocessed meeting data, the method further includes: Keyword extraction is performed based on the preprocessed meeting data, and the extracted keywords are displayed in a graphical interface.

16. A data processing device for a conference setting, characterized in that, include: The meeting data acquisition module is used to acquire meeting data, the types of which include single-modal meeting data and multimodal meeting data; The preprocessing module is used to preprocess the meeting data according to the type of the meeting data to obtain preprocessed meeting data; The large model determination module is used to determine the large model corresponding to the preprocessed meeting data; the large model includes a meeting large language model and a multimodal visual language large model, the meeting large language model is used to process the single-modal meeting data, and the multimodal visual language large model is used to process the multimodal meeting data; The target task generation module is used to input the preprocessed meeting data into the large model and output the data results corresponding to the target task through the large model. The target task includes at least one of intelligent chapter, meeting minutes, speaker summary and meeting to-do list.

17. An electronic device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Meeting hot spot processing method and device, terminal equipment and storage medium

    CN110211590A

  • Task processing method, automatic question answering method and task processing system

    CN117193960A

  • Methods and apparatus to controllable multimodal meeting summarization with semantic entities augmentation

    US20230343331A1

Cited By

  • Real-time voice interaction method based on large model and electronic equipment

    CN121415784A

  • Conference process resource processing result acquisition method, electronic equipment and program product

    CN121579537A