An intelligent conference system based on multi-modal data
By designing an intelligent conference system based on multimodal data, using the interaction and fusion of multimodal data and artificial intelligence technology, the problem of insufficient intelligence of the existing conference system is solved, a wider application scenario and more efficient functions are achieved, and the reliability and compatibility of the conference system are improved.
Patent Information
- Application Number
- CN202510156521.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The existing conference system has poor intelligence, limited application scenarios and fewer functions, which limits the practical application and development of information technology in conference systems.
An intelligent conference system based on multimodal data is designed, including data acquisition module, data transmission module, participant identification module, content visualization module, conference voting module, knowledge recommendation module, speech transcription module, archive processing module and auxiliary decision-making module. Through the interaction and fusion of multimodal data, artificial intelligence technology is used to improve the adaptability of the system.
It enhances the reliability, compatibility and business capabilities of the intelligent conference system, broadens application scenarios, provides more efficient, scientific and intelligent functions, supports post-conference solutions and knowledge recommendations, and improves decision-making efficiency.
Smart Images

Figure CN119652689B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to fields such as computer science and digital signal processing, and in particular to an intelligent conference system based on multimodal data. Background Art
[0002] A meeting is an organized, led, and purposeful deliberative activity. It is conducted at a limited time and place according to a certain procedure. There are meetings in almost every organized place. The main functions of meetings include decision-making, control, coordination, and education. In recent years, with the rapid development of information technology, the form of meetings has gradually evolved from traditional meeting modes to meetings in various terminal forms. Offline meetings, online meetings, and online and offline combined meetings are becoming more and more common. Meetings generally include three elements: discussion, decision, and action. Therefore, it is necessary to have discussions during meetings, decisions during discussions, and actions after decisions.
[0003] Traditional conference systems serve as auxiliary media for holding meetings and provide basic convenience for meetings. Although existing conference systems have initially realized the function of voice tracking, they still have problems such as poor intelligence, many limitations on application scenarios, and few application functions, which have restricted the actual application and development of information technology in conference systems to varying degrees.
[0004] Therefore, making full use of the various information data of the conference system, effectively realizing basic functions such as identity recognition and data collection, using information technology to broaden application scenarios, and developing more efficient, scientific, and intelligent functions, and building an intelligent conference system based on big data analysis has become an engineering problem that needs to be solved urgently. Summary of the invention
[0005] The present application provides an intelligent conference system based on multimodal data to enhance the reliability, compatibility and business capabilities of the intelligent conference system.
[0006] The technical solution adopted in this application is:
[0007] An intelligent conference system based on multimodal data, comprising: a data acquisition module, a data transmission module, a participant identification module, a content visualization module, a conference voting module, a knowledge recommendation module, a voice transcription module, an archiving processing module and a decision support module;
[0008] The data collection module collects data from the target area and / or target object based on the received data collection instruction or the input user operation instruction, and encapsulates the collected data to obtain a collection data packet, and then uploads it to the data transmission module, wherein the collection data packet includes a service identifier of the collected data for the data transmission module to distribute the collected data to the corresponding module of the intelligent conference system;
[0009] The data transmission module parses the uploaded collected data packets, transmits the collected data to the corresponding module based on the service identifier, and stores the collected data in the corresponding storage device of the data transmission module;
[0010] The participant identification module is used to identify and confirm the participants based on the collected data;
[0011] The content visualization module is used to visually output and display the meeting content configured by the user and the information related to the set meeting process based on the output terminals deployed in the meeting venue;
[0012] The meeting voting module provides a voting function including various voting forms, collects voting information and counts the voting situation to form the meeting voting result;
[0013] The knowledge recommendation module, based on the deployed knowledge recommendation model, generates recommended knowledge information of the meeting-related auxiliary knowledge of the current meeting according to the theme and content of the current meeting, and pushes the recommended knowledge information to the designated meeting participants;
[0014] The speech transcription module performs speech transcription processing on the speech data collected by the data collection module based on the deployed machine learning model to obtain the transcribed meeting process text information;
[0015] The archiving processing module automatically generates meeting minutes based on the deployed large language model based on the meeting-related information; and classifies and organizes the meeting-related materials to form meeting archiving materials and saves them;
[0016] The auxiliary decision-making module, based on the deployed auxiliary decision-making model, extracts the designated information in the meeting archiving materials as the input of its model, and generates the text information of the specific measures of the meeting through the auxiliary decision-making model.
[0017] Furthermore, an intelligent meeting system based on multi-modal data of the present application further includes a meeting initiation notification module, which initiates meeting notification information to the designated notified objects based on the meeting notification information configured by the user.
[0018] Furthermore, the meeting initiation notification module forms a meeting notification and a confirmed list of participants based on the participation information feedback by the notified objects.
[0019] Furthermore, the service identifier is set as the identifier of the module that issues the data collection instruction and / or the identifier of the module corresponding to the user operation instruction.
[0020] Furthermore, the collection devices of the data collection module include but are not limited to: audio collection devices, video collection devices, image collection devices, and text collection devices.
[0021] Further, the data transmission module regularly clears and deletes the collected data stored in the storage device according to the configured cleaning period.
[0022] Further, the participant identification module specifically includes:
[0023] Based on the data collection module, collect the basic information of the person to be identified;
[0024] Match the basic information of the person to be identified collected with the information of the participants recorded in the preset identification database. If the current person to be identified is in the database, confirm that they are a participant in this meeting; if not, add the basic information of the current person to be identified to the identification database and regard the current person to be identified as a participant in this meeting.
[0025] Based on the set participant statistics time or the meeting host starts the participant statistics operation, the participant identification module conducts statistics on the participants in this meeting and matches them with the list of participants determined at the start of the meeting. If they are consistent, the meeting starts; if not, send a confirmation message to the meeting host asking whether the meeting can start.
[0026] If the meeting host agrees to start, the meeting is started; if not, re-determine the participants in this meeting based on the participant information feedback by the host.
[0027] Further, the identification data of the participant identification module includes but is not limited to: face, fingerprint, and voiceprint.
[0028] Further, the knowledge recommendation module specifically includes:
[0029] Extract the meeting content in multiple modalities and perform feature extraction on each of them to obtain the feature vectors of the meeting content in each modality; usually, the modalities include video, image, and text. Different modality feature vectors can be extracted based on the large language model to obtain the feature vectors of the meeting content in each modality.
[0030] Extract the feature vector of the text of the meeting theme to obtain the feature vector of the meeting theme.
[0031] In each modality, concatenate the feature vector of the meeting content and the feature vector of the meeting theme as the input of the causal transfer module based on the attention mechanism, and obtain the fused feature in each modality based on the output of the causal transfer module.
[0032] Concatenate the fused features in all modalities as the input of the preset generation prompt network model, and obtain the text prompt of the meeting-related auxiliary knowledge of the current meeting based on the output of the generation prompt network model.
[0033] Then, input the text prompt of the meeting-related auxiliary knowledge into the large language model, and obtain the meeting-related auxiliary knowledge of the current meeting based on its output.
[0034] Furthermore, the voting forms of the meeting voting module include: oral, action, text, and other forms.
[0035] Furthermore, the meeting-related materials include: information of participants, basic meeting materials, meeting-related auxiliary knowledge, text information of the meeting process, meeting voting results, and meeting minutes.
[0036] Furthermore, the meeting-related materials extracted by the archiving processing module also include the multi-modal acquisition data saved by the data transmission module; and after completing the meeting archiving materials for the corresponding meeting, the archiving processing module notifies the data transmission module to clean and delete the acquisition data corresponding to the meeting.
[0037] Furthermore, based on the deployed meeting news generation template, the archiving processing module generates the corresponding meeting news through the large language model and pushes it to the configured meeting news notification users.
[0038] The technical solution provided by this application at least brings the following beneficial effects:
[0039] An intelligent meeting system based on multi-modal data provided by this application strengthens the interaction and integration of multi-modal data by organically integrating information technology means with the three elements of the meeting (discussion, decision-making, action), uses artificial intelligence technology to improve the adaptability of the system, provides solutions after the meeting, enhances the reliability and compatibility of the intelligent meeting system, and creates better conditions for the practical application of intelligent and smart scenarios.
[0040] In order to better utilize causal information at the feature level, the knowledge recommendation module of this application adopts a causal relationship enhancement method, which can extract features of multi-modal information such as images, texts, and languages, connect them with the meeting theme feature information respectively, and then input them into the causal transfer module respectively. In this module, the fused features are calculated through the attention mechanism to guide the large model to generate relevant knowledge information, so as to realize the real-time push of knowledge blind spots and reference materials in the meeting process by using relevant technologies of the large language model, and provide knowledge reference for meeting decision-makers.
[0041] The archiving processing module of this application provides a text generation function for users. According to the meeting convening situation, it uses big data technology for machine learning of text writing, can automatically generate meeting minutes, meeting news, etc., and can also develop corresponding text generation functions according to needs.
[0042] The auxiliary decision-making module of this application provides an auxiliary solution function. By using related technologies of large language models, it can automatically generate notices, solutions, methods, etc. for solving problems in the next step according to the situation of the meeting, providing support for solving problems in the next step. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0044] Figure 1 is a schematic structural diagram of an intelligent conference system based on multimodal data provided by an embodiment of this application;
[0045] Figure 2 is a schematic hierarchical architecture diagram of an intelligent conference system based on multimodal data provided by an embodiment of this application;
[0046] Figure 3 is a schematic diagram of the participant identification process in an embodiment of this application;
[0047] Figure 4 is a schematic diagram of the processing process of the knowledge recommendation module in an embodiment of this application;
[0048] Figure 5 is a schematic diagram of the processing process of the conference voting module in an embodiment of this application.
[0049] Figure 6 is a working flow chart of an intelligent conference system based on multimodal data provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this application will be described in detail and completely below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described by referring to the drawings are exemplary and are intended to explain this application, and should not be construed as limiting this application.
[0051] In order to build an intelligent platform for meetings based on the information technology field, by organically integrating information technology means with the three elements of meetings (discussion, decision-making, and action), strengthening the interaction and integration of multimodal data, using artificial intelligence technology to improve the adaptability of the system, providing solutions after meetings, and enhancing the reliability and compatibility of the intelligent conference system, the embodiments of this application provide an intelligent conference system based on multimodal data.
[0052] As Figure 1As shown in the figure, an intelligent conference system based on multimodal data provided by an embodiment of the present application includes: a data acquisition module, a data transmission module, a participant identification module, a content visualization module, a meeting voting module, a knowledge recommendation module, a speech transcription module, an archiving processing module, and an auxiliary decision-making module;
[0053] Among them, the data acquisition module performs data acquisition on the target area and / or target object based on data acquisition instructions issued by other modules of the intelligent conference system or input user operation instructions, and performs data encapsulation on the acquired data to obtain an acquisition data packet, and then uploads it to the data transmission module. Among them, the acquisition data packet includes task / business identifiers of the acquisition data (such as the identifier of the module that issues the data acquisition instruction, the working module identifier corresponding to the user operation instruction, etc.), so that the data transmission module can distribute the acquisition data to the corresponding functional modules for corresponding processing; among them, the data modalities of the acquisition data include but are not limited to voice data, image / video data, text data, etc. Specifically, the data acquisition module in this embodiment can be specifically set as: an audio acquisition device, a video acquisition device, an image acquisition device, and a text acquisition device.
[0054] The data transmission module performs data packet parsing processing on the uploaded acquisition data packet, and then transmits the acquisition data to the corresponding module; for example, if the current acquisition data is the data to be identified of the participant (such as the face to be identified, the fingerprint to be identified, the sound wave to be identified, etc.), it is transmitted to the participant identification module; another example is that if the current acquisition data is the data related to the meeting vote, it is transmitted to the meeting voting module; if it is the data related to speech transcription, it is transmitted to the speech transcription module; if it is the data related to the meeting process, it is transmitted to the archiving processing module and the knowledge recommendation module, etc. In addition, a corresponding storage space can also be set for the data transmission module to facilitate the storage of the parsed acquisition data (using the meeting number as the storage data index). In order to improve the utilization rate of the storage space, the acquisition data corresponding to the meeting archiving materials that have been archived or the meeting number corresponding to the completed meeting archiving materials issued by the archiving processing module can also be cleared / deleted regularly.
[0055] The participant identification module is used to identify and confirm the participants based on the acquisition data (such as face, fingerprint, voiceprint, etc.).
[0056] Content visualization module, used to visually output and display the meeting content configured by the user based on the output terminals deployed in the meeting venue; that is, this module usually configures a user interaction interface for the user to configure the meeting content to be output and displayed; at the same time, it can also achieve personalized display for different meeting participants, such as the output display of basic meeting content for participants, and the separate presentation of additional information such as auxiliary materials for meeting hosts / speakers. In addition, the content visualization module is also used to output and display the voting process and voting results of the meeting voting module.
[0057] Meeting voting module, used to provide voting functions in various voting forms (such as oral, action, text, and other forms), collect voting information and count the voting situation to form the meeting voting result; in addition, the meeting voting result can be output and displayed to designated personnel through the content visualization module.
[0058] Knowledge recommendation module, based on the deployed knowledge recommendation model (which can be implemented based on large language models), generates recommended knowledge information of meeting-related auxiliary knowledge for the current meeting according to the theme and content of the current meeting, and pushes the recommended knowledge information to designated meeting participants to provide basic support for the scientific judgment of meeting decision-makers.
[0059] Speech transcription module, which performs speech transcription processing on the speech data collected by the data collection module based on the deployed machine learning model to obtain the text information of the meeting process after transcription for subsequent processing by the archiving processing module, auxiliary decision-making module, etc.
[0060] Archiving processing module, automatically generates meeting minutes based on the deployed large language model based on meeting-related information; and classifies and organizes meeting-related materials such as participant information, basic meeting materials, meeting-related auxiliary knowledge, meeting process text information, meeting voting results, and meeting minutes to form meeting archival materials and save them. In addition, when forming the meeting archival materials, the archiving processing module can also extract the multi-modal collected data saved by the data transmission module. After completing the meeting archival materials for the corresponding meeting, it can notify the data transmission module to delete the collected data corresponding to the meeting. Furthermore, the archiving processing module can also generate the corresponding meeting news based on the deployed meeting news generation template through the large language model and push it to the configured meeting news notification users. It can also store the meeting news as one item in the meeting archival materials.
[0061] The auxiliary decision-making module, based on the deployed auxiliary decision-making model (which can be implemented based on large language models), extracts specified information (such as meeting content and text information during the meeting process, etc.) from the meeting archived materials as the input of its model, and generates text information on specific measures for the meeting through the auxiliary decision-making model, that is, based on the artificial intelligence technology of large models, according to the situation discussed in the meeting, automatically forms notices, plans, methods, measures, etc. for implementation, to assist decision-makers to carry out actual work more efficiently.
[0062] In one embodiment, an intelligent meeting system based on multi-modal data of the present application further includes a meeting initiation notification module, which initiates meeting notification information to specified notified objects based on the meeting notification information configured by the user (such as time, place, host, topic, etc.). Further, it can also form a meeting notification and a confirmed list of participants based on the participation information feedback by the notified objects.
[0063] Further, in actual application implementation, the architecture of an intelligent meeting system based on multi-modal data provided by the embodiments of the present application can be divided into five layers, such as Figure 2 shown, which includes a physical layer, a data layer, a control layer, an application layer, and an expansion layer. Each layer is specifically as follows:
[0064] The physical layer is mainly data information collection devices. This layer is the bottom layer of the system, deploying a large number of sensors, devices, and apparatuses. The devices include but are not limited to mobile phones, computers, microphones, speakers, cameras, image capture devices, intelligent meeting devices, etc., which are used to collect and transmit basic materials required for the meeting such as audio, images, videos, and texts, and at the same time provide basic guarantees for the convening of the meeting.
[0065] The data layer is mainly data recording and storage. This layer stores data of different categories, different times, different devices, and different objects, and makes marks during the data recording process, fully considering efficient storage and access of different types of data.
[0066] The control layer is mainly data processing and interaction. During the multi-modal data collection process, data synchronization and interaction will be involved. Situations such as different times for one modality and the same time for different modalities may all have synchronization or interaction, and some data-level marks need to be made through the control layer. At the same time, there will also be situations such as redundant information and data noise in the data, and data control and supervision need to be strengthened. The feature extraction and application of data can be constructed through digital signal processing technology.
[0067] The application layer is mainly for developing basic applications. Combining with the meeting scenario, using the reserved database to develop applications in the actual scenario. For the participant positioning function, according to the pre - scheduled meeting plan, identify the participants, collect the face image data of the participants, compare with the existing personnel information database, conduct face recognition, and lock the situation of the participants; the voice transcription function transcribes the voice data in real - time through natural language processing methods; the meeting voting function records and counts the voting results according to the actual situation through trigger voting, raising hands to vote, oral voting, etc.
[0068] The expansion layer mainly provides interfaces for data - interaction applications. For the text generation function, according to the meeting situation, use big data technology for machine learning of text writing to automatically generate meeting minutes, meeting news, etc., and corresponding text generation functions can also be developed according to needs; for the knowledge recommendation function, use related technologies of large language models to push the knowledge blind spots and reference materials during the meeting in real - time to provide knowledge reference for meeting decision - makers; for the auxiliary solution function, use related technologies of large language models to automatically generate notices, solutions, methods, etc. for solving problems in the next step according to the meeting situation to provide support for solving problems in the next step.
[0069] In one embodiment, as Figure 3 shown, the specific processing process of the participant recognition module of this application is as follows: When the meeting is initiated, the scope of participants in this meeting has been determined. Before the meeting, personnel recognition is carried out. If the person is in the known database (which can be the database configured by the intelligent meeting system, such as the participant recognition database, which can be set based on the confirmed list of participants) and is confirmed as a participant in this meeting, then it is confirmed that the person participates in the meeting. If the person is not in the known database, the basic information (face, fingerprint, voiceprint, etc.) of the person needs to be added and the database is updated, and finally it is judged whether the person participates in the meeting. After all the expected participants arrive, the host reviews and approves, and then this meeting is held.
[0070] In the embodiment of the present application, the basic information of the person to be identified can be first collected based on the data collection module (which can be a basic information collection device for personnel deployed in a designated area such as the entrance of the venue, such as collecting face, fingerprint, voiceprint, etc.), and transmitted to the participant identification module through the data transmission module; the module performs personnel matching (such as face matching, fingerprint matching, voiceprint matching, etc.) based on the basic information of the personnel and the participant information recorded in the pre-configured identification database. If the current person to be identified is in the database, it is confirmed as a participant of this meeting; if not in the database, the basic information of the current person to be identified is added to the identification database, and the identification is realized. The identification database is now updated, and the current person to be identified is regarded as a participant of this meeting; based on the set participant counting time (which can be set based on the meeting start time) or the participant counting operation initiated by the meeting host, the participant identification module counts the participants of this meeting and matches them with the list of participants determined when the meeting was initiated. If they are consistent, the meeting starts; if they are inconsistent, a confirmation message on whether the meeting has started is sent to the meeting host, and the meeting is started based on the feedback of the meeting host; if the host disagrees, the participants of this meeting are re-determined based on the processing information of their feedback.
[0071] The theme of a conference often guides the discussion direction of the entire conference. The content of the conference is generally presented in the form of video, image, text, voice and other visual information. The multimodal information in the conference theme and content contains causal relationships. In order to better utilize causal information at the feature level, the knowledge recommendation module of this application adopts a causal enhancement method. Taking the three modal information of image, text and language as an example, the three modal information is extracted and connected with the conference theme feature information respectively, and then respectively input into the causal migration module (denoted as module T), in which the fusion features are calculated through the attention mechanism:
[0072]
[0073]
[0074]
[0075] in, and Represent the input features and output fusion features / data respectively, represents the output of the causal transfer module, represents the softmax function, represents the attention score, which is expressed as:
[0076]
[0077] in, , , respectively represent three fully connected layers, that is, , , are the outputs of the query linear layer, the key linear layer, and the value linear layer respectively. In the embodiments of the present application, they are all set as fully connected layers. represents a preset constant, and the specific value can be set based on the actual application scenario.
[0078] In one embodiment, the specific processing process of the knowledge recommendation module of the present application based on the causal transfer module is as Figure 4 shown: Using the feature vectors of three modal information (image data, text data, speech data) (denoted as image feature , text feature , and speech feature ) and the feature vector corresponding to the meeting theme information (denoted as theme feature ) as the input of the causal transfer module, corresponding to three causal transfer modules , , respectively, that is, each modality corresponds to one; the process of obtaining the corresponding fusion feature based on the corresponding causal transfer module can be expressed as:
[0079]
[0080]
[0081]
[0082]
[0083] Among them, , , are the fusion features of image and theme, the fusion features of text and theme, and the fusion features of speech and theme respectively, represents the concatenation operation, represents the generated text prompt, represents the generated prompt network model. Then, is used as the input of the large model to guide the large model to generate relevant knowledge information, that is, the meeting-related auxiliary knowledge.
[0084] In one embodiment, a pre-trained machine learning model (also referred to as a speech transcription model) can be directly used as the initial structure, and transfer learning training is performed on it in combination with manually verified text and then deployed in the speech transcription module. Among them, the specific process of transfer learning training includes: based on the selected initial structure, preprocessing the speech data to be transcribed, such as segmenting, so that the preprocessed speech data matches the input of the machine learning model. Then, the preprocessed speech data is input into the model, and the corresponding text information is obtained based on its output, and it is matched with the manually verified text corresponding to the input. If they are inconsistent, the model parameters of the machine learning model are learned and optimized based on the set loss function (such as the cross-entropy loss between the text information output by the model and the manually verified text) to complete the model correction and deep learning; if they are consistent, the transfer learning training is stopped, and thus the final transcribed text information is output based on the current model.
[0085] In addition, after the transfer training is completed, the machine learning model can also be periodically updated based on the processed speech data and manually verified text as training data to further improve the transcription accuracy of the model.
[0086] In one embodiment, as Figure 5 shown, the voting forms supported by the meeting voting module of the present application include: oral, action, text, and other forms. Among them, oral voting corresponds to speech recognition, action voting (such as raising a hand) corresponds to image recognition, and text voting corresponds to text recognition. Oral voting can be set as: agree, pass, approve, etc., all of which can indicate the tendency towards the topic or matter. Action voting: raising a hand, nodding, etc., can all express agreement. Text voting is more direct, through ticking, liking, etc., and the text record will be more direct and clear. There are also other voting methods, such as clicking on a specific identifier on the screen, pressing a specific button, etc., which are also a form of voting. Based on the voting form selected by the meeting host, the corresponding data collection device is activated, and the statistical result of the meeting voting is obtained based on the recognition processing of the data collection device, that is, after voting, the voting situation is statistically analyzed according to the collected and recognized voting information, and finally the meeting voting result is formed.
[0087] Figure 6 The working flowchart of an intelligent conference system based on multi-modal data provided by the embodiments of the present application is shown. In the figure, it is described according to the session of the meeting, which is mainly divided into three parts, namely before the meeting, during the meeting, and after the meeting. The specific process is as follows:
[0088] Before the meeting is officially held, the host will initiate the meeting arrangement based on the meeting initiation notification module of this system, including the basic information of the meeting: time, place, host, participants, agenda, etc. After the meeting is initiated, the preparatory work for the meeting will be carried out: preparation of meeting materials, preparation of hardware facilities, preparation of software facilities. It should be noted that the intelligent conference system of this application has added multimedia system management on the basis of traditional conferences, and at the same time has higher requirements for the update of databases, knowledge bases, etc., and conducts machine learning and knowledge reserves according to the theme of the meeting.
[0089] During the meeting, participants are first confirmed based on the participant identification module. The confirmation method is selected according to the different conference media: face recognition, fingerprint recognition, sound wave recognition and other recognition methods.
[0090] After the participants are confirmed, the meeting is formally held. The meeting is conducted based on an intelligent conference system, whose content visualization module is connected to multimedia devices to facilitate the transmission and reception of multimedia information, including: video, images, text, voice and other visual information.
[0091] It is worth mentioning that this application introduces an auxiliary decision-making system during the meeting, that is, the knowledge recommendation module of this application can use artificial intelligence technology based on the large language model (Large Language Model, referred to as the large model) to recommend related knowledge information according to the theme and content of the meeting, providing basic support for the scientific judgment of decision makers. At the same time, the text information of the meeting process can also be obtained based on the voice transcription module of this application, and its content includes but is not limited to: the speech information of the speaker of the meeting, the discussion information of the participants, etc.
[0092] The end of a meeting usually requires judgment and decision-making, the most common of which is approval or rejection. At this time, a variety of voting methods are introduced: oral voting, raising hands, text voting and other voting methods. Based on the selected voting method, the corresponding meeting voting results are obtained through the meeting voting module of this application.
[0093] After the meeting, it is necessary to organize the situation of the meeting and arrange the next steps. Traditional meeting archives only classify and store the materials during the meeting. The information in the archiving processing module provided by the intelligent conference system of this application is classified and organized, and after processing by methods such as noise reduction and cleaning, it is numbered and stored. At the same time, according to the actual situation of the meeting, artificial intelligence means are used to automatically generate meeting minutes as an archiving material. As one of the results of the meeting, specific implementation measures are indispensable. Artificial intelligence technology based on large models automatically forms implementation notices, plans, methods, measures, etc. according to the discussion of the meeting, assisting decision makers to carry out actual work more efficiently.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
[0095] The above are only some embodiments of the present application. For those of ordinary skill in the art, without departing from the creative concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. An intelligent conference system based on multimodal data, characterized in that: include: Data collection module, data transmission module, participant identification module, content visualization module, meeting voting module, knowledge recommendation module, voice transcription module, archiving processing module and decision support module; The data collection module collects data from the target area and / or target object based on the received data collection instruction or the input user operation instruction, and encapsulates the collected data to obtain a collection data packet, and then uploads it to the data transmission module, wherein the collection data packet includes a service identifier of the collected data for the data transmission module to distribute the collected data to the corresponding module of the intelligent conference system; The data transmission module performs data packet parsing processing on the uploaded collected data packets, transmits the collected data to the corresponding module based on the service identifier, and stores the collected data in the corresponding storage device of the data transmission module; A participant identification module is used to identify and confirm the participants based on the collected data; The content visualization module is used to visualize the conference content configured by the user based on the output terminals deployed at the conference venue, and to visualize the conference process related information; The meeting voting module provides voting functions including various voting forms, collects voting information and counts voting results to form meeting voting results; The knowledge recommendation module generates recommended knowledge information of the current meeting's associated auxiliary knowledge based on the deployed knowledge recommendation model according to the current meeting's theme and content, and pushes the recommended knowledge information to designated meeting participants; A speech transcription module, which performs speech transcription processing on the speech data collected by the data collection module based on the deployed machine learning model to obtain transcribed text information of the meeting process; The archiving processing module automatically generates meeting minutes based on the meeting-related information based on the deployed large language model; and classifies and organizes the meeting-related materials to form and save the meeting archive materials; The decision-making support module, based on the deployed decision-making support model, extracts the specified information in the meeting archive as its model input, and generates text information of specific measures of the meeting through the decision-making support model; The knowledge recommendation module is as follows: Extract conference contents of multiple modes, and perform feature extraction on them respectively to obtain feature vectors of conference contents of each mode; Extract the feature vector of the conference theme text to obtain the feature vector of the conference theme; In each mode, the feature vector of the conference content and the feature vector of the conference theme are concatenated as the input of the causal transfer module based on the attention mechanism. Based on the output of the causal transfer module, the fusion features in each mode are obtained: f out =T(f in ) =softmax(w f )i V (f in ) Among them, f in and f out They represent the fusion features of the input features and the output, T() represents the output of the causal migration module, softmax() represents the softmax function, and w f represents the attention score, which is expressed as: Among them, θ Q (),θ K (),θ V () are the outputs of the query linear layer, key linear layer and value linear layer respectively, all of which are set as fully connected layers, d f Indicates a preset constant; The fusion features under all modalities are spliced as the input of the preset generative prompt network model, and the text prompt of the meeting-related auxiliary knowledge of the current meeting is obtained based on the output of the generative prompt network model: Corresponding to three causal migration modules T i , T t , T v , that is, each modality corresponds to one; the process of obtaining the corresponding fusion feature based on the corresponding causal migration module can be expressed as: Among them, f th represents the topic feature, f i 、f t 、f v are image features, text features and speech features respectively, and f ith 、f tth 、f vth Represented as the corresponding image features, text features and speech features respectively; T i , T t , T v are the output features of the causal transfer module corresponding to images, text, and speech respectively; Indicates the connection operation, P th represents the generated text prompt, θ represents the network model for generating prompts; Then the text prompt P of the conference related auxiliary knowledge th The large language model is input, and based on its output, the conference-related auxiliary knowledge of the current conference is obtained.
2. The intelligent conference system based on multimodal data according to claim 1, characterized in that: It also includes a conference initiation notification module, which initiates conference notification information to designated notified objects based on the conference notification information configured by the user.
3. The intelligent conference system based on multimodal data according to claim 1, characterized in that: The service identifier is set to the identifier of the module that issues the data collection instruction and / or the identifier of the module corresponding to the user operation instruction.
4. The intelligent conference system based on multimodal data according to claim 1, characterized in that: The acquisition device of the data acquisition module includes an audio acquisition device, a video acquisition device, an image acquisition device and a text acquisition device.
5. The intelligent conference system based on multimodal data according to claim 1, characterized in that: The data transmission module regularly cleans and deletes the collected data stored in the storage device according to the configured cleanup cycle.
6. The intelligent conference system based on multimodal data according to claim 1, characterized in that: The specific module for identifying participants is as follows: Collect basic information of the person to be identified based on the data collection module; The collected basic information of the person to be identified is matched with the information of the participants recorded in the preset identification database. If the current person to be identified is in the database, it is confirmed as a participant of this meeting; if not in the database, the basic information of the current person to be identified is added to the identification database, and the current person to be identified is regarded as a participant of this meeting; Based on the set participant counting time or the participant counting operation initiated by the conference host, the participant identification module counts the participants of the current meeting and matches them with the participant list determined when the meeting was initiated. If they are consistent, the meeting starts; if they are inconsistent, a confirmation message is sent to the conference host to confirm whether the meeting has started; If the meeting host agrees to start the meeting, the meeting will be started; if the host disagrees to start the meeting, the participants of this meeting will be re-determined based on the participant information fed back by the host.
7. The intelligent conference system based on multimodal data according to claim 1, characterized in that: Meeting-related materials include: information about participants, basic meeting materials, meeting-related auxiliary knowledge, text information about the meeting process, meeting voting results and meeting minutes.
8. The intelligent conference system based on multimodal data according to claim 1, characterized in that: The meeting-related data extracted by the archiving processing module also includes the multimodal collected data saved by the data transmission module; and after completing the meeting archiving data of the corresponding meeting, the archiving processing module notifies the data transmission module to clean up and delete the collected data corresponding to the meeting.
9. The intelligent conference system based on multimodal data according to claim 1, characterized in that: The archiving processing module generates corresponding conference news based on the deployed conference news generation template through a large language model and pushes it to the configured conference news notification users.
Citation Information
Patent Citations
AI intelligent conference system based on voice and semantics and implementation method thereof
CN109474763A
Intelligent conference recording method
CN116187949A
Intelligent conference assistant system based on AIGC
CN117573621A
Conference decision support condition prediction method, device, equipment, medium and product
CN118134049A