Conference system, intelligent agent and conference processing method

Through the collaborative work of the intelligent body and cloud server, the problems of low access efficiency, insufficient information utilization and inaccurate recording of traditional conference systems are solved, automated and intelligent conference management is realized, and meeting efficiency and quality are improved.

CN120297933AActive Publication Date: 2025-07-11LINGYANGE SEMICONDUCTOR, INC

Patent Information

Application Number
CN202510771762.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The problems of inefficient access to traditional conference systems, inadequate use of pre-meeting information, and inaccurate meeting notes and analysis lead to inefficient meetings and delayed decision-making.

Method used

The intelligent body and cloud server work together, and through the information acquisition module, edge equipment, clock and network status monitoring module, knowledge graph construction module and conference record and analysis module, automatic access to conference, in-depth processing of pre-meeting information, real-time translation and intelligent record analysis are realized.

Benefits of technology

It improves the speed and accuracy of meeting access, provides clear conference focus before the meeting, translate and record and analyze in real time, generates high-quality conference summary documents, and improves meeting efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297933A_ABST
    Figure CN120297933A_ABST
Patent Text Reader

Abstract

The invention provides a conference system, an intelligent agent and a conference processing method, and the system comprises an edge device which is used for carrying out the preliminary screening and feature extraction of a conference link file, and determining a first feature; the cloud server is used for analyzing the first feature after receiving the first feature sent by the edge device, and determining a conference link and conference starting time; the clock and network state monitoring module is used for evaluating the quality of a network environment where the intelligent agent is located within a first preset time before the conference starting time after receiving the conference link and the conference starting time sent by the cloud server, and selecting an optimal network access strategy to access the conference; the knowledge graph construction module is used for constructing a first knowledge graph according to the conference data; the conference recording and analyzing module is used for performing real-time translation according to the source language and the target language of the spokesman during the conference; and analyzing the conference content by adopting the optimized pre-trained large model and a multi-modal fusion algorithm to generate a conference summary document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of conference technology, and in particular to a conference system, an intelligent agent, and a conference processing method. Background Art

[0002] In today's era of rapid development of digital office, meetings are an important way for enterprises, organizations and other users to communicate, collaborate and make decisions, and improving their efficiency and quality is becoming increasingly critical. However, in the long-term use of traditional conference room systems, the following problems have gradually been exposed that are difficult to meet the needs of modern office: 1. Low efficiency of conference access: Traditional conference rooms usually rely on manual operation when accessing conference systems. The manual access to conference systems seriously affects the efficiency of conference startup. According to relevant surveys, in daily meetings of some large companies, due to problems with participants manually accessing the conference system, the meeting is delayed 2 to 3 times a week on average, with each delay ranging from 5 to 15 minutes. This not only wastes a lot of working time and reduces work efficiency, but may also lead to delays in important decisions, bringing potential economic losses to the company.

[0003] 2. Insufficient use of pre-meeting information: Due to the lack of in-depth processing of pre-meeting materials, it is difficult for participants to quickly form a comprehensive understanding of the meeting content before the meeting, and they are unable to quickly enter the state during the meeting to conduct efficient discussions and decisions. This increases the communication costs in the early stages of the meeting, and the meeting process is easily hindered, making it difficult to achieve the expected results. For example, in some meetings involving complex projects, since participants do not have a deep understanding of the pre-meeting materials, a lot of time is needed to re-sort out the project background and key issues during the meeting, resulting in a longer meeting time, while the time actually used to discuss solutions is relatively insufficient.

[0004] 3. Inaccurate meeting records and analysis: Traditional meeting records and analysis methods cannot meet the requirements of modern meetings for information accuracy and comprehensiveness. The subjectivity and incompleteness of manual records make it possible for meeting records to fail to truly reflect the full picture of the meeting and easily miss important viewpoints and decision-making information. Relying solely on recording equipment for recording, the subsequent sorting and analysis work is cumbersome and time-consuming, and cannot provide strong support for meeting decisions in a timely manner. In multilingual meetings, the lack of real-time and effective translation methods has severely restricted the communication between participants of different languages ​​and hindered the smooth transmission of information.

[0005] Based on the above problems, there is an urgent need to provide a conference system to provide users with a more convenient, efficient and intelligent conference experience. Summary of the invention

[0006] Embodiments of the present disclosure provide a conference system, an intelligent agent, and a conference processing method to solve the problems of low efficiency of existing conference access, insufficient utilization of pre-conference information, and inaccurate conference recording and analysis.

[0007] Based on the above problems, in a first aspect, a conference system provided by embodiments of the present disclosure includes: an intelligent agent and a cloud server. The intelligent agent includes: an information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module, and a conference recording and analysis module; The information acquisition module is used to acquire a conference link file and conference materials; The edge device is used to perform preliminary screening and feature extraction on the conference link file by using a compressed and quantized convolutional neural network model or a recurrent neural network model to determine a first feature; The cloud server is used to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device to determine a conference link and a conference start time; The clock and network status monitoring module is used to evaluate the network environment quality where the intelligent agent is located within a first preset time before the conference start time after receiving the conference link and the conference start time sent by the cloud server, and select an optimal network access strategy to access the conference by using a deep Q-network; The knowledge graph construction module is used to construct a first knowledge graph according to the conference materials; The conference recording and analysis module is used to perform real-time translation according to the source language and target language of the speaker during the conference; and generate a conference summary document according to the conference content.

[0008] In combination with the first aspect, in a possible implementation manner, the formula for the clock and network status monitoring module to select an optimal network access strategy to access the conference by using a deep Q-network is:

[0009] where represents the network status where the intelligent agent is located Under this condition, by selecting a network access strategy the expected reward obtained represents the reward, represents the discount factor, represents the network access strategy selected at the next time step the maximum expected reward obtained.

[0010] In combination with the first aspect, in a possible implementation, the clock and network status monitoring module is further configured to, when the network environment quality indicates that the meeting cannot be accessed, send the shared network status information and a request for the meeting access method to other agents, and access the meeting according to the shared network status information and the meeting access method sent by the received other agents.

[0011] In combination with the first aspect, in a possible implementation, the knowledge graph construction module is configured to optimize a pre-trained large model according to the meeting materials by using an online learning algorithm; Analyze the meeting materials according to the optimized pre-trained large model to extract second features; Predict the meeting type according to the meeting materials by using a random forest algorithm to determine the meeting type; Determine the entities in the meeting materials and the corresponding relationships between the entities according to the second features and the meeting type by using a knowledge graph construction tool, and construct a first knowledge graph.

[0012] In combination with the first aspect, in a possible implementation, the meeting content includes: voice records, video records, and environmental sensor data; The meeting recording and analysis module is configured to optimize the translation model parameters in real time according to the speech content of the speaker during the meeting; According to the source language and target language of the speaker and the analysis results of the audio records, video records, and environmental sensor data by using a multimodal fusion algorithm, perform real-time translation on the speech content of the speaker by using the optimized translation model.

[0013] In combination with the first aspect, in a possible implementation, the meeting content includes: voice records, video records, and environmental sensor data; The meeting recording and analysis module is configured to use a speech recognition model according to the voice record to convert the voice into text data in real time; Optimize a pre-trained large model according to the text data by using an online learning algorithm; Generate a meeting record document according to the analysis results of the optimized pre-trained large model on the text data and the analysis results of the audio records, video records, and environmental sensor data by using a multimodal fusion algorithm; Generate a meeting summary document according to the meeting record document.

[0014] In combination with the first aspect, in a possible implementation, the knowledge graph construction module is further configured to analyze the audio records, video records, and environmental sensor data by using a multimodal fusion algorithm during the meeting; Optimize a pre-trained large model according to the meeting content by using an online learning algorithm; Analyze the meeting content based on the analysis results of audio records, video records, and environmental sensor data by the optimized pre-trained large model and multi-modal fusion algorithm; and update the first knowledge graph according to the analysis results of the meeting content to generate a second knowledge graph.

[0015] The meeting record and analysis module is also used to generate a preliminary meeting summary report using the T5 model according to the meeting record document; Submit the preliminary meeting summary report to the meeting chair for adjudication; Generate a meeting summary document using a natural language generation model according to the adjudicated preliminary meeting summary report, the meeting record document, and the second knowledge graph.

[0016] Combined with the first aspect, in a possible implementation, the clock and network status monitoring module is also used to send a meeting notice to meeting participants after receiving the meeting link and meeting start time sent by the cloud server. The meeting notice includes: the meeting link, the meeting start time, and the first knowledge graph; The knowledge graph construction module is also used to introduce the meeting content according to the first knowledge graph using a pre-trained speech synthesis model; The meeting record and analysis module is also used to generate an AR plan according to the meeting summary document and the meeting room environment features recognized by computer vision.

[0017] In a second aspect, an intelligent agent is provided, including: an information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module, and a meeting record and analysis module; The information acquisition module is used to acquire a meeting link file and meeting materials; The edge device is used to perform preliminary screening and feature extraction on the meeting link file using a compressed and quantized convolutional neural network model or a recurrent neural network model to determine the first feature; and send the first feature to the cloud server; the cloud server is used to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device to determine the meeting link and meeting start time; The clock and network status monitoring module is used to evaluate the network environment quality of the intelligent agent within a first preset time before the meeting start time after receiving the meeting link and meeting start time sent by the cloud server, and select the optimal network access strategy to access the meeting using a deep Q network; The knowledge graph construction module is used to construct a first knowledge graph according to the meeting materials; The meeting record and analysis module is used to perform real-time translation according to the source language and target language of the speaker during the meeting; and generate a meeting summary document according to the meeting content.

[0018] Thirdly, a meeting processing method is provided, including: Obtain a meeting link file and meeting materials; Use a convolutional neural network model or a recurrent neural network model after compression and quantization processing to perform preliminary screening and feature extraction on the meeting link file, and determine the first feature; Dynamically allocate computing resources to analyze the first feature, and determine the meeting link and meeting start time; According to the meeting link and meeting start time, evaluate the network environment quality within a first preset time before the meeting start time, and use a deep Q network to select the optimal network access strategy to access the meeting; Construct a first knowledge graph according to the meeting materials; Perform real-time translation according to the source language and target language of the speaker during the meeting; and generate a meeting summary document according to the meeting content.

[0019] The beneficial effects of the embodiments of the present disclosure include: The conference system, intelligent agent and conference processing method provided by the embodiment of the present disclosure include: an intelligent agent and a cloud server, wherein the intelligent agent includes: an information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module and a conference recording and analysis module; the information acquisition module is used to obtain conference link files and conference materials; the edge device is used to use a convolutional neural network model or a recurrent neural network model after compression and quantization to perform preliminary screening and feature extraction on the conference link files to determine a first feature; the cloud server is used to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device to determine the conference link and the conference start time; the clock and network status monitoring module is used to evaluate the quality of the network environment where the intelligent agent is located within a first preset time before the conference start time after receiving the conference link and the conference start time sent by the cloud server, and use a deep Q network to select the optimal network access strategy to access the conference; the knowledge graph construction module is used to construct a first knowledge graph based on the conference materials; the conference recording and analysis module is used to perform real-time translation according to the speaker's source language and target language during the meeting; and generate a conference summary document based on the content of the meeting. The disclosure uses intelligent agents and cloud servers to identify the link formats of various conference systems through algorithms, and can automatically and accurately access the conference system according to the meeting time in a complex network environment, greatly simplifying the conference access process. Before the meeting, the intelligent agent uses natural language processing technology and machine learning algorithms to deeply analyze the previous meeting summary, the theme of this meeting and other materials provided by the user, extract key information and construct the first knowledge graph, laying a solid data foundation for conference interaction and decision-making. During the meeting, using advanced speech recognition, semantic understanding and sentiment analysis technology, the intelligent agent records the speeches of participants in real time, accurately summarizes the points of view, clearly distinguishes between pros and cons, and simultaneously realizes real-time speech recognition, speaker annotation and on-demand multi-language translation functions. After the meeting, the intelligent agent quickly generates a meeting summary document. With the functions of intelligent automatic access, deep processing of pre-meeting information, and intelligent recording and analysis of speeches, the full process of the meeting is fully automated and intelligently managed, and the efficiency and quality of the meeting are significantly improved. It provides enterprises, organizations and other users with efficient, convenient and intelligent conference solutions that meet the needs of digital office, helping users to reduce costs and increase efficiency in conference scenarios, and promote more scientific and accurate decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A schematic diagram of the structure of a conference system provided by an embodiment of the present disclosure; Figure 2 A schematic diagram of the structure of edge-cloud collaborative computing provided by an embodiment of the present disclosure; Figure 3 A schematic diagram of implementing functions of a conference system provided by an embodiment of the present disclosure; Figure 4Schematic diagram of the meeting processing method provided by the embodiments of the present disclosure. Detailed implementation manners

[0021] The embodiments of the present disclosure provide a meeting system, an intelligent agent, and a meeting processing method. The preferred embodiments of the present disclosure will be described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present disclosure, and are not used to limit the present disclosure. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0022] The embodiments of the present disclosure provide a meeting system, as Figure 1 shown, including: an intelligent agent 100 and a cloud server 200. The intelligent agent 100 includes: an information acquisition module 101, an edge device 102, a clock and network status monitoring module 103, a knowledge graph construction module 104, and a meeting record and analysis module 105; The information acquisition module 101 is used to acquire a meeting link file and meeting materials; The edge device 102 is used to perform preliminary screening and feature extraction on the meeting link file by using a compressed and quantized convolutional neural network model or a recurrent neural network model to determine a first feature; The cloud server 200 is used to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device 102 to determine a meeting link and a meeting start time; The clock and network status monitoring module 103 is used to evaluate the network environment quality where the intelligent agent 100 is located within a first preset time before the meeting start time after receiving the meeting link and the meeting start time sent by the cloud server 200, and select an optimal network access strategy to access the meeting by using a deep Q network; The knowledge graph construction module 104 is used to construct a first knowledge graph according to the meeting materials; The meeting record and analysis module 105 is used to perform real-time translation according to the source language and target language of the speaker during the meeting, and generate a meeting summary document according to the meeting content.

[0023] The disclosed embodiments are applied to the field of conference technology. Traditional conference systems usually rely on manual operation when accessing conferences. Participants need to open the corresponding conference software by themselves before the meeting starts, and enter information such as the conference ID and password to complete the conference access process. This method not only consumes a lot of time and energy of participants in large-scale meetings or frequent meeting arrangements, but is also prone to input errors due to human negligence, thereby affecting the punctual start of the meeting. According to relevant surveys, in the daily meetings of some large enterprises, the delay in the start of the meeting occurs an average of 2 to 3 times a week due to problems with the manual access of the conference system by participants, and each delay time ranges from 5 to 15 minutes. This not only wastes a lot of working time and reduces work efficiency, but also may lead to delays in important decisions, bringing potential economic losses to the company. In the meeting preparation stage, users often simply distribute meeting materials such as the summary of the previous meeting and the theme of this meeting to the participants. These meeting materials lack effective organization and analysis, and participants need to read and understand them by themselves before the meeting, making it difficult to quickly grasp the key points and key information of the meeting. Although some existing conference systems have simple document sharing functions, they are almost unable to achieve in-depth processing of materials, such as extracting key information and building knowledge associations. Due to the lack of in-depth processing of pre-meeting materials, it is difficult for participants to quickly form a comprehensive understanding of the meeting content before the meeting, and it is difficult to quickly enter the state during the meeting to conduct efficient discussions and decisions. This increases the communication cost in the early stage of the meeting, and the meeting process is easily hindered, making it difficult to achieve the expected results. For example, in some meetings involving complex projects, since the participants do not have a deep understanding of the pre-meeting materials, a lot of time is needed to re-sort out the project background and key issues during the meeting, resulting in extended meeting time, while the time actually used to discuss solutions is relatively insufficient. During the meeting, traditional recording methods mainly rely on manual recording or simple recording equipment. Manual recording is not only inefficient and difficult to fully record the speech content of each participant, but also prone to omissions or errors during the recording process, so that the meeting minutes may not truly reflect the full picture of the meeting and are prone to omitting important views and decision-making information. Although the recording equipment can record the conference audio, it takes a lot of time for manual listening and sorting to convert it into a usable text record, which cannot provide strong support for the meeting decision in a timely manner. The analysis of speech content, such as summarizing arguments and distinguishing viewpoints, relies on human subjective judgment and lacks accuracy and comprehensiveness. In multilingual conference scenarios, translation work usually requires professional translators, which is costly and has poor real-time performance. This severely limits the communication between participants of different languages ​​and hinders the smooth transmission of information.

[0024] In the embodiments of the present disclosure, the conference system includes an agent 100 and a cloud server 200. The agent 100 can be a software entity, program module or related device developed based on artificial intelligence technology, capable of autonomously executing tasks, perceiving the environment and making decisions. The cloud server 200 can be a remote server built based on cloud computing technology, used for storing, processing and transmitting conference-related data and information. The cloud server 200 is the core infrastructure for the conference system to achieve intelligent and cross-regional collaboration, undertaking key tasks such as connecting participants, managing the conference process, and ensuring the stable operation of the conference. The agent 100 includes: an information acquisition module 101, an edge device 102, a clock and network status monitoring module 103, a knowledge graph construction module 104, and a conference record and analysis module 105. In the conference preparation stage before the conference starts, the user can provide the conference link file and conference materials to the information acquisition module 101. The conference materials can include the summary of the previous conference and the theme of the current conference. The information acquisition module 101 can also be integrated with office systems (such as OA systems, document management platforms), and automatically obtain the conference link file and conference materials through the API interface, improving the efficiency and accuracy of information collection. Such as Figure 2As shown, the edge device 102 can be an intelligent gateway device local to the meeting room. The edge device 102 can use lightweight models such as a convolutional neural network model (CNN) or a recurrent neural network model (RNN) to perform preliminary screening and feature extraction on the meeting link file, so as to relieve the computing pressure on the cloud server 200 and obtain the first feature. For the edge device 102, when using a lightweight model to perform preliminary screening and feature extraction on the meeting link file, the direct memory access (DMA) method can be adopted. DMA can efficiently transfer data from the memory to the algorithm processing unit without CPU intervention, so as to improve the computing efficiency of the edge device 102. Compressing and quantizing the convolutional neural network model or the recurrent neural network model can reduce the model size and improve the inference speed. For example, unimportant neurons in the model can be deleted through pruning techniques, or the weights and activation values of the model can be converted from floating-point numbers to low-precision integers. The quantization technique can reduce the numerical precision of the model from 32 bits to 16 bits or even lower, thereby reducing the computing amount and memory occupancy of the edge device 102. The cloud server 200 is used to perform in-depth analysis and classification on the first features uploaded by the edge device 102 by using powerful computing resources, improve the accuracy and efficiency of meeting link format recognition, and obtain the meeting link and the meeting start time. The cloud server 200 can communicate with the edge device 102. Multiple first features obtained by the edge device 102 and then deeply analyzed and classified by the cloud server 200 can relieve the pressure on the cloud server 200. The edge device 102 performs edge computing locally and processes the meeting link file in real time, while migrating complex or large-scale computing tasks to the cloud server 200. The cloud server 200 can provide large-scale storage resources and can store a large amount of training data and model files. The elastic computing ability allows the cloud server 200 to dynamically allocate computing resources according to needs. During the peak period of model training or inference, computing nodes can be quickly expanded, and resources can be released during low load to reduce costs. A meeting link file including a large number of different meeting links can be used to train the convolutional neural network or the recurrent neural network. The edge device 102 determines the first feature by performing preliminary screening and feature extraction on the meeting link file by using the compressed and quantized convolutional neural network model or recurrent neural network model, and its formula is expressed as:

[0025] where represents the iA meeting link file, where CNN and RNN respectively represent a convolutional neural network model and a recurrent neural network model applied to the compressed and quantized processing of meeting data. The cloud server 200 can further analyze the first feature through a deep learning model, and its formula is expressed as:

[0026] Among them, represents the first feature extracted by the edge device 102, represents the weighting coefficient, represents the k th deep learning model, such as the GPT-4 model, which is used to deeply analyze the first feature. Through the edge-cloud collaborative computing method, the recognition efficiency of the meeting link can be improved in a complex network environment, the meeting access speed can be accelerated, the requirement for network bandwidth can be reduced, and the response speed and stability of the system can be enhanced.

[0027] After receiving the meeting link and the meeting start time sent by the cloud server 200, the clock and network status monitoring module 103 automatically initiates a meeting access request near the meeting start time, such as at the first preset time before the meeting start time. The clock and network status monitoring module 103 uses a reinforcement learning algorithm, such as the Deep Q-Network (DNQ), to evaluate the quality of the network environment where the intelligent agent 100 is located in real time, and selects the optimal network access strategy to access the meeting according to the evaluation result. For example, it can intelligently switch between office Wi-Fi and mobile 5G to avoid meeting access failures or meeting interruptions caused by personnel movement.

[0028] The knowledge graph construction module 104 is used to construct the first knowledge graph according to the meeting materials. For example, the knowledge graph construction module 104 optimizes and fine-tunes a pre-trained large model (such as GPT-4) using the meeting materials through an online learning algorithm, and dynamically updates the parameters of the pre-trained large model to adapt to the language characteristics of the meeting scenario. The optimized pre-trained large model is used to analyze the meeting materials, extract the second feature, and classify and predict the meeting materials in combination with the random forest algorithm to determine the meeting type. A random forest model is customized according to different meeting types to improve the classification accuracy. Finally, according to the second feature and the meeting type, a knowledge graph construction tool (such as Stardog) is used to identify the entities and their relationships in the meeting materials, and the first knowledge graph is constructed. This first knowledge graph can clearly present the meeting theme, objectives, and key issues, helping the participants to clarify the meeting focus in advance and prepare materials targeted, thereby improving the meeting efficiency.

[0029] Furthermore, the meeting record and analysis module 105 can implement real-time multilingual translation using the Google Cloud Translation Model based on the Transformer architecture. The Google Cloud Translation Model is trained on a massive multilingual parallel corpus and supports translation of numerous language pairs. The meeting record and analysis module 105 translates the source language of the speech or text into the target language in real time according to the source language and target language of the speech, and performs target language speech synthesis through the Lark model to output the translation result in the form of speech or text. The meeting content can include: voice records, video records, and environmental sensor data. The meeting record and analysis module 105 is used to optimize the pre-trained large model according to the meeting content using an online learning algorithm. For example, the voice record is converted into text data, and the pre-trained large model is optimized according to the text data using an online learning algorithm. The optimized pre-trained large model (such as the GPT-4 model) analyzes and infers the text data. The multimodal fusion algorithm can analyze the voice record, video record, and environmental sensor data. The voice record can include acoustic features such as voice content, intonation, speech rate, and background noise, which can be used for semantic recognition and sentiment analysis. The video record can include visual features such as visual images, human movements, scene structures, facial expressions, and body language, which are used for object detection and behavior recognition. The environmental sensor data can include temperature, humidity, air quality, light intensity, and noise level for environmental state perception. The voice record, video record, and environmental sensor data can serve as multimodal data for the multimodal fusion algorithm. The core goal of the multimodal fusion algorithm is to make up for the limitations of a single modality (such as when the speech semantics are ambiguous, combining the video picture to clarify the referent object). Align the time and space of the multimodal data. The multimodal fusion algorithm can capture the internal dependencies within a single modality through self-attention (such as the association between different objects in a video), and model cross-modal interactions through cross-attention (such as the correspondence between speech keywords and video picture regions). A sentiment analysis model (such as a model based on BERT) can be used to judge the sentiment tendency of the speaker. Thus, the accuracy and comprehensiveness of the analysis results are improved. The speech recognition, semantic understanding, and sentiment analysis technologies cooperate with each other to achieve in-depth mining of the meeting content and generate a meeting summary document.

[0030] In the embodiments of the present application, the intelligent agent 100 and the cloud server 200 can accurately parse the meeting link. Even in a weak network or complex network environment, they can accurately access the meeting according to the meeting start time, avoiding the cumbersome steps of manually entering the link and repeatedly debugging. Before the meeting, the intelligent agent 100 uses natural language processing and machine learning technologies to deeply interpret the historical meeting minutes and the current topic documents, extract the core points from the massive information, and construct the first knowledge graph. The first knowledge graph can not only assist the participants in quickly grasping the background, but also provide solid data support for the meeting discussion and decision-making. During the meeting, according to advanced speech recognition, semantic understanding and sentiment analysis technologies, the intelligent agent 100 can not only transcribe the speech content in real time, but also intelligently extract the core viewpoints, and accurately distinguish between the approval and opposition positions. At the same time, the multi-language real-time translation function breaks down the language barriers. After the meeting, the intelligent agent 100 can generate a well-organized and key-pointed meeting summary document in a short time. From automatic access, pre-meeting preparation to post-meeting summary, the intelligent agent 100 runs through the entire meeting process, greatly improving the meeting efficiency and quality with an intelligent and automated management mode, providing an efficient and convenient digital meeting solution for enterprises and organizations, helping to reduce the operation cost, and promoting more scientific and efficient decision-making.

[0031] In another embodiment of the present disclosure, the formula for the clock and network status monitoring module to use the deep Q-network to select the optimal network access strategy to access the meeting is:

[0032] Among them, represents the network status where the intelligent agent is located Under the condition, by selecting the network access strategy , the expected return represents the reward represents the discount factor represents the maximum expected return obtained by selecting the network access strategy at the next time step .

[0033] In the embodiments of the present disclosure, the formula for the deep Q-network (DNQ, Deep Q-Network) to optimize the network access strategy is:

[0034] Among them, represents the value of taking the action in the state . At time step t, in the network status where the intelligent agent 100 is located, by selecting the access strategy , the expected return (including the discount of the immediate reward and future rewards) represents the reward represents a discount factor, represents the access strategy selected at the next time step, such as Wi-Fi, mobile 5G, represents the access strategy selected at the next time step , and the maximum expected return. For example, the network where the current agent 100 is located is a Wi-Fi network with poor network environment quality. The agent 100 uses a deep Q-network to update the Q value at each time step to gradually select the optimal network access strategy to access the meeting.

[0035] In another embodiment of the present disclosure, the clock and network status monitoring module 103 is further configured to send the shared network status information and the access meeting method request to other agents 100 when the network environment quality indicates that the meeting cannot be accessed, and access the meeting according to the shared network status information and the access meeting method sent by the received other agents 100.

[0036] In the embodiment of the present disclosure, in a complex network environment, through the innovation of the distributed collaborative agent architecture, multiple agents 100 cooperate to play a role in the network. When an agent 100 in a meeting room encounters difficulties in access, for example, when the network environment quality indicates that the meeting cannot be accessed, the clock and network status monitoring module 103 can ask other agents 100 for help, send the shared network status information and the access meeting method request to other agents 100, and access the meeting according to the shared network status information and the access meeting method sent by the received other agents 100. The shared network information may include: basic network metrics, such as Wi-Fi signal strength, bandwidth, latency; real-time load status, etc. The access meeting methods may include: Wi-Fi, 5G network, etc. It realizes collaborative access and improves the success rate and stability of access. In terms of knowledge processing, when an agent 100 encounters difficulties in understanding a specific term, the agent 100 can also send a shared knowledge request to other agents 100, and analyze the specific term according to the shared knowledge sent by the received other agents 100, improving the overall system's processing ability and adaptability to complex problems.

[0037] In another embodiment of the present disclosure, the knowledge graph construction module 104 is configured to optimize the pre-trained large model according to the meeting materials by using an online learning algorithm; analyze the meeting materials according to the optimized pre-trained large model to extract second features; predict the meeting type according to the meeting materials by using a random forest algorithm to determine the meeting type; determine the corresponding relationship between entities and entities in the meeting materials by using a knowledge graph construction tool according to the second features and the meeting type, and construct a first knowledge graph.

[0038] In the embodiments of the present disclosure, through an online learning algorithm, a pre-trained large model is optimized according to meeting materials, and then a first knowledge graph is constructed. The knowledge graph construction module 104 is used to optimize the pre-trained large model according to the meeting materials by using the online learning algorithm. The pre-trained large model can be a GPT-4 model. First, a large amount of meeting materials related to the meeting scenario are used to optimize and fine-tune GPT-4. During the optimization and fine-tuning process, the innovation of the dynamic adaptive learning mechanism is applied, and the parameters of the pre-trained large model are optimized in real time through the online learning algorithm. For example, when the information acquisition module 101 obtains new meeting materials, the knowledge graph construction module 104 will, based on the online learning algorithm of gradient descent, update the word vector representation and semantic understanding parameters of the pre-trained large model in real time according to the new terms, specific industry vocabulary, etc. in the meeting materials, so that the pre-trained large model can better adapt to the language characteristics of the meeting scenario. Analyze the meeting materials according to the optimized pre-trained large model to extract the second features. The second features can include basic information (such as time, location, participants, main topics), deep semantic features, association features, or structured information. Use the random forest algorithm to classify and predict the meeting materials to determine the meeting type. For example, use the text features of the meeting materials (such as keyword frequency, sentence length, etc.) as input to train a random forest model to predict the theme category, importance level, etc. of the materials. Through personalized learning customization, random forest models are constructed for different types of meetings (such as enterprise daily office meetings, cross-regional project collaboration meetings, academic discussion and exchange meetings) to improve the accuracy and pertinence of classification. According to the second features and the meeting type, use a knowledge graph construction tool (such as Stardog) to construct the first knowledge graph. The knowledge graph construction tool can identify the entities (such as people, organizations, events) in the meeting materials and the corresponding relationships between the entities (such as causal relationships, subordination relationships), and convert them into the nodes and edges of the knowledge graph, and then construct the first knowledge graph. Constructing the first knowledge graph can help organizers and participants clearly sort out the theme and objectives of the meeting. By presenting the core topics, related sub-topics, and expected goals of the meeting in the form of the first knowledge graph, everyone can have a more intuitive understanding of the focus of the meeting. For example, in a meeting on new product development, the first knowledge graph can clearly show key elements such as the functional requirements, market positioning, and technical difficulties of the product, so that the participants can clearly understand the core issues to be solved by the meeting before the meeting starts, and thus prepare relevant materials and viewpoints more pertinently, improving the efficiency of the meeting.

[0039] In another embodiment of the present disclosure, the meeting content includes: voice recordings, video recordings, and environmental sensor data; The meeting recording and analysis module 105 is used to optimize the translation model parameters in real time according to the speech content of the speaker during the meeting; Based on the source language and target language of the speaker, and the analysis results of the audio recording, video recording, and environmental sensor data by the multimodal fusion algorithm, the speech content of the speaker is translated in real time using the translation model with optimized parameters.

[0040] In the embodiments of the present disclosure, for the meeting record and analysis module 105, the translation model can be the Google Cloud Translation Model. Real-time multilingual translation can be achieved using the Google Cloud Translation Model. During the translation process, combined with the innovation of the dynamic adaptive learning mechanism, the parameters of the translation model are optimized in real time according to the new terms and specific industry vocabulary constantly appearing in the speech content of the speaker during the meeting, improving the accuracy and professionalism of the translation. The analysis results of the multimodal fusion algorithm for the audio recording, video recording, and environmental sensor data consider factors such as the emotion and tone of the speaker, making the translation result more natural and accurate. Based on the source language and target language of the speaker, and the analysis results of the multimodal fusion algorithm for the audio recording, video recording, and environmental sensor data, the speech content of the speaker is translated in real time using the translation model with optimized parameters. The deep interaction innovation of the multimodal fusion algorithm improves the accuracy and robustness of speaker annotation and multilingual translation. The dynamic adaptive learning mechanism ensures that the translation model can be continuously optimized according to the real-time situation of the meeting, improving the translation quality.

[0041] In another embodiment of the present disclosure, the meeting content includes: voice recording, video recording, and environmental sensor data; The meeting record and analysis module 105 is used to convert the voice into text data in real time according to the voice recording using a speech recognition model; According to the text data, an online learning algorithm is used to optimize the pre-trained large model; According to the analysis results of the optimized pre-trained large model for the text data and the analysis results of the multimodal fusion algorithm for the audio recording, video recording, and environmental sensor data, a meeting record document is generated; According to the meeting record document, a meeting summary document is generated.

[0042] In the embodiments of the present disclosure, the meeting record and analysis module 105 is used to convert speech into text data in real time according to the speech record by using a speech recognition model. The speech recognition model can be the DeepSpeech2 model for real-time speech recognition. This speech recognition model is trained on a large amount of speech data in meeting scenarios to adapt to the complex acoustic environment of the meeting room. The analysis results of audio records, video records, and environmental sensor data can be utilized by a multimodal fusion algorithm to improve the accuracy of speech recognition. For example, by analyzing the lip movements and facial expressions of the participants, the speech recognition model can be assisted in correcting misrecognized content. In some embodiments, a speaker recognition model can also be used to label the speaker identity of the text data. While recognizing speech, according to the output of the speaker recognition model (such as the i-vector model), the speaker identity of each speech segment in the text data is labeled. The analysis results of audio records, video records, and environmental sensor data, such as video image features (such as face recognition and limb movement analysis), are utilized by the multimodal fusion algorithm to further improve the accuracy and robustness of speaker labeling. The pre-trained large model can be the GPT-4 model. According to the text data, an online learning algorithm is used to optimize the pre-trained large model. Through the innovation of a dynamic adaptive learning mechanism, the pre-trained large model continuously optimizes the semantic understanding parameters according to the real-time meeting content being discussed, improving the ability to understand complex semantics. The optimized pre-trained large model is used to analyze the text data of the real-time speech recognition, and the analysis results of the text data are obtained. According to the analysis results of audio records, video records, and environmental sensor data by the multimodal fusion algorithm, the meaning, theme, and sentiment tendency of the sentences are recognized, and a meeting record document is generated. Based on the meeting record document, refinement and summary are carried out to generate a meeting summary document. Combining video image records and environmental sensor data to assist speech recognition significantly improves the accuracy of speech recognition, can better adapt to complex acoustic environments in meeting rooms and scenarios such as multiple people speaking simultaneously, and improves the accuracy of the meeting summary document.

[0043] In another embodiment of the present disclosure, the knowledge graph construction module 104 is further used to analyze audio records, video records, and environmental sensor data by using a multimodal fusion algorithm during the meeting; According to the meeting content, an online learning algorithm is used to optimize the pre-trained large model; According to the optimized pre-trained large model and the analysis results of audio records, video records, and environmental sensor data by the multimodal fusion algorithm, the meeting content is analyzed; and the first knowledge graph is updated according to the analysis results of the meeting content to generate a second knowledge graph; The meeting record and analysis module 105 is further used to generate a preliminary meeting collation report according to the meeting record document by using the T5 model; Submit the preliminary meeting collation report to the meeting chairperson for adjudication; Generate a meeting summary document using a natural language generation model based on the preliminary organized report after the ruling, the meeting record document, and the second knowledge graph.

[0044] In an embodiment of the present disclosure, the knowledge graph construction module 104 is further configured to, during the meeting, analyze audio records, video records, and environmental sensor data using a multimodal fusion algorithm. For example, the multimodal fusion algorithm can utilize the deep interaction innovation of multimodal fusion to fuse multimodal information such as audio records, video records, and environmental sensor data, further enriching the content of the first knowledge graph. For example, by analyzing features such as the facial expressions and actions of people in the video record, and the intonation and speech rate in the voice record, add emotional and behavioral features to the person nodes in the first knowledge graph. Environmental sensor data can include temperature, humidity, air quality, light intensity, and noise level. Environmental sensor data directly affects the comfort and attention of the participants, and thus affects the meeting efficiency. Optimize and fine-tune a pre-trained large model (such as GPT-4) according to the meeting content during the meeting. Through an online learning algorithm, update the parameters of the pre-trained large model in real time according to the meeting content during the meeting (such as the transcribed text of new meeting audio records, professional terms in specific industries, etc.). Its formula is expressed as:

[0045] where represents the learning rate, represents the gradient of the loss function, represents the loss of the model, so that the pre-trained large model continuously adapts to new terms, scenarios, and topics in the meeting.

[0046] For example, when new technical terms appear in the meeting, the pre-trained large model can adjust the word vector representation and semantic understanding parameters through a dynamic adaptive learning mechanism, so as to better understand the specific meanings of these terms in the meeting. According to the analysis results of the optimized pre-trained large model and the multimodal fusion algorithm on audio records, video records, and environmental sensor data, analyze the meeting content, and update the first knowledge graph in real time according to the analysis results of the meeting content. For example, update the key information and decision results in the meeting. Exemplarily, key information such as project budget figures and product technical parameters identified in the audio record, and the voting results of the participants on a certain plan observed in the video. Then generate a second knowledge graph to ensure that the second knowledge graph is consistent with the actual situation of the meeting. Its formula is expressed as:

[0047] where, represents the weight of each entity and relationship, reflecting its importance; and respectively represent the real-time data of entities in the meeting content and the relationships between entities.

[0048] The second knowledge graph more accurately and comprehensively reflects the content and structure of the meeting. For example, by analyzing the associations between different topics in the meeting, the connection relationships between topic nodes in the knowledge graph are optimized; the second knowledge graph with real-time updates can immediately reflect the discussion content, decisions, and action items in the meeting. For example, when a new idea or decision is proposed in the meeting, the knowledge graph can immediately add relevant nodes and relationships to ensure that all participants can timely understand the latest information. This immediacy reduces the time cost of post-meeting collation and update, and avoids misunderstandings or repeated discussions caused by information lag. The second knowledge graph can provide more powerful support for subsequent meeting summaries, decision-making support, and knowledge sharing. For example, through the second knowledge graph, the core content and key points of the meeting can be quickly understood, providing a basis for formulating subsequent action plans; the second knowledge graph can also be used as a knowledge sharing platform to facilitate team members to access key information and decision-making results in the meeting.

[0049] The meeting record and analysis module 105 is also used to preliminarily organize the meeting record document using a text summarization generation tool based on the T5 model. The T5 model can compress and refine long texts through an encoder-decoder architecture based on the attention mechanism, extract key information and core viewpoints, and obtain a preliminary meeting organization report. The preliminary meeting organization report is submitted to the meeting chairperson in a structured report form for adjudication. According to the adjudicated preliminary meeting organization report, the meeting record document, and the second knowledge graph, a detailed meeting summary document is generated using a natural language generation model. The meeting summary document can include the meeting topic, participants, meeting agenda, main discussion content, reached consensus, unresolved issues, and the next action plan, etc. Through innovation of the dynamic adaptive learning mechanism, according to the actual situation of the meeting and the adjudication result of the meeting chairperson on the preliminary meeting organization report, the parameters of the natural language generation model are continuously optimized to make the language expression of the meeting summary document more natural and fluent, and the content more accurate and comprehensive.

[0050] In another embodiment of the present disclosure, the clock and network status monitoring module 103 is also used to send a meeting notice to meeting participants after receiving the meeting link and meeting start time sent by the cloud server. The meeting notice includes: the meeting link, the meeting start time, and the first knowledge graph; The knowledge graph construction module 104 is also used to introduce the meeting content using a pre-trained speech synthesis model according to the first knowledge graph; The meeting record and analysis module 105 is also used to generate an AR plan according to the meeting summary document and the meeting room environment characteristics recognized by computer vision.

[0051] In the embodiments of the present disclosure, the clock and network status monitoring module 103 is further configured to send a meeting notice to meeting participants after receiving the meeting link and meeting start time sent by the cloud server. The list of meeting participants can be obtained from the cloud server or specified by the meeting chairperson. The notification channels may include: email, mobile phone number, etc. The meeting notice includes: the meeting link, the meeting start time, and the first knowledge graph. The first knowledge graph is constructed by extracting key information based on meeting materials such as the summary of the previous meeting and the theme of the current meeting. The first knowledge graph can help meeting participants master the meeting background in a short time, reduce the understanding cost, and improve the discussion efficiency.

[0052] The knowledge graph construction module 104 uses a pre-trained speech synthesis model (such as the Lark model) to introduce the meeting content and clarify the meeting theme according to the first knowledge graph. The speech synthesis model is trained based on a large amount of speech data and can simulate human speech intonation and emotional expression to make the opening more vivid and professional.

[0053] The meeting record and analysis module 105 can use the Vuforia augmented reality engine to convert the information such as tasks, responsible persons, and time nodes determined in the meeting discussion in the meeting summary document into an AR plan. The AR plan can use computer vision technology to identify the environmental characteristics of the meeting room and superimpose the task-related information in an intuitive and visual form (such as virtual tags, task progress bars, etc.) on the field of vision of the corresponding responsible person (through AR glasses or mobile phone AR applications). For example, the task progress bar information is displayed as:

[0054] Among them, represents the progress percentage of the task, represents the scheduled time of the task, represents the responsible person of the task. Each task in the virtual space can be described by the position vector ∈ and the virtual information as:

[0055] Among them, represents the current timestamp and updates the task progress in real time. For example, the meeting record and analysis module 105, according to In the location within the meeting room space, task information is superimposed in real time in the form of progress bar information and other forms within the field of vision of the person in charge. During the generation process of the AR plan, based on the meeting summary document and the analysis results of audio records, video records, and environmental sensor data by the multi-modal fusion algorithm, the AR plan can be made more in line with the actual situation of the meeting, improving the accuracy and efficiency of task execution. Its formula is expressed as:

[0056] Among them, represents the output after the emotional analysis of the audio record, represents the output after the face recognition of the video record, represents the environmental sensor data, and the weight can be dynamically adjusted according to the importance of the task and the real-time situation. For example, the weight of is set to a larger value when discussing emotion-related tasks. According to the generated AR plan, the meeting record and analysis module 105 can track the actions and positions of the responsible personnel. For example, when a certain responsible person approaches the task target, the meeting record and analysis module 105 can automatically push the relevant task progress or reminder. This interactive design enables the AR plan to be not only statically displayed, but also updated in real time according to the task progress, ensuring the accuracy and efficiency of task execution.

[0057] Figure 3 is a schematic diagram of the functions implemented by the meeting system. As Figure 3 shown, it includes: S301: Pre-meeting preparation stage: Edge-cloud collaborative computing accesses the meeting; network status information is shared among multiple intelligent agents; a first knowledge graph is constructed; a meeting notice is sent; S302: During the meeting: The meeting content is introduced; real-time translation is performed according to the source language and target language of the speaker, combined with the multi-modal fusion algorithm; the first knowledge graph is updated using the optimized pre-trained large model and the multi-modal fusion algorithm to generate a second knowledge graph; S303: After the meeting: The meeting content is analyzed using the optimized pre-trained large model and the multi-modal fusion algorithm to generate a meeting summary document; an AR plan is generated.

[0058] In the pre-meeting preparation stage, the recognition efficiency of conference links is improved through edge-cloud collaborative computing, which improves the response speed and stability of the system. Network status information and access methods are shared between multiple intelligent agents to improve the efficiency of accessing meetings. The first knowledge graph is constructed to help participants clarify the focus of the meeting in advance and improve the efficiency of the meeting. During the meeting, real-time translation is performed based on the source language and target language of the speaker in combination with the multimodal fusion algorithm to improve the translation accuracy. The first knowledge graph is updated using the optimized pre-trained large model and multimodal fusion algorithm to generate the second knowledge graph, which can improve the accuracy of the second knowledge graph construction, instantly reflect the discussion content, decisions and action items in the meeting, and improve the efficiency of the meeting. After the meeting, the meeting content is analyzed using the optimized pre-trained large model and multimodal fusion algorithm to generate a meeting summary document, which can improve the accuracy of the meeting summary. The whole process of intelligent management of the meeting from preparation to post-meeting processing is realized, which improves the user experience, improves the overall efficiency and quality of the meeting, and provides enterprises and organizations with a more convenient and efficient meeting solution.

[0059] Based on the same disclosed concept, the disclosed embodiment also provides an intelligent agent. Since the principle of solving the problem by the intelligent agent is similar to that of the aforementioned conference system, the implementation of the intelligent agent can refer to the implementation of the aforementioned conference system, and the repeated parts will not be repeated.

[0060] The present disclosure provides an intelligent agent 100, such as Figure 1 As shown, it includes: an information acquisition module 101, an edge device 102, a clock and network status monitoring module 103, a knowledge graph construction module 104 and a meeting record and analysis module 105; The information acquisition module 101 is used to acquire conference link files and conference materials; The edge device 102 is used to use a convolutional neural network model or a recurrent neural network model after compression and quantization to perform preliminary screening and feature extraction on the conference link file to determine a first feature; and send the first feature to the cloud server 200; the cloud server 200 is used to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device 102, and determine the conference link and the conference start time; The clock and network status monitoring module 103 is used to evaluate the quality of the network environment where the agent 100 is located within a first preset time before the meeting start time after receiving the meeting link and meeting start time sent by the cloud server 200, and use a deep Q network to select the optimal network access strategy to access the meeting; The knowledge graph construction module 104 is used to construct a first knowledge graph based on the conference materials; The meeting record and analysis module 105 is used to perform real-time translation according to the source language and target language of the speaker during the meeting; and generate a meeting summary document according to the meeting content.

[0061] In another embodiment of the present disclosure, the formula for the clock and network status monitoring module to select the optimal network access strategy to access the meeting using the deep Q network is:

[0062] Wherein, represents the network status where the agent is located Under the condition, by selecting the network access strategy , the expected return represents the reward represents the discount factor represents selecting the network access strategy at the next time step , the maximum expected return.

[0063] In another embodiment of the present disclosure, the clock and network status monitoring module 103 is further used to send a shared network status information and a request for a meeting access method to other agents 100 when the network environment quality characterization cannot access the meeting, and access the meeting according to the received shared network status information and meeting access method sent by other agents 100.

[0064] In another embodiment of the present disclosure, the knowledge graph construction module 104 is used to optimize the pre-trained large model using an online learning algorithm according to the meeting materials; Analyze the meeting materials according to the optimized pre-trained large model to extract second features; Predict the meeting type using a random forest algorithm according to the meeting materials to determine the meeting type; Determine the entities in the meeting materials and the corresponding relationships between the entities using a knowledge graph construction tool according to the second features and the meeting type, and construct a first knowledge graph.

[0065] In another embodiment of the present disclosure, the meeting content includes: voice records, video records, and environmental sensor data; The meeting record and analysis module 105 is used to optimize the translation model parameters in real time according to the speech content of the speaker during the meeting; According to the source language and target language of the speaker, and the analysis results of the audio record, video record, and environmental sensor data by the multi-modal fusion algorithm, use the optimized translation model to perform real-time translation on the speech content of the speaker.

[0066] In another embodiment of the present disclosure, the meeting content includes: voice recordings, video recordings, and environmental sensor data; The meeting recording and analysis module 105 is configured to use a speech recognition model according to the voice recording to convert the voice into text data in real time; According to the text data, an online learning algorithm is used to optimize the pre-trained large model; According to the analysis results of the optimized pre-trained large model for the text data and the analysis results of the audio recording, video recording, and environmental sensor data by the multi-modal fusion algorithm, a meeting record document is generated; According to the meeting record document, a meeting summary document is generated.

[0067] In another embodiment of the present disclosure, the knowledge graph construction module 104 is further configured to, during the meeting, use a multi-modal fusion algorithm to analyze the audio recording, video recording, and environmental sensor data; According to the meeting content, an online learning algorithm is used to optimize the pre-trained large model; According to the analysis results of the optimized pre-trained large model and the multi-modal fusion algorithm for the audio recording, video recording, and environmental sensor data, the meeting content is analyzed; and according to the analysis results of the meeting content, the first knowledge graph is updated to generate a second knowledge graph; The meeting recording and analysis module 105 is further configured to generate a preliminary meeting collation report according to the meeting record document using the T5 model; Submit the preliminary meeting collation report to the meeting chairperson for adjudication; According to the adjudicated preliminary meeting collation report, the meeting record document, and the second knowledge graph, a natural language generation model is used to generate a meeting summary document.

[0068] In another embodiment of the present disclosure, the clock and network status monitoring module 103 is further configured to, after receiving the meeting link and meeting start time sent by the cloud server, send a meeting notice to the meeting participants, where the meeting notice includes: the meeting link, the meeting start time, and the first knowledge graph; The knowledge graph construction module 104 is further configured to introduce the meeting content according to the first knowledge graph using a pre-trained speech synthesis model; The meeting recording and analysis module 105 is further configured to generate an AR plan according to the meeting summary document and the meeting room environment features recognized by computer vision.

[0069] Based on the same general inventive concept, embodiments of the present disclosure further provide a conference processing method. Since the principle of the problem solved by this conference processing method is similar to that of the foregoing conference system, the implementation of this conference processing method can refer to the implementation of the foregoing conference system, and repeated parts will not be elaborated.

[0070] Embodiments of the present disclosure provide a conference processing method, as Figure 4 shown, including the following steps: S401. Obtain a conference link file and conference materials; S402. Use a compressed and quantized convolutional neural network model or a recurrent neural network model to perform preliminary screening and feature extraction on the conference link file to determine a first feature; S403. Dynamically allocate computing resources to analyze the first feature to determine the conference link and the conference start time; S404. According to the conference link and the conference start time, evaluate the network environment quality within a first preset time before the conference start time, and use a deep Q-network to select an optimal network access strategy to access the conference; S405. Construct a first knowledge graph according to the conference materials; S406. Perform real-time translation according to the source language and target language of the speaker during the conference; and generate a conference summary document according to the conference content.

[0071] In another embodiment of the present disclosure, the formula for using a deep Q-network to select an optimal network access strategy to access the conference is expressed as:

[0072] where represents the network state where the agent is located under which, by selecting a network access strategy , the expected obtained reward, represents the reward, represents the discount factor, represents the maximum expected obtained reward by selecting a network access strategy at the next time step.

[0073] In another embodiment of the present disclosure, the method further includes: in the case where the network environment quality indicates that the conference cannot be accessed, send the shared network state information and a request for a conference access method to other agents, and access the conference according to the received shared network state information and conference access method sent by other agents.

[0074] In another embodiment of the present disclosure, constructing a first knowledge graph according to the conference materials includes: According to the conference materials, use an online learning algorithm to optimize a pre-trained large model; Analyze the meeting materials according to the optimized pre-trained large model and extract the second feature; Predict the meeting type according to the meeting materials by using the random forest algorithm and determine the meeting type; According to the second feature and the meeting type, use a knowledge graph construction tool to determine the entities in the meeting materials and the corresponding relationships between entities, and construct the first knowledge graph.

[0075] In another embodiment of the present disclosure, the meeting content includes: voice records, video records, and environmental sensor data; Perform real-time translation according to the source language and target language of the speaker during the meeting, including: Optimize the translation model parameters in real time according to the speech content of the speaker during the meeting; According to the source language and target language of the speaker and the analysis results of the audio record, video record, and environmental sensor data by the multi-modal fusion algorithm, use the optimized translation model to perform real-time translation on the speech content of the speaker.

[0076] In another embodiment of the present disclosure, the meeting content includes: voice records, video records, and environmental sensor data; Generate a meeting summary document according to the meeting content, including: According to the voice record, use a speech recognition model to convert the voice into text data in real time; Optimize the pre-trained large model according to the text data by using an online learning algorithm; Generate a meeting record document according to the analysis results of the optimized pre-trained large model on the text data and the analysis results of the audio record, video record, and environmental sensor data by the multi-modal fusion algorithm; Generate a meeting summary document according to the meeting record document.

[0077] In another embodiment of the present disclosure, the method further includes: During the meeting, analyze the audio record, video record, and environmental sensor data by using the multi-modal fusion algorithm; Optimize the pre-trained large model according to the meeting content by using an online learning algorithm; Analyze the meeting content according to the analysis results of the optimized pre-trained large model and the audio record, video record, and environmental sensor data by the multi-modal fusion algorithm; and update the first knowledge graph according to the analysis results of the meeting content to generate a second knowledge graph; Generate a meeting summary document according to the meeting record document, including: Generate a preliminary meeting collation report according to the meeting record document by using the T5 model; Submit the preliminary meeting summary report to the meeting chairperson for adjudication; Generate a meeting summary document using a natural language generation model based on the adjudicated preliminary meeting summary report, the meeting record document, and the second knowledge graph.

[0078] In another embodiment of the present disclosure, the method further includes: sending a meeting notice to the meeting participants, where the meeting notice includes: a meeting link, a meeting start time, and a first knowledge graph; Introduce the meeting content using a pre-trained speech synthesis model based on the first knowledge graph; Generate an AR plan based on the meeting summary document and the meeting room environment features recognized by computer vision.

[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present disclosure.

[0080] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present disclosure.

[0081] Those skilled in the art can understand that the modules in the device in the embodiments can be distributed in the device in the embodiments according to the description of the embodiments, or can be changed accordingly and located in one or more devices different from the present embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.

[0082] The serial numbers of the above embodiments of the present disclosure are only for description and do not represent the advantages and disadvantages of the embodiments.

[0083] Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and variations.

Claims

1. A conference system, characterized in that, Including: An agent and a cloud server, where the agent includes: an information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module, and a meeting record and analysis module; The information acquisition module is used to acquire a meeting link file and meeting materials; The edge device is used to perform preliminary screening and feature extraction on the meeting link file using a compressed and quantized convolutional neural network model or a recurrent neural network model to determine the first feature; The cloud server is used to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device to determine the meeting link and meeting start time; The clock and network status monitoring module is used to evaluate the network environment quality where the agent is located within a first preset time before the meeting start time after receiving the meeting link and meeting start time sent by the cloud server, and use a deep Q-network to select the optimal network access strategy to access the meeting; The knowledge graph construction module is used to construct a first knowledge graph based on the meeting materials; The meeting record and analysis module is used to perform real-time translation according to the source language and target language of the speaker during the meeting; and generate a meeting summary document according to the meeting content.

2. The system according to claim 1, wherein The formula for the clock and network status monitoring module to use a deep Q-network to select the optimal network access strategy to access the meeting is: Among them, represents the network state where the agent is located Under this condition, by selecting a network access policy the expected obtained return represents the reward represents the discount factor represents the maximum expected return obtained by selecting a network access policy at the next time step ​ 3. The system according to claim 1, wherein The clock and network status monitoring module is further used to send shared network status information and a request for a meeting access method to other agents in the case where the network environment quality indicates that the meeting cannot be accessed, and access the meeting according to the shared network status information and the meeting access method sent by the received other agents.

4. The system according to claim 1, wherein The knowledge graph construction module is used to optimize a pre-trained large model using an online learning algorithm according to the meeting materials; Analyze the meeting materials according to the optimized pre-trained large model to extract the second feature; Predict the meeting type using a random forest algorithm according to the meeting materials to determine the meeting type; Determine the entities in the meeting materials and the corresponding relationships between the entities using a knowledge graph construction tool according to the second feature and the meeting type, and construct a first knowledge graph.

5. The system according to claim 1, wherein The meeting content includes: voice records, video records, and environmental sensor data; The meeting record and analysis module is used to optimize the translation model parameters in real time according to the speaker's speech content during the meeting; According to the source language and target language of the speaker and the analysis results of the audio record, video record, and environmental sensor data by a multimodal fusion algorithm, use the optimized translation model to perform real-time translation on the speaker's speech content.

6. The system according to claim 1, wherein The meeting content includes: voice records, video records, and environmental sensor data; The meeting record and analysis module is used to use a speech recognition model according to the voice record to convert the voice into text data in real time; Optimize a pre-trained large model using an online learning algorithm according to the text data; Generate a meeting record document based on the analysis results of the text data by the optimized pre-trained large model and the analysis results of the audio records, video records, and environmental sensor data by the multimodal fusion algorithm; Generate a meeting summary document based on the meeting record document.

7. The system according to claim 6, wherein The knowledge graph construction module is further configured to analyze the audio records, video records, and environmental sensor data using the multimodal fusion algorithm during the meeting; Optimize the pre-trained large model using an online learning algorithm according to the meeting content; Analyze the meeting content based on the analysis results of the optimized pre-trained large model and the multimodal fusion algorithm for the audio records, video records, and environmental sensor data; and update the first knowledge graph according to the analysis results of the meeting content to generate a second knowledge graph; The meeting record and analysis module is further configured to generate a preliminary meeting arrangement report using the T5 model according to the meeting record document; Submit the preliminary meeting arrangement report to the meeting chairperson for adjudication; Generate a meeting summary document using a natural language generation model according to the adjudicated preliminary meeting arrangement report, the meeting record document, and the second knowledge graph.

8. The system according to claim 1, characterized in that The clock and network status monitoring module is further configured to send a meeting notice to the meeting participants after receiving the meeting link and meeting start time sent by the cloud server, where the meeting notice includes: the meeting link, the meeting start time, and the first knowledge graph; The knowledge graph construction module is further configured to introduce the meeting content using a pre-trained speech synthesis model according to the first knowledge graph; The meeting record and analysis module is further configured to generate an AR plan according to the meeting summary document and the meeting room environment features recognized by computer vision.

9. An agent, characterized in that, Comprising: An information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module, and a meeting record and analysis module; The information acquisition module is configured to acquire a meeting link file and meeting materials; The edge device is configured to perform preliminary screening and feature extraction on the meeting link file using a compressed and quantized convolutional neural network model or a recurrent neural network model to determine a first feature; And send the first feature to the cloud server; The cloud server is configured to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device to determine the meeting link and meeting start time; The clock and network status monitoring module is configured to evaluate the network environment quality of the agent within a first preset time before the meeting start time after receiving the meeting link and meeting start time sent by the cloud server, and select an optimal network access strategy to access the meeting using a deep Q network; The knowledge graph construction module is configured to construct a first knowledge graph according to the meeting materials; The meeting record and analysis module is configured to perform real-time translation according to the source language and target language of the speaker during the meeting; and generate a meeting summary document according to the meeting content.

10. A conference processing method, characterized in that, Comprising: Acquire a meeting link file and meeting materials; The convolutional neural network model or recurrent neural network model after compression and quantization processing is used to preliminarily screen and extract features from the meeting link file to determine the first feature; Computing resources are dynamically allocated to analyze the first feature to determine the meeting link and the meeting start time; According to the meeting link and the meeting start time, the network environment quality is evaluated within the first preset time before the meeting start time, and the deep Q-network is used to select the optimal network access strategy to access the meeting; According to the meeting materials, a first knowledge graph is constructed; Real-time translation is performed according to the source language and target language of the speaker during the meeting; and a meeting summary document is generated according to the meeting content.

Citation Information

Patent Citations

  • Method for quickly joining conference through applet card

    CN110601863A

  • Heterogeneous wireless network access selection method and system based on SDN

    CN111586809A

  • Multi-modal conference recording method and system based on edge calculation

    CN116405635A

  • Conference access method and device, equipment and storage medium

    CN118509272A

  • Mobile video conference interaction system based on smart learning tablet

    CN119277012A

Cited By

  • Method and device for realizing intelligent quality valve in demand submission process, equipment and medium

    CN120610686A

  • Methods, devices, equipment, and media for implementing intelligent quality valves in the demand submission process.

    CN120610686B

  • Intelligent conference implementation method and device, equipment and storage medium

    CN120750685A

  • Multi-mode environment perception adaptive calibration system and method of intelligent projector

    CN120786043A

  • Multi-agent collaborative cross-language system translation method and system based on loAs

    CN121189343A