A conference system, intelligent device and conference processing method
Through the collaborative work of smart devices and cloud servers, the problems of low efficiency and insufficient information utilization of traditional conference systems have been solved, automatic access, in-depth pre-meeting processing and real-time translation have been achieved, and the overall efficiency and quality of the meeting have been improved.
Patent Information
- Application Number
- CN202510771762.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Traditional conferencing systems are inefficient when connecting to meetings, with insufficient utilization of pre-meeting information, inaccurate meeting records and analysis, and poor real-time multi-language translation, all of which affect meeting efficiency and decision-making quality.
By adopting smart devices and cloud servers, through information acquisition modules, edge devices, clock and network status monitoring modules, knowledge graph construction modules and meeting recording and analysis modules, automatic access to meetings, in-depth processing of pre-meeting information, real-time translation and accurate recording and analysis can be achieved.
It realizes the automated and intelligent management of the conference system, improves the efficiency of conference access and the utilization rate of pre-meeting information, ensures the accuracy of meeting records and the real-time translation, and improves the efficiency and quality of meetings.
Smart Images

Figure CN120297933B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of conference technology, and in particular to a conference system, an intelligent device, and a conference processing method. Background Art
[0002] In today's rapidly developing digital office environment, meetings are an important way for businesses, organizations, and other users to communicate, collaborate, and make decisions. Improving their efficiency and quality is becoming increasingly critical. However, traditional conference room systems have gradually exposed the following problems in their long-term use, making them difficult to meet the needs of modern offices:
[0003] 1. Inefficient conference access: Traditional conference rooms typically rely on manual access to conference systems. This manual access severely impacts meeting start-up efficiency. Surveys show that in some large enterprises, daily meetings are delayed an average of two to three times per week due to participants' manual access to the conference system, with each delay lasting between five and 15 minutes. This not only wastes significant work time and reduces efficiency, but can also delay important decisions, resulting in potential financial losses for the company.
[0004] 2. Insufficient use of pre-meeting information: Due to a lack of in-depth processing of pre-meeting materials, participants struggle to quickly form a comprehensive understanding of the meeting content beforehand, unable to quickly engage in effective discussions and decision-making. This increases communication costs in the early stages of the meeting, easily hinders the progress of the meeting, and makes it difficult to achieve the desired results. For example, in some meetings involving complex projects, due to participants' limited understanding of pre-meeting materials, a considerable amount of time is spent re-examining project background and key issues, resulting in extended meeting times and relatively insufficient time for actual solution discussion.
[0005] 3. Inaccurate meeting records and analysis: Traditional methods of recording and analyzing meetings cannot meet the requirements of modern meetings for accuracy and comprehensiveness. The subjectivity and incompleteness of manual recording mean that meeting records may not fully reflect the meeting, easily omitting important insights and decision-making information. Relying solely on recording equipment for recording, the subsequent compilation and analysis is tedious and time-consuming, failing to provide timely and effective support for meeting decisions. In multilingual meetings, the lack of real-time and effective translation severely limits communication between participants speaking different languages and hinders the smooth flow of information.
[0006] Based on the above problems, there is an urgent need to provide a conference system that can provide users with a more convenient, efficient and intelligent conference experience. Summary of the Invention
[0007] The embodiments of the present disclosure provide a conference system, an intelligent device, and a conference processing method to solve the existing problems of low conference access efficiency, insufficient use of pre-conference information, and inaccurate conference records and analysis.
[0008] Based on the above problems, in a first aspect, an embodiment of the present disclosure provides a conference system, comprising: an intelligent device and a cloud server, wherein the intelligent device comprises: an information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module, and a conference recording and analysis module;
[0009] The information acquisition module is used to obtain conference link files and conference materials;
[0010] The edge device is configured to perform preliminary screening and feature extraction on the conference link file using a convolutional neural network model or a recurrent neural network model that has been compressed and quantized, to determine a first feature;
[0011] The cloud server is configured to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device, and determine the conference link and conference start time;
[0012] The clock and network status monitoring module is used to evaluate the network environment quality of the smart device within a first preset time before the meeting start time after receiving the meeting link and meeting start time sent by the cloud server, and use the deep Q network to select the optimal network access strategy to access the meeting;
[0013] The knowledge graph construction module is used to construct a first knowledge graph based on the conference materials;
[0014] The conference record and analysis module is used to perform real-time translation according to the speaker's source language and target language during the conference; and to generate a conference summary document based on the conference content.
[0015] In conjunction with the first aspect, in one possible implementation, the clock and network status monitoring module is used to select an optimal network access strategy for accessing a conference using a deep Q network. The formula is expressed as follows:
[0016]
[0017] in, Indicates the network status of the smart device Next, select Network Access Policy , expected return, Indicates reward, represents the discount factor, Indicates the network access strategy selected in the next time step , the maximum expected return.
[0018] In combination with the first aspect, in a possible implementation, the clock and network status monitoring module is also used to send shared network status information and conference access method requests to other smart devices when the network environment quality indicates that the meeting cannot be accessed, and to access the meeting based on the shared network status information and conference access method received from other smart devices.
[0019] In conjunction with the first aspect, in one possible implementation, the knowledge graph construction module is configured to optimize the pre-trained large model using an online learning algorithm based on the conference materials;
[0020] Analyze the conference data according to the optimized pre-trained large model to extract the second feature;
[0021] Based on the meeting information, a random forest algorithm is used to predict the meeting type and determine the meeting type;
[0022] According to the second feature and the conference type, a knowledge graph construction tool is used to determine the entities in the conference materials and the corresponding relationships between the entities to construct a first knowledge graph.
[0023] In conjunction with the first aspect, in one possible implementation, the conference content includes: voice recording, video recording, and environmental sensor data;
[0024] The conference recording and analysis module is used to optimize the translation model parameters in real time according to the speaker's speech content during the conference;
[0025] Based on the speaker's source language and target language, and the analysis results of the multimodal fusion algorithm on voice recordings, video recordings and environmental sensor data, a translation model with optimized parameters is used to translate the speaker's speech in real time.
[0026] In conjunction with the first aspect, in one possible implementation, the conference content includes: voice recording, video recording, and environmental sensor data;
[0027] The conference recording and analysis module is used to convert the voice into text data in real time using a voice recognition model based on the voice recording;
[0028] Based on the text data, an online learning algorithm is used to optimize the pre-trained large model;
[0029] Generate meeting record documents based on the analysis results of text data by the optimized pre-trained large model and the analysis results of voice recordings, video recordings and environmental sensor data by the multimodal fusion algorithm;
[0030] Generate a meeting summary document based on the meeting record document.
[0031] In conjunction with the first aspect, in one possible implementation, the knowledge graph construction module is further configured to analyze voice recordings, video recordings, and environmental sensor data using a multimodal fusion algorithm during the meeting;
[0032] Based on the meeting content, an online learning algorithm is used to optimize the pre-trained large model;
[0033] The meeting content is analyzed based on the analysis results of voice recordings, video recordings and environmental sensor data using the optimized pre-trained large model and multimodal fusion algorithm; and the first knowledge graph is updated based on the analysis results of the meeting content to generate a second knowledge graph.
[0034] The meeting record and analysis module is further used to generate a preliminary meeting summary report based on the meeting record document using the T5 model;
[0035] Submit the preliminary report of the meeting to the chairman of the meeting for decision;
[0036] Based on the preliminary summary report of the meeting after the ruling, the meeting record document and the second knowledge graph, a natural language generation model is used to generate a meeting summary document.
[0037] In conjunction with the first aspect, in one possible implementation, the clock and network status monitoring module is further configured to, after receiving the meeting link and meeting start time sent by the cloud server, send a meeting notification to the meeting participants, the meeting notification including: the meeting link, the meeting start time, and the first knowledge graph;
[0038] The knowledge graph construction module is further configured to introduce the content of the meeting using a pre-trained speech synthesis model based on the first knowledge graph;
[0039] The meeting recording and analysis module is further used to generate an AR plan based on the meeting summary document and the meeting room environment features identified by computer vision.
[0040] In a second aspect, an intelligent device is provided, comprising: an information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module, and a meeting recording and analysis module;
[0041] The information acquisition module is used to obtain conference link files and conference materials;
[0042] The edge device is configured to perform preliminary screening and feature extraction on the conference link file using a convolutional neural network model or a recurrent neural network model that has undergone compression and quantization processing to determine a first feature; and transmit the first feature to a cloud server; and the cloud server is configured to dynamically allocate computing resources to analyze the first feature after receiving the first feature transmitted by the edge device, and determine the conference link and the meeting start time.
[0043] The clock and network status monitoring module is used to evaluate the network environment quality of the smart device within a first preset time before the meeting start time after receiving the meeting link and meeting start time sent by the cloud server, and use the deep Q network to select the optimal network access strategy to access the meeting;
[0044] The knowledge graph construction module is used to construct a first knowledge graph based on the conference materials;
[0045] The conference record and analysis module is used to perform real-time translation according to the speaker's source language and target language during the conference; and to generate a conference summary document based on the conference content.
[0046] A third aspect provides a conference processing method, including:
[0047] Obtain conference link files and conference materials;
[0048] Using a convolutional neural network model or a recurrent neural network model that has been compressed and quantized to perform preliminary screening and feature extraction on the conference link file to determine a first feature;
[0049] Dynamically allocating computing resources to analyze the first feature and determine a conference link and a conference start time;
[0050] According to the conference link and the conference start time, the network environment quality is evaluated within a first preset time before the conference start time, and the optimal network access strategy is selected using the deep Q network to access the conference;
[0051] Constructing a first knowledge graph based on the conference materials;
[0052] During the meeting, real-time translation is performed based on the speaker's source and target languages; and a meeting summary document is generated based on the meeting content.
[0053] The beneficial effects of the embodiments of the present disclosure include:
[0054] The conference system, smart device, and conference processing method provided by the embodiments of the present disclosure include: a smart device and a cloud server, wherein the smart device includes: an information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module, and a conference recording and analysis module; the information acquisition module is used to obtain conference link files and conference materials; the edge device is used to use a convolutional neural network model or a recurrent neural network model that has been compressed and quantized to perform preliminary screening and feature extraction on the conference link files to determine a first feature; the cloud server is used to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device to determine the conference link and the meeting start time; the clock and network status monitoring module is used to, after receiving the conference link and the meeting start time sent by the cloud server, evaluate the quality of the network environment in which the smart device is located within a first preset time before the meeting start time, and use a deep Q network to select the optimal network access strategy for accessing the meeting; the knowledge graph construction module is used to construct a first knowledge graph based on the conference materials; the conference recording and analysis module is used to perform real-time translation according to the speaker's source language and target language during the meeting; and generate a meeting summary document based on the meeting content. This disclosure leverages smart devices and cloud servers to identify various conference system link formats through algorithms. This allows for automatic and precise access to conference systems based on meeting times in complex network environments, significantly simplifying the conference access process. Before a meeting, the smart device uses natural language processing technology and machine learning algorithms to deeply analyze user-provided information, such as summaries of previous meetings and the current meeting topic, extracting key information and constructing a primary knowledge graph, laying a solid data foundation for meeting interaction and decision-making. During the meeting, using advanced speech recognition, semantic understanding, and sentiment analysis technologies, the smart device records participant speeches in real time, accurately summarizes points, and clearly distinguishes between pros and cons. It also enables real-time speech recognition, speaker annotation, and on-demand multilingual translation. After the meeting, the smart device quickly generates a meeting summary document. With features such as intelligent automatic access, in-depth pre-meeting information processing, and intelligent speech recording and analysis, it fully realizes automated and intelligent management of the entire meeting process, significantly improving meeting efficiency and quality. This provides businesses, organizations, and other users with an efficient, convenient, and intelligent conference solution that meets the needs of digital office work, helping users reduce costs and increase efficiency in conference scenarios and promote more scientific and accurate decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 A schematic diagram of the structure of a conference system provided in an embodiment of the present disclosure;
[0056] Figure 2 A schematic diagram of the structure of edge-cloud collaborative computing provided by an embodiment of the present disclosure;
[0057] Figure 3 A schematic diagram illustrating the functions of the conference system provided in an embodiment of the present disclosure;
[0058] Figure 4 A schematic diagram of a conference processing method provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0059] The present disclosure provides a conference system, intelligent device, and conference processing method. Preferred embodiments of the present disclosure are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are intended only to illustrate and explain the present disclosure and are not intended to limit the present disclosure. Furthermore, the embodiments and features of the embodiments may be combined with each other unless there is a conflict.
[0060] The present disclosure provides a conference system. Figure 1 As shown, it includes: a smart device 100 and a cloud server 200, the smart device 100 includes: an information acquisition module 101, an edge device 102, a clock and network status monitoring module 103, a knowledge graph construction module 104 and a meeting recording and analysis module 105;
[0061] Information acquisition module 101, used to obtain conference link files and conference materials;
[0062] The edge device 102 is configured to perform preliminary screening and feature extraction on the conference link file using a convolutional neural network model or a recurrent neural network model that has been compressed and quantized, to determine a first feature;
[0063] The cloud server 200 is configured to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device 102 to determine the meeting link and meeting start time;
[0064] The clock and network status monitoring module 103 is configured to, after receiving the conference link and conference start time from the cloud server 200, evaluate the quality of the network environment where the smart device 100 is located within a first preset time before the conference start time, and select the optimal network access strategy for accessing the conference using a deep Q network;
[0065] A knowledge graph construction module 104 is configured to construct a first knowledge graph based on conference materials;
[0066] The conference record and analysis module 105 is used to perform real-time translation according to the speaker's source language and target language during the conference, and generate a conference summary document based on the conference content.
[0067] The disclosed embodiments are applied to the field of conference technology. Traditional conference systems typically rely on manual operation to access conferences. Before the meeting begins, participants must open the corresponding conference software and enter information such as the conference ID and password to complete the conference access process. This approach not only consumes a significant amount of participants' time and energy when dealing with large-scale or frequently scheduled meetings, but is also prone to input errors due to human negligence, which in turn affects the on-time start of meetings. According to relevant surveys, in daily meetings at some large enterprises, problems with participants manually accessing the conference system cause meeting start delays an average of 2 to 3 times per week, with each delay ranging from 5 to 15 minutes. This not only wastes significant working time and reduces work efficiency, but can also delay important decisions, resulting in potential financial losses for the company. During the meeting preparation phase, users often simply distribute meeting materials such as summaries of previous meetings and the current meeting theme to participants. These meeting materials lack effective organization and analysis, forcing participants to read and understand them on their own before the meeting, making it difficult to quickly grasp the meeting's key points and key information. While some existing conferencing systems offer basic document sharing capabilities, they are nearly incapable of in-depth data processing, such as extracting key information and building knowledge connections. This lack of in-depth pre-meeting material processing makes it difficult for participants to quickly develop a comprehensive understanding of the meeting content, effectively engaging in discussions and making decisions during the meeting. This increases communication costs in the early stages of the meeting, hinders progress, and makes it difficult to achieve the desired results. For example, in meetings involving complex projects, participants often lack a thorough understanding of pre-meeting materials, leading to significant time spent revisiting project background and key issues. This prolongs the meeting and leaves relatively little time for actual solution discussion. Traditionally, meeting recorders rely on manual recording or simple recording devices. Manual recording is not only inefficient and difficult to fully capture every participant's speech, but is also prone to omissions and errors. Consequently, meeting minutes may not fully reflect the meeting and may miss important insights and decision-making information. While recording devices can capture meeting audio, subsequent manual listening and editing required for usable transcripts is time-consuming and ineffective, preventing them from providing timely and effective support for decision-making. Analysis of speech content, such as summarizing arguments and distinguishing viewpoints, relies heavily on subjective human judgment, lacking accuracy and comprehensiveness. In multilingual conferences, translation typically requires specialized translators, which is costly and lacks real-time performance. This severely limits communication between participants speaking different languages and hinders the smooth flow of information.
[0068] In the disclosed embodiments, the conference system includes a smart device 100 and a cloud server 200. The smart device 100 can be a software entity, program module, or related device developed based on artificial intelligence technology that can autonomously execute tasks, perceive the environment, and make decisions. The cloud server 200 can be a remote server built using cloud computing technology, used to store, process, and transmit conference-related data and information. The cloud server 200 is the core infrastructure for the conference system to achieve intelligent and cross-regional collaboration, and is responsible for key tasks such as connecting participants, managing meeting processes, and ensuring stable conference operations. The smart device 100 includes an information acquisition module 101, an edge device 102, a clock and network status monitoring module 103, a knowledge graph construction module 104, and a meeting recording and analysis module 105. During the meeting preparation phase, users can provide the information acquisition module 101 with meeting links and meeting materials. These materials can include summaries of previous meetings and the current meeting theme. The information acquisition module 101 can also be integrated with office systems (such as OA systems or document management platforms) to automatically acquire meeting links and meeting materials through APIs, improving the efficiency and accuracy of information collection. like Figure 2As shown, edge device 102 can be a local intelligent gateway device in the conference room. Edge device 102 can use lightweight models such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs) to perform preliminary screening and feature extraction on conference link files, reducing the computational burden on cloud server 200 and obtaining the first feature. For edge device 102, using lightweight models for preliminary screening and feature extraction can employ direct memory access (DMA). DMA efficiently transfers data from memory to the algorithm processing unit without CPU intervention, improving computational efficiency on edge device 102. Compression and quantization of convolutional neural network or recurrent neural network models can reduce model size and improve inference speed. For example, pruning can be used to remove unimportant neurons in the model, or the model's weights and activation values can be converted from floating-point numbers to low-precision integers. Quantization can reduce the model's numerical precision from 32 bits to 16 bits or even lower, thereby reducing computational effort and memory usage on edge device 102. Cloud server 200 utilizes powerful computing resources to perform in-depth analysis and classification of the first features uploaded by edge device 102, improving the accuracy and efficiency of conference link format recognition and obtaining the conference link and meeting start time. Cloud server 200 can communicate with edge device 102. Multiple first features are obtained by edge device 102, and then in-depth analysis and classification by cloud server 200 can reduce the pressure on cloud server 200. Edge device 102 performs local edge computing to process conference link files in real time, while migrating complex or large-scale computing tasks to cloud server 200. Cloud server 200 provides large-scale storage resources for storing large amounts of training data and model files. Elastic computing capabilities allow cloud server 200 to dynamically allocate computing resources based on demand. During peak model training or inference periods, computing nodes can be rapidly expanded, while resources can be released during low-load periods to reduce costs. Conference link files containing a large number of different conference links can be used to train multiplication neural networks or recurrent neural networks. The edge device 102 performs preliminary screening and feature extraction on the conference link file by using a convolutional neural network model or a recurrent neural network model after compression and quantization processing to determine a first feature, which is expressed as follows:
[0069]
[0070] in, Indicates the iconference link files, CNN and RNN respectively represent the convolutional neural network model and recurrent neural network model applied to the compressed and quantized conference data. The cloud server 200 can further analyze the first feature through the deep learning model, which is expressed as follows:
[0071]
[0072] in, represents the first feature extracted by the edge device 102, represents the weighting coefficient, Indicates the k A deep learning model, such as the GPT-4 model, is used to deeply analyze the first feature. Through edge-cloud collaborative computing, the recognition efficiency of conference links can be improved in complex network environments, speeding up conference access, reducing network bandwidth requirements, and improving system response speed and stability.
[0073] After receiving the meeting link and meeting start time from the cloud server 200, the clock and network status monitoring module 103 automatically initiates a meeting access request near the meeting start time, for example, at a first preset time before the meeting start time. The clock and network status monitoring module 103 uses a reinforcement learning algorithm, such as a Deep Q-Network (DNQ), to evaluate the quality of the network environment of the smart device 100 in real time. Based on the evaluation results, it selects the optimal network access strategy for meeting access, such as intelligently switching between office Wi-Fi and mobile 5G, to avoid meeting access failures or interruptions caused by personnel movement.
[0074] The knowledge graph construction module 104 is used to construct a first knowledge graph based on the conference materials. For example, the knowledge graph construction module 104 uses the conference materials through an online learning algorithm to optimize and fine-tune a pre-trained large model (such as GPT-4), and dynamically updates the parameters of the pre-trained large model to adapt to the language characteristics of the conference scenario. The optimized pre-trained large model is used to analyze the conference materials, extract the second feature, and combine it with the random forest algorithm to classify and predict the conference materials to determine the conference type. The random forest model is customized according to different conference types to improve classification accuracy. Finally, based on the second feature and the conference type, a knowledge graph construction tool (such as Stardog) is used to identify entities and their relationships in the conference materials to construct a first knowledge graph. This first knowledge graph can clearly present the conference theme, objectives, and key issues, helping participants to clarify the conference focus in advance and prepare materials in a targeted manner, thereby improving meeting efficiency.
[0075] Furthermore, the meeting recording and analysis module 105 can implement real-time multilingual translation using the Google Cloud Translator model, based on the Transformer architecture. The Google Cloud Translator model is trained on a massive multilingual parallel corpus and supports translation of numerous language pairs. Based on the source and target languages of the speech, the meeting recording and analysis module 105 translates the source language of speech or text into the target language in real time. It then uses the Skylark model to perform speech synthesis in the target language, and outputs the translation results in speech or text form. Meeting content can include voice recordings, video recordings, and environmental sensor data. The meeting recording and analysis module 105 is configured to optimize a pre-trained large model based on the meeting content using an online learning algorithm. For example, it converts voice recordings into text data and optimizes the pre-trained large model using an online learning algorithm based on the text data. The optimized pre-trained large model (such as the GPT-4 model) performs analysis and inference on the text data. Multimodal fusion algorithms can analyze voice recordings, video recordings, and environmental sensor data. Voice recordings can include acoustic features such as speech content, intonation, speaking rate, and background noise, which can be used for semantic recognition and sentiment analysis. Video recordings can include visual features such as visual imagery, person movement, scene structure, facial expressions, and body language for object detection and behavior recognition. Environmental sensor data can include temperature, humidity, air quality, light intensity, and noise levels for environmental state perception. Voice recordings, video recordings, and environmental sensor data can serve as multimodal data for multimodal fusion algorithms. The core goal of multimodal fusion algorithms is to overcome the limitations of a single modality (for example, when speech semantics are ambiguous, clarify the object reference by combining the video image). They also aim to align multimodal data temporally and spatially. Multimodal fusion algorithms can use self-attention to capture dependencies within a single modality (such as the association between different objects in a video) and cross-attention to model cross-modal interactions (such as the correspondence between speech keywords and video image regions). Sentiment analysis models (such as those based on BERT) can be used to determine the speaker's emotional leanings, thereby improving the accuracy and comprehensiveness of analysis results. Speech recognition, semantic understanding, and sentiment analysis technologies work together to deeply mine meeting content and generate meeting summary documents.
[0076] In the embodiment of the present application, the smart device 100 and the cloud server 200 can accurately parse the conference link. Even in a weak network or complex network environment, they can accurately access the meeting based on the meeting start time, avoiding the tedious steps of manually entering the link and repeated debugging. Before the meeting, the smart device 100 uses natural language processing and machine learning technology to deeply interpret historical meeting minutes and current agenda documents, extract core points from massive information, and construct a first knowledge graph. The first knowledge graph can not only assist participants in quickly grasping the background, but also provide solid data support for meeting discussions and decision-making. During the meeting, the smart device 100 uses advanced speech recognition, semantic understanding and sentiment analysis technology. It can not only transcribe the speech content in real time, but also intelligently refine the core ideas and accurately distinguish between pro and con positions. At the same time, the multi-language real-time translation function breaks down language barriers. After the meeting, the smart device 100 can generate a clear and focused meeting summary document in a short period of time. From automatic access, pre-meeting preparations to post-meeting summaries, Smart Device 100 runs through the entire meeting process, significantly improving meeting efficiency and quality with intelligent and automated management models, providing enterprises and organizations with efficient and convenient digital meeting solutions, helping to reduce operating costs and promote more scientific and efficient decision-making.
[0077] In another embodiment of the present disclosure, the clock and network status monitoring module is used to select the optimal network access strategy for accessing a conference using a deep Q network as follows:
[0078]
[0079] in, Indicates the network status of the smart device Next, select Network Access Policy , expected return, Indicates reward, represents the discount factor, Indicates the network access strategy selected in the next time step , the maximum expected return.
[0080] In the embodiment of the present disclosure, the formula for optimizing the network access strategy of the Deep Q-Network (DNQ) is expressed as follows:
[0081]
[0082] in, Indicates that the status Take action The value of, at time step t, the network state of the smart device 100 Next, select Access Strategy , expected rewards (including immediate rewards and discounted future rewards), Indicates reward, represents the discount factor, Indicates the access strategy selected in the next time step, such as Wi-Fi, mobile 5G, Indicates the access strategy selected for the next time step , the expected maximum reward. For example, the current network where the smart device 100 is located is a Wi-Fi network with poor network quality. The smart device 100 uses a deep Q network to update the Q value at each time step to gradually select the optimal network access strategy and access the meeting.
[0083] In another embodiment of the present disclosure, the clock and network status monitoring module 103 is also used to send shared network status information and a conference access method request to other smart devices 100 when the network environment quality indicates that the conference cannot be accessed, and access the conference based on the shared network status information and conference access method received from the other smart devices 100.
[0084] In the disclosed embodiments, in complex network environments, through innovative distributed collaborative intelligent device architectures, multiple intelligent devices 100 collaborate to form a network. If a smart device 100 in a conference room encounters access difficulties, for example, if the network quality indicates that it cannot access a meeting, the clock and network status monitoring module 103 can seek help from other intelligent devices 100, sending shared network status information and a request for conference access to the other intelligent devices 100. The module then accesses the meeting based on the shared network status information and conference access methods received from the other intelligent devices 100. Shared network information can include basic network metrics such as Wi-Fi signal strength, bandwidth, and latency, as well as real-time load status. Conference access methods can include Wi-Fi and 5G networks. This enables collaborative access, improving access success rate and stability. Regarding knowledge processing, if a smart device 100 encounters difficulty understanding specific terminology, it can also send a request for shared knowledge to other intelligent devices 100 and analyze the specific terminology based on the shared knowledge received from the other intelligent devices 100. This improves the overall system's ability to handle complex problems and its adaptability.
[0085] In another embodiment of the present disclosure, the knowledge graph construction module 104 is used to optimize the pre-trained large model using an online learning algorithm based on conference materials;
[0086] Analyze the meeting data based on the optimized pre-trained large model and extract the second feature;
[0087] Based on the meeting data, the random forest algorithm is used to predict the meeting type and determine the meeting type;
[0088] According to the second feature and the conference type, a knowledge graph construction tool is used to determine the corresponding relationships between entities in the conference materials and construct the first knowledge graph.
[0089] In the disclosed embodiment, a pre-trained large model is optimized based on conference materials using an online learning algorithm to construct a first knowledge graph. The knowledge graph construction module 104 is configured to optimize the pre-trained large model using an online learning algorithm based on the conference materials. The pre-trained large model can be a GPT-4 model. First, GPT-4 is optimized and fine-tuned using a large amount of conference materials relevant to the conference scenario. During the optimization and fine-tuning process, an innovative dynamic adaptive learning mechanism is employed to optimize the parameters of the pre-trained large model in real time using an online learning algorithm. For example, when the information acquisition module 101 acquires new conference materials, the knowledge graph construction module 104, using a gradient descent-based online learning algorithm, updates the word vector representation and semantic understanding parameters of the pre-trained large model in real time based on new terms and industry-specific vocabulary in the conference materials, enabling the pre-trained large model to better adapt to the linguistic characteristics of the conference scenario. The optimized pre-trained large model is used to analyze the conference materials and extract second features. These second features can include basic information (such as time, location, participants, and main topics), deep semantic features, association features, or structured information. A random forest algorithm is then used to classify and predict the conference materials to determine the conference type. For example, a random forest model is trained using text features of meeting materials (such as keyword frequency and sentence length) as input to predict the materials' topic categories and importance levels. Through personalized learning, random forest models are constructed for different types of meetings (such as daily corporate office meetings, cross-regional project collaboration meetings, and academic seminars and exchanges), improving classification accuracy and relevance. Based on the second feature and the meeting type, a knowledge graph construction tool (such as Stardog) is used to construct a primary knowledge graph. This knowledge graph construction tool identifies entities (such as people, organizations, and events) and the corresponding relationships between entities (such as causal and subordinate relationships) in meeting materials, converting them into nodes and edges in the knowledge graph to construct the primary knowledge graph. This primary knowledge graph helps organizers and participants clearly understand the meeting's theme and objectives. By presenting the meeting's core topics, related subtopics, and expected goals in the primary knowledge graph, participants gain a more intuitive understanding of the meeting's focus. For example, in a meeting about new product development, the first knowledge graph can clearly display key elements such as the product's functional requirements, market positioning, and technical difficulties, so that participants can clearly understand the core issues to be addressed before the meeting begins, and thus prepare relevant materials and opinions more specifically, thereby improving the efficiency of the meeting.
[0090] In yet another embodiment of the present disclosure, the conference content includes: voice recordings, video recordings, and environmental sensor data;
[0091] The conference recording and analysis module 105 is used to optimize the translation model parameters in real time according to the content of the speakers' speeches during the conference;
[0092] Based on the speaker's source language and target language, and the analysis results of the multimodal fusion algorithm on voice recordings, video recordings and environmental sensor data, a translation model with optimized parameters is used to translate the speaker's speech in real time.
[0093] In the disclosed embodiment, the conference record and analysis module 105, the translation model can be a Google Cloud Translation model. The Google Cloud Translation model can be used to achieve real-time multilingual translation. During the translation process, combined with the innovation of the dynamic adaptive learning mechanism, the translation model parameters are optimized in real time according to the new terms and specific industry vocabulary that continue to appear in the speaker's speech content during the meeting, thereby improving the accuracy and professionalism of the translation. The multimodal fusion algorithm analyzes the results of voice recordings, video recordings and environmental sensor data, taking into account factors such as the speaker's emotions and tone, making the translation results more natural and accurate. According to the speaker's source language and target language, and the multimodal fusion algorithm's analysis results of voice recordings, video recordings and environmental sensor data, the translation model with optimized parameters is used to translate the speaker's speech content in real time. The deep interactive innovation of the multimodal fusion algorithm improves the accuracy and robustness of speaker annotation and multilingual translation. The dynamic adaptive learning mechanism ensures that the translation model can be continuously optimized according to the real-time situation of the meeting to improve the quality of translation.
[0094] In yet another embodiment of the present disclosure, the conference content includes: voice recordings, video recordings, and environmental sensor data;
[0095] The conference recording and analysis module 105 is used to convert the speech into text data in real time using a speech recognition model based on the speech recording;
[0096] Based on text data, online learning algorithms are used to optimize the pre-trained large model;
[0097] Generate meeting record documents based on the analysis results of text data by the optimized pre-trained large model and the analysis results of voice recordings, video recordings and environmental sensor data by the multimodal fusion algorithm;
[0098] Generate a meeting summary document based on the meeting minutes document.
[0099] In the disclosed embodiments, the meeting recording and analysis module 105 is configured to convert speech recordings into text data in real time using a speech recognition model. The speech recognition model can be the DeepSpeech2 model, which performs real-time speech recognition. This speech recognition model is trained on a large amount of speech data from conference scenes and adapts to the complex acoustic environment of conference rooms. A multimodal fusion algorithm can be used to analyze the results of speech recordings, video recordings, and environmental sensor data to improve speech recognition accuracy. For example, by analyzing the lip movements and facial expressions of participants, the speech recognition model can be assisted in correcting misidentified content. In some embodiments, a speaker recognition model can also be used to annotate the text data with the speaker's identity. While recognizing speech, the speaker's identity is annotated for each speech segment in the text data based on the output of the speaker recognition model (e.g., an i-vector model). The accuracy and robustness of speaker annotation can be further improved by using a multimodal fusion algorithm to analyze the results of speech recordings, video recordings, and environmental sensor data, such as video image features (e.g., facial recognition and body movement analysis). The pre-trained large model can be a GPT-4 model, which is optimized using an online learning algorithm based on the text data. Through innovative dynamic adaptive learning mechanisms, the pre-trained large model continuously optimizes semantic understanding parameters based on real-time meeting content, improving its ability to understand complex semantics. The optimized pre-trained large model analyzes text data from real-time speech recognition to generate analysis results. Using a multimodal fusion algorithm to analyze voice recordings, video recordings, and environmental sensor data, the model identifies the meaning, theme, and sentiment of sentences and generates meeting minutes. Meeting minutes are then summarized and refined to produce a meeting summary document. Combining video and environmental sensor data to assist speech recognition significantly improves speech recognition accuracy, enabling better adaptation to complex meeting room acoustics and scenarios with multiple people speaking simultaneously, enhancing the accuracy of meeting summary documents.
[0100] In another embodiment of the present disclosure, the knowledge graph construction module 104 is further configured to analyze voice recordings, video recordings, and environmental sensor data using a multimodal fusion algorithm during the meeting;
[0101] Based on the meeting content, online learning algorithms are used to optimize the pre-trained large model;
[0102] Analyze the meeting content based on the analysis results of voice recordings, video recordings, and environmental sensor data using the optimized pre-trained large model and multimodal fusion algorithm; and update the first knowledge graph based on the analysis results of the meeting content to generate a second knowledge graph;
[0103] The meeting record and analysis module 105 is also used to generate a preliminary meeting summary report based on the meeting record document using the T5 model;
[0104] Submit the preliminary report of the meeting to the chairman of the meeting for decision;
[0105] Based on the preliminary summary report of the meeting after the ruling, the meeting minutes document and the second knowledge graph, a natural language generation model is used to generate a meeting summary document.
[0106] In the disclosed embodiment, the knowledge graph construction module 104 is also used to analyze voice recordings, video recordings, and environmental sensor data using a multimodal fusion algorithm during the meeting. For example, the multimodal fusion algorithm can utilize the deep interactive innovation of multimodal fusion to fuse multimodal information such as voice recordings, video recordings, and environmental sensor data to further enrich the content of the first knowledge graph. For example, by analyzing the facial expressions and movements of characters in video recordings, and the intonation and speaking speed in voice recordings, emotional and behavioral features are added to the character nodes in the first knowledge graph. Environmental sensor data may include temperature, humidity, air quality, light intensity, and noise level. Environmental sensor data directly affects the comfort and attention of participants, and thus affects the efficiency of the meeting. The pre-trained large model (such as GPT-4) is optimized and fine-tuned according to the content of the meeting during the meeting. Through the online learning algorithm, the parameters of the pre-trained large model are updated in real time according to the content of the meeting during the meeting (such as the transcription of new meeting recordings, professional terms for specific industries, etc.). Its formula is expressed as follows:
[0107]
[0108] in represents the learning rate, represents the gradient of the loss function, The loss of the representation model allows the pre-trained large model to continuously adapt to new terms, scenarios, and topics in the meeting.
[0109] For example, when new technical terms appear in a meeting, the pre-trained large model can adjust the word vector representation and semantic understanding parameters through a dynamic adaptive learning mechanism to better understand the specific meaning of these terms in the meeting. According to the analysis results of the voice recordings, video recordings and environmental sensor data by the optimized pre-trained large model and the multimodal fusion algorithm, the content of the meeting is analyzed, and the first knowledge graph is updated in real time based on the analysis results of the meeting content, such as updating the key information and decision results in the meeting. For example, key information such as project budget figures and product technical parameters identified in the voice recordings, as well as the voting results of participants on a certain plan observed in the video, are included. Then a second knowledge graph is generated to ensure that the second knowledge graph is consistent with the actual situation of the meeting. The formula is expressed as follows:
[0110]
[0111] in, Indicates the weight of each entity and relationship, reflecting its importance; and They represent the real-time data of entities and the relationships between entities in the meeting content respectively.
[0112] The secondary knowledge graph more accurately and comprehensively reflects the content and structure of the meeting. For example, by analyzing the relationships between different topics in the meeting, the connections between topic nodes in the knowledge graph are optimized. The real-time updated secondary knowledge graph can instantly reflect the discussion content, decisions, and action items in the meeting. For example, when a new idea or decision is raised in a meeting, the knowledge graph can immediately add relevant nodes and relationships, ensuring that all participants are kept up to date with the latest information. This immediacy reduces the time cost of post-meeting review and updates, avoiding misunderstandings or repeated discussions caused by information lags. The secondary knowledge graph can provide stronger support for subsequent meeting summaries, decision support, and knowledge sharing. For example, the secondary knowledge graph can quickly understand the core content and key points of the meeting, providing a basis for developing subsequent action plans. It can also be used as a knowledge sharing platform, allowing team members to easily access key information and decision results from the meeting.
[0113] The meeting record and analysis module 105 is also used to use a text summary generation tool based on the T5 model to preliminarily organize the meeting record documents. The T5 model can be an encoder-decoder architecture based on the attention mechanism to compress and refine long texts, extract key information and core ideas, and obtain a preliminary meeting report. The preliminary meeting report is submitted to the meeting chairman for adjudication in the form of a structured report. Based on the adjudicated preliminary meeting report, meeting record documents and the second knowledge graph, a natural language generation model is used to generate a detailed meeting summary document. The meeting summary document may include the meeting theme, participants, meeting agenda, main discussion content, consensus reached, unresolved issues and next action plan, etc. Through innovation of dynamic adaptive learning mechanism, the parameters of the natural language generation model can be continuously optimized according to the actual situation of the meeting and the adjudication results of the meeting chairman on the preliminary meeting report, so that the language expression of the meeting summary document is more natural and fluent, and the content is more accurate and comprehensive.
[0114] In another embodiment of the present disclosure, the clock and network status monitoring module 103 is further configured to send a meeting notification to the meeting participants after receiving the meeting link and meeting start time sent by the cloud server, the meeting notification including: the meeting link, the meeting start time and the first knowledge graph;
[0115] The knowledge graph construction module 104 is further configured to introduce the meeting content using a pre-trained speech synthesis model based on the first knowledge graph;
[0116] The meeting record and analysis module 105 is further configured to generate an AR plan based on the meeting summary document and the meeting room environment features identified by computer vision.
[0117] In the disclosed embodiment, the clock and network status monitoring module 103 is also used to send a meeting notification to the meeting participants after receiving the meeting link and meeting start time sent by the cloud server. The list of meeting participants can be obtained from the cloud server or specified by the meeting chair. Notification channels can include: email, mobile phone number, etc. The meeting notification includes: meeting link, meeting start time and the first knowledge graph. The first knowledge graph is constructed by extracting key information based on meeting materials such as the summary of the previous meeting and the theme of this meeting. The first knowledge graph can help meeting participants grasp the background of the meeting in a short time, reduce understanding costs, and improve discussion efficiency.
[0118] The knowledge graph construction module 104 uses a pre-trained speech synthesis model (e.g., the Skylark model) based on the first knowledge graph to introduce the meeting content and clarify the meeting theme. The speech synthesis model is trained on a large amount of speech data and can simulate human voice intonation and emotional expression, making the opening more lively and professional.
[0119] The meeting recording and analysis module 105 can utilize the Vuforia augmented reality engine to convert the tasks, responsible personnel, time points, and other information determined during the meeting discussion in the meeting summary document into an AR plan. The AR plan can use computer vision technology to identify the characteristics of the meeting room environment and overlay task-related information in an intuitive and visual form (such as virtual labels, task progress bars, etc.) in the field of view of the corresponding person in charge (via AR glasses or mobile AR applications). For example, the task progress bar information display is represented as follows:
[0120]
[0121] in, Indicates the progress percentage of the task. Indicates the scheduled time of the task. Indicates the person responsible for the task. Each task The representation in the virtual space can be represented by the position vector in the virtual space ∈ and virtual information Describe it as:
[0122]
[0123] in, Indicates the current timestamp, and updates the task progress in real time. For example, the meeting record and analysis module 105, according to In the conference room, task information is overlaid in real time in the form of progress bars in the person in charge's field of view. During the AR plan generation process, the results of the analysis of voice recordings, video recordings, and environmental sensor data based on the meeting summary document and the multimodal fusion algorithm can be used to make the AR plan more consistent with the actual meeting situation, improving the accuracy and efficiency of task execution. The formula is expressed as:
[0124]
[0125] in, represents the output of the speech recording after sentiment analysis, Represents the output of the video record after facial recognition, Represents environmental sensor data, weight It can be adjusted dynamically according to the importance of the task and the real-time situation, for example, The weight of is set to a larger value when discussing emotion-related tasks. Based on the generated AR plan, the meeting record and analysis module 105 can track the movements and locations of the responsible personnel. For example, if a responsible personnel approaches the task target, the meeting record and analysis module 105 can automatically push relevant task progress or reminders. This interactive design allows the AR plan to not only be displayed statically, but also be dynamically updated in real time based on task progress, ensuring the accuracy and efficiency of task execution.
[0126] Figure 3 A schematic diagram of the conference system to realize its functions, such as Figure 3 Shown, including:
[0127] S301: Pre-meeting preparation stage: edge-cloud collaborative computing access meeting; sharing network status information between multiple smart devices; building a first knowledge graph; sending meeting notifications;
[0128] S302: During the meeting: introduce the meeting content; perform real-time translation based on the speaker's source language and target language using a multimodal fusion algorithm; update the first knowledge graph using the optimized pre-trained large model and multimodal fusion algorithm to generate a second knowledge graph;
[0129] S303: After the meeting: Use the optimized pre-trained large model and multimodal fusion algorithm to analyze the meeting content and generate a meeting summary document; generate an AR plan.
[0130] During the pre-meeting preparation phase, edge-cloud collaborative computing improves the recognition efficiency of meeting links, enhancing system responsiveness and stability. Network status information and meeting access methods are shared across multiple smart devices, improving access efficiency. A primary knowledge graph is constructed to help attendees identify meeting focus in advance, improving meeting efficiency. During the meeting, real-time translation is performed based on the speaker's source and target languages, combining a multimodal fusion algorithm to enhance translation accuracy. An optimized pre-trained large model and multimodal fusion algorithm are used to update the primary knowledge graph and generate a secondary knowledge graph. This improves the accuracy of the secondary knowledge graph, instantly reflecting discussion content, decisions, and action items, and enhancing meeting efficiency. After the meeting, an optimized pre-trained large model and multimodal fusion algorithm are used to analyze the meeting content and generate a meeting summary document, improving its accuracy. This achieves intelligent management of the entire meeting process, from preparation to post-meeting processing, enhancing the user experience, improving overall meeting efficiency and quality, and providing businesses and organizations with a more convenient and efficient meeting solution.
[0131] Based on the same disclosed concept, the embodiment of the present disclosure also provides an intelligent device. Since the principle of solving the problem by the intelligent device is similar to that of the aforementioned conference system, the implementation of the intelligent device can refer to the implementation of the aforementioned conference system, and the repeated parts will not be repeated.
[0132] The present disclosure provides a smart device 100, such as Figure 1 As shown, it includes: information acquisition module 101, edge device 102, clock and network status monitoring module 103, knowledge graph construction module 104 and meeting recording and analysis module 105;
[0133] The information acquisition module 101 is used to obtain conference link files and conference materials;
[0134] The edge device 102 is configured to perform preliminary screening and feature extraction on the conference link file using a convolutional neural network model or a recurrent neural network model that has undergone compression and quantization processing to determine a first feature; and transmit the first feature to the cloud server 200; the cloud server 200 is configured to dynamically allocate computing resources to analyze the first feature after receiving the first feature transmitted by the edge device 102, and determine the conference link and the meeting start time;
[0135] The clock and network status monitoring module 103 is configured to, after receiving the conference link and conference start time sent by the cloud server 200, evaluate the network environment quality of the smart device 100 within a first preset time before the conference start time, and select the optimal network access strategy using a deep Q network to access the conference;
[0136] The knowledge graph construction module 104 is used to construct a first knowledge graph based on the conference materials;
[0137] The conference recording and analysis module 105 is used to perform real-time translation according to the speaker's source language and target language during the conference; and generate a conference summary document based on the conference content.
[0138] In another embodiment of the present disclosure, the clock and network status monitoring module is used to select the optimal network access strategy for accessing a conference using a deep Q network. The formula is expressed as follows:
[0139]
[0140] in, Indicates the network status of the smart device Next, select Network Access Policy , expected return, Indicates reward, represents the discount factor, Indicates the network access strategy selected in the next time step , the maximum expected return.
[0141] In another embodiment of the present disclosure, the clock and network status monitoring module 103 is also used to send shared network status information and a request for accessing the conference to other smart devices 100 when the network environment quality indicates that the conference cannot be accessed, and to access the conference based on the shared network status information and access conference method received from the other smart devices 100.
[0142] In another embodiment of the present disclosure, the knowledge graph construction module 104 is used to optimize the pre-trained large model using an online learning algorithm based on the conference materials;
[0143] Analyze the conference data according to the optimized pre-trained large model to extract the second feature;
[0144] Based on the meeting information, a random forest algorithm is used to predict the meeting type and determine the meeting type;
[0145] According to the second feature and the conference type, a knowledge graph construction tool is used to determine the entities in the conference materials and the corresponding relationships between the entities to construct a first knowledge graph.
[0146] In yet another embodiment of the present disclosure, the conference content includes: voice recordings, video recordings, and environmental sensor data;
[0147] The conference recording and analysis module 105 is used to optimize the translation model parameters in real time according to the speaker's speech content during the conference;
[0148] Based on the speaker's source language and target language, and the analysis results of the multimodal fusion algorithm on voice recordings, video recordings and environmental sensor data, a translation model with optimized parameters is used to translate the speaker's speech in real time.
[0149] In yet another embodiment of the present disclosure, the conference content includes: voice recordings, video recordings, and environmental sensor data;
[0150] The conference recording and analysis module 105 is used to convert the speech into text data in real time using a speech recognition model based on the speech recording;
[0151] Based on the text data, an online learning algorithm is used to optimize the pre-trained large model;
[0152] Generate meeting record documents based on the analysis results of text data by the optimized pre-trained large model and the analysis results of voice recordings, video recordings and environmental sensor data by the multimodal fusion algorithm;
[0153] Generate a meeting summary document based on the meeting record document.
[0154] In another embodiment of the present disclosure, the knowledge graph construction module 104 is further configured to analyze voice recordings, video recordings, and environmental sensor data using a multimodal fusion algorithm during the meeting;
[0155] Based on the meeting content, an online learning algorithm is used to optimize the pre-trained large model;
[0156] Analyze the meeting content based on the analysis results of the voice recordings, video recordings, and environmental sensor data using the optimized pre-trained large model and the multimodal fusion algorithm; and update the first knowledge graph based on the analysis results of the meeting content to generate a second knowledge graph;
[0157] The meeting record and analysis module 105 is further configured to generate a preliminary meeting summary report based on the meeting record document using the T5 model;
[0158] Submit the preliminary report of the meeting to the chairman of the meeting for decision;
[0159] Based on the preliminary summary report of the meeting after the ruling, the meeting record document and the second knowledge graph, a natural language generation model is used to generate a meeting summary document.
[0160] In another embodiment of the present disclosure, the clock and network status monitoring module 103 is further configured to send a meeting notification to the meeting participants after receiving the meeting link and meeting start time sent by the cloud server, wherein the meeting notification includes: the meeting link, the meeting start time, and the first knowledge graph;
[0161] The knowledge graph construction module 104 is further configured to introduce the meeting content using a pre-trained speech synthesis model based on the first knowledge graph;
[0162] The meeting recording and analysis module 105 is further configured to generate an AR plan based on the meeting summary document and the meeting room environment features identified by computer vision.
[0163] Based on the same disclosed concept, the embodiment of the present disclosure also provides a conference processing method. Since the principle of the problem solved by the conference processing method is similar to that of the aforementioned conference system, the implementation of the conference processing method can refer to the implementation of the aforementioned conference system, and the repeated parts will not be repeated.
[0164] The present disclosure provides a conference processing method, such as Figure 4 As shown, the following steps are included:
[0165] S401. Obtain conference link files and conference materials;
[0166] S402: Using a compressed and quantized convolutional neural network model or a recurrent neural network model to perform preliminary screening and feature extraction on the conference link file to determine a first feature;
[0167] S403: Dynamically allocate computing resources to analyze the first feature and determine the conference link and conference start time;
[0168] S404: Evaluate the network environment quality within a first preset time before the meeting start time based on the meeting link and the meeting start time, and use the Deep Q network to select the optimal network access strategy to access the meeting;
[0169] S405. Construct a first knowledge graph based on the conference materials;
[0170] S406. During the meeting, real-time translation is performed based on the speaker's source language and target language; and a meeting summary document is generated based on the meeting content.
[0171] In another embodiment of the present disclosure, the formula for selecting the optimal network access strategy for accessing a conference using a deep Q network is expressed as:
[0172]
[0173] in, Indicates the network status of the smart device Next, select Network Access Policy , expected return, Indicates reward, represents the discount factor, Indicates the network access strategy selected in the next time step , the maximum expected return.
[0174] In another embodiment of the present disclosure, the method further includes: when the network environment quality indicates that the conference cannot be accessed, sending shared network status information and a request for accessing the conference to other smart devices, and accessing the conference based on the shared network status information and accessing the conference received from the other smart devices.
[0175] In another embodiment of the present disclosure, constructing a first knowledge graph based on conference materials includes:
[0176] Based on the conference materials, online learning algorithms were used to optimize the pre-trained large model;
[0177] Analyze the meeting data based on the optimized pre-trained large model and extract the second feature;
[0178] Based on the meeting data, the random forest algorithm is used to predict the meeting type and determine the meeting type;
[0179] According to the second feature and the conference type, a knowledge graph construction tool is used to determine the corresponding relationships between entities in the conference materials and construct the first knowledge graph.
[0180] In yet another embodiment of the present disclosure, the conference content includes: voice recordings, video recordings, and environmental sensor data;
[0181] Real-time translation during meetings based on the speaker's source and target languages, including:
[0182] Optimize translation model parameters in real time based on the speaker's speech content during the meeting;
[0183] Based on the speaker's source language and target language, and the analysis results of the multimodal fusion algorithm on voice recordings, video recordings and environmental sensor data, a translation model with optimized parameters is used to translate the speaker's speech in real time.
[0184] In yet another embodiment of the present disclosure, the conference content includes: voice recordings, video recordings, and environmental sensor data;
[0185] Generate a meeting summary document based on the meeting content, including:
[0186] Use speech recognition models based on voice recordings to convert speech into text data in real time;
[0187] Based on text data, online learning algorithms are used to optimize the pre-trained large model;
[0188] Generate meeting record documents based on the analysis results of text data by the optimized pre-trained large model and the analysis results of voice recordings, video recordings and environmental sensor data by the multimodal fusion algorithm;
[0189] Generate a meeting summary document based on the meeting minutes document.
[0190] In another embodiment of the present disclosure, the method further includes:
[0191] During the meeting, multimodal fusion algorithms were used to analyze voice recordings, video recordings, and environmental sensor data;
[0192] Based on the meeting content, online learning algorithms are used to optimize the pre-trained large model;
[0193] Analyze the meeting content based on the analysis results of voice recordings, video recordings, and environmental sensor data using the optimized pre-trained large model and multimodal fusion algorithm; and update the first knowledge graph based on the analysis results of the meeting content to generate a second knowledge graph;
[0194] Generate a meeting summary document based on the meeting minutes, including:
[0195] Based on the meeting record documents, use the T5 model to generate a preliminary meeting report;
[0196] Submit the preliminary report of the meeting to the chairman of the meeting for decision;
[0197] Based on the preliminary summary report of the meeting after the ruling, the meeting minutes document and the second knowledge graph, a natural language generation model is used to generate a meeting summary document.
[0198] In another embodiment of the present disclosure, the method further includes: sending a meeting notice to the meeting participants, the meeting notice including: a meeting link, a meeting start time, and the first knowledge graph;
[0199] Based on the first knowledge graph, a pre-trained speech synthesis model is used to introduce the meeting content;
[0200] Generate an AR plan based on the meeting summary document and the meeting room environment features identified by computer vision.
[0201] Through the above description of the embodiments, those skilled in the art will clearly understand that the embodiments of the present disclosure can be implemented through hardware or through software plus the necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in the various embodiments of the present disclosure.
[0202] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes in the accompanying drawings are not necessarily required for implementing the present disclosure.
[0203] Those skilled in the art will appreciate that the modules in the devices of the embodiments may be distributed in the devices of the embodiments as described in the embodiments, or may be located in one or more devices different from the embodiments with corresponding changes. The modules of the above embodiments may be combined into one module or further split into multiple submodules.
[0204] The serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.
[0205] Obviously, those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A conference system, characterized in that: include: Smart devices and cloud servers, the smart devices including: an information acquisition module, an edge device, a clock and network status monitoring module, a knowledge graph construction module, and a meeting recording and analysis module; The information acquisition module is used to obtain conference link files and conference materials; The edge device is configured to perform preliminary screening and feature extraction on the conference link file using a convolutional neural network model or a recurrent neural network model that has been compressed and quantized, to determine a first feature; The cloud server is configured to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device, and determine the conference link and conference start time; The clock and network status monitoring module is used to evaluate the network environment quality of the smart device within a first preset time before the meeting start time after receiving the meeting link and meeting start time sent by the cloud server, and use the deep Q network to select the optimal network access strategy to access the meeting; The knowledge graph construction module is used to construct a first knowledge graph based on the conference materials; The knowledge graph construction module is also used to analyze voice recordings, video recordings, and environmental sensor data using a multimodal fusion algorithm during the meeting; An online learning algorithm is used to optimize the pre-trained large model based on the meeting content, which includes voice recordings, video recordings, and environmental sensor data. Analyze the meeting content based on the analysis results of the multimodal fusion algorithm on the voice recordings, video recordings, and environmental sensor data and the optimized pre-trained large model; and update the first knowledge graph based on the analysis results of the meeting content to generate a second knowledge graph; The meeting recording and analysis module is used to perform real-time translation according to the speaker's source language and target language during the meeting; and generate a meeting summary document based on the meeting content; The meeting record and analysis module is also used to generate a preliminary meeting summary report based on the meeting record document using the T5 model; Submit the preliminary report of the meeting to the chairman of the meeting for decision; Based on the preliminary summary report of the meeting after the ruling, the meeting record document and the second knowledge graph, a natural language generation model is used to generate a meeting summary document.
2. The system according to claim 1, wherein The clock and network status monitoring module is used to select the optimal network access strategy to access the conference using the deep Q network. The formula is expressed as: in, Indicates the network status of the smart device Next, select Network Access Policy , expected return, Indicates reward, represents the discount factor, Indicates the network access strategy selected in the next time step , the maximum expected return.
3. The system according to claim 1, wherein The clock and network status monitoring module is also used to send shared network status information and conference access method requests to other smart devices when the network environment quality indicates that the meeting cannot be accessed, and access the meeting based on the shared network status information and conference access method received from other smart devices.
4. The system according to claim 1, wherein: The knowledge graph construction module is used to optimize the pre-trained large model using an online learning algorithm based on the conference materials; Analyze the conference data according to the optimized pre-trained large model to extract the second feature; Based on the meeting information, a random forest algorithm is used to predict the meeting type and determine the meeting type; According to the second feature and the conference type, a knowledge graph construction tool is used to determine the entities in the conference materials and the corresponding relationships between the entities to construct a first knowledge graph.
5. The system according to claim 1, wherein: The conference recording and analysis module is used to optimize the translation model parameters in real time according to the speaker's speech content during the conference; Based on the speaker's source language and target language, and the analysis results of the multimodal fusion algorithm on voice recordings, video recordings and environmental sensor data, a translation model with optimized parameters is used to translate the speaker's speech in real time.
6. The system according to claim 1, wherein: The conference recording and analysis module is used to convert the voice into text data in real time using a voice recognition model based on the voice recording; Based on the text data, an online learning algorithm is used to optimize the pre-trained large model; Meeting record documents are generated based on the analysis results of text data by the optimized pre-trained large model and the analysis results of voice recordings, video recordings and environmental sensor data by the multimodal fusion algorithm.
7. The system according to claim 1, wherein: The clock and network status monitoring module is further configured to send a meeting notification to the meeting participants after receiving the meeting link and meeting start time sent by the cloud server, wherein the meeting notification includes: the meeting link, the meeting start time and the first knowledge graph; The knowledge graph construction module is further configured to introduce the content of the meeting using a pre-trained speech synthesis model based on the first knowledge graph; The meeting recording and analysis module is further used to generate an AR plan based on the meeting summary document and the meeting room environment features identified by computer vision.
8. A smart device, characterized in that: include: Information acquisition module, edge device, clock and network status monitoring module, knowledge graph construction module and meeting recording and analysis module; The information acquisition module is used to obtain conference link files and conference materials; The edge device is configured to perform preliminary screening and feature extraction on the conference link file using a convolutional neural network model or a recurrent neural network model that has been compressed and quantized, to determine a first feature; and sending the first feature to a cloud server; The cloud server is configured to dynamically allocate computing resources to analyze the first feature after receiving the first feature sent by the edge device, and determine the conference link and conference start time; The clock and network status monitoring module is used to evaluate the network environment quality of the smart device within a first preset time before the meeting start time after receiving the meeting link and meeting start time sent by the cloud server, and use the deep Q network to select the optimal network access strategy to access the meeting; The knowledge graph construction module is used to construct a first knowledge graph based on the conference materials; The knowledge graph construction module is also used to analyze voice recordings, video recordings, and environmental sensor data using a multimodal fusion algorithm during the meeting; An online learning algorithm is used to optimize the pre-trained large model based on the meeting content, which includes voice recordings, video recordings, and environmental sensor data. Analyze the meeting content based on the analysis results of the multimodal fusion algorithm on the voice recordings, video recordings, and environmental sensor data and the optimized pre-trained large model; and update the first knowledge graph based on the analysis results of the meeting content to generate a second knowledge graph; The meeting recording and analysis module is used to perform real-time translation according to the speaker's source language and target language during the meeting; and generate a meeting summary document based on the meeting content; The meeting record and analysis module is also used to generate a preliminary meeting summary report based on the meeting record document using the T5 model; Submit the preliminary report of the meeting to the chairman of the meeting for decision; Based on the preliminary summary report of the meeting after the ruling, the meeting record document and the second knowledge graph, a natural language generation model is used to generate a meeting summary document.
9. A conference processing method, characterized in that: include: Obtain conference link files and conference materials; Performing preliminary screening and feature extraction on the conference link file using a convolutional neural network model or a recurrent neural network model after compression and quantization processing by an edge device to determine a first feature; Dynamically allocate computing resources using a cloud server to analyze the first feature and determine the meeting link and meeting start time; According to the conference link and the conference start time, the network environment quality is evaluated within a first preset time before the conference start time, and the optimal network access strategy is selected using the deep Q network to access the conference; Constructing a first knowledge graph based on the conference materials; During the meeting, multimodal fusion algorithms were used to analyze voice recordings, video recordings, and environmental sensor data; An online learning algorithm is used to optimize the pre-trained large model based on the meeting content, which includes voice recordings, video recordings, and environmental sensor data. Analyze the meeting content based on the analysis results of the multimodal fusion algorithm on the voice recordings, video recordings, and environmental sensor data and the optimized pre-trained large model; and update the first knowledge graph based on the analysis results of the meeting content to generate a second knowledge graph; Provide real-time translation between the speaker's source and target languages during the meeting; and generate a meeting summary document based on the meeting content; The meeting summary document is generated based on the meeting content, including: Based on the meeting record documents, use the T5 model to generate a preliminary meeting report; Submit the preliminary report of the meeting to the chairman of the meeting for decision; Based on the preliminary summary report of the meeting after the ruling, the meeting record document and the second knowledge graph, a natural language generation model is used to generate a meeting summary document.
Citation Information
Patent Citations
Method for quickly joining conference through applet card
CN110601863A
Heterogeneous wireless network access selection method and system based on SDN
CN111586809A