Processing system for intelligently analyzing and recording video conference content

By leveraging deep learning networks and natural language processing technologies, this system addresses the problem of video conferencing systems failing to accurately capture key discussion and decision-making content. It enables intelligent analysis and recording of video and audio data, generating detailed meeting minutes, improving analysis efficiency, and ensuring data security.

CN121887947AInactive Publication Date: 2026-04-17WUHAN JIXUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-04-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing video conferencing systems cannot accurately capture key discussions and decisions during meetings, lack intelligent analysis capabilities, cannot analyze video and audio content in real time, and cannot effectively identify important information and potential risks.

Method used

It employs deep learning networks and natural language processing methods to intelligently analyze video and audio data, including a video conferencing module, a network service module, a video conferencing monitoring module, a data storage module, and an information display module. It acquires data through video cameras and audio acquisition devices, and performs data transformation, intelligent analysis, encrypted storage, and visualization.

Benefits of technology

It enables intelligent analysis of video and audio data, accurately captures visual and auditory features in meetings, generates detailed meeting minutes, improves the efficiency of the analysis and recording process, ensures data security and reliability, and supports security guarantees for commercial use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887947A_ABST
    Figure CN121887947A_ABST
Patent Text Reader

Abstract

The invention provides a processing system for intelligently analyzing and recording video conference content. Relates to the technical field of video conferences, and comprises a video conference module used for acquiring video data during a conference through a video camera and acquiring audio data during the conference through an audio acquisition device, and a network service module used for performing data conversion on the video data and the audio data, and obtaining standardized data through the data conversion. According to the processing system for intelligently analyzing and recording the video conference content, intelligent analysis and recording of the video conference content, acquisition of video and audio data, data conversion and intelligent analysis are realized by utilizing a deep learning network and a natural language processing technology, and the system accurately captures and analyzes visual and auditory features in a conference, so that the video conference content can be intelligently analyzed and recorded. Comprising facial expressions, body languages, speech intonations and speed, generating a complete conference record comprising emotion and dialogue content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video conferencing technology, specifically to a processing system for intelligent analysis and recording of video conferencing content. Background Technology

[0002] Existing video conferencing systems are widely used in various scenarios, possessing basic functions such as audio and video transmission, participant management, and meeting scheduling. Video conferencing transmits audio and video signals to all participants via a network, ensuring real-time interaction. The system supports the collaborative work of video terminals, meeting servers, and network services, ensuring participants can share audio, video, files, presentations, and other content. The system provides authentication and meeting room creation functions for convenient meeting security management. With technological advancements, video conferencing systems are gradually incorporating technologies such as cloud computing and big data analytics, enhancing the system's scalability and flexibility.

[0003] However, existing technologies have failed to address the challenges of intelligent analysis and recording of video conferencing content. Traditional systems rely on manual recording, which cannot accurately capture key discussions and decisions made during meetings. Discussion points and decision-making information from participants are not recorded in a timely and comprehensive manner. Furthermore, existing technologies lack intelligent analysis capabilities. Systems fail to analyze video and audio content in real time, making it difficult to effectively identify important information and potential risks within the meeting. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a processing system for intelligent analysis and recording of video conferencing content. The technical problem this invention aims to solve is: how to intelligently analyze video and audio data using deep learning networks and natural language processing methods, thereby solving the problem that traditional video conferencing systems cannot accurately capture key discussion and decision-making content.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a processing system for intelligent analysis and recording of video conference content, comprising: Video conferencing module: used to acquire video data during the meeting through a video camera and to acquire audio data during the meeting through an audio acquisition device; Network service module: used to perform data conversion on the video and audio data, and obtain standardized data through the data conversion; Video conferencing monitoring module: used to perform intelligent analysis on the standardized data, and generate meeting record data through the intelligent analysis. The intelligent analysis includes deep learning networks and natural language processing methods. Data storage module: used to store the meeting record data and encrypt the meeting record data to obtain encrypted record data; Information display module: used to decrypt the encrypted record data to form a meeting content record, and to visualize the meeting content record.

[0006] Preferably, the video camera has a resolution of 1080p, and the audio acquisition device captures multi-party audio signals during the meeting through a microphone array, and performs signal filtering on the multi-party audio signals to obtain audio data.

[0007] Preferably, the data conversion performs compression encoding on the video data to obtain compressed video data, the compression encoding adopts the H.265 standard, performs audio encoding on the audio data to obtain compressed audio data, the audio encoding adopts the AACs standard, and encapsulates the compressed video data and compressed audio data to obtain standardized data.

[0008] Preferably, the deep learning network includes a convolutional neural network and a recurrent neural network. The convolutional neural network extracts spatial features from standardized data to obtain visual features, and the recurrent neural network extracts time series data from standardized data to obtain speech features. The visual features include facial expression features and body language, and the speech features include speech tone and speech rate.

[0009] Preferably, the natural language processing method classifies facial expression features and body language, outputs facial expression action labels and posture action labels, calculates the frequency of occurrence of the facial expression action labels and posture action labels to obtain a preset weight table, assigns values ​​to the facial expression action labels and posture action labels according to the preset weight table, obtains the emotion intensity value through the assignment, and calculates the speech intensity of the speech tone and speech rate.

[0010] Preferably, the speech intensity calculation is based on extracting speech signals from standardized data, dividing the speech signals into frames to obtain framed speech segments, performing short-time energy calculation on the framed speech signals to obtain short-time energy sequences, linearly combining the short-time energy sequences, the changes in speech intonation, and the changes in speech rate to obtain speech emotion intensity values, and performing weighted average fusion of the emotion intensity values ​​and speech emotion intensity values ​​to generate meeting record data.

[0011] Preferably, the encryption obtains a key through a key derivation function, performs symmetric encryption on the meeting record data based on the key, and obtains encrypted record data through the symmetric encryption.

[0012] Preferably, the data decryption involves key verification to obtain a key decryption service, followed by symmetric decryption of the encrypted record data based on the key decryption service, and obtaining the meeting content record through the symmetric decryption. The visualization includes meeting summary display and decision record visualization.

[0013] This invention provides a processing system for intelligent analysis and recording of video conference content. It has the following beneficial effects: This intelligent video conferencing content analysis and recording system utilizes deep learning networks and natural language processing technology to achieve intelligent analysis and recording of video conferencing content. Through the acquisition, data conversion, and intelligent analysis of video and audio data, the system accurately captures and analyzes the visual and auditory features in the meeting, including facial expressions, body language, tone of voice, and speaking speed, generating a complete meeting record that includes emotions and dialogue content.

[0014] By combining convolutional neural networks, recurrent neural networks, and natural language processing techniques, this invention improves the overall efficiency of the analysis and recording process. The system generates detailed emotion intensity values ​​and voice sentiment intensity, integrating this information into the final meeting transcript. Data encryption ensures the security of sensitive information, enabling the system to efficiently capture meeting content and providing reliable security for commercial use. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the system structure of the present invention; Figure 2 This is a flowchart illustrating the analysis process of the video conferencing monitoring module of the present invention. Figure 3 This is a flowchart illustrating the decryption and display process of the information display module of this invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1

[0018] like Figure 1-3 As shown, this embodiment of the invention provides a processing system for intelligent analysis and recording of video conference content, including: Video conferencing module: Used to acquire video data during meetings via a webcam and audio data via an audio acquisition device. The webcam has a resolution of 1080p, and the audio acquisition device captures audio signals from multiple parties during the meeting through a microphone array, filtering the signals to obtain the audio data.

[0019] Network service module: Used for data conversion of video and audio data to obtain standardized data. Data conversion includes compressing video data using the H.265 standard, encoding audio data using the AACs standard, and encapsulating the compressed video and audio data to obtain standardized data.

[0020] The video conferencing monitoring module performs intelligent analysis on standardized data, generating meeting minutes. This intelligent analysis utilizes deep learning networks and natural language processing (NLP). The deep learning networks include convolutional neural networks (CNNs) and recurrent neural networks (RNNs). CNNs extract spatial features from the standardized data to obtain visual features, while RNNs extract time-series features to obtain speech features. Visual features include facial expressions and body language, while speech features include tone of voice and speech rate. NLP classifies facial expressions and body language, outputting facial expression / gesture labels. It then calculates the frequency of these labels to obtain a pre-defined weight table, assigns values ​​to the labels according to this table, and obtains emotion intensity values. Finally, it calculates speech intensity based on tone of voice and speech rate. Speech intensity calculation is based on extracting speech signals from standardized data, dividing the speech signals into frames to obtain framed speech segments, performing short-time energy calculation on the framed speech signals to obtain short-time energy sequences, linearly combining the short-time energy sequences, the changes in speech intonation and speech rate, and obtaining speech emotion intensity values ​​through linear combination, and then performing weighted average fusion of the emotion intensity values ​​and speech emotion intensity values ​​to generate meeting record data.

[0021] Data storage module: Used to store meeting minutes and encrypt them to obtain encrypted meeting minutes. Encryption is achieved by obtaining a key through a key derivation function, and then using that key to perform symmetric encryption on the meeting minutes to obtain the encrypted meeting minutes.

[0022] Information Display Module: This module decrypts encrypted meeting records to create a visual representation of the meeting content. Data decryption involves key verification to obtain a key decryption service. Based on this service, the encrypted data is symmetrically decrypted to obtain the meeting content. The visualization includes a meeting summary and a visualization of decision records.

[0023] By using deep learning models to intelligently analyze video and audio data, features such as facial expressions, body language, tone of voice, and speech rate are extracted to generate accurate and comprehensive meeting records, enhancing the ability to capture details of meeting content.

[0024] By analyzing facial expressions and body language, and combining them with vocal emotional features, the emotional state of participants can be determined, and an emotional intensity value can be generated, which helps to reflect the meeting atmosphere and decision-making process.

[0025] Automatically generate meeting minutes and summaries, reducing manual intervention and improving efficiency. Automated processing reduces the possibility of errors in manual recording, ensuring more accurate extraction of meeting content.

[0026] The data storage module encrypts meeting records to ensure the security of meeting content, prevent the leakage of sensitive information, and protect meeting content with commercial value or confidentiality requirements.

[0027] Visualizing meeting summaries and decision records makes it easier for users to quickly access key information and improves the readability and usability of meeting data.

[0028] By comprehensively processing video, audio, and natural language data, and fully analyzing meeting content, we can ensure the efficient integration of various types of information and improve the accuracy and completeness of meeting records.

[0029] Example 2

[0030] This embodiment is a processing system based on intelligent analysis and recording of video conference content. It collects and intelligently analyzes the audio and video data of participants through a video conference monitoring module, extracts visual and voice features, generates accurate meeting records, and ensures secure data storage and visualization. The specific implementation method is as follows: 1. Data Acquisition and Preprocessing In a real-world video conference, a video conferencing monitoring system was used to capture and process audio and video data. The meeting, attended by five participants, focused on project progress and team collaboration. It lasted 60 minutes and covered topics such as task allocation, project progress, and member feedback.

[0031] Video data collection: The meeting used a 1080p resolution camera with a frame rate of 30fps. Based on actual testing results of the equipment configuration and meeting environment, the video data size was determined to be 185MB per minute.

[0032] The video data size is 185MB per minute, and the total meeting duration is 60 minutes. The total video data size of the meeting is 185MB × 60 minutes = 11.1GB.

[0033] Audio data acquisition: Audio data was acquired using a 4-microphone array with a sampling rate of 48kHz. Equipment testing showed that the audio data size was 6MB per minute, reflecting the actual data acquisition volume in a standard voice conferencing environment.

[0034] The audio data size is 6MB per minute, and the total meeting duration is 60 minutes. The total audio data size of the meeting is 6MB × 60 minutes = 360MB.

[0035] 2. Data Conversion

[0036] Video data compression: Using the H.265 video encoding standard, the video data compression ratio is determined to be 1:3 based on the actual compression performance of the device and test results. This embodiment of the video encoding standard achieves high compression efficiency while ensuring video quality.

[0037] The video data compression ratio is 1:3, and the compressed video data size is 11.1GB ÷ 3 = 3.7GB.

[0038] Audio data encoding: Audio data is compressed using AAC encoding at a compression ratio of 1:3.

[0039] AAC encoding reduces data volume while maintaining audio quality, making it suitable for compressing high-quality audio data.

[0040] The audio data compression ratio is 1:3, and the size of the compressed audio data is 360MB ÷ 3 = 120MB.

[0041] Standardized data: The compressed audio and video data are merged into a standardized data stream, and the size of the generated standardized data is: 3.7GB + 120MB = 3.82GB.

[0042] 3. Intelligent Analysis and Feature Extraction

[0043] Visual feature extraction: The system uses a convolutional neural network to process video data and extract facial expression and body language features.

[0044] Based on the meeting video data, the system identified the following facial expressions and actions: The smile appeared 10 times. Smiles generally indicate positive emotions, suggesting that participants may have a positive attitude towards the meeting. Therefore, the weight of smiles is relatively high and set at 0.7.

[0045] Frowning appeared 4 times. Frowning is usually associated with emotions such as doubt, confusion, or dissatisfaction. Frowning can reflect emotions, but compared to smiling, it has a smaller impact on the intensity of emotions, so its weight is set to 0.3.

[0046] The nodding occurred 8 times. Nodding usually indicates agreement or understanding. In a meeting, the appearance of a nodding expression may indicate the participant's approval and understanding, and is assigned a weight of 0.4.

[0047] Calculate the emotional intensity of each feature: Smile intensity: 10 times × 0.7 = 7.0; frown intensity: 4 times × 0.3 = 1.2; nod intensity: 8 times × 0.4 = 3.2.

[0048] Speech feature extraction: The system uses a recurrent neural network to analyze audio data and extract speech pitch and rate features. Based on actual speech data from the meeting, the speech rate and pitch fluctuation rates are: With a speech intonation fluctuation rate of 5% and a speech rate of 130 words per minute, the system calculates the emotional intensity value of the speech:

[0049] Experimental analysis suggests that intonation fluctuations have a significant impact on the intensity of emotion; therefore, the following settings were established. , .

[0050] The above data were normalized, and a 5% intonation fluctuation rate was directly represented by 0.05. Speech rate was standardized based on the typical speech rate in a conversational environment; the average speech rate in a meeting is 120 words per minute, so a speech rate of 130 words per minute is expressed as... .

[0051]

[0052] 4. Sentiment Analysis and Tag Generation

[0053] Facial expressions and body language classification: Facial expressions and body language were classified using natural language processing methods, and weights were assigned to each label. The final visual emotion intensity was 11.4.

[0054] Voice sentiment analysis: Through audio data analysis, the emotional intensity of the participants' voices was found to be 0.462.

[0055] Weighted fusion: The system performs a weighted average of visual emotion intensity and vocal emotion intensity. The weight for visual emotion intensity is 0.6, and the weight for vocal emotion intensity is 0.4. The final comprehensive emotion intensity value is calculated as follows:

[0056] 5. Generate meeting minutes

[0057] The system generates detailed meeting minutes based on extracted visual features, speech features, and emotion intensity. The minutes include: The visual emotion intensity was 11.4, the verbal emotion intensity was 0.462, and the overall emotion intensity was 7.0248.

[0058] Meeting summary: Covers discussions on project progress, teamwork, and task allocation, and includes a sentiment analysis of the participants.

[0059] In summary, this embodiment demonstrates that the video conferencing monitoring module supports sentiment analysis and decision recording in practical applications, while ensuring data security and privacy protection.

[0060] Example 3

[0061] This embodiment is a video conferencing content intelligent analysis and recording processing system. Through its data storage and information display modules, this system aims to ensure secure storage, encryption protection, and efficient display of meeting data. The specific implementation method is as follows: 1. Data Acquisition and Transformation During the video conference, the system acquires video and audio data in real time through a 1080p resolution video camera and a microphone array audio acquisition device.

[0062] Video data is compressed using the H.265 encoding standard, and audio data is compressed using AAC encoding. The converted data is then encapsulated by the network service module and transformed into a standardized data format for subsequent analysis.

[0063] The system converts the video data of participants' facial expressions and body language, as well as the audio data of their speech tone and speed, into standard format data for subsequent processing.

[0064] 2. Sentiment Analysis Data Storage

[0065] The system uses deep learning networks to analyze video and audio data and extract the emotional intensity of each participant.

[0066] Zhang San's emotional intensity value during the meeting was 0.82, Li Si's emotional intensity value during the meeting was 0.78, and Wang Wu's emotional intensity value during the meeting was 0.90.

[0067] The aforementioned emotional intensity values ​​were obtained by analyzing the participants' facial expressions such as smiles and frowns, as well as the speech characteristics of changes in speech rate and tone.

[0068] All sentiment analysis data will be stored in the system database.

[0069] 3. Meeting minutes data storage

[0070] Basic information about the meeting, including the meeting ID, participants, meeting time, and meeting topic, as well as sentiment analysis data, will be stored in a standard format. To ensure data security, all meeting records will be encrypted using the AES-256 symmetric encryption algorithm to guarantee confidentiality. The encrypted data will be stored in a database to prevent unauthorized access.

[0071] 4. Data encryption

[0072] Meeting minutes are encrypted using a generated encryption key before storage. Encryption employs the AES-256 algorithm to ensure data security. The encrypted data is then stored in a database, typically in binary format or Base64 encoded form.

[0073] 5. Data Decryption

[0074] When it is necessary to view meeting minutes, the system will retrieve the encrypted meeting minutes data from storage and verify the validity of the key through the key decryption service.

[0075] After successful verification, the encrypted data is decrypted into the original meeting transcript using the AES-256 decryption algorithm. The decrypted data includes the meeting ID, attendee list, sentiment analysis data, and other information.

[0076] 6. Meeting Summary Presentation

[0077] The decrypted data will be displayed on the front-end interface in text or table format, including information such as meeting topic, time, attendees, and meeting summary. The displayed content includes: The meeting, themed "Annual Performance Summary," was to be held on December 22, 2025, at 2:00 PM. Participants included Zhang San, Li Si, and Wang Wu.

[0078] Meeting Summary: This meeting summarized the performance from Q1 to Q3 of 2025 and discussed the contributions of each department.

[0079] 7. Sentiment Analysis Demonstration

[0080] The system uses visual charts to display the emotional intensity of each participant, usually represented by bar charts or pie charts. The bar chart of emotional intensity shows the emotional state of each participant.

[0081] In the above analysis, Wang Wu, a participant with higher emotional intensity, is represented by a taller bar, while Li Si, a participant with lower emotional intensity, is represented by a shorter bar. This visualization helps managers and participants understand the overall atmosphere of the meeting.

[0082] 8. Decision Record Display

[0083] Key decisions and discussion outcomes from the meeting will be presented in list format. Decision records may include: raising the Q4 sales target; requiring Department A to submit a more detailed performance report.

[0084] Decision items will be presented in the form of lists or charts, making it easier for participants to review the important decisions made during the meeting and helping to improve the transparency and effectiveness of meeting decisions.

[0085] Through the above steps, secure storage and encryption of meeting minutes are achieved, ensuring data confidentiality and integrity. The system intuitively displays meeting summaries, sentiment analysis results, and decision records, providing users with tools for meeting review and decision tracking. The sentiment analysis function helps assess meeting atmosphere, promotes transparency and execution of decisions, and improves the usability and operational efficiency of meeting data while ensuring data security. It is suitable for various scenarios requiring efficient management and analysis of meeting data.

[0086] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A processing system for intelligent analysis and recording of video conference content, characterized in that, include: Video conferencing module: used to acquire video data during the meeting through a video camera and to acquire audio data during the meeting through an audio acquisition device; Network service module: used to perform data conversion on the video and audio data, and obtain standardized data through the data conversion; Video conferencing monitoring module: used to perform intelligent analysis on the standardized data, and generate meeting record data through the intelligent analysis. The intelligent analysis includes deep learning networks and natural language processing methods. Data storage module: used to store the meeting record data and encrypt the meeting record data to obtain encrypted record data; Information display module: used to decrypt the encrypted record data to form a meeting content record, and to visualize the meeting content record.

2. The intelligent analysis and recording system for video conferencing content according to claim 1, characterized in that: The video camera has a resolution of 1080p, and the audio acquisition device captures multi-party audio signals during the meeting through a microphone array, and performs signal filtering on the multi-party audio signals to obtain audio data.

3. The intelligent analysis and recording system for video conferencing content according to claim 1, characterized in that: The data conversion process involves compressing and encoding the video data to obtain compressed video data, using the H.265 standard; performing audio encoding on the audio data to obtain compressed audio data, using the AACs standard; and then encapsulating the compressed video data and compressed audio data to obtain standardized data.

4. The intelligent analysis and recording system for video conferencing content according to claim 1, characterized in that: The deep learning network includes a convolutional neural network and a recurrent neural network. The convolutional neural network extracts spatial features from standardized data to obtain visual features, and the recurrent neural network extracts time series features from standardized data to obtain speech features. The visual features include facial expression features and body language, and the speech features include speech tone and speech rate.

5. The intelligent analysis and recording system for video conferencing content according to claim 4, characterized in that: The natural language processing method classifies facial expression features and body language, outputs facial expression action labels and posture action labels, calculates the frequency of occurrence of the facial expression action labels and posture action labels to obtain a preset weight table, assigns values ​​to the facial expression action labels and posture action labels according to the preset weight table, obtains the emotion intensity value through the assignment, and calculates the speech intensity of the speech tone and speech rate.

6. The intelligent analysis and recording system for video conferencing content according to claim 5, characterized in that: The speech intensity calculation is based on extracting speech signals from standardized data, dividing the speech signals into frames to obtain framed speech segments, performing short-time energy calculation on the framed speech signals to obtain short-time energy sequences, linearly combining the short-time energy sequences, the changes in speech tone, and the changes in speech rate to obtain speech emotion intensity values, and performing weighted average fusion of the emotion intensity values ​​and speech emotion intensity values ​​to generate meeting record data.

7. The intelligent analysis and recording system for video conferencing content according to claim 1, characterized in that: The encryption process obtains a key through a key derivation function, performs symmetric encryption on the meeting record data based on the key, and obtains encrypted record data through the symmetric encryption.

8. The intelligent analysis and recording system for video conferencing content according to claim 7, characterized in that: The data decryption process verifies the key to obtain a key decryption service. Based on the key decryption service, the encrypted record data is symmetrically decrypted to obtain the meeting content record. The visualization includes meeting summary display and decision record visualization.