Method for generating highlights in the scenario of business meetings based on artificial intelligence
By using multimodal consistency detection and long-term memory network model to verify the authenticity of behind-the-scenes videos in business meeting scenarios, the problem that AI editing may generate false content is solved, ensuring the authenticity and credibility of behind-the-scenes videos.
Patent Information
- Application Number
- CN202510336193.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The prior art can generate false or misleading content when AI automatically edits and optimizes business meeting highlights, affecting the credibility of the meeting and leading to false strategic decisions based on false information.
The behind-the-scenes generation method based on artificial intelligence is used to verify the authenticity of the behind-the-scenes fragments through the multimodal consistency detection model and the long-term and short-term memory network model, detect the synchronization of speech and facial expressions, quantify the speech recognition abnormality index and facial expression variation index, and predict the authenticity of the behind-the-scenes fragments through machine learning models.
It effectively solves the problems of audio and video out-of-synchronization, Deepfake tampering and false information dissemination caused by AI editing, ensuring that the generated behind-the-scenes videos are highly authentic, with enhanced content credibility, and avoiding wrong strategic decisions.
Smart Images

Figure CN119854441B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method for generating highlights in the scenario of business meetings based on artificial intelligence. Background Art
[0002] Generating highlights in the scenario of business meetings based on artificial intelligence refers to using AI technology to intelligently extract, edit, and automatically generate content from key segments, exciting moments, or interesting interactions during the meeting. These highlights can include excellent speeches by speakers, audience interactions, data visualization displays, and even AI-generated captions, background music, or stylized videos for meeting review, promotion, or social media sharing. This technology typically combines natural language processing (NLP), computer vision, and video editing algorithms to intelligently identify and screen out the most valuable segments by analyzing meeting video, audio, and text data. For example, AI can automatically extract high-frequency discussion topics, moments of high emotion, or the formation process of key decisions and integrate them into short highlight videos to make the meeting content more communicable and valuable for review.
[0003] The existing technology has the following deficiencies:
[0004] During the process of AI automatically editing and optimizing highlights, some algorithms (such as Deepfake or GANs) may generate false or misleading content. For example, due to audio-visual synchronization errors, AI may cause the speech of a speaker not to match their facial expression, or even generate false facial expressions or gestures, making the video content appear tampered with and affecting the credibility of the meeting. In addition, if AI wrongly edits or tampers with the meeting content, it may cause executives or decision-makers to make wrong judgments based on false information, thus triggering wrong strategic decisions. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for generating highlights in the scenario of business meetings based on artificial intelligence to solve the deficiencies in the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for generating highlights in the scenario of business meetings based on artificial intelligence, including the following steps:
[0007] S1: Obtain the video, audio, and text data of the business meeting, and preprocess the data, including noise removal, speech recognition, and text transcription;
[0008] S2: Use natural language processing and computer vision technologies to identify high-frequency discussion topics, moments of high emotion, and the formation process of key decisions in the meeting, and screen out potential highlight segments;
[0009] S3: Based on the multi-modal consistency detection model, verify the selected highlight clips. By comparing the speaker's voice features with the features of their facial expression changes, detect whether there is forged content generated by Deepfake or GANs, and determine the authenticity of the highlight clips;
[0010] S4: Edit and optimize the highlight clips that have passed the authenticity verification, including automatic subtitle generation, background music matching, and visual effect enhancement;
[0011] S5: Further verify the authenticity of the optimized highlight clips through the long short-term memory network model. By tracking the time series relationship between video frames and speech, filter out incorrectly edited or tampered videos, and set permission controls;
[0012] S6: Export the finally generated highlight video to the enterprise internal platform, social media, or storage system, and support user feedback and secondary editing optimization.
[0013] Preferably, in S3, after analyzing the extracted voice features, generate a voice recognition anomaly index. The method for obtaining the voice recognition anomaly index is as follows:
[0014] Define the input text: T r : Manually transcribed text, T a : ASR speech recognition output text; Calculate the Levenshtein distance D(T r , T a ), The Levenshtein distance is defined as the minimum number of edit operations required to convert one string to another string. The expression is:
[0015]
[0016] Where: D(i,j) represents the minimum edit distance from the first i characters of T r to the first j characters of T a , T r [i]≠T a [j] is an indicator function. If T r [i] and T a [j] are different, the value is 1; otherwise the value is 0; D(i - 1,j)+1 represents deleting a character, D(i,j - 1)+1 represents inserting a character, D(i - 1,j - 1)+1(T r [i]≠T a [j]) represents replacing a character; Calculate the normalized edit distance NED. The expression is: Where: |T r | and |T a | are the number of words in the manually transcribed text and the ASR recognized text respectively; max(|Tr , T a ), so that the normalized value is always between 0 and 1. The speech recognition anomaly index SRAI reflects the deviation degree of ASR recognition, and the expression is: SRAI = 1 - e -λ·NED ; where: λ is the smoothing parameter.
[0017] Preferably, set an anomaly threshold. If SRAI < 0.2, the ASR recognition quality is good and the error is acceptable; if 0.2 ≤ SRAI < 0.5, there are some recognition errors and further evaluation is needed; if SRAI ≥ 0.5, the recognition quality is poor, including serious misidentifications or ASR model anomalies.
[0018] Preferably, in S3, after analyzing the extracted facial expression change features, a facial expression variation index is generated. The acquisition method of the facial expression variation index is:
[0019] It is assumed that the conference video has been segmented into multiple highlight clips, and each clip contains consecutive video frames. For each frame f t Perform facial expression feature extraction. Input: V: input video clip, containing N frames; F t : the facial expression feature vector detected in the t-th frame; the facial expression feature of each frame f t is represented as: F t ={L t , E t}; where: L t represents the two-dimensional coordinates of facial key points, and E t represents the deep learning feature of facial muscle movement;
[0020] In consecutive video frames, calculate the expression change rate dF of adjacent frames t as the instantaneous expression variation degree, and the expression is: Calculate the facial expression variation index, and the expression is: N is the total number of frames of the video clip, and FEVI is the facial expression variation index.
[0021] Preferably, if FEVI < 0.2, the expression change is normal and there is no anomaly; if 0.2 ≤ FEVI < 0.5, there is an anomaly in the expression change and further analysis is needed; if FEVI ≥ 0.5, the expression change is highly abnormal, which is Deepfake or unnatural facial expression synthesis.
[0022] Preferably, in S3, the speech recognition anomaly index and the facial expression variation index are converted into a comprehensive feature vector, and the comprehensive feature vector is used as the input of the machine learning model. The machine learning model takes predicting the authenticity value label of the selected flower video clip for each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of the authenticity value labels of all the selected flower video clips as the training target to train the machine learning model until the sum of the prediction errors reaches convergence and then stops the model training. The authenticity value of the selected flower video clip is determined according to the model output result, where the machine learning model is a polynomial regression model.
[0023] Preferably, the authenticity value of the obtained selected flower video clip is compared with the authenticity value reference threshold preset according to historical data. If the authenticity value of the obtained selected flower video clip is greater than or equal to the preset authenticity value reference threshold, it indicates that there is no forged content generated by Deepfake or GANs, and the authenticity of the flower video clip is high, and the corresponding flower video clip is marked as a low-risk clip; if the authenticity value of the obtained selected flower video clip is less than the preset authenticity value reference threshold, it indicates that there is forged content generated by Deepfake or GANs, and the authenticity of the flower video clip is low, and the corresponding flower video clip is marked as a high-risk clip; the low-risk clips automatically pass the verification and are stored in the database; the high-risk clips are submitted to manual review to further confirm the authenticity.
[0024] Preferably, in S5, the authenticity of the optimized flower video clip is further verified through a long short-term memory network model, specifically:
[0025] An input sequence for the LSTM model is constructed through the speech recognition anomaly index and the facial expression variation index extracted from the flower video clip;
[0026] The LSTM processes the time series data and predicts the video authenticity score for each time step. The LSTM processes the input features at time step s, calculates the hidden state and the authenticity score, and the model outputs the authenticity prediction value T for each time step s s : The mean value of the authenticity scores for all time steps s of the entire flower video clip is calculated as the overall authenticity score of the flower video clip, and the value range is [0,1].
[0027] Preferably, according to the authenticity score output by the model, the permission level of the flower video clip is determined, and an authenticity threshold is set; if the authenticity score output by the model is greater than or equal to the set authenticity threshold, it indicates that the video authenticity is high, the audio and video are synchronized, and there are no obvious tampering traces, and it is marked as a low-risk authenticity clip and is allowed to be automatically published or stored in the database; if the authenticity score output by the model is less than the set authenticity threshold, it indicates that the video has forged faces, voice tampering or editing errors, and it is marked as a high-risk non-authenticity clip and requires manual review.
[0028] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:
[0029] 1. The present invention deeply verifies the authenticity of the highlight clips through multi-modal consistency detection and long short-term memory network (LSTM), effectively solving problems such as audio-visual asynchrony, Deepfake tampering, and false information dissemination that may be caused by AI editing in the prior art. Quantify the authenticity of the meeting content through the Speech Recognition Anomaly Index (SRAI) and Facial Expression Variation Index (FEVI), combine with a polynomial regression model for prediction, and track the time series relationship between speech and video frames through LSTM to further detect editing errors or forged content. In addition, the system supports intelligent editing optimization, including automatic subtitle generation, soundtrack matching, and visual effect enhancement, improving the viewing experience and dissemination value of the video.
[0030] 2. While ensuring the authenticity of the highlight clips, the present invention also improves the automation and security of content generation. By setting permission control, low-risk clips can be directly published or stored in the database, while high-risk clips require manual review to ensure the credibility of the disseminated content. The finally generated highlight video can be exported to the enterprise internal platform or social media, and supports user feedback and secondary editing optimization, enabling enterprises to efficiently and securely manage and promote business meeting content, thereby enhancing brand influence and decision-making transparency. Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0032] Figure 1 It is the flowchart of the method of the present invention. Detailed Embodiments
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0034] Embodiment, please refer to Figure 1 As shown, the method for generating highlights in the business meeting scenario based on artificial intelligence described in this embodiment includes the following steps:
[0035] S1: Obtain the video, audio, and text data of a business meeting, and preprocess the data, including noise removal, speech recognition, and text transcription;
[0036] S2: Utilize natural language processing and computer vision technologies to identify the high-frequency discussion topics, emotionally intense moments, and the formation process of key decisions in the meeting, and screen out potential highlight clips;
[0037] S3: Based on a multi-modal consistency detection model, verify the screened highlight clips. By comparing the speaker's voice characteristics with the changes in their facial expression features, detect whether there is forged content generated by Deepfake or GANs, and determine the authenticity of the highlight clips;
[0038] S4: Edit and optimize the highlight clips that have passed the authenticity verification, including automatic subtitle generation, music score matching, and visual effect enhancement;
[0039] S5: Further verify the authenticity of the optimized highlight clips through a long short-term memory network model. By tracking the time series relationship between video frames and speech, filter out incorrectly edited or tampered videos, and set permission controls;
[0040] S6: Export the finally generated highlight video to an enterprise internal platform, social media, or storage system, and support user feedback and secondary editing optimization.
[0041] In S1, before the business meeting starts, the system deploys multiple data acquisition devices, including high-definition cameras, omnidirectional microphone arrays, and meeting recording software, to ensure comprehensive capture of audio-visual information during the meeting.
[0042] Video Capture: A high-definition camera is used to record the meeting scene, including the expressions, gestures of the speakers, and the projected content. The camera supports auto-focus and can intelligently track the position of the speakers. Audio Capture: A omnidirectional microphone array is used to obtain high-quality voice data, and beamforming technology is combined to enhance the voices of the speakers and reduce the interference of ambient noise. Text Data Capture: If the meeting is equipped with a real-time caption or note-taking system, the text data can be directly obtained; otherwise, the audio needs to be transcribed into text through speech recognition technology. The collected meeting data is cleaned, optimized, and formatted to improve the accuracy and stability of subsequent analysis. Adaptive Noise Cancellation (ANC) technology is used to filter background noise, such as keyboard typing sounds, air conditioner noise, or echoes from remote participants in the meeting room. Through speech enhancement algorithms, such as the Wavenet noise reduction model based on deep learning, the clarity of the human voice is improved, enabling the speech recognition system to transcribe text more accurately.
[0043] Use an end-to-end speech recognition model (such as DeepSpeech or Whisper) to convert the meeting audio into text. Speaker Diarization technology is adopted, based on the x-vector or Deep Speaker Embeddings model, to distinguish the voices of different speakers and label the identities of the speakers. Grammar correction and semantic optimization are performed on the transcribed text, and automatic error correction is carried out using BERT or GPT language models to improve the transcription quality.
[0044] Combined with natural language processing (NLP) technology, keyword extraction and topic analysis are performed on the transcribed text to identify the high-frequency discussion topics in the meeting. Sentiment analysis algorithms are used to evaluate the tone of the speakers and identify the moments of high emotion in the meeting for subsequent editing. Through text segmentation, long texts are split into logically complete segments to provide structured data for subsequent editing.
[0045] The preprocessed audio-visual data, text transcription results, and related metadata (such as speaker identities, timestamps, etc.) are uniformly stored in the database for subsequent analysis and retrieval. The data format is structured using JSON or XML to make subsequent AI editing and highlight generation more efficient.
[0046] In S2, first, the system receives and processes the preprocessed video, audio, and text data, and extracts the key features for screening highlight segments.
[0047] Text data: Extract keywords, themes, and sentiment features based on the results of conference speech transcription.
[0048] Audio data: Analyze speech pitch, speech rate, and speech sentiment to identify moments of high emotional fluctuations.
[0049] Video data: Based on facial expression recognition and body movement analysis, detect the emotional states of speakers or audiences, such as behaviors like smiling, being surprised, or applauding.
[0050] Utilize natural language processing techniques to extract high-frequency discussion topics from conference text data to determine the most valuable content segments.
[0051] Keyword extraction: Adopt algorithms such as TF-IDF and TextRank to extract core terms in the conference, and combine pre-trained language models (such as BERT and GPT) for context understanding to identify the most important topics.
[0052] Topic modeling: Use LDA (Latent Dirichlet Allocation) or BERT topic clustering methods to perform semantic classification on the conference content and extract the main discussion directions.
[0053] Co-occurrence analysis: Analyze the co-occurrence relationships between keywords, construct a topic network to discover important discussion content related to the core issues of the conference.
[0054] Combine audio-visual data to detect critical moments related to the emotional fluctuations of the audience or speakers during the conference to screen out interesting highlights with dissemination value.
[0055] Audio emotion analysis: Based on speech emotion recognition models such as CNN + BiLSTM or Wav2Vec, analyze the changes in speech pitch, speech rate, and energy to identify exciting, enthusiastic, or emphasized speech segments.
[0056] Facial expression recognition: Use deep learning models such as FaceNet, OpenCV DNN, or EmotionNet to detect the emotions of speakers and audiences, such as smiling, being surprised, and thinking seriously, etc., to screen out content with strong interactivity.
[0057] Body movement analysis: Through pose estimation techniques, such as OpenPose or MediaPipe, detect the gesture changes and standing movements of speakers to identify highly dynamic moments.
[0058] In a conference, key decisions are often the most valuable interesting highlights. Therefore, it is necessary to comprehensively identify the decision-making process based on text, speech, and visual data.
[0059] Text Decision Recognition: Based on semantic analysis models such as BERT and RoBERTa, detect statements in the meeting that involve decision-making terms such as "decide", "approve", "vote", and "take action", and analyze the decision-making process in combination with the context.
[0060] Speech Emphasis Detection: Use audio feature extraction models such as MFCC (Mel Frequency Cepstral Coefficients) and VGGish to analyze the stress, pauses, and intonation changes in the speech to determine the speech patterns of the speakers in key decision-making links.
[0061] Voting and Interaction Detection: Based on computer vision technology, detect actions such as raising hands, applauding, and nodding in the meeting, and combine with the speech recognition results to judge whether the decision is unanimously approved and screen the corresponding segments.
[0062] Score all the identified potential highlight segments and screen the content with the most dissemination value. Content Value Scoring: Calculate the comprehensive score of each segment based on the popularity of the discussion topic, the intensity of emotions, and the importance of the decision. Visual Analysis: Use data visualization technology to display the emotional fluctuation curves, keyword clouds, and video timelines of each highlight segment to assist manual adjustment and optimization. Final Screening: Set a threshold to screen high-value segments and provide personalized recommendations according to user needs, such as "the most interactive segment" and "core decision-making segment".
[0063] The generated highlight segments are stored in the enterprise database and archived in a structured format (such as JSON, XML, etc.) to support subsequent calls. They can be exported to standard video formats (such as MP4, MOV, etc.) and automatically added with subtitles, background music, transition effects and other optimization elements. Support one-click publishing to social media, enterprise internal platforms or meeting record systems for quick dissemination and review.
[0064] In S3, after screening out the candidate highlight segments, extract multi-modal features for authenticity verification from video, audio, and text data.
[0065] Speech Feature Extraction, including:
[0066] Adopt acoustic feature extraction methods such as Mel-spectrogram and MFCC (Mel Frequency Cepstral Coefficients) to obtain the key features of the speakers' speech.
[0067] Use deep learning models such as Wav2Vec2.0 or Deep Speaker Embeddings to perform feature encoding on the speakers' speech identities to ensure the authenticity and reliability of the speech sources.
[0068] Combine with BiLSTM or Transformer models to analyze the prosody, stress, and pause patterns of speech, ensuring that the speech content is natural without signs of speech splicing or GANs generation.
[0069] The extraction of facial expression change features includes:
[0070] Adopt face recognition models such as FaceNet, Dlib, or MediaPipe to extract the facial key point coordinates of the speaker, including the dynamic changes of the mouth, eyes, and eyebrows.
[0071] Use models such as LipNet or SyncNet to analyze the mouth movement trajectory of the speaker, perform time alignment with the speech waveform, and determine whether there is a problem of lip movement being out of sync with the speech.
[0072] Adopt models such as CNN+LSTM or VisionTransformer to perform temporal analysis on facial expression changes and detect whether there are unnatural facial movements (such as abnormal blinking frequency, stiff facial expressions, etc.).
[0073] After analyzing the extracted speech features, generate a speech recognition anomaly index. The method for obtaining the speech recognition anomaly index is as follows:
[0074] Define the input text: T r : Manually transcribed text (GroundTruth), T a : ASR speech recognition output text; calculate the Levenshtein distance D(T r , T a ), where the Levenshtein distance is defined as the minimum number of edit operations required to convert one string to another, and the expression is:
[0075]
[0076] where: D(i,j) represents the minimum edit distance from the first i characters of T r to the first j characters of T a , T r [i]≠T a [j] is an indicator function. If T r [i] and T a [j] are different, the value is 1 (i.e., a replacement operation); otherwise, the value is 0 (no edit operation). D(i-1,j)+1 represents deleting a character, D(i,j-1)+1 represents inserting a character, and D(i-1,j-1)+1 (T r [i]≠T a [j]) represents replacing a character;
[0077] To standardize the edit distance for texts of different lengths, the normalized edit distance NED is calculated, and the expression is: where: |T r | and |T a | are the number of words (or characters) in the manually transcribed text and the ASR recognized text respectively. max(|T r |, |T a |) makes the normalized value always between 0 and 1. The speech recognition anomaly index SRAI reflects the deviation degree of ASR recognition, and the expression is: SRAI = 1 - e -λ·NED ; where: λ is the smoothing parameter (usually taking values from 3 to 5, depending on the sensitivity of the system to errors).
[0078] The exponential function e -λ·NED makes a smaller NED correspond to a lower anomaly index, while a larger NED makes the anomaly index approach 1.
[0079] Set the anomaly threshold. If SRAI < 0.2, the ASR recognition quality is good and the error is acceptable. If 0.2 ≤ SRAI < 0.5, there may be some recognition errors and further evaluation is needed. If SRAI ≥ 0.5, the recognition quality is poor and there may be serious misidentifications or ASR model anomalies.
[0080] After analyzing the extracted facial expression change features, a facial expression variation index is generated. The acquisition method of the facial expression variation index is as follows:
[0081] Assume that the conference video has been segmented into multiple highlight clips, and each clip contains consecutive video frames. For each frame f t (at time t), facial expression feature extraction is performed. Input: V: the input video clip, containing N frames; F t : the facial expression feature vector detected in the t-th frame;
[0082] Facial feature extraction method: Use FaceNet, OpenFace, Dlib or MediaPipe to extract 68 facial key points (Facial Landmarks) to form the key point vector L t ; Use ResNet + LSTM or Facial Action Coding System (FACS) to calculate the expression embedding vector E t ; Finally, the facial expression feature of each frame f t is represented as: F t = {L t , E t}; where: L t represents the two-dimensional coordinates of the facial key points (such as eyes, mouth, eyebrows, etc.), and E tDeep learning features representing facial muscle movements;
[0083] In consecutive video frames, calculate the rate of change of expression dF between adjacent frames t As the instantaneous expression variability, the expression is: Calculate the facial expression variability index, the expression is: N is the total number of frames in the video clip, and FEVI is the facial expression variability index.
[0084] If FEVI < 0.2, the expression change is normal and there is no abnormality; if 0.2 ≤ FEVI < 0.5, there is a certain abnormality in the expression change and further analysis is required; if FEVI ≥ 0.5, the expression change is highly abnormal and may be Deepfake or unnatural facial expression synthesis.
[0085] Convert the speech recognition anomaly index and the facial expression variability index into a comprehensive feature vector, and use the comprehensive feature vector as the input of the machine learning model. The machine learning model takes predicting the authenticity value label of the selected highlight clip for each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of the authenticity value labels of all the selected highlight clips as the training target. Train the machine learning model until the sum of the prediction errors reaches convergence and then stop the model training. Determine the authenticity value of the selected highlight clip according to the model output result, where the machine learning model is a polynomial regression model.
[0086] The method for obtaining the authenticity value of the selected highlight clip is: obtain the corresponding function expression from the comprehensive feature vector training data of the trained machine learning model: LR = F(SRAI, FEVI); where F is the output function of the model, SRAI is the speech recognition anomaly index, FEVI is the facial expression variability index, and LR is the authenticity value of the selected highlight clip.
[0087] Compare the obtained authenticity value of the selected highlight clip with the authenticity value reference threshold preset according to historical data. If the obtained authenticity value of the selected highlight clip is greater than or equal to the preset authenticity value reference threshold, it indicates that there is no forged content generated by Deepfake or GANs, and the authenticity of the highlight clip is high, and mark the corresponding highlight clip as a low-risk clip; if the obtained authenticity value of the selected highlight clip is less than the preset authenticity value reference threshold, it indicates that there is forged content generated by Deepfake or GANs, and the authenticity of the highlight clip is low, and mark the corresponding highlight clip as a high-risk clip; low-risk clips automatically pass the verification and are stored in the database; high-risk clips (such as detecting traces of GAN generation or identity mismatch) are submitted to manual review to further confirm the authenticity.
[0088] The verified behind-the-scenes clips are stored in the enterprise database and marked as "verified". If a behind-the-scenes clip fails the authenticity test, it can be re-screened, edited, or directly discarded to prevent the spread of misleading information. A detailed authenticity test log is generated, recording the risk score and test results of each clip for subsequent traceability and optimization.
[0089] S4: Edit and optimize the behind-the-scenes clips that have passed the authenticity verification, including automatic subtitle generation, background music matching, and visual effect enhancement.
[0090] To edit and optimize the behind-the-scenes clips that have passed the authenticity verification, it is first necessary to automatically generate subtitles to improve the readability and dissemination effect of the content. Use speech recognition (ASR) technology to extract the audio content and combine it with natural language processing (NLP) for grammar correction and sentence segmentation to ensure that the subtitles are fluent and contextually appropriate. In addition, adopt time synchronization technology (such as DTW or VAD) to accurately align the subtitles and speech, avoiding subtitle delay or misalignment. At the same time, support multilingual translation and provide high-quality cross-language subtitles through models such as Transformer or BERT to meet the needs of international conferences.
[0091] In terms of background music matching, the system will automatically recommend appropriate background music based on the mood and rhythm of the meeting content. Through audio emotion analysis (Speech Emotion Recognition, SER), detect the speaker's tone, intonation, and emotion type, and select background music that matches the meeting atmosphere. For example, in a formal business decision-making scenario, match low-key and steady background music, while in a team interaction session, you can choose faster-paced and positive background music. In addition, use a dynamic volume adjustment algorithm to ensure that the background music does not interfere with the speech content, and at the same time support AI automatic noise reduction and audio equalization to optimize the overall auditory experience.
[0092] To enhance the visual appeal of the video, visual effect enhancement can also be carried out, including automatic lens switching, highlighting key content, and overlaying information graphics. Based on computer vision technology (such as YOLO, OpenPose, etc.), automatically detect and track the speaker, and intelligently adjust the camera angle to make the picture more dynamic. At the same time, use augmented reality (AR) technology to overlay key information in the video, such as meeting points, real-time data analysis, or voting results, so that the audience can more intuitively understand the content. In addition, smooth transition animations, brand logos, and dynamic titles can be added to enhance the professionalism and brand recognition of the video.
[0093] S5: Further verify the authenticity of the optimized behind-the-scenes clips through a long short-term memory network model, filter out incorrectly edited or tampered videos by tracking the time series relationship between video frames and speech, and set permission controls.
[0094] Construct an input sequence for the LSTM model by extracting the speech recognition anomaly index and the facial expression variation index from the behind-the-scenes footage clips;
[0095] Use LSTM to process the time-series data and predict the video authenticity score for each time step. LSTM processes the input features at time step s, calculates the hidden state and the authenticity score. Finally, the model outputs the authenticity prediction value T for each time step s. s : Calculate the mean of the authenticity scores for all time steps s of the entire behind-the-scenes footage clip as the overall authenticity score of the behind-the-scenes footage clip, with a value range of [0, 1].
[0096] According to the authenticity score output by the model, determine the permission level of the behind-the-scenes footage clip. Set the authenticity threshold Tth (e.g., Tth = 0.7): If the authenticity score output by the model is greater than or equal to the set authenticity threshold, it indicates that the video has high authenticity, the audio and video are synchronized, and there are no obvious signs of tampering. Mark it as a low-risk authenticity clip and allow automatic publishing or storage in the database. If the authenticity score output by the model is less than the set authenticity threshold, it indicates that the video may have forged faces, voice tampering, or editing errors. Mark it as a high-risk non-authenticity clip and require manual review. It may trigger access permission restrictions and only allow administrators or reviewers to view.
[0097] S6: Export the finally generated behind-the-scenes video to the enterprise internal platform, social media, or storage system, and support user feedback and secondary editing optimization.
[0098] After the finally generated behind-the-scenes video passes the authenticity verification and editing optimization, it can be automatically exported to the enterprise internal platform, social media, or cloud storage system for quick access and sharing by different audiences. The enterprise internal platform can be used for archiving and employee training, while social media platforms (such as LinkedIn, YouTube, Twitter) are helpful for brand promotion and marketing. In addition, the system supports multiple video formats (MP4, MOV, AVI) and resolution adaptation to ensure smooth playback on different devices, and integrates a security encryption mechanism to prevent unauthorized access and tampering.
[0099] To improve the content quality and interactivity, the behind-the-scenes video also supports user feedback and secondary editing optimization. Users can provide opinions through the like, comment, or rating system, and the AI can improve the editing strategy based on the feedback analysis, such as adjusting the clip order, optimizing the subtitles, or enhancing the visual effects. At the same time, it supports video editing permission management, allowing authorized users to directly perform secondary editing in the cloud, such as supplementing key content, replacing the background music, or adding enterprise brand elements, to ensure that the final video meets the dissemination requirements and achieves the best effect.
[0100] In this embodiment, first, the video, audio, and text data of a business meeting are obtained and preprocessed, including noise removal, speech recognition, and text transcription, to ensure data quality. Subsequently, natural language processing and computer vision technologies are used to analyze the meeting content, identify high-frequency discussion topics, moments of high emotion, and the formation process of key decisions, and filter out potential highlight clips. To ensure authenticity, a multi-modal consistency detection model is used to compare the speech features of the speakers with the facial expression change features to detect whether there is any forged content generated by Deepfake or GANs. The highlight clips that pass the authenticity verification will be edited and optimized, including automatically generating subtitles, matching suitable background music, and enhancing visual effects, to improve the viewing experience. On this basis, a long short-term memory network model is further used to track the time series relationship between video frames and speech, detect whether there is any incorrect editing or tampering, and set corresponding access permissions. Finally, the optimized highlight video will be exported to the enterprise internal platform, social media, or storage system, and user feedback and secondary editing optimization are supported to ensure content quality and dissemination effect.
[0101] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0102] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0103] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this text can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0104] As described above, this is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.
Claims
1. A method for generating highlights in a business meeting scenario based on artificial intelligence, characterized in that: The following steps are involved: S1: Obtain video, audio and text data of a business meeting, and pre-process the data, including noise removal, speech recognition and text transcription; S2: Use natural language processing and computer vision technology to identify frequently discussed topics, emotional moments, and key decision-making processes in meetings, and screen out potential highlights; S3: Based on the multimodal consistency detection model, the selected highlights are verified. By comparing the speaker's voice features with the changes in his or her facial expressions, it is detected whether there is forged content generated by Deepfake or GANs, and the authenticity of the highlights is determined; S4: Edit and optimize the authenticity-verified footage, including automatic subtitle generation, music matching, and visual effects enhancement; S5: The authenticity of the optimized highlights is further verified through the long short-term memory network model. By tracking the time series relationship between video frames and voice, incorrect editing or tampering of videos is filtered out, and permission control is set; S6: Export the final generated behind-the-scenes video to the company's internal platform, social media or storage system to support user feedback and secondary editing optimization.
2. The method for generating highlights in a business conference scenario based on artificial intelligence according to claim 1 is characterized in that: In S3, the extracted speech features are analyzed to generate a speech recognition anomaly index, and the method for obtaining the speech recognition anomaly index is as follows: Define input text: T r : Manual transcription of text, T a :ASR speech recognition output text; calculate Levenshtein distance D(T r ,T a ), the Levenshtein distance is defined as the minimum number of edit operations required to transform one string into another, and the expression is: Where: D(i,j) represents the r The first i characters to T a The minimum edit distance of the first j characters, T r [i]≠T a [j] is the indicator function, if T r [i] and T a [j] is different, the value is 1; otherwise, the value is 0; D(i-1,j)+1 means deleting a character, D(i,j-1)+1 means inserting a character, and D(i-1,j-1)+1(T r [i]≠T a [j]) means replacing a character; the normalized edit distance NED is calculated as follows: Where: |T r | and |T a | are the number of words in the manually transcribed text and the ASR recognized text respectively; max(|T r |,|T a |) makes the normalized value always between 0 and 1. The speech recognition anomaly index SRAI reflects the degree of deviation of ASR recognition, and the expression is: SRAI = 1-e -λ·NED ; Where: λ is the smoothing parameter.
3. The method for generating highlights in a business conference scenario based on artificial intelligence according to claim 2 is characterized in that: The abnormal threshold is set. If SRAI < 0.2, the ASR recognition quality is good and the error is acceptable. If 0.2 ≤ SRAI < 0.5, some recognition errors exist and further evaluation is required. If SRAI ≥ 0.5, the recognition quality is poor, including serious misrecognition or ASR model abnormality.
4. The method for generating highlights in a business conference scenario based on artificial intelligence according to claim 3 is characterized in that: In S3, the facial expression variation index is generated after analyzing the extracted facial expression variation features. The facial expression variation index is obtained by: Assume that the conference video has been divided into multiple highlights segments, each segment contains continuous video frames, and for each frame f t To extract facial expression features, input: V: input video clip, containing N frames; F t : The facial expression feature vector detected in the tth frame; each frame f t The facial expression feature is expressed as: F t ={L t ,E t }; where: L t Represents the two-dimensional coordinates of facial key points, E t Deep learning features representing facial muscle movements; In continuous video frames, calculate the expression change rate dF of adjacent frames t As the instantaneous expression variability, the expression is: Calculate the facial expression variability index, the expression is: N is the total number of frames in the video clip, and FEVI is the facial expression variability index.
5. The method for generating highlights in a business conference scenario based on artificial intelligence according to claim 4 is characterized in that: If FEVI<0.2, the expression changes are normal and there is no abnormality; if 0.2≤FEVI<0.5, the expression changes are abnormal and further analysis is required; if FEVI≥0.5, the expression changes are highly abnormal, which is Deepfake or unnatural facial expression synthesis.
6. The method for generating highlights in a business conference scenario based on artificial intelligence according to claim 5 is characterized in that: In S3, the speech recognition anomaly index and the facial expression variation index are converted into a comprehensive feature vector, and the comprehensive feature vector is used as the input of the machine learning model. The machine learning model predicts the authenticity value label of the screened highlights with each group of comprehensive feature vectors as the prediction target, and minimizes the sum of the prediction errors of the authenticity value labels of all the screened highlights as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence, and the model training is stopped. The authenticity value of the screened highlights is determined according to the model output results, wherein the machine learning model is a polynomial regression model; From the comprehensive feature vector training data of the trained machine learning model, the corresponding function expression is obtained: LR=F(SRAI, FEVI); where F is the output function of the model, SRAI is the speech recognition anomaly index, FEVI is the facial expression variation index, and LR is the authenticity value of the screened highlights.
7. The method for generating highlights in a business conference scenario based on artificial intelligence according to claim 6 is characterized in that: The authenticity value of the obtained screened footage is compared with the authenticity value reference threshold pre-set according to historical data. If the authenticity value of the obtained screened footage is greater than or equal to the pre-set authenticity value reference threshold, it means that there is no forged content generated by Deepfake or GANs, the authenticity of the footage is high, and the corresponding footage is marked as a low-risk footage; if the authenticity value of the obtained screened footage is less than the pre-set authenticity value reference threshold, it means that there is forged content generated by Deepfake or GANs, the authenticity of the footage is low, and the corresponding footage is marked as a high-risk footage; low-risk footage automatically passes the verification and is stored in the database; high-risk footage is submitted for manual review to further confirm the authenticity.
8. The method for generating highlights in a business conference scenario based on artificial intelligence according to claim 7 is characterized in that: In S5, the authenticity of the optimized highlights was further verified by the long short-term memory network model, specifically: The input sequence for the LSTM model is constructed by extracting the speech recognition anomaly index and facial expression variation index from the highlights; LSTM is used to process time series data and predict the video authenticity score at each time step. LSTM processes the input features of time step s, calculates the hidden state and authenticity score, and the model outputs the authenticity prediction value T for each time step s. s : Take the average authenticity score of all time steps s of the entire behind-the-scenes clip as the overall authenticity score of the behind-the-scenes clip, with a value range of [0,1].
9. The method for generating highlights in a business conference scenario based on artificial intelligence according to claim 8, characterized in that: Based on the authenticity score output by the LSTM model, the permission level of the behind-the-scenes clip is determined and the authenticity threshold is set; if the authenticity score output by the LSTM model is greater than or equal to the set authenticity threshold, it means that the video is highly authentic, the audio and video are synchronized, and there are no obvious signs of tampering. It will be marked as a low-risk authenticity clip and allowed to be automatically published or stored in the database; if the authenticity score output by the LSTM model is less than the set authenticity threshold, it means that the video contains fake faces, voice tampering or editing errors, and it will be marked as a high-risk non-authentic clip and require manual review.
Citation Information
Patent Citations
Method for intelligently generating splendid moment video of children in juvenile garden based on artificial intelligence
CN113792694A
Deepfake video detection method and system and readable storage medium
CN115661725A