Live broadcast data processing method and device, and storage medium
By analyzing and feature extraction of multimodal live broadcast data of the target game, combined with the clip detection model, the accurate and efficient identification and generation of highlight live video clips are achieved, and the problem of low manual screening efficiency is solved.
Patent Information
- Application Number
- CN202510452330.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the screening of highlight live video clips mainly relies on manual labor, is inefficient and cannot meet the actual promotion needs.
By determining the multimodal live broadcast data of the target game, data analysis is performed to obtain the initial characteristics of multiple modalities, then feature analysis is performed to obtain the target characteristics, and fragment detection is performed using the fragment detection model to generate the target live video clip.
It realizes accurate and efficient identification and generation of highlight live video clips, avoids the inefficiency of manual screening, and meets the needs of promotion and publicity.
Smart Images

Figure CN119996721A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to a method for processing live data. One or more embodiments of this specification also relate to a live data processing device and a computer-readable storage medium. Background Art
[0002] With the continuous development of Internet live streaming, game live streaming has become a more mainstream live streaming content. In order to increase the influence of game live streaming content, you can select highlight live video clips from game live streaming videos for promotion.
[0003] Currently, highlight live video clips are mostly determined from game live videos by manual screening. However, the efficiency of manual screening is low and cannot meet the needs of actual promotion and publicity. Therefore, how to accurately and efficiently determine highlight live video clips has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] In view of this, an embodiment of this specification provides a live data processing method. One or more embodiments of this specification also relate to a live data processing device, a computer-readable storage medium and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a live broadcast data processing method is provided, including: Determine multimodal live broadcast data of a target game, and perform data analysis on the multimodal live broadcast data to obtain initial features of multiple modes, wherein the multimodal live broadcast data includes target game data, a game live broadcast video, and live broadcast interaction data associated with the game live broadcast video; Performing feature analysis on the initial features of the multiple modes to obtain target features of the multiple modes; Using a segment detection model to perform segment detection on target features of the multiple modalities to obtain target video segment parameters; Based on the target video segment parameters and the game live video, a target live video segment corresponding to the target game is generated.
[0006] According to a second aspect of an embodiment of this specification, a live broadcast data processing device is provided, including: A data determination module is configured to determine multimodal live broadcast data of a target game, and perform data analysis on the multimodal live broadcast data to obtain initial features of multiple modes, wherein the multimodal live broadcast data includes target game data, a game live broadcast video, and live broadcast interaction data associated with the game live broadcast video; A feature analysis module is configured to perform feature analysis on the initial features of the multiple modes to obtain target features of the multiple modes; A segment detection module is configured to perform segment detection on target features of the multiple modalities using a segment detection model to obtain target video segment parameters; The segment generation module is configured to generate a target live video segment corresponding to the target game based on the target video segment parameters and the game live video.
[0007] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the above-mentioned live data processing method are implemented.
[0008] According to a fourth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which implements the steps of the above-mentioned live data processing method when executed by a processor.
[0009] One or more embodiments of the present specification provide a live data processing method. First, data analysis can be performed on multimodal live data including target game data, game live video, and live interactive data associated with the game live video to obtain initial features of multiple modalities. Secondly, a more detailed in-depth analysis can be performed on the initial features of multiple modalities to obtain relatively accurate target features of multiple modalities, thereby improving the accuracy of the generated target live video clips. Finally, a clip detection model is used to efficiently and accurately identify highlight clips of target features of multiple modalities to obtain target video clip parameters. Based on the target video clip parameters and the game live video, a target live video clip corresponding to the target game is generated, thereby achieving accurate and efficient determination of the target live video clip and avoiding the problem that the efficiency of manual screening is low and cannot meet the needs of actual promotion and publicity. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is an application diagram of a live broadcast data processing method provided by an embodiment of this specification; Figure 2 is a flow chart of a live broadcast data processing method provided by an embodiment of this specification; Figure 3 It is a processing flow chart of an AI analysis layer architecture in a live broadcast data processing method provided by an embodiment of this specification; Figure 4 It is a system architecture diagram of a live broadcast data processing method provided by an embodiment of this specification; Figure 5is a processing flow chart of a live broadcast data processing method provided by an embodiment of this specification; Figure 6 This is an architecture diagram of an AI analysis layer in a live data processing method provided by an embodiment of this specification; Figure 7 It is a structural diagram of a live data processing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0011] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0012] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0013] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0014] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0015] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.
[0016] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0017] First, the terms involved in one or more embodiments of this specification are explained.
[0018] Multimodal Analysis: An analysis method that can simultaneously process and fuse multiple different forms of data. In this method, it mainly processes data in multiple modes such as video, audio, and text, and extracts features and fusion analysis through deep learning models, thereby achieving a comprehensive understanding of the live broadcast content.
[0019] Automatic Speech Recognition (ASR): is a technology that automatically converts human speech into text. This method can use ASR technology to convert the host's speech into text subtitles in real time.
[0020] PR editing script: An automated editing script based on Adobe Premiere software (a video editing software) that can automatically complete video editing according to preset parameters and templates. This method uses this technology to achieve automatic editing of highlight clips.
[0021] Freeze Frame Effect: A processing technique that selects a specific frame in a video and pauses it briefly, usually used to emphasize important scenes. This method automatically determines the freeze frame time point and duration through AI analysis. Dot log: A log format that records key system events and data in chronological order. This method includes the correspondence between multi-dimensional information such as video timestamps, game events, and barrage density. Large Language Model (LLM): is an advanced natural language processing system based on deep learning technology. Through training on massive text data, it has acquired powerful language understanding and generation capabilities. In this method, the large language model is mainly used to conduct comprehensive analysis and decision-making on multimodal feature data, and accurately identify the highlights in the live content by understanding the correlation between different features. The large language model based on the Transformer architecture has excellent long sequence processing capabilities and context understanding capabilities, which enables it to effectively handle complex scene analysis tasks.
[0022] AE presets: Effect preset files in Adobe After Effects software, which contain parameter settings for special effects such as animations and transitions. This method uses these presets to achieve standardized special effects processing for videos.
[0023] VLM (Vision Language Model): A visual language model is an AI model that can understand both image and text information. In this method, it is mainly used to understand the content in the video, including the analysis of static features such as the host's expression and game screen.
[0024] Deep Feature Extraction: The process of extracting high-dimensional features from raw data using a deep learning model. These features contain the essential characteristics of the data and can better express the inherent laws of the data. In this system, it is used to extract valuable feature information from raw data such as video and audio.
[0025] Near Real-time Processing: A processing method between real-time processing and offline processing, which allows the system to have a small delay in exchange for better processing results. The deep analysis layer of this system adopts this processing mechanism to ensure the quality of analysis while meeting the timeliness requirements.
[0026] Speech Emotion Recognition (SER): A technology that recognizes and analyzes the emotional information contained in speech. The system uses this technology to analyze the host's speech and identify the emotional features contained in it, such as excitement, surprise, etc.
[0027] Pipeline Processing Architecture: A system architecture that breaks down complex tasks into multiple sequential processing stages. In this system, data is processed in the order of basic feature extraction, deep analysis, and decision analysis to form a complete processing pipeline.
[0028] Freeze Frame Text: It is a common post-processing technique in the process of short video packaging, mainly used to enhance content communication or create visual effects through text information. Freeze Frame Text refers to inserting a short still picture (usually 0.5-3 seconds) into a dynamic video and superimposing text content on the picture. This technology combines the two techniques of Freeze Frame and Typography.
[0029] PR project file: refers to a file that can be used for video editing in Adobe Premiere Pro (PR for short), which saves all settings and operations in the editing process. PR project files can contain multiple media elements such as video, audio, subtitles, special effects, as well as editing operations such as timeline, editing, effects, transitions, etc.
[0030] With the continuous development of Internet live broadcasting, game live broadcasting has become a more mainstream live broadcasting content. In order to increase the influence of game live broadcasting content, highlight live broadcasting video clips can be selected from game live broadcasting videos for promotion. Currently, highlight live broadcasting video clips are mostly selected from game live broadcasting videos by manual screening, but the efficiency of manual screening is low and cannot meet the actual promotion needs.
[0031] For example, game live broadcasting has become an important content marketing channel. During the live broadcasting process, a large number of exciting content clips with dissemination value are often generated. These contents can not only attract new users, but also enhance brand influence. The distribution department recommends games to users by cooperating with well-known anchors. However, due to the long duration and fixed time of live broadcasting, the user group that can watch in real time is relatively limited. In order to expand the dissemination range of content, it is necessary to edit the live broadcast content into slices for secondary dissemination on major platforms, which can not only break through the time and scene restrictions, but also significantly increase the exposure opportunities of game content. In the process of identifying and editing the highlights of live broadcasting, there are the following outstanding problems: 1. The efficiency of slice production is low; specifically, it includes: the need for operators to follow the broadcast to record the live broadcast content and analyze the exciting moments; the editor manually finds the marked exciting moments according to the operation records and captures the key pictures; subtitles, titles and other materials need to be manually written, which is time-consuming and labor-intensive, and the cross-departmental collaboration process is lengthy, resulting in reduced content timeliness. 2. The cost of content processing is high; specifically, it includes: professional editors are required to perform a large amount of repetitive post-processing work; the manual editing cycle is long, which is difficult to meet the needs of fast-iteration short video marketing; the human cost investment is large, and it is difficult to support large-scale content distribution on multiple platforms. 3. It is difficult to control the content quality; specifically, there are no unified subjective judgment standards, resulting in uneven content quality, lack of a systematic content evaluation mechanism, and difficulty in establishing content screening standards suitable for short video dissemination.
[0032] In addition, current live content automatic editing products are mainly used in live streaming scenarios, and such products have obvious deficiencies in video processing. First, in terms of content integrity, the system often cuts before the host finishes speaking, resulting in content being taken out of context. Secondly, video packaging lacks a systematic strategy, and only adopts simple processing methods such as randomly inserting emoticons, and the generated titles and tags are also relatively mechanical. In addition, since only the final video file is output without retaining editable engineering files, professional editors cannot perform post-optimization. Moreover, in the field of game live streaming, the existing live content processing system relies too much on game screen analysis, ignoring multi-dimensional information such as host performance and audience interaction. This single analysis method is particularly insufficient when dealing with chess and card strategic games, and cannot accurately identify the exciting decision-making moments in the game and the key commentary of the host, resulting in the lack of viewing value of the generated content.
[0033] Based on this, in this specification, a live broadcast data processing method is provided. One or more embodiments of this specification also involve a live broadcast data processing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0034] Considering the huge number of model parameters of the large model and the limited computing resources of the mobile terminal, the live broadcast data processing method provided in the embodiment of the present application can be applied to Figure 1 The application scenarios shown are not limited to these. Figure 1 In the application scenario shown, the large model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client device 20 here may include but is not limited to: smart phones, tablet computers, laptops, PDAs, personal computers, smart home devices, vehicle-mounted devices, etc. The client device 20 can interact with the user through a graphical user interface to implement the call of the large model, thereby implementing the method provided in the embodiments of this specification.
[0035] In the embodiment of the present specification, the system composed of the client device 20 and the server 10 can perform the following steps: the client device 20 is a device for executing game live broadcast, and the image user interface displays the game live broadcast screen of the game live broadcast; the server 10 performs the steps of basic feature extraction, deep feature extraction, highlight segment prediction and generation of highlight video segments; wherein, basic feature extraction refers to: extracting basic features from the collected game live broadcast video, audience interaction data stream, game event data stream, anchor commentary text stream and audio change data stream to obtain basic feature vectors (i.e., initial vectors of multiple modes); deep feature extraction refers to: extracting deep features from the basic feature vectors to obtain deep feature vectors (i.e., target vectors of multiple modes); highlight segment prediction refers to: comprehensively analyzing the deep feature vectors, identifying highlight segments, and outputting the editing configuration information of the highlight segments (i.e., target video segment parameters); generating highlight video segments refers to: generating highlight video segments (i.e., target live video segments) according to the editing configuration information and the game live broadcast video. This significantly improves the conversion efficiency of live broadcast content to short videos and expands the dissemination range and influence of game live broadcast content.
[0036] See also Figure 2 , Figure 2 A flowchart of a live broadcast data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0037] Step 202: Determine the multimodal live broadcast data of the target game, and perform data analysis on the multimodal live broadcast data to obtain initial features of multiple modalities, wherein the multimodal live broadcast data includes target game data, game live broadcast video, and live broadcast interaction data associated with the game live broadcast video.
[0038] Among them, the target game can be understood as an online game, or a real-person competitive game in reality; the online game includes but is not limited to online chess and card games, role-playing games, etc., and the real-person competitive game includes but is not limited to mahjong, go, chess, etc. Multimodal live broadcast data can be understood as multimodal data collected during the live broadcast of the target game. The target game data can be understood as game data generated during the target game, such as game scores, game coins, game critical events, etc.; it should be noted that the target game data can be obtained by analyzing the live game video screen, for example, by performing data analysis on the live game video screen through a neural network model to obtain the target game data. The game live video can be understood as the game live video collected during the live broadcast of the target game. The live interactive data associated with the game live video can be understood as interactive data generated during the target game, such as interactive data between the game anchor and the audience, or interactive data generated by users during the game live broadcast; the live interactive data is not limited to: barrage data, gift giving data, live likes or comments data, etc. The initial features of multiple modalities can be understood as features obtained by extracting features from live broadcast data of multiple modalities; it should be noted that the initial features of the multiple modalities can be obtained through real-time processing (such as real-time feature extraction operations).
[0039] In one or more embodiments provided in the present specification, the method of determining multimodal live broadcast data of a target game and performing data analysis on the multimodal live broadcast data to obtain initial features of multiple modalities includes: determining multiple live broadcast data sources, wherein the multiple live broadcast data sources include a game server, a live broadcast video storage unit, and an interactive data storage unit; obtaining the target game data of the target game from the game server, obtaining the game live broadcast video of the target game from the live broadcast video storage unit, and obtaining the live broadcast interactive data from the interactive data storage unit; performing data analysis on the target game data, the game live broadcast video, and the live broadcast interactive data to obtain initial features of the multiple modalities.
[0040] Among them, the multiple live data sources can be understood as data sources for storing live data of multiple modes; the live video storage unit can be understood as a device for storing live game videos, for example, the live video storage unit can be a server, a database, etc.; the interactive data storage unit can be understood as a unit for storing live interactive data, for example, the interactive data storage unit can be a live server, a live platform, a database, etc. Among them, the live game video is obtained by collecting video of the live game of the target game in a real-time data collection manner, the live interactive data is obtained by collecting interactive data of the live game in a real-time data collection manner, and the target game data is obtained by querying the game server according to the time tag of the live game video. Among them, the time tag can be understood as a tag used to indicate the live broadcast time corresponding to the live game video. For example, the time tag can be the time information, recording time, acquisition time, etc. of the live game video.
[0041] Taking the application of the live broadcast data processing method provided in this specification in the scenario of live broadcast highlight intelligent editing as an example, the live broadcast data processing method is explained, wherein the live broadcast data processing method in this specification is applied to an intelligent processing system, which is an intelligent processing system from content discovery to automatic editing (hereinafter referred to as the system), and the live broadcast data processing method can be applied to the intelligent processing system for live video processing. Based on this, in the process of data collection, this method can first obtain the game live broadcast screen recording (i.e., game live broadcast video) from the server that stores the game live broadcast video, and obtain the audience interaction information (i.e., live broadcast interaction data) from the live broadcast platform. Specifically: When the anchor starts the game live broadcast, the system will automatically start two parallel tasks.
[0042] 1. Automatic screen recording: used to capture the complete live broadcast screen and audio (i.e., live game voice), and upload the recorded content to the server every hour. It should be noted that this method can mark the collected screen recording data, which increases the availability and analyzability of the data. For example, this system can mark the host ID and time information (i.e., time tag) in the recorded video, providing a basis for subsequent game data association. Alternatively, mark the time information of the start of recording and the host ID in the file name of the screen recording. The file name of the screen recording contains the time information of the start of recording, which is convenient for subsequent query of game data within a specific time period according to the timestamp (such as the start and end time of the game, gold flow changes, etc.); and these fixed game data can be obtained through SQL queries. 2. Audience information capture: Real-time collection of audience interaction information in the live broadcast room, which includes but is not limited to audience behaviors such as barrage and gift giving.
[0043] In addition to the two parallel tasks mentioned above, this system can obtain game data (i.e., target game data) from the game server, such as gold flow and critical hit data obtained from the game server. The way to obtain game data can be to query the game server through SQL to obtain game data, such as querying the background database according to the marked information (such as the anchor ID, recording time information) to obtain the corresponding specific game data, such as gold flow changes, critical hit status, etc.
[0044] After obtaining the game live screen recording, audience interaction information, and game data, this system can perform preliminary feature extraction on these multimodal live data, obtain basic feature vectors of multiple modalities, and provide a data basis for subsequent analysis.
[0045] Based on the above embodiments, it can be seen that the method automatically starts the screen recording function during live broadcast, and can capture the complete live broadcast screen and audio without manual intervention. This not only improves the efficiency of data collection, but also ensures the integrity of the data. The design of automatically uploading the recorded content to the server every hour effectively avoids the risk of data loss caused by long-term recording, and simplifies the subsequent query operation by including the time information of the start recording in the file name. In addition, the system can collect the audience interaction information in the live broadcast room in real time, which is convenient for accurately determining the highlight fragments according to the audience interaction. At the same time, the system can also obtain game data within a specific time period from the game server through SQL queries and other methods. This method ensures the accuracy and comprehensiveness of the data and provides a solid foundation for subsequent in-depth analysis. In other words, the system obtains multimodal data from multiple data sources through automated means, improves the data collection efficiency of live content, and uses multi-source data integration to enhance the accuracy and comprehensiveness of data analysis, ultimately providing strong support for improving user experience, optimizing service quality and increasing commercial value.
[0046] In one or more embodiments provided in the present specification, the target game data, the game live video and the live interactive data are subjected to data analysis to obtain the initial features of the multiple modalities, including: determining the game live voice of the game live video, and converting the game live voice into game live voice text; time-aligning the game live video, the game live voice, the game live voice text, the live interactive data and the target game data according to the video time information of the game live video, the interaction time information of the live interactive data and the game time information of the target game data to obtain aligned multimodal live data; performing feature extraction on the aligned multimodal live data to obtain initial game screen features, initial live video features, initial live voice features, initial voice text features and initial interactive data features; determining the initial game screen features, the initial live video features, the initial live voice features, the initial voice text features and the initial interactive data features as the initial features of the multiple modalities.
[0047] Among them, the initial game screen features can be understood as the features obtained after feature extraction of the game screen in the game live video; the initial live video features can be understood as the features obtained by feature extraction of the game live video; the initial live voice features can be understood as the features obtained by feature extraction of the audio data corresponding to the game live video; the initial voice text features can be understood as the features obtained by text feature extraction of the game live voice text; the initial interactive data features can be understood as the features obtained by feature extraction of the live interactive data. It should be noted that the operation of feature extraction of the multimodal live data of game live video, game live voice, game live voice text, live interactive data and target game data can be implemented by a neural network model, a deep learning model, or an encoder. Among them, the aligned multimodal live data can be understood as a file that records game live events and data in chronological order. For example, the aligned multimodal live data can be a dot log.
[0048] Specifically, the present application can synchronously record the game live voice during the process of recording the game live video, and convert the game live voice into the game live voice text during the data analysis process. For example, the game live voice can be converted into the game live voice text by using voice recognition technology. After obtaining the game live video, game live voice, game live voice text, live interactive data and target game data, in order to ensure the accuracy of the feature extraction operation, it is necessary to time-align the multiple data to ensure that the multiple data can be aligned in time order to ensure the consistency of the data. After obtaining the aligned multimodal live data, feature extraction can be performed on the aligned multimodal live data to obtain initial game screen features, initial live video features, initial live voice features, initial voice text features, and initial interactive data features.
[0049] Using the above example, after the data collection operation, this system can convert the data collected from multiple data sources, increasing the availability and analyzability of the data. Specifically, this system uses speech recognition technology to convert the host's voice (i.e., game live broadcast voice) into text subtitles (i.e., game live broadcast voice text) to facilitate information retrieval and analysis in text form.
[0050] After the above conversion operation is completed, this method can deeply integrate and synchronize all processed data, and finally form a highly valuable dot log, which provides a basis for subsequent content analysis, editing and prediction of highlight videos. Specifically, the dot log generation includes the following two steps: 1. Data synchronization: Determine the time axis (i.e., video time information, interaction time information, and game time information) corresponding to video data, audio data, audience interaction data (such as bullet screen), subtitle text, and game data (such as game event data); and synchronize the video data, audio data, audience interaction data, subtitle text, and game event data obtained from the game server according to the time axis (i.e., time alignment), and obtain the video data, audio data, audience interaction data, subtitle text, and game data after time synchronization. It should be noted that due to the variety of data sources (screen recording, platform interface, OCR recognition, etc.), there are certain deviations, so a certain degree of data alignment is required.
[0051] 2. Generate a dot log: Based on the synchronized data above, generate a dot log containing complete information. These logs not only contain the key moments in the video, but also combine the audience's interactive information and the events that occurred in the game, providing solid data support for subsequent highlight moment detection.
[0052] After obtaining the dot log, the synchronized audience interaction data stream, game event data stream, anchor commentary text stream (i.e. subtitle text) and audio change data stream (i.e. audio data) are processed for preliminary feature extraction to generate basic feature vectors, providing a data basis for subsequent analysis.
[0053] Based on the above embodiments, it can be seen that this method ensures that data from different sources can be accurately aligned in the time dimension by corresponding video data, audio data, audience interaction data, subtitle text and game data to a unified timeline. This process not only needs to process the time information of the video and audio, but also needs to consider the timestamp of the audience interaction data and the time stamp of the game event data obtained from the game server. Since the data comes from multiple data sources, it is inevitable that there will be certain deviations in the time synchronization process. In order to overcome these deviations, the system adopts a timeline alignment method to arrange the data from different sources as accurately as possible in the order of their actual occurrence. This provides an accurate and coherent data basis for subsequent in-depth analysis.
[0054] Step 204: Perform feature analysis on the initial features of the multiple modes to obtain target features of the multiple modes.
[0055] Among them, the target features of multiple modes can be understood as the features obtained by performing feature extraction and multi-dimensional analysis on the initial features of multiple modes. It should be noted that the target features of multiple modes can be obtained through quasi-real-time processing (for example, performing feature extraction operations in a quasi-real-time manner).
[0056] Specifically, after obtaining the initial features of multiple modalities, the method can perform deep feature extraction on the initial features of multiple modalities, thereby extracting valuable target features of multiple modalities.
[0057] In one or more embodiments provided in this specification, the feature analysis of the initial features of the multiple modalities to obtain the target features of the multiple modalities includes: determining the feature analysis modules corresponding to the initial features of the multiple modalities, wherein one feature analysis module corresponds to the initial features of at least one modality; using multiple feature analysis modules, the feature analysis of the initial features of the multiple modalities is performed to obtain the target features of the multiple modalities. The feature analysis module can be understood as a functional module for performing deep feature extraction on the initial features of different modalities; the feature analysis module can be integrated with a deep learning model or a large model for deep feature extraction of the initial features.
[0058] Continuing with the above example, the live broadcast data processing method provided in this specification proposes an AI analysis layer architecture based on multimodal analysis. The architecture adopts a hierarchical and progressive technical route to achieve intelligent processing of live broadcast content through multi-level feature extraction and analysis. The technical implementation and processing flow of the architecture are described in detail below. The AI analysis layer adopts a three-level architecture design, including a basic feature layer, a deep analysis layer, and a large language model analysis layer. Through different processing timings and analysis depths, this architecture realizes a complete technical link from data acquisition to highlight segment recognition. The processing flow of the AI analysis layer architecture can be referred to Figure 3 , Figure 3 It is a processing flow chart of an AI analysis layer architecture in a live broadcast data processing method provided in an embodiment of this specification.
[0059] based on Figure 3 It can be seen that before the AI analysis layer architecture performs analysis, real-time data collection is required to obtain live broadcast related data (i.e. multi-modal live broadcast data); then, the basic feature layer, as the data acquisition and primary processing unit of the system, can perform basic feature extraction on the collected data. Specifically, after simultaneously analyzing multiple data sources through a parallel processing mechanism and obtaining audience interaction data streams, game event data streams, anchor commentary text streams, and audio change data streams, these data can be processed through preliminary feature extraction to generate basic feature vectors (i.e., initial features of multiple modalities).
[0060] Then, the deep analysis layer is used to extract deep features from the basic feature vectors. The deep analysis layer can adopt a quasi-real-time processing mechanism and integrate multiple professional models for deep feature extraction. The deep analysis layer can include multiple functional modules (i.e., feature analysis modules), and each functional module can perform deep feature extraction and multi-dimensional analysis on the basic feature vectors of the corresponding modality through a parallel computing architecture to obtain valuable features (i.e., target features of multiple modalities). Subsequently, the large language model (LLM) analysis layer can analyze the valuable features and select highlight clips from the live game video.
[0061] Based on the above embodiments, it can be seen that in the process of multimodal analysis, this method provides an AI analysis layer architecture including a basic feature layer, a deep analysis layer and a large language model analysis layer; the basic feature layer, as a data acquisition and primary processing layer, can simultaneously analyze data streams from multiple data sources through an efficient parallel processing mechanism. These data sources include audience interaction data, game event data, anchor commentary text, and audio change data. Through this design, the system can perform a preliminary understanding and feature extraction of the live content at the first time. At the same time, during the preliminary feature extraction process, the system converts the above-mentioned multiple types of data into their respective basic feature vectors. This step not only simplifies the complexity of the original data, but also provides structured input for subsequent deep analysis. Each modality (video, audio, text, etc.) has its corresponding basic feature vector, which allows different types of raw data to be processed and analyzed in the same framework. The deep analysis layer adopts a quasi-real-time processing mechanism, which means that the system can perform in-depth analysis of the live content with almost no delay. This is crucial for capturing dynamic changes during the live broadcast process, such as timely identification of highlight moments or changes in audience emotions. The deep analysis layer integrates multiple professional models for different modalities to perform deep feature extraction tasks. Each functional module focuses on a specific type of data (such as video, audio, text), and improves processing speed and efficiency through a parallel computing architecture. This design allows the system to extract more valuable feature representations from the basic feature vectors, thereby more accurately reflecting the key information and internal patterns of the live content, making it easier for the subsequent large language model (LLM) to further analyze the live content and select highlight clips.
[0062] In one or more embodiments provided in this specification, the multiple feature analysis modules include an image analysis module, a video analysis module, a text analysis module, an interactive data analysis module, and a voice analysis module, and the initial features of the multiple modalities include initial game screen features, initial live video features, initial live voice features, initial voice text features, and initial interactive data features; The method uses multiple feature analysis modules to perform feature analysis on the initial features of the multiple modalities to obtain target features of the multiple modalities, including: using the image analysis module, the video analysis module, the text analysis module, the interactive data analysis module and the voice analysis module to perform deep feature extraction on the initial game screen features, the initial live video features, the initial voice text features, the initial interactive data features and the initial live voice features to obtain target game screen features, target live video features, target text emotion features, target interactive data features and target voice emotion features; determining the target game screen features, the target live video features, the target voice emotion features, the target text emotion features and the target interactive data features as the target features of the multiple modalities.
[0063] Specifically, the method can use the image analysis module to perform screen content analysis on the initial game screen features to obtain target game screen features; use the video analysis module to perform video feature analysis on the initial live video features to obtain target live video features; use the voice analysis module to perform voice emotion analysis on the initial live voice features to obtain target voice emotion features; use the text analysis module to perform text emotion analysis on the initial voice text features to obtain target text emotion features; use the interactive data analysis module to perform interactive data analysis on the initial interactive data features to obtain target interactive data features; determine the target game screen features, the target live video features, the target voice emotion features, the target text emotion features and the target interactive data features as the target features of the multiple modalities.
[0064] Among them, the image analysis module can be understood as a functional module for analyzing and processing the game screen. For example, the image analysis module can be a VLM-based screen content understanding module, which can use the Visual Language Model (VLM) to understand and interpret the content of the video screen; the VLM can link the image or video frame with the corresponding text description to achieve a high-level understanding of the screen content; the VLM-based screen content understanding module can be used to identify and classify objects, scenes and activities in the game live video. For example, in a game live broadcast, it can identify game characters, action sequences or specific game events to provide a basis for subsequent analysis.
[0065] The video analysis module can be understood as a module for analyzing dynamic features in live game videos. For example, the video analysis module can be a video dynamic feature analysis module based on deep learning; the video dynamic feature analysis module based on deep learning can be used to extract dynamic features (that is, features of the video that change over time) from live game video data. This usually involves the use of deep learning models such as convolutional neural networks (CNNs), recurrent neural networks (RNNs) or their variants such as LSTM, GRU, etc. The video dynamic feature analysis module based on deep learning can capture dynamic information such as movement patterns and behavior trajectories in live game videos. For live game broadcasts, it can help analyze important moments such as key actions and tactical changes during the game, and provide a basis for the selection of highlight clips.
[0066] The text analysis module can be understood as a module used to perform text sentiment analysis on the initial speech text features. For example, the text analysis module can be an NLP-based text sentiment analysis module; the NLP-based text sentiment analysis module can use natural language processing (NLP) technology to analyze the emotional tendency of the subtitle text; the NLP-based text sentiment analysis module can understand the emotional reaction of the anchor during the live broadcast, such as whether it is positive or negative, through sentiment analysis.
[0067] The interactive data analysis module can be understood as a module that performs statistical analysis on the initial interactive data features. The interactive data analysis module can be understood as an interactive data analysis module based on statistical learning. The interactive data analysis module based on statistical learning can use statistical learning methods to analyze interactive data from the audience, including but not limited to likes, gift giving and other behaviors; by quantitatively analyzing these interactive data, it can reveal which content is more popular with the audience, thereby helping to select highlight clips.
[0068] The speech analysis module can be understood as a module that performs speech emotion analysis based on the initial live speech features. For example, the language analysis module can be understood as a speech emotion recognition module based on a deep neural network. The speech emotion recognition module based on a deep neural network uses deep neural network (DNNs) technology to identify the emotional state in human speech. By analyzing the host's voice characteristics, such as intonation and volume changes, the emotions expressed by the host during the live broadcast, such as excitement and nervousness, can be judged. This is very helpful in selecting highlight clips.
[0069] It should be noted that the above-mentioned multiple functional modules all achieve efficient data processing through parallel computing architecture. This means that different data types can be processed simultaneously in their respective modules, greatly improving the processing speed and efficiency of the entire system, so that even in the face of large and complex data flows, it can maintain a fast response.
[0070] Continuing with the above example, in the process of using the AI analysis layer architecture for multimodal analysis, the deep analysis layer can be used to perform deep feature extraction on the basic feature vector. The deep analysis layer can include multiple functional modules, namely: a picture content understanding module based on VLM, a video dynamic feature analysis module based on deep learning, a text sentiment analysis module based on NLP, an interactive data analysis module based on statistical learning, and a speech emotion recognition module based on deep neural network.
[0071] Among them, the VLM-based screen content understanding module can use the visual language model (VLM) to perform deep feature extraction on the game screen feature vector (i.e., the initial game screen feature) extracted from the game live screen by the basic feature layer, and obtain the target game screen feature vector that can represent important information such as game characters, action sequences, or specific game events; The video dynamic feature analysis module based on deep learning can perform deep feature extraction on the live video feature vectors (i.e., initial live video features) extracted from the basic feature layer based on the deep learning module, and capture dynamic feature vectors (i.e., target live video features) such as motion patterns and behavior trajectories in the live game video.
[0072] The NLP-based text sentiment analysis module can use natural language processing technology to perform deep feature extraction on the subtitle text feature vector (i.e., initial speech text features) extracted from the basic feature layer, analyze the sentiment tendency of the subtitle text, and thus obtain a text sentiment feature vector (target text sentiment feature) that can represent the text sentiment.
[0073] The interactive data analysis module based on statistical learning can perform deep feature extraction on the audience interaction feature vector (i.e., initial interactive data features) extracted from the basic feature layer using statistical learning methods, and obtain quantitative interactive feature vectors (i.e., target interactive data features) by performing quantitative analysis on these interactive data.
[0074] The speech emotion recognition module based on deep neural network can use deep neural network technology to perform deep feature extraction on the sound feature data extracted by the basic feature layer (i.e., the initial live speech features), so as to extract the emotion feature vector (i.e., the target speech emotion features) expressed through speech during the live broadcast. It should be noted that the basic feature layer will collect basic sound feature data such as volume and speech speed changes, and the deep analysis layer can analyze the host's voice data and audio change data stream. This part of the analysis is based on the sound feature data detected by the basic feature layer. Multiple factors can be used to make the analysis effect better.
[0075] Based on the contents of the above embodiments, it can be seen that this method constructs a comprehensive and efficient deep analysis architecture by integrating a variety of advanced artificial intelligence technologies. Through the above multiple functional modules, the basic feature vectors are subjected to deep feature extraction and multi-dimensional analysis, so that the system can understand the core elements of the game screen content from a visual perspective, capture key actions and content in the live video, understand the emotional response of the anchor during the live broadcast, and reveal which content is more popular with the audience through quantitative analysis of these interactive data, so as to deeply explore the value of the live content from multiple dimensions, so that the subsequent large language model (LLM) analysis layer can analyze the valuable features and accurately select highlight clips from the live game video.
[0076] Step 206: Use the segment detection model to perform segment detection on the target features of the multiple modalities to obtain target video segment parameters.
[0077] Among them, the segment detection model can be understood as a model for detecting the target live segment in the game live video. For example, the segment detection model can be a large language model, a large model or a deep learning model, and no specific restrictions are made here. Among them, the target live segment can be a highlight live segment or a non-highlight live segment in the game live video; the highlight live segment can refer to an indicator segment in which the player performs well, triggers an important game plot or an exciting event occurs during the game live broadcast; for example, a live segment with a significant increase in the number of viewers can be a highlight live segment because the game content at this moment attracts a significant increase in the number of viewers; or, the interaction rate of interactive behaviors such as likes, comments, and shares is greater than the preset interaction rate threshold, which usually means that something exciting has happened, such as a wonderful kill or a reversal victory, and the live segment can be a highlight live segment. Alternatively, when the anchor has a violent reaction to a game event (such as an increase in voice or emotional excitement), it can also be a highlight live segment. Non-highlight live segments can be other live segments in the game live video that are not highlight live segments. Among them, the target video segment parameters can be understood as parameters used to generate a target live video segment, and the target live video segment includes a video segment display text and video segment time information.
[0078] In one or more embodiments provided in this specification, the segment detection model is a large language model; the use of the segment detection model to perform segment detection on the target features of the multiple modalities to obtain target video segment parameters includes: using the large language model to perform feature conversion on the target features of the multiple modalities to obtain live data features corresponding to the multimodal live data, and performing segment detection on the live data features to obtain the target video segment parameters. The feature conversion can be understood as an operation of converting target features of multiple dimensions into the same dimension, and the live data features can be understood as feature vectors under the same dimension.
[0079] Following the above example, the system adopts a pipeline processing architecture during the data flow process. After the basic features are extracted, the original data stream generates a feature vector. The data fragments with research value are transmitted to the deep analysis layer for feature enhancement. Finally, the large language model performs decision analysis and generates editing configuration information based on the target features after feature enhancement by the deep analysis layer. The editing configuration information includes the generated title, label, freeze time and / or freeze text. Specifically, first, the target game screen features, target live video features, target voice emotion features, target text emotion features and target interactive data features are input into the large language model, and the large language model is used to perform feature conversion on multiple features to obtain data features of the same dimension; for example, using the encoder in the large language model, multiple features are encoded into data features of the same dimension; or, through the feature conversion network layer configured in the large language model, multiple features are converted into data features of the same dimension. Then, the large language model performs decision analysis on the data features of the same dimension to generate editing configuration information.
[0080] Based on the above embodiments, it can be seen that this method not only ensures the real-time response capability of the system, but also ensures the accuracy of the analysis results through this progressive processing mechanism. In addition, the large language model (LLM) can further analyze the live content of the high-value features extracted by the deep analysis layer and accurately select the highlight clips from the live game video. The advantage of LLM lies in its powerful semantic understanding and context association capabilities, which enables it to accurately locate the most attractive highlight clips in a complex data environment.
[0081] In one or more embodiments provided in the present specification, the method of utilizing the large language model to perform feature conversion on the target features of the multiple modalities to obtain live data features corresponding to the multimodal live data, and performing segment detection on the live data features to obtain the target video segment parameters includes: inputting the target features of the multiple modalities into the large language model, mapping the target features of the multiple modalities to a semantic space in the large language model for feature conversion to obtain the live data features corresponding to the multimodal live data; and performing a multi-dimensional evaluation on the live data features to obtain the target video segment parameters, wherein the target video segment parameters include video segment display text and video segment time information.
[0082] The video clip display text can be understood as the text information that needs to be displayed in the target live video clip, and the video clip display text includes but is not limited to: video title, video tag and / or freeze frame text. The video clip time information can be understood as the time information or time range for cutting the video clip from the live game video, for example, the video clip time information can be the freeze frame time.
[0083] Continuing with the above example, the AI analysis layer architecture includes a large language model analysis layer. The large language model analysis layer serves as a decision-making unit and uses a post-processing mechanism to conduct a comprehensive analysis of the deep features (i.e., target features of multiple modalities) extracted by the deep analysis layer. Specifically, first, the large language model analysis layer uses a large language model to understand features, maps multimodal features (i.e., target features of multiple modalities) to a unified semantic space for feature conversion, and obtains multimodal features under the same data dimension; secondly, contextual analysis is performed on the multimodal features under the same data dimension to evaluate the temporal coherence of the content; finally, through a multi-dimensional scoring mechanism, highlight clips are identified and screened, and the editing configuration information of the highlight clips is output.
[0084] Based on the above embodiments, this method establishes an innovative two-layer multimodal analysis framework. The system uses lightweight algorithms in the basic feature layer to achieve real-time processing, including audience interaction analysis, game event processing, subtitle generation, and volume detection; in the deep analysis layer, it integrates advanced AI technologies such as VLM, video understanding model, and sentiment analysis to achieve quasi-real-time analysis; finally, through the large language model, feature fusion and decision analysis are performed to achieve efficient and accurate highlight segment recognition. This innovative layered architecture not only ensures the real-time responsiveness of the system, but also provides high-quality analysis results through deep learning models.
[0085] Through this technical solution, the system realizes intelligent processing of live broadcast content and can accurately identify and extract highlight clips with dissemination value. This technical route based on multimodal analysis significantly improves the efficiency and quality of secondary creation of live broadcast content. At the same time, the layered architecture design of the system also provides a good technical foundation for subsequent function expansion and performance optimization.
[0086] Step 208: Based on the target video segment parameters and the game live video, generate a target live video segment corresponding to the target game.
[0087] Among them, the target live video clip can be understood as the video corresponding to the highlight live clip and the wonderful live clip in the game live video; the target live video clip can be a highlight video, a slice video, a PR project file, etc.
[0088] In one or more embodiments provided in this specification, the target video segment parameters are video segment display text and video segment time information; the generating the target live video segment corresponding to the target game based on the target video segment parameters and the game live video includes: cutting the game live video based on the video segment time information to obtain a cut video segment; generating the target live video segment corresponding to the target game based on the video segment display text and the cut video segment. The cut video segment can be understood as a video segment cut from the game live video.
[0089] Continuing with the above example, this method uses PR software to cut the live game video based on the freeze-frame time output by the large language model to obtain a live game video clip corresponding to the freeze-frame time; then, the PR software is used to generate a highlight video based on the title, label and / or freeze-frame text and the live game video clip output by the large language model to obtain a highlight video or PR project file (i.e., the target live video clip) corresponding to the target game.
[0090] Based on the above embodiments, it can be seen that this method constructs a complete intelligent editing workflow; the system innovatively converts the AI analysis results directly into standardized editing parameters, which not only realizes the automatic cutting and packaging of videos, but also generates engineering files that can be further optimized by professional editors, truly opening up the entire process from content recognition to work generation.
[0091] In one or more embodiments provided in this specification, the generating the target live video clip corresponding to the target game based on the video clip display text and the cut video clip includes: determining a target special effect template from a plurality of candidate special effect templates, and performing video clip rendering based on the target special effect template, the video clip display text and the cut video clip, to obtain the target live video clip corresponding to the target game. The candidate special effect template can be understood as a special effect animation template required to generate a target live video clip, for example, the candidate special effect template can be an AE preset file; the target special effect template can be any one of a plurality of candidate feature templates.
[0092] Continuing with the above example, the template system in this method can provide professional material support, which includes special effects animation templates (AE preset files); in the process of using PR software to generate highlight videos based on video titles, video tags and / or freeze-frame text, and live game video clips, special effects animation templates can be selected from the materials, and based on the video titles, video tags and / or freeze-frame text, live game video clips, and special effects animation templates, highlight videos or PR project files (i.e., target live video clips) can be generated.
[0093] Based on the above embodiments, it can be seen that in terms of content integrity, the system uses multimodal analysis technology to fully understand the live content, and combines the text data of voice recognition for semantic analysis to ensure the integrity of video cutting. Through the deep learning model to analyze the game process data, anchor commentary and barrage interaction, the system can accurately identify the highlight moments of the content, effectively avoiding the problem of content out of context. In terms of video packaging quality, the system has established a scene-based freeze strategy. By analyzing the key nodes of the game, the changes in the anchor's expression and the peak of the audience's interaction, it intelligently determines the best freeze time and generates the corresponding copy (i.e., the video clip display text). At the same time, the system will also select appropriate packaging effects (such as special effects animation templates) from the professional template library, which significantly improves the video viewing experience. Especially in the processing of chess and card games, the system can accurately identify strategic highlights by analyzing the background data of the game, combining the anchor's commentary and the audience's reaction. The system also generates standardized PR engineering files, enabling professional editors to further optimize and improve the efficiency and flexibility of content creation. After actual testing, the system has achieved significant improvements in key indicators such as content integrity, packaging quality, and communication effect, fully demonstrating the system's technological innovation value in the field of intelligent editing of live content.
[0094] According to the contents of one or more of the above embodiments, in order to improve the dissemination effect of live broadcast content, the live broadcast data processing method in this specification provides an intelligent processing system, which is an intelligent processing system from content discovery to automatic editing. The live broadcast data processing method can be applied to the intelligent processing system to perform live broadcast clip editing. Figure 4 is a system architecture diagram of a live broadcast data processing method provided in an embodiment of this specification; based on Figure 4 It can be seen that the intelligent processing system includes: infrastructure layer, data collection layer, data processing layer, AI analysis layer, editing resource layer, content generation layer, and output layer. Among them, the infrastructure layer includes Win Server and AI analysis server; WinServer: as the basic server of the system, is responsible for the following functions: capturing live broadcast information and recording the screen, capturing live broadcast barrage interactive information; AI analysis server is the core computing unit of the system, and is responsible for the following functions: querying and analyzing data from the background in real time according to user ID and live broadcast time, and providing computing resources for multimodal analysis.
[0095] The functions of the data collection layer include: 1. Video stream collection: Real-time recording of live broadcast images; thereby ensuring the clarity of video quality, synchronization of audio and video, stability of the recording process, and ensuring that the anchor ID and recording time information are marked in the title of the recorded video (such as the file name). 2. Barrage collection: Obtain real-time interactive data in the live broadcast room; the interactive data includes: real-time comments from viewers, gift reward information, interactive heat data, etc. 3. Speech recognition: Convert the anchor's voice into text; the focus of this speech recognition is on: real-time conversion efficiency, Chinese recognition accuracy, dialect adaptability, speaker recognition, etc. 4. Game background log: Used to record game process data. The game background log can be a record of the game server during live broadcast. The game background log can include: game process status and corresponding timestamps, game cash flow data.
[0096] The data processing layer includes the original data processing module and the log processing module; the original data processing module is responsible for integrating various types of data, including mp4 video files, srt subtitle files, and ass bullet screen files. The mp4 video file is used to store the complete live broadcast content; the srt subtitle file is used to record the speech-to-text results; the ass bullet screen file is used to save the audience interaction information. The log processing module is used for data fusion and analysis, including dot log generation and information extraction; the dot log generation refers to integrating multi-source information, establishing time series associations, and generating easy-to-read dot logs; information extraction refers to screening key data to provide support for highlight detection.
[0097] The AI analysis layer can be understood as the AI analysis layer architecture in the above-mentioned embodiment, which includes a basic feature layer, a deep analysis layer and a large language model analysis layer; wherein the basic feature layer and the deep analysis layer constitute a two-layer multimodal analysis engine (i.e., a two-layer multimodal analysis framework); the large language model analysis layer is used for highlight moment detection and editing configuration information (i.e., target video clip parameters) generation. Specifically, the system uses a lightweight algorithm in the basic feature layer to achieve real-time processing, including audience interaction analysis, game event processing, subtitle generation, and volume detection; in the deep analysis layer, VLM, video understanding model, sentiment analysis and other advanced AI technologies are integrated to achieve quasi-real-time analysis; finally, feature fusion and decision analysis are performed through a large language model to achieve efficient and accurate highlight clip recognition. This layered architecture not only ensures the real-time responsiveness of the system, but also provides high-quality analysis results through a deep learning model.
[0098] The editing resource layer includes a template system and a configuration system; the template system is used to provide professional material support; the material includes: 1. PR project template: preset editing style; 2. AE preset file: special effect animation template; 3. Script library: automatic processing script. The configuration system is used to ensure output specifications. The configurations set by the configuration system include: 1. Freeze configuration: freeze timestamp and corresponding text; 2. Style configuration: select editing template; 3. Title configuration: standardized information presentation.
[0099] The content production layer includes an automatic editing module, a rendering queue, and output control. The automatic editing module is used to realize intelligent material selection, automatic special effects addition, and content rhythm control; the rendering queue is used for task priority management, resource scheduling optimization, and parallel rendering processing. The output control is used for quality monitoring, format specification, and error handling. Among them, the quality monitoring refers to the quality monitoring of the output highlight video or PR project file. The specific monitoring rules include: when the output highlight video is less than 5 seconds in length, the clarity is less than 360p, or the video resolution and frame rate do not meet the standard, it is determined that the highlight video does not meet the quality requirements. The error handling can be understood as handling errors in the highlight video, for example, analyzing the video metadata to observe whether there are problems such as missing data and encoding errors that cause the video to be unable to open.
[0100] The output layer is used to output the data produced by the content production layer, including but not limited to: 1. MP4 video: directly usable finished video; 2. PR project file: project file that can be further optimized; 3. Material library storage: use a unified resource management system to store the data produced by the content production layer.
[0101] Through the above-mentioned layered architecture, the intelligent processing system realizes full automation of live content processing and significantly improves content production efficiency. In particular, the innovative application of the AI analysis layer enables the system to accurately identify highlight moments and generate high-quality video content. At the same time, the modular design of the system also provides a good foundation for future functional expansion and optimization.
[0102] One or more embodiments of the present specification provide a live data processing method. First, data analysis can be performed on multimodal live data including target game data, game live video, and live interactive data associated with the game live video to obtain initial features of multiple modalities. Secondly, a more detailed in-depth analysis can be performed on the initial features of multiple modalities to obtain relatively accurate target features of multiple modalities, thereby improving the accuracy of the generated target live video clips. Finally, a clip detection model is used to efficiently and accurately identify highlight clips of target features of multiple modalities to obtain target video clip parameters. Based on the target video clip parameters and the game live video, a target live video clip corresponding to the target game is generated, thereby achieving accurate and efficient determination of the target live video clip and avoiding the problem that the efficiency of manual screening is low and cannot meet the needs of actual promotion and publicity.
[0103] The following combination Figure 5 , taking the application of the live broadcast data processing method provided in this specification in the live broadcast highlight intelligent editing scene as an example, the live broadcast data processing method is further explained. Figure 5 A processing flow chart of a live broadcast data processing method provided by an embodiment of the present specification is shown, which specifically includes the following steps.
[0104] Step 502: Data collection.
[0105] Specifically, when the anchor starts the game live broadcast, the system will automatically trigger two parallel tasks.
[0106] 1. Automatic screen recording: used to capture the complete live broadcast screen and audio (i.e. game live broadcast voice), and upload the recorded content to the server every hour. 2. Audience information capture: real-time collection of audience interaction information in the live broadcast room; the audience interaction information includes but is not limited to audience behaviors such as bullet comments and gift giving. Based on this, the audience information capture includes bullet comment capture or gift giving information capture.
[0107] Step 504: Data labeling and conversion.
[0108] Specifically, this method can process the collected raw data, and the specific processing method is as follows; 1. Mark the host ID and time information in the recorded video for subsequent data association.
[0109] This method can mark the collected screen recording data, increasing the availability and analyzability of the data; for example, the system can mark the host ID and time information in the recorded video, providing a basis for subsequent game data association. Alternatively, the time information of the start of recording and the host ID are marked in the file name of the screen recording. The file name of the screen recording contains the time information of the start of recording, which is convenient for subsequent query of game data within a specific time period according to the timestamp (such as the start and end time of the game, gold flow changes, etc.).
[0110] 2. Convert the host’s voice into text subtitles through speech recognition technology.
[0111] Speech recognition technology is used to convert the host's voice into text subtitles to facilitate information retrieval and analysis in text form.
[0112] 3. Query the background according to the marked information and obtain the corresponding game data.
[0113] Game data obtained from the game server, such as gold flow and critical hit data obtained from the game server, can be obtained by querying the game server through SQL, such as querying the backend database based on marked information (such as anchor ID, recording time information) to obtain the corresponding specific game data.
[0114] Step 506: Generate a log.
[0115] Specifically, this method integrates and processes all acquired data, and the specific processing method is as follows: 1. Synchronize the time of video data, barrage interaction data, subtitle text and game data.
[0116] Determine the timeline corresponding to the video data, audio data, audience interaction data, subtitle text, and game data.
[0117] The video data, audio data, audience interaction data, subtitle text, and game event data obtained from the game server are synchronized according to the time axis to obtain the video data, audio data, audience interaction data, subtitle text, and game data after time synchronization. It should be noted that due to the variety of data sources (screen recording, platform interface, OCR recognition, etc.), there are certain deviations, so a certain degree of data alignment is required.
[0118] 2. Generate a dot log containing complete information to provide a data basis for highlight detection.
[0119] Based on the synchronized data, a complete log is generated. These logs not only include the key moments in the video, but also combine the audience's interactive information and the events that occurred in the game, providing solid data support for subsequent highlight moment detection.
[0120] Step 508: Multimodal highlight detection.
[0121] Specifically, this method performs in-depth analysis on the integrated data through an AI analysis module, which can be understood as an AI analysis layer architecture. The AI analysis layer architecture can be used to comprehensively evaluate the integrated multi-dimensional data, intelligently identify highlight moments in the video, and generate video titles, labels, freeze times, text and other information based on the video model's (i.e., large language model) understanding of the video. Figure 6 yes Figure 6 This is an architecture diagram of the AI analysis layer in a live data processing method provided in an embodiment of this specification, based on Figure 6 It can be seen that the AI analysis layer architecture in this system includes the basic feature layer, the deep analysis layer, and the large language model analysis layer. This architecture realizes the complete technical link from data collection to highlight segment recognition through different processing timings and analysis depths.
[0122] Among them, the basic feature layer, as the data acquisition and primary processing unit of the system, can extract basic features from the collected data. Specifically, after simultaneously analyzing multiple data sources through a parallel processing mechanism to obtain audience interaction data streams, game event data streams, anchor commentary text streams (i.e. anchor commentary subtitles), and audio change data streams, these data can be processed through preliminary feature extraction to generate basic feature vectors. The audience interaction data stream is obtained by performing audience interaction analysis in this method, and the audio change data stream is obtained by detecting volume changes in the anchor audio in this method.
[0123] Among them, the deep analysis layer is used to perform deep feature extraction on the basic feature vector. The deep analysis layer can include multiple functional modules, namely: a picture content understanding module based on VLM, a video dynamic feature analysis module based on deep learning, a text sentiment analysis module based on NLP, an interactive data analysis module based on statistical learning, and a speech emotion recognition module based on deep neural network.
[0124] The VLM-based screen content understanding module can use VLM analysis to perform deep feature extraction on the game screen feature vector extracted from the game live screen by the basic feature layer, and obtain the target game screen feature vector that can represent important information such as game characters, action sequences or specific game events. The deep learning-based video dynamic feature analysis module can perform video understanding based on the deep learning module, thereby performing deep feature extraction on the live video feature vector extracted from the basic feature layer, and capturing dynamic feature vectors such as movement patterns and behavior trajectories in the game live video. The NLP-based text sentiment analysis module can use natural language processing technology to perform deep feature extraction on the subtitle text feature vector extracted from the basic feature layer, analyze the emotional tendency of the subtitle text, and obtain a text emotional feature vector that can represent the text emotion. The statistical learning-based interactive data analysis module can use statistical learning methods to perform deep feature extraction on the audience interaction feature vector extracted from the basic feature layer, and obtain a quantitative interactive feature vector by quantitatively analyzing these interactive data. The speech emotion recognition module based on deep neural network can use deep neural network technology to perform deep feature extraction on the sound feature data extracted by the basic feature layer, so as to extract the emotion feature vector expressed by the voice during the live broadcast. It should be noted that the basic feature layer will collect basic sound feature data such as volume and speech speed changes, and the deep analysis layer can analyze the host's voice data and audio change data stream. This part of the analysis is based on the sound feature data detected by the basic feature layer. Only through multiple factors can the analysis effect be better.
[0125] Among them, the large language model analysis layer serves as a decision-making unit and adopts a post-processing mechanism to comprehensively analyze the deep feature vectors extracted by the deep analysis layer.
[0126] Specifically, first, the large language model analysis layer performs feature understanding through the large language model, maps the multimodal features to a unified semantic space for feature conversion, and obtains multimodal features under the same data dimension; secondly, context analysis is performed on the multimodal features under the same data dimension to evaluate the temporal coherence of the content; finally, highlight clips are identified and screened through a multi-dimensional scoring mechanism, and the editing configuration information of the highlight clips is output, which includes video title, video tag, freeze time, and freeze text.
[0127] Step 510: Video processing.
[0128] Specifically, this method can process the video according to the result of highlight detection. The specific processing methods include: first, cutting the video based on the freeze time to accurately cut the highlight segment from the original screen recording. Secondly, the information generated by the detection (such as video title, video label, freeze text) is configured into the editing parameters. Finally, the PR editing script is used for automatic processing to generate a highlight video or PR project file.
[0129] Step 512: Finished product output.
[0130] Specifically, this method can generate two types of outputs: sliced video and PR project file; the sliced video (i.e., highlight video) is a directly usable sliced video that can be stored in the material library for use; the PR project file is used for editors to make further optimization adjustments.
[0131] Based on the above steps, it can be seen that the live broadcast data processing of this specification provides a live broadcast highlight intelligent editing system based on game data and multimodal analysis, and based on this system, a live broadcast highlight intelligent editing method based on game data and multimodal analysis is implemented. The system realizes automatic identification of wonderful content through intelligent log analysis, ensures processing efficiency through standardized workflow, and improves content quality through algorithm assistance, thereby significantly improving the conversion efficiency of live broadcast content to short videos and expanding the dissemination range and influence of game content.
[0132] This system implements an intelligent editing solution based on multimodal analysis; the system establishes an innovative two-layer multimodal analysis framework, and uses lightweight algorithms in the basic feature layer to achieve real-time processing, including audience interaction analysis, game event processing, subtitle generation, and volume detection; in the deep analysis layer, it integrates VLM, video understanding models, sentiment analysis and other advanced AI technologies to achieve quasi-real-time analysis; finally, through the large language model, feature fusion and decision analysis are performed to achieve efficient and accurate highlight segment recognition. This innovative layered architecture not only ensures the real-time responsiveness of the system, but also provides high-quality analysis results through deep learning models. At the same time, this system builds a complete intelligent editing workflow; by directly converting AI analysis results into standardized editing parameters, it not only realizes automatic cutting and packaging of videos, but also generates engineering files that can be further optimized by professional editors, truly opening up the entire process from content recognition to work generation.
[0133] In addition, in terms of data collection, in addition to the real-time recording method adopted by this system, it can also be implemented in the following ways; for example, the system can directly capture and store the data of the streaming server; or by implanting a collection module in the live broadcast client, data can be collected at the source; even the live broadcast content can be saved through web recording technology.
[0134] In the multimodal analysis framework, this system can also adopt different data fusion strategies. For example, it can first extract single-modal features and then fuse them; it can also design an end-to-end multimodal model to directly process the original data; it can also adopt a hierarchical analysis method, first processing the more important modal data, and then gradually integrating information from other dimensions.
[0135] For the recognition of highlight moments, in addition to deep learning algorithms, this system can also consider: rule-based judgment systems, which identify key moments by setting multiple thresholds and conditions; or adopt traditional machine learning methods, such as random forests, support vector machines and other algorithms; or use sequence models based on attention mechanisms to analyze time series data.
[0136] In terms of video editing, in addition to generating PR project files, this system can also: directly output editing decision parameters (i.e. target video clip parameters) for use by other editing software; or build an independent editing engine that does not rely on specific commercial software; and even design a cloud-based online editing system.
[0137] Corresponding to the above method embodiment, this specification also provides a live data processing device embodiment. Figure 7 FIG. 1 is a schematic diagram showing the structure of a live broadcast data processing device provided by an embodiment of the present specification. Figure 7 As shown, the device comprises: The data determination module 702 is configured to determine the multimodal live broadcast data of the target game, and perform data analysis on the multimodal live broadcast data to obtain initial features of multiple modes, wherein the multimodal live broadcast data includes target game data, game live broadcast video, and live broadcast interaction data associated with the game live broadcast video; The feature analysis module 704 is configured to perform feature analysis on the initial features of the multiple modes to obtain target features of the multiple modes; The segment detection module 706 is configured to perform segment detection on the target features of the multiple modalities using the segment detection model to obtain target video segment parameters; The segment generation module 708 is configured to generate a target live video segment corresponding to the target game based on the target video segment parameters and the game live video.
[0138] Optionally, the feature analysis module 704 is further configured to: determine the feature analysis modules corresponding to the initial features of the multiple modalities, wherein one feature analysis module corresponds to the initial features of at least one modality; and use multiple feature analysis modules to perform feature analysis on the initial features of the multiple modalities to obtain target features of the multiple modalities.
[0139] Optionally, the multiple feature analysis modules include an image analysis module, a video analysis module, a text analysis module, an interactive data analysis module, and a voice analysis module, and the initial features of the multiple modalities include initial game screen features, initial live video features, initial live voice features, initial voice text features, and initial interactive data features; The feature analysis module 704 is further configured to: utilize the image analysis module, the video analysis module, the text analysis module, the interactive data analysis module and the voice analysis module to perform deep feature extraction on the initial game screen features, the initial live video features, the initial voice text features, the initial interactive data features and the initial live voice features to obtain target game screen features, target live video features, target text emotion features, target interactive data features and target voice emotion features; determine the target game screen features, the target live video features, the target voice emotion features, the target text emotion features and the target interactive data features as the target features of the multiple modalities.
[0140] Optionally, the segment detection model is a large language model; the segment detection module 706 is further configured to: utilize the large language model to perform feature conversion on the target features of the multiple modalities, obtain the live data features corresponding to the multimodal live data, and perform segment detection on the live data features to obtain the target video segment parameters.
[0141] Optionally, the segment detection module 706 is further configured to: input the target features of the multiple modalities into the large language model, and in the large language model, map the target features of the multiple modalities to the semantic space for feature conversion to obtain the live broadcast data features corresponding to the multimodal live broadcast data; perform multi-dimensional evaluation on the live broadcast data features to obtain the target video segment parameters, wherein the target video segment parameters include video segment display text and video segment time information.
[0142] Optionally, the data determination module 702 is further configured to: determine multiple live data sources, wherein the multiple live data sources include a game server, a live video storage unit, and an interactive data storage unit; obtain the target game data of the target game from the game server, obtain the game live video of the target game from the live video storage unit, and obtain the live interactive data from the interactive data storage unit; perform data analysis on the target game data, the game live video, and the live interactive data to obtain the initial features of the multiple modalities.
[0143] Optionally, the data determination module 702 is further configured to: determine the game live voice of the game live video, and convert the game live voice into game live voice text; time-align the game live video, the game live voice, the game live voice text, the live interactive data and the target game data according to the video time information of the game live video, the interactive time information of the live interactive data and the game time information of the target game data to obtain aligned multimodal live data; perform feature extraction on the aligned multimodal live data to obtain initial game screen features, initial live video features, initial live voice features, initial voice text features and initial interactive data features; determine the initial game screen features, the initial live video features, the initial live voice features, the initial voice text features and the initial interactive data features as the initial features of the multiple modalities.
[0144] Optionally, the target video segment parameters are video segment display text and video segment time information; the segment generation module 708 is further configured to: perform video cutting on the game live video based on the video segment time information to obtain a cut video segment; and generate the target live video segment corresponding to the target game based on the video segment display text and the cut video segment.
[0145] The fragment generation module 708 is also configured to: determine a target special effect template from multiple candidate special effect templates, and render the video fragment based on the target special effect template, the video fragment display text and the cut video fragment to obtain the target live video fragment corresponding to the target game.
[0146] Optionally, the game live broadcast video is obtained by collecting video of the game live broadcast of the target game in a real-time data collection manner, the live interactive data is obtained by collecting interactive data of the game live broadcast in a real-time data collection manner, and the target game data is obtained by querying the game server according to the time tag of the game live broadcast video.
[0147] One or more embodiments of the present specification provide a live data processing device. First, data analysis can be performed on multimodal live data including target game data, game live video, and live interactive data associated with the game live video to obtain initial features of multiple modalities. Secondly, feature analysis can be performed on the initial features of multiple modalities in a more detailed manner to obtain relatively accurate target features of multiple modalities, thereby improving the accuracy of the generated target live video clips. Finally, a clip detection model is used to efficiently and accurately identify highlight clips of target features of multiple modalities to obtain target video clip parameters. Based on the target video clip parameters and the game live video, a target live video clip corresponding to the target game is generated, thereby achieving accurate and efficient determination of the target live video clip and avoiding the problem that the efficiency of manual screening is low and cannot meet the needs of actual promotion and publicity.
[0148] The above is a schematic scheme of a live data processing device of this embodiment. It should be noted that the technical scheme of the live data processing device and the technical scheme of the live data processing method described above belong to the same concept, and the details of the technical scheme of the live data processing device that are not described in detail can all be referred to the description of the technical scheme of the live data processing method described above.
[0149] An embodiment of the present specification also provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the above-mentioned live data processing method when executed by a processor.
[0150] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to the live data processing method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the live data processing method embodiment.
[0151] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned live data processing method when executed by a processor.
[0152] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the live data processing method described above are of the same concept, and the details not described in detail in the technical scheme of the computer program product can be found in the description of the technical scheme of the live data processing method described above.
[0153] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0154] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0155] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0156] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0157] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A live broadcast data processing method, characterized in that: include: Determine multimodal live broadcast data of a target game, and perform data analysis on the multimodal live broadcast data to obtain initial features of multiple modes, wherein the multimodal live broadcast data includes target game data, a game live broadcast video, and live broadcast interaction data associated with the game live broadcast video; Performing feature analysis on the initial features of the multiple modes to obtain target features of the multiple modes; Using a segment detection model to perform segment detection on target features of the multiple modalities to obtain target video segment parameters; Based on the target video segment parameters and the game live video, a target live video segment corresponding to the target game is generated.
2. The live broadcast data processing method according to claim 1, characterized in that: The performing feature analysis on the initial features of the multiple modes to obtain target features of the multiple modes includes: Determining feature analysis modules corresponding to the initial features of the multiple modalities, wherein one feature analysis module corresponds to the initial feature of at least one modality; A plurality of feature analysis modules are used to perform feature analysis on the initial features of the plurality of modes to obtain target features of the plurality of modes.
3. The live broadcast data processing method according to claim 2, characterized in that: The multiple feature analysis modules include an image analysis module, a video analysis module, a text analysis module, an interactive data analysis module, and a voice analysis module, and the initial features of the multiple modalities include initial game screen features, initial live video features, initial live voice features, initial voice text features, and initial interactive data features; The method of using multiple feature analysis modules to perform feature analysis on the initial features of the multiple modes to obtain target features of the multiple modes includes: Utilizing the image analysis module, the video analysis module, the text analysis module, the interactive data analysis module and the voice analysis module, deep feature extraction is performed on the initial game screen features, the initial live video features, the initial voice text features, the initial interactive data features and the initial live voice features to obtain target game screen features, target live video features, target text emotion features, target interactive data features and target voice emotion features; The target game screen features, the target live video features, the target voice emotion features, the target text emotion features and the target interactive data features are determined as target features of the multiple modalities.
4. The live broadcast data processing method according to any one of claims 1 to 3, characterized in that: The segment detection model is a large language model; The using the segment detection model to perform segment detection on the target features of the multiple modalities to obtain target video segment parameters includes: The large language model is used to perform feature conversion on the target features of the multiple modalities to obtain live data features corresponding to the multimodal live data, and segment detection is performed on the live data features to obtain the target video segment parameters.
5. The live broadcast data processing method according to claim 4, characterized in that: The using of the large language model to perform feature conversion on the target features of the multiple modalities to obtain live data features corresponding to the multimodal live data, and performing segment detection on the live data features to obtain the target video segment parameters includes: Inputting the target features of the multiple modalities into the large language model, mapping the target features of the multiple modalities to a semantic space in the large language model for feature conversion, and obtaining the live broadcast data features corresponding to the multimodal live broadcast data; A multi-dimensional evaluation is performed on the live broadcast data features to obtain the target video segment parameters, wherein the target video segment parameters include video segment display text and video segment time information.
6. The live broadcast data processing method according to any one of claims 1 to 3, characterized in that: The step of determining the multimodal live broadcast data of the target game and performing data analysis on the multimodal live broadcast data to obtain initial features of multiple modes includes: Determine a plurality of live data sources, wherein the plurality of live data sources include a game server, a live video storage unit, and an interactive data storage unit; Acquire the target game data of the target game from the game server, acquire the game live video of the target game from the live video storage unit, and acquire the live interactive data from the interactive data storage unit; Data analysis is performed on the target game data, the game live video, and the live interactive data to obtain initial features of the multiple modalities.
7. The live broadcast data processing method according to claim 6, characterized in that: The performing data analysis on the target game data, the game live video, and the live interactive data to obtain initial features of the multiple modalities includes: Determine the game live broadcast voice of the game live broadcast video, and convert the game live broadcast voice into game live broadcast voice text; According to the video time information of the game live video, the interaction time information of the live interactive data and the game time information of the target game data, the game live video, the game live voice, the game live voice text, the live interactive data and the target game data are time-aligned to obtain aligned multimodal live data; Performing feature extraction on the aligned multimodal live broadcast data to obtain initial game screen features, initial live broadcast video features, initial live broadcast voice features, initial voice text features, and initial interactive data features; The initial game screen features, the initial live video features, the initial live voice features, the initial voice text features, and the initial interactive data features are determined as initial features of the multiple modalities.
8. The live broadcast data processing method according to any one of claims 1 to 3, characterized in that: The target video segment parameters include video segment display text and video segment time information; The generating a target live video segment corresponding to the target game based on the target video segment parameters and the game live video includes: Performing video cutting on the live game video based on the video segment time information to obtain a cut video segment; Based on the video segment display text and the cut video segment, the target live video segment corresponding to the target game is generated.
9. The live broadcast data processing method according to claim 8, characterized in that: The step of generating the target live video segment corresponding to the target game based on the video segment display text and the cut video segment includes: A target special effect template is determined from a plurality of candidate special effect templates, and video segment rendering is performed based on the target special effect template, the video segment display text, and the cut video segment to obtain the target live video segment corresponding to the target game.
10. The live broadcast data processing method according to any one of claims 1 to 3, characterized in that: The game live video is obtained by collecting video of the game live of the target game in a real-time data collection manner, the live interactive data is obtained by collecting interactive data of the game live in a real-time data collection manner, and the target game data is obtained by querying the game server according to the time tag of the game live video.
11. A live broadcast data processing device, characterized in that: include: A data determination module is configured to determine multimodal live broadcast data of a target game, and perform data analysis on the multimodal live broadcast data to obtain initial features of multiple modes, wherein the multimodal live broadcast data includes target game data, a game live broadcast video, and live broadcast interaction data associated with the game live broadcast video; A feature analysis module is configured to perform feature analysis on the initial features of the multiple modes to obtain target features of the multiple modes; A segment detection module is configured to perform segment detection on target features of the multiple modalities using a segment detection model to obtain target video segment parameters; The segment generation module is configured to generate a target live video segment corresponding to the target game based on the target video segment parameters and the game live video.
12. A computer-readable storage medium, characterized in that: It stores a computer program / instruction, which implements the steps of the method described in any one of claims 1 to 10 when executed by a processor.
Citation Information
Patent Citations
Video processing method and device
CN112738557A
Video recommendation method and device, computer equipment and storage medium
CN115604510A
Live broadcast monitoring method and device and electronic equipment
CN117097921A
Live video key point marking method and device, equipment and storage medium
CN119071520A
Automatic video montage generation
US20220108727A1
Cited By
Video outline generation method and device, computer equipment and medium
CN120281995A
Video outline generation method, device, computer equipment and medium
CN120281995B
Highlight combination set generation method and system
CN120321474A
A high photosynthetic yield generation method and system
CN120321474B
Direct broadcasting room explanation content generation method and system, readable storage medium and equipment
CN121309861A