Generation method and device of commentary video, electronic equipment and computer storage medium
By determining the information of the target scene during the video generation process and tuning it using a large language model to generate commentary information that fits the video scene, the problems of poor coherence and high manual production cost are solved, and efficient and smooth commentary video generation is achieved.
Patent Information
- Application Number
- CN202410036728.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
The traditional commentary video generation model has insufficient coherence and fluency in the commentary information, and the production of manual commentary videos is expensive and time-consuming.
By determining the information of the target scene in the original video, targeted prompt information is generated, and tuning is used to use a large language model to generate an explanatory video.
Improve the pertinence and consistency of the information on understanding, reduce production costs, and improve the fluency and quality of video commentary.
Smart Images

Figure CN120281987A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of video processing, and in particular, to a method for generating an explanatory video, a device for generating an explanatory video, an electronic device, and a computer-readable storage medium. Background Art
[0002] An explanatory video refers to a video containing explanatory information, which can enable users to more intuitively understand the video content. In a related technology, explanatory information for a related video can be generated based on a machine learning model. Specifically, the machine learning model generally used is a traditional language model, such as an n-gram model, a Hidden Markov Model (HMM), or a conditional random field model, etc. However, the explanatory information generated by traditional language models has the problem of poor coherence. Summary of the Invention
[0003] The present application provides a method for generating an explanatory video, a device for generating an explanatory video, an electronic device, and a computer-readable storage medium, which at least to some extent improve the coherence of the explanatory information.
[0004] In a first aspect, the present application provides a method for generating an explanatory video, the method including: determining target prompt information according to information reflecting a target scene in the original video, where the target scene is any one of the scenes included in the original video; inputting the target prompt information into a large language model, and determining explanatory information for the target scene according to the output of the large language model; and generating an explanatory video according to the explanatory information for the target scene and the original video.
[0005] In some embodiments, based on the above solution, the determining target prompt information according to information reflecting a target scene in the original video includes: determining a target detection object in the original video; determining a target key event corresponding to the target detection object in the original video, where the information reflecting the target scene includes the target detection object and the target key event corresponding to the target detection object; and performing a structured process on the target detection object and the target key event corresponding to the target detection object to obtain target prompt information for the target scene.
[0006] In some embodiments, based on the above solution, the determining a target detection object in the original video includes: determining a target detection object in the original video based on the processing of the original video by a trained target detection model;
[0007] Determining the target key event corresponding to the target detection object in the original video as described above includes: after determining the target detection object, determining the target motion trajectory of the target detection object based on the trained trajectory determination model; and, based on the trained behavior analysis model, determining the target key event according to the target detection object and the target motion trajectory.
[0008] In some embodiments, based on the above solution, structuring the target detection object and the target key event corresponding to the target detection object to obtain target prompt information for the target scenario includes: determining the target time information of the target scenario in the original video; and, structuring the target time information, the target detection object, and the target key event to obtain target prompt information for the target scenario.
[0009] In some embodiments, based on the above solution, generating an explanatory video according to the explanatory information for the target scenario and the original video includes: splitting the original video according to the target time information to obtain multiple sub-videos; and, adding the explanatory information of the target scenario to the sub-video corresponding to the target scenario to obtain an explanatory video.
[0010] In some embodiments, based on the above solution, after splitting the original video according to the target time information to obtain multiple sub-videos, the method further includes: performing personalized processing on the explanatory information corresponding to the target sub-video; and / or, adding special effects processing between the first explanatory information and the second explanatory information respectively corresponding to two adjacent sub-videos.
[0011] In some embodiments, based on the above solution, inputting the target prompt information into a large language model and determining the explanatory information for the target scenario according to the output of the large language model includes: inputting the prompt information of the target scenario into the large language model to optimize the large language model; and, determining the explanatory information about the target scenario according to the output information of the large language model.
[0012] In some embodiments, based on the above solution, the explanatory information includes pre-game explanatory information about a competition; the method further includes: obtaining basic information about the detection object in the competition; determining first prompt information according to the basic information; and, inputting the first prompt information into the large language model and determining the pre-game explanatory information about the competition according to the output of the large language model;
[0013] Generating an explanatory video according to the explanatory information for the target scenario and the original video as described above includes: determining an explanatory video according to the pre-game explanatory information, the explanatory information for the target scenario, and the original video.
[0014] In some embodiments, based on the above solution, the above-mentioned commentary information includes post-game commentary information about the event; the above method further includes: obtaining summary information about the event; determining second prompt information according to the above summary information; and inputting the above second prompt information into a large language model, and determining post-game commentary information about the above event according to the output of the above large language model; the generating of the commentary video according to the commentary information for the above target scenario and the above original video includes: determining a commentary video according to the above post-game commentary information, the commentary information for the above target scenario, and the above original video.
[0015] The commentary video generation method provided by the embodiments of the present application can make the commentary information fit the displayed video scenario, and can also improve the coherence and fluency of the video commentary.
[0016] In a second aspect, the present application provides a device for generating a commentary video, the device includes: a prompt information determination module, a commentary information determination module, and an end video generation module;
[0017] Wherein, the above prompt information determination module is used to determine target prompt information according to the information reflecting the target scenario in the original video, where the above target scenario is any one of the scenarios included in the above original video; the above commentary information determination module is used to input the above target prompt information into a large language model, and determine commentary information for the above target scenario according to the output of the above large language model; and the above commentary video generation module is used to generate a commentary video according to the commentary information for the above target scenario and the above original video.
[0018] In some embodiments, based on the above solution, the above prompt information determination module includes: a first determination unit, a second determination unit, and a structuring unit;
[0019] Wherein, the above first determination unit is used to: determine the target detection object in the above original video; the above second determination unit is used to: determine the target key event corresponding to the above target detection object in the above original video, wherein the information reflecting the target scenario includes the above target detection object and the target key event corresponding to the above target detection object; and the above structuring unit is used to: perform structuring processing on the above target detection object and the target key event corresponding to the above target detection object to obtain target prompt information for the above target scenario.
[0020] In some embodiments, based on the above solution, the above first determination unit is specifically used to: determine the target detection object in the above original video based on the processing of the above original video by a trained target detection model;
[0021] The above-mentioned second determination unit is specifically configured to: after determining the target detection object, determine the target motion trajectory of the target detection object based on the trained trajectory determination model; and, based on the trained behavior analysis model, determine the target key event according to the target detection object and the target motion trajectory.
[0022] In some embodiments, based on the above solution, the above-mentioned structuring unit is specifically configured to: determine the target time information of the target scene in the original video; and perform structuring processing on the target time information, the target detection object, and the target key event to obtain target prompt information for the target scene.
[0023] In some embodiments, based on the above solution, the above-mentioned commentary video module includes: a sub-video segmentation unit and an information addition unit;
[0024] Among them, the above-mentioned sub-video segmentation unit is used to: segment the original video according to the target time information to obtain a plurality of sub-videos; and the above-mentioned information addition unit is used to: add the process commentary information of the target scene to the sub-video corresponding to the target scene to obtain a commentary video.
[0025] In some embodiments, based on the above solution, the above-mentioned device for generating the commentary video further includes: a personalized processing module;
[0026] Among them, the above-mentioned personalized processing module is used to: after the sub-video segmentation unit segments the original video according to the target time information to obtain a plurality of sub-videos, perform personalized processing on the commentary information corresponding to the target sub-video; and / or add special effect processing between the first commentary information and the second commentary information respectively corresponding to two adjacent sub-videos.
[0027] In some embodiments, based on the above solution, the above-mentioned commentary information determination module is specifically configured to: input the prompt information of the target scene into a large language model to optimize the large language model; and determine the commentary information about the target scene according to the output information of the large language model.
[0028] In some embodiments, based on the above solution, the above-mentioned commentary information includes pre-game commentary information about the event; the above-mentioned device for generating the commentary video further includes: a first information acquisition module;
[0029] Among them, the above-mentioned first information acquisition module is used to: acquire basic information about the detection object in the event; the above-mentioned prompt information determination module is further used to: determine the first prompt information according to the above-mentioned basic information; the above-mentioned commentary information determination module is further used to: input the above-mentioned first prompt information into a large language model, and determine the pre-game commentary information about the above-mentioned event according to the output of the above-mentioned large language model; the above-mentioned commentary video generation module is further used to: determine the commentary video according to the above-mentioned pre-game commentary information, the commentary information for the above-mentioned target scenario, and the above-mentioned original video.
[0030] In some embodiments, based on the above solution, the above-mentioned commentary information includes post-game commentary information about the event; the generating device of the above-mentioned commentary video further includes: a second information acquisition module;
[0031] Among them, the above-mentioned second information acquisition module is used to: acquire summary information about the event; the above-mentioned prompt information determination module is further used to: determine the second prompt information according to the above-mentioned summary information; the above-mentioned commentary information determination module is further used to: input the above-mentioned second prompt information into a large language model, and determine the post-game commentary information about the above-mentioned event according to the output of the above-mentioned large language model; the above-mentioned commentary video generation module is further used to: determine the commentary video according to the above-mentioned post-game commentary information, the commentary information for the above-mentioned target scenario, and the above-mentioned original video.
[0032] The commentary video generating device provided by the embodiments of the present application can make the commentary information fit the displayed video scene, and can also improve the coherence and fluency of the video commentary.
[0033] In a fifth aspect, an electronic device is provided, including a processor and a memory. The above-mentioned memory is used to store a computer program, and the above-mentioned processor is used to call and run the computer program stored in the above-mentioned memory to execute the method for generating a commentary video provided in the first aspect and its various implementation manners.
[0034] In a sixth aspect, a chip is provided for implementing the method in any one of the first aspects or its various implementation manners. Specifically, the above-mentioned chip includes: a processor, which is used to call and run a computer program from a memory, so that a device equipped with the above-mentioned chip executes the method for generating a commentary video provided in the first aspect and its various implementation manners.
[0035] In a seventh aspect, a computer-readable storage medium is provided for storing a computer program, and the above-mentioned computer program causes a computer to execute the method for generating a commentary video provided in the first aspect and its various implementation manners.
[0036] In an eighth aspect, there is provided a computer program product including computer program instructions, and the computer program instructions cause a computer to execute the method for generating an explanatory video provided in the first aspect and its various implementation manners above.
[0037] In a ninth aspect, there is provided a computer program which, when running on a computer, causes the computer to execute the method for generating an explanatory video provided in the first aspect and its various implementation manners above.
[0038] In summary, in the solution provided in the embodiments of the present application, information reflecting a target scene in the original video is determined, and then target prompt information corresponding to the target scene is determined according to the information reflecting the target scene. By inputting the target prompt information for the target scene into a large language model, tuning of the large language model for the target scene can be achieved, so that the large language model outputs explanatory information about the target scene. An explanatory video is generated according to the explanatory information about each scene and the original video. Since the prompt information corresponding to different scene information is different, relevant explanatory information for the relevant scene can be determined specifically based on different prompt information, thereby increasing the pertinence of the explanatory information and making the explanatory information fit the video scene being displayed; at the same time, the large language model adopted has strong content understanding ability and can generate smooth and coherent explanatory information. It can be seen that the explanatory video generation solution provided in the embodiments of the present application can make the explanatory information fit the video scene being displayed and can also improve the coherence and smoothness of the video explanation. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention of this application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 It is a schematic diagram of an application scenario of the explanatory video provided in the embodiments of the present application;
[0041] Figure 2 It is a schematic flowchart of the method for generating an explanatory video provided in the embodiments of the present application;
[0042] Figure 3 It is a schematic diagram of the acquisition process of detection objects and their key events for a target scene provided in the embodiments of the present application;
[0043] Figure 4 It is a schematic diagram of a detection target provided in the embodiments of the present application;
[0044] Figure 5Schematic diagram of the process for obtaining prompt information and commentary information regarding a target scenario provided by an embodiment of the present application;
[0045] Figure 6A Schematic diagram of an explanatory video provided by an embodiment of the present application;
[0046] Figure 6B Schematic diagram of an explanatory video provided by another embodiment of the present application;
[0047] Figure 7 Schematic flowchart of a method for generating an explanatory video provided by another embodiment of the present application;
[0048] Figure 8 Schematic block diagram of a device for generating an explanatory video provided by an embodiment of the present application;
[0049] Figure 9 Schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In the embodiments of the present invention and the present application, "B corresponding to A" means that B is associated with A. In one implementation manner, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two.
[0052] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0053] The technical solution proposed in the embodiments of the present application can be applied to the video playback process of various platform applications (Application, APP) and official accounts. The usage scenarios are, for example, the commentary on sports events, the commentary on games, and the commentary on film and television videos, etc. It is mainly used to make the commentary information fit the displayed video playback, and improve the coherence and fluency of the video commentary.
[0054] Figure 1 It is a schematic diagram of the commentary video application scenario 001 provided by the embodiments of the present application. Refer to Figure 1 , the terminal 110 plays a video about basketball. To facilitate the user's understanding of the video content, commentary information can be provided according to the played scenario. For example Figure 1 The commentary information corresponding to the shown scenario can be that athlete A is shooting and athlete B is defending. Thus, even if the audience does not understand the basketball rules, they can correctly understand the content of the scenario shown on the terminal according to the above commentary information. In addition, when the above commentary information is in the form of voice, even if the audience does not watch the video image, they can accurately understand the progress of the event. It can be seen that the commentary information that fits the scenario and is coherent and smooth can help the audience deeply and accurately understand the original video.
[0055] In the commentary video generation scheme provided by the first related technology for generating commentary videos, as described in the background technology, when the traditional speech model is used to process natural language tasks, it is limited by the scale and quality of its training data. Therefore, there are problems such as low refinement degree of the generated commentary information, low fitting degree between the commentary information and the video display content, and further poor coherence and low fluency in the video commentary.
[0056] Another related technology for generating commentary videos is to determine the video commentary information manually. For example, the commentary information of a sports event is determined manually, and then a series of editing and other processes are used to determine the relevant commentary video. However, this related technology has problems of high cost, time-consuming and laborious. Especially for small-scale competitions and non-professional competitions, it is not suitable to produce artificial commentary videos at a high cost.
[0057] To solve the above technical problems, in the embodiments of the present application, the information reflecting the scene in the original video is first determined, and then the prompt information for optimizing the large language model is determined according to the scene information. For example Figure 1 in the scene information, it can be that athlete a is shooting a basket, and the generated prompt information can be: You are a professional basketball game commentator, and the following is some information about a basketball game: Athlete a is shooting a basket. Please summarize the game scene in about 50 words based on the above information. Among them, different scenes correspond to different prompt information. Further, the above prompt information can be used to optimize the large language model, so as to determine the commentary information about the scene according to the output of the above large language model. According to the commentary information about the above scene and the original video, a commentary video is generated. Since different scene information corresponds to different prompt information, relevant scene commentary information can be determined specifically based on different prompt information, thus increasing the pertinence of the commentary information. At the same time, the large language model adopted has the advantages of strong content understanding ability and strong language generation ability, and can generate smooth and coherent commentary information. It can be seen that the commentary video generation solution provided by the embodiments of the present application can make the commentary information fit the video scene being displayed, and can improve the coherence and fluency of the video commentary. On this basis, the embodiments of the present application do not require a large amount of labor cost, which is beneficial to saving the manufacturing cost of the commentary video.
[0058] It can be understood that the above large language model can be deployed on the terminal 110 or on a server connected to the terminal 110 through a network. Among them, the network can be a wired communication link, a wireless communication link or an optical fiber cable, etc. The embodiments of the present application do not make any restrictions here. For example, it can be various types of communication media that can provide a communication link between the terminal and the server. In addition, the above terminal 110 can be a computer, a smart phone, a tablet, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a wearable smart device, a medical device, etc. The device is often configured with a display device, and the display device can also be a monitor, a display screen, a touch screen, etc. The touch screen can also be a touch panel, a touch panel, etc. But it is not limited to this. The above server can be a cloud server, which can specifically provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms; in addition, the server can also be an independent physical server, or a server cluster or distributed system composed of multiple physical servers.
[0059] It should be noted that the application scenarios of the embodiments of the present application are not limited to Figure 1The scene 001 shown may also be other application scenarios, such as video commentary about other sports, video commentary about games, and commentary of film and television videos, etc., which is not limited in this embodiment of the present application.
[0060] The technical solutions of the embodiments of the present application are described in detail below through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0061] Figure 2 The flowchart of the method P200 for generating a narration video provided in the embodiment of the present application is shown in FIG. 1 . The execution subject of the method P200 may be a server. Figure 2 , method P200 includes: S210 to S230.
[0062] In S210, the server determines target prompt information according to information reflecting the target scene in the original video, wherein the target scene is any one of the scenes included in the original video.
[0063] The original video is the video to which the commentary information is to be added. The original video may be composed of a series of scenes, such as Figure 1 The scenes related to shooting in the original video are shown. In this embodiment of the present application, any scene (target scene) in the original video is used as an example for description.
[0064] The above prompt information refers to the information used as "Prompt" to tune the large language model. Since the original video is composed of multiple continuous and different scenes, different prompts are determined for different scenes. Therefore, the large language model is tuned from different angles based on different prompts, and the information output by the large language model is different, which can increase the pertinence of the commentary information and ultimately make the commentary information fit the displayed video scene.
[0065] The above large language model is a natural language processing model built based on deep learning technology, which can generate, understand, and process natural language texts by learning the statistical laws and semantic features of language. Large language models can be widely applied to tasks such as language modeling, machine translation, text summarization, and sentiment analysis. The large language model generates coherent text by predicting the probability of the next word or character. Due to its huge number of parameters, the large language model has better performance than general language models. The large language model adopted in the embodiments of this application has the advantages of strong content understanding ability and strong language generation ability, and can generate smooth and coherent commentary information. It can be seen that since the commentary video generation solution provided by the embodiments of this application is implemented based on the above large language model, it is beneficial to generate coherent and smooth commentary information, and thus beneficial to improving the coherence and fluency of the commentary video.
[0066] In an exemplary embodiment, as a specific implementation manner of S210, it includes S210-1 to S210-3.
[0067] In S210-1, determine the target detection object in the original video.
[0068] Exemplarily, Figure 3 is a schematic diagram of the acquisition process 003 of the detection object for the target scene provided by the embodiments of this application. Refer to Figure 3 , the target detection can be implemented by using a pre-trained target detection model 310. Among them, for example, the target detection model 310 can adopt the You Only Look Once (Yolo) series of models, such as Yolo-V5. Input the original video 30 into the target detection model 310, and through the processing of the original video 30 by this model, the target detection object 32 can be determined.
[0069] Exemplarily, the target detection model 310 identifies the target of interest in each video frame of the original video 30, such as players, basketballs, basketball hoops, etc. Specifically, the target detection model 300 detects the bounding box and category of the target in the video frame, so as to determine the position information of the detection target. For example, refer to Figure 4 the schematic diagram of the detection target 004 shown, the detection objects can be determined by the target detection model 310 to include: the basketball hoop corresponding to the detection box 400, the basketball corresponding to the detection box 410, the athlete A corresponding to the detection box 420, and the athlete B corresponding to the detection box 430. Furthermore, the position information of the above detection targets can also be determined. For example, athlete A faces the basketball hoop, athlete B has his back to the basketball hoop and is between the basketball hoop and athlete A, etc.
[0070] In S210-2, determine the target key events corresponding to the target detection object in the above original video.
[0071] Among them, the information reflecting the target scenario includes the above-mentioned target detection object and its corresponding target key event. Therefore, in the embodiments of the present application, the explanatory information of the target scenario will be determined according to the above-mentioned target detection object and its corresponding target key event. Specifically:
[0072] After determining the target detection object, the target motion trajectory of the target detection object is determined based on the trained trajectory determination model. Exemplarily, Figure 3 The schematic diagram of the acquisition process 003 of the key event corresponding to the detection object of the target scenario provided by the embodiments of the present application is also provided. Refer to Figure 3 , after determining the target detection object, continuously identify and capture these targets in subsequent consecutive video frames. The trajectory determination model 320 can be used to implement the continuous identification and capture of the target detection object 32. Among them, the above-mentioned target detection model 320 can be a target continuous identification algorithm based on a Kalman filter, an optical flow method, or deep learning. The trajectory determination model 320 specifically identifies the motion trajectory of the same target detection object 32 in subsequent video frames, so as to capture the dynamic changes of the target detection object 32. For example, refer to Figure 4 , taking the above-mentioned target detection object 32 as an example of a basketball, after determining that the target detection object basketball appears in video frame x through the target detection model 310, the trajectory determination model is used to identify the motion trajectory of the basketball in a series of subsequent video frames of video frame x.
[0073] After determining the motion trajectory of the target detection object, based on the trained behavior analysis model, according to the above-mentioned target detection object and its corresponding target motion trajectory, the target key event of the target detection object can be determined. Exemplarily refer to Figure 3 , the behavior analysis model 330 can be used to determine the target key event 36 based on the target detection object 32 and its corresponding target motion trajectory 34. Among them, the above-mentioned behavior analysis model 36 can adopt a Temporal Shift Module (TSM). The behavior analysis model 36 can determine key events by analyzing the mutual relationship and motion pattern between targets. For example, according to the motion trajectory of the target detection object basketball in the above-mentioned embodiment, and analyzing the positional relationship between the motion trajectory and the basketball hoop, it can be determined whether the basketball successfully enters the basketball hoop, so that the key event regarding the target detection object basketball can be determined: scoring a goal.
[0074] It should be noted that the models used for target detection, target continuous capture, and event monitoring in the embodiments of the present application may not be limited to the models listed in the above embodiments, and may also be other models for implementing the same functions. The embodiments of the present application do not make any limitations in this regard.
[0075] S210-3: Structurally process the above-mentioned object to be detected and the target key event corresponding to the object to be detected to obtain target prompt information for the above-mentioned target scenario.
[0076] Among them, the structured information is beneficial to improving the processing standardization and processing efficiency. In the embodiments of the present application, in order to improve the accuracy of the prompt information Prompt and provide explanatory information that fits the current scenario, the embodiments of the present application will also generate structured prompt information.
[0077] In an exemplary embodiment, since the target key event occurs in a series of video frames of the original video, in order to subsequently combine the explanatory information with a high degree of correspondence with its corresponding video frames, when generating the structured improvement information in the embodiments of the present application, it also includes time information related to the scenario, such as a time period or a time point. Exemplarily, Figure 5 is a schematic diagram of the acquisition process 005 of the prompt information for the target scenario provided by the embodiments of the present application. Refer to Figure 5 , determine the target time information of the above-mentioned target scenario in the above-mentioned original video, for example, the time stamp is 05:02-06:12; structurally process the above-mentioned target time information 50, the above-mentioned object to be detected 32, and the corresponding target key event 36 to obtain target prompt information 52 for the above-mentioned target scenario. For example, taking the scenario of athlete A operating a basketball as an example, the objects to be detected are the basketball and athlete A, and the corresponding key events are athlete A shooting and the basketball scoring. If the time stamp of the video segment corresponding to the basketball movement trajectory is 03:01-03:52, then "03:01-03:52", "athlete A", "basketball", and "scoring" can be structurally processed. In the embodiments of the present application, the structured information extracted by target detection, target continuous recognition and capture, and event detection is integrated together in the order of events to form a complete scene description. For example, the player's position, the ball's trajectory, and the game events are organized in chronological order into a structured data table or JSON format of prompt information Prompt for subsequent generation of explanatory information.
[0078] Exemplarily, the structured prompt information is as follows:
[0079] (1). Time point m - time point n: Player A's action - moving from the center of the court to the right side of the court; Player B's action - staying still on the right side of the court;
[0080] (2). Time point n - time point o: Player A's action - having a physical collision; Player B's action - having a physical collision;
[0081] (3). Time point o - time point p: Player C's action - shooting; Ball's action - scoring;
[0082] (4). Time point p - Time point q: Referee's action - Whistle blowing.
[0083] The above - structured prompt information can extract useful information from the original video. The structured prompt information will help the large - language model generate richer and more targeted commentary text, thereby improving the quality of the commentary video.
[0084] Continue to refer to Figure 2 In S220, the server inputs the above - mentioned target prompt information into the large - language model and determines the commentary information for the above - mentioned target scenario according to the output of the large - language model.
[0085] Exemplarily, Figure 5 There is also provided a schematic diagram of the acquisition process 005 of the commentary information for the target scenario provided by the embodiment of the present application. Refer to Figure 5 The prompt information about the target scenario is input as a Prompt into the large - language model 500, so that the large - language model is tuned for the above - mentioned target scenario based on this Prompt, and then outputs the commentary information 54 for the target scenario.
[0086] The embodiment of the present application makes full use of the strong text - understanding ability of the large - language model, which is beneficial to generating smooth and logical commentary information and is conducive to efficiently completing the task of generating commentary information.
[0087] In an exemplary embodiment, the above - mentioned large - language model can be an open - source pre - trained large - language model, a self - developed large - language model, or a closed - source large - language model. The embodiment of the present application does not make any limitations in this regard. Exemplarily, the above - mentioned large - language model can adopt the open - source pre - trained model baichuan - 13b.
[0088] In an exemplary embodiment, the type of the commentary information output by the above - mentioned large - language model is text. In other embodiments, the above - mentioned large - language model can also directly output commentary information of the voice type. The embodiment of the present application does not make any limitations in this regard.
[0089] In S230, the server generates a commentary video according to the commentary information for the above - mentioned target scenario and the above - mentioned original video.
[0090] In order to increase the refinement and pertinence of the commentary information so that the commentary information is more in line with each link of the original video, the embodiment of the present application can divide the original video into multiple sub - videos, and then add corresponding commentary information to each sub - video, which is beneficial to improving the overall coherence and smoothness of the commentary video. Figure 6A There is provided a schematic diagram of the commentary video 60 provided by an embodiment of the present application. Refer to Figure 6A The commentary video 60 can be divided into multiple sub - videos 600 - 608 and their respectively corresponding commentary information.
[0091] In an exemplary embodiment, the original video may be segmented according to a preset duration to obtain a plurality of sub-videos, thereby increasing the refinement of the explanation information in the explanation video.
[0092] In another exemplary embodiment, the original video can also be segmented according to the target time information about the target scene to obtain multiple sub-videos. Thus, the commentary information for each scene is added to the corresponding sub-video. For example, the commentary information 1 for scene 1 is determined through the above embodiment, and the content displayed by sub-video 600 is the content of scene 1, then the commentary information 1 can be added to sub-video 600. Similarly, the commentary information 2 for scene 2 is determined through the above embodiment, and the content displayed by sub-video 602 is the content of scene 2, then the commentary information 2 can be added to sub-video 602. And so on, it can be determined Figure 6A The commentary information corresponding to each sub-video in the video is obtained, and then the commentary information corresponding to each sub-video is added to the corresponding position, so that the commentary video corresponding to the above-mentioned original video can be generated.
[0093] In an exemplary embodiment, a specific implementation of adding the commentary information to the corresponding sub-video may include: S11 and S12; in another exemplary embodiment, a specific implementation of adding the commentary information to the corresponding sub-video may also include: S11 to S13.
[0094] S11, converting the commentary information from a text type to an audio type.
[0095] For example, the commentary text may be converted into commentary speech through a text to speed (TTS) model, so as to facilitate the audience to listen.
[0096] S12. Add the commentary information to the corresponding sub-video according to the time information corresponding to the commentary information.
[0097] It is understandable that, since the structured prompt information Prompt contains time information, the corresponding time information can be retained in the commentary information determined based on Prompt. For example, the timestamp information about the basketball goal scene is: 03:01-03:52, and the corresponding timestamp information 03:01-03:52 is retained in the commentary information m about the basketball goal scene generated by the large language model. Further, the commentary information m is added to 03:01-03:52 in the original video according to the timestamp information, and the commentary information m is added to a series of video frames showing the basketball entering.
[0098] S13. Personalize the commentary information corresponding to the target sub-video; and / or, add special effects processing between the first commentary information and the second commentary information corresponding to two adjacent sub-videos respectively.
[0099] Exemplarily, the personalization of the target sub-video can be highlighting the music, animation, etc. of the scene of the target sub-video. It can also be when there is an intersection of at least two scenes in the target sub-video, fusing the commentary information corresponding to the two scenes. By personalizing the commentary information corresponding to the target sub-video, it is beneficial to increase the diversity of the commentary video, and can also enhance the empathy of the audience, thereby improving the viewing experience of the commentary video and increasing the ratings of the commentary video. Exemplarily, if the above-mentioned target sub-video is a segment about a basketball scoring scene, then the soundtrack of cheering for the goal can be added to the corresponding commentary information m. Exemplarily, if the above-mentioned target sub-video is a segment about the scene of team Y after team X scores a basketball goal, then the animation of cheering can be added to the corresponding commentary information m. Exemplarily, if the above-mentioned target sub-video is a segment about a basketball scoring scene, and this segment also includes the scene of athlete A shooting, therefore, there is an intersection between the commentary information n for athlete A's shooting scene and the above-mentioned commentary information m. Therefore, through the way of personalization, the commentary information m and the commentary information n can be fused into "Athlete A is extremely powerful and achieves a dunk to get two points for team X", thereby increasing the coherence and smoothness of the commentary of the commentary video.
[0100] Exemplarily, the special effects processing between adjacent sub-videos can be the music, animation, and connecting commentary used to connect the two videos. When the commentary information corresponding to two adjacent sub-videos is the first commentary information and the second commentary information respectively, in order to increase coherence, connecting words, etc. can be added between the first commentary information and the second commentary information. By this way of special effects processing, it is also beneficial to further increase the smoothness of the commentary. For example, if the previous sub-video is a sub-video about team X scoring a goal, and the next sub-video is a sub-video about the counterattack of the opponent team Y, then connecting commentary such as "Next, let's see the reaction of team Y" can be inserted between the two sub-videos. Another example, if the previous sub-video is a sub-video about player a of team Y passing the ball to player b, and the next sub-video is a sub-video about player c of team X stealing the pass, then music with a plot twist can be inserted between the two sub-videos to highlight the intense game atmosphere and at the same time make the connection between the two sub-videos coherent.
[0101] In an exemplary embodiment, Figure 6B is a schematic diagram of the commentary video 60' provided by another embodiment of this application. Refer to Figure 6B , compared with the commentary video 60 shown in Figure 6A , the commentary video 60' also includes a start stage 62 and a summary stage 64.
[0102] Exemplarily, in the commentary video about the event, before the official event is broadcast, an introduction to the basic information of the event can be provided so that the audience can understand the relevant objects in advance, thereby increasing the readability of the commentary video.
[0103] Exemplarily, the method for determining the commentary information corresponding to the above-mentioned start stage 62 includes: S21 to S23.
[0104] S21. Obtain the basic information about the detection object in the event;
[0105] For example, the basic information of each athlete in the event, the introduction to the event rules, the commentary on historical related events, etc.
[0106] S22. Determine the first prompt information according to the above basic information; and, S23. Input the above first prompt information into the large language model, and determine the pre-game commentary information about the above event according to the output of the large language model.
[0107] The above basic information can be structured to generate a structured first prompt information. Then, through the above first prompt information as a Prompt input into the large language model, the large language model can be tuned specifically to obtain the pre-game commentary information for the above event.
[0108] In an exemplary embodiment, the pre-game commentary information for the above event can be combined with the video corresponding to the above basic information to obtain the start stage 62 of the commentary video 60'.
[0109] Exemplarily, the method for determining the commentary information corresponding to the above-mentioned summary stage 64 includes: S31 to S33.
[0110] S31. Obtain the summary information about the event;
[0111] For example, the performance of each athlete in this event, the commentary on the exciting scenes in this event, the commentary on the regrettable points in this event, the evaluation information of relevant news and social media about this event, etc.
[0112] S32. Determine the second prompt information according to the above summary information; and, S23. Input the above second prompt information into the large language model, and determine the post-game commentary information about the above event according to the output of the large language model.
[0113] The above summary information can be structured to generate a structured second prompt information. Then, through the above second prompt information as a Prompt input into the large language model, the large language model can be tuned specifically to obtain the post-game commentary information for the above event.
[0114] In an exemplary embodiment, the post-match commentary video for the above-mentioned event may be combined with the video corresponding to the above-mentioned summary information to obtain the summary stage 64 of the commentary video 60 ′.
[0115] The generation method of the explanation video provided in the embodiment of the present application first determines the information reflecting the target scene in the original video, and then determines the target prompt information corresponding to the target scene according to the target scene information. By inputting the target prompt information of the target scene into the large language model, the large language model can be tuned for the target scene, so that the large language model outputs the explanation information about the target scene. Finally, the explanation video can be generated according to the explanation information and the original video about each scene. In the embodiment of the present application, due to the different prompt information corresponding to the different scene information of the original video, the explanation information of the relevant scene can be determined specifically based on different prompt information, thereby increasing the pertinence of the explanation information, and the explanation information can be made to fit the displayed video scene; at the same time, the large language model used has a strong ability to understand content and can generate fluent and coherent explanation information. It can be seen that the explanation video generation scheme provided in the embodiment of the present application can make the explanation information fit the displayed video scene, and can also improve the coherence and fluency of the video explanation. It can be seen that the explanation video generation scheme provided in the embodiment of the present application can make the explanation information fit the displayed video scene, and can improve the coherence and fluency of the video explanation. On this basis, the embodiment of the present application does not require a large manpower cost, which is helpful to save the production expenses of the commentary video.
[0116] The above is an overall introduction to the method for generating a narration video provided in the embodiment of the present application. The following is a further introduction to the method for generating a narration video provided in the embodiment of the present application through a specific embodiment.
[0117] Figure 7 The flowchart of the method P400 for generating a commentary video provided by another embodiment of the present application is shown in FIG. Figure 7 , the embodiment shown in the figure includes: S410 to S440.
[0118] In S410, the competition data is collected and organized.
[0119] S410 can collect rich game information to determine targeted prompts. Then the large language model is tuned based on the targeted prompts, which is conducive to generating expected commentary information. Exemplarily, the collected game data may include:
[0120] 1) Original video of the competition;
[0121] On the one hand, structured information is extracted from the original video in S420 to determine Prompts for different scenarios. On the other hand, it is used as the basis for generating an explanatory video in S440, that is, by adding relevant explanatory information to the relevant time periods of the original video, an explanatory video corresponding to the game can be generated.
[0122] 2) Player information, which belongs to the basic game data;
[0123] Collect the basic information of the participating players, such as name, age, affiliated team, on-field position, etc. In addition, historical performance data of the players can also be collected, such as the number of goals, assists, appearances, etc. Exemplarily, the above information can be obtained from player profiles, team official websites or sports databases.
[0124] Among them, the above player information, as the basic game information, can be used to generate Prompts for the pre-game start stage (such as Figure 6B in 62), that is, for generating the above first prompt information.
[0125] 3) Game background information, which belongs to the basic game data;
[0126] Understand the background information of the game, such as the name of the event, game time, game location, team rankings, historical head-to-head records, etc. Exemplarily, the above information can be obtained from the official event website, news reports or sports repositories.
[0127] Among them, the above game background information, as the basic game information, can be used to generate Prompts for the pre-game start stage (such as Figure 6B in 62), that is, for generating the above first prompt information.
[0128] 4) Game statistical data, which belongs to the game summary data;
[0129] Collect statistical data related to the game, such as the score, number of shots, possession time, passing success rate, number of fouls, etc. Exemplarily, the above data can be obtained from official statistical agencies, game organizers or third-party data providers.
[0130] Among them, the above game statistical information, as the game summary data, can be used to generate Prompts for the post-game summary stage (such as Figure 6B in 64), that is, for generating the above second prompt information.
[0131] 5) Related news and social media information, which belongs to the game summary data;
[0132] Pay attention to news reports and social media dynamics related to the game to understand emergencies, highlights of players' performances, audience reactions, etc. during the game. Exemplarily, the above information can be obtained from news websites, social media platforms or professional sports forums.
[0133] Among them, the above-mentioned relevant news and social media information, as game summary data, can be used to generate Prompts for the post-game summary stage (such as Figure 6B in 64), that is, to generate the above-mentioned second prompt information.
[0134] Exemplarily, in order to ensure data quality and improve data processing efficiency, for the above-mentioned game data collected, preprocessing operations such as data cleaning, deduplication, and format conversion are also required.
[0135] In S420, structured data extraction.
[0136] Among them, structured information such as player movements and game progress is extracted from the above-mentioned collected game videos. Specifically, it is implemented through an object detection model, a trajectory determination model, and an event monitoring model. The specific implementation method can refer to the corresponding embodiment of S210. Among them, in the embodiment of the present application, the structured information of the above-mentioned sports event includes: at a certain time point or a certain time - detection object (such as, athlete, coach, referee, football, etc.) - behavior - result. For example: time point n - time point o: player A's action - physical collision occurs; player B's action - physical collision occurs.
[0137] In S430, commentary information generation.
[0138] On the one hand, the embodiment of the present application generates rich prompt information Prompts for different scenarios, and on the other hand, by making full use of the capabilities of the large language model, it is beneficial to generate highly fitting commentary information for different scenarios.
[0139] In an exemplary embodiment, referring to Figure 6B , it is necessary to generate commentary information corresponding to three stages, specifically: pre-game background commentary corresponding to the start stage 62, in-game process commentary corresponding to the process, and post-game summary commentary corresponding to the summary stage.
[0140] Among them, regarding the generation process of pre-game background commentary information corresponding to the start stage 62:
[0141] The commentary information generated in this stage includes the background introduction of the game, player introductions, historical head-to-head records between the two sides, etc. The corresponding first prompt can be as follows: "You are a professional football game commentator. The following are the basic information of a football game: 'Basic game data'. Please briefly introduce the background of this game, the strength comparison between the two sides, and the historical head-to-head records between the two sides in about 500 words before the game." Here, the 'Basic game data' can specifically refer to the basic game data collected in the first step.
[0142] Regarding the process of generating post-match summary commentary information corresponding to the summary stage 64:
[0143] The commentary information generated in this stage includes the post-match summary of this game, a review of outstanding players, and a recap of exciting moments, etc. The corresponding second prompt is designed as follows: "You are a professional football game commentator. The following are some information I have collected about a football game: 'Game summary information'. Please summarize this game in about 500 words based on the above information, with a focus on introducing the highlights and outstanding players of this game." Here, the 'Game summary information' can specifically refer to the game statistical data and news social media information collected in the first step.
[0144] Regarding the process of generating commentary information for the game process corresponding to the overall process:
[0145] Since the game duration is long and a large amount of text needs to be generated, the original video is segmented into multiple sub-videos. Specifically, the segmentation method can refer to the detailed introduction in the above embodiments. It should be noted that the two segmentation methods provided in the above embodiments can also be combined, that is, the original video is segmented according to a preset duration (such as about one minute), but if the segmentation point is within the same scene, then this segmentation point is cancelled. Thus, it can be ensured that exciting moments are not split into two sub-videos. Specifically, according to the time information contained in the structured prompt information, the segmentation points are avoided from exciting moments such as goals and fouls, which helps to improve the fluency of the commentary video.
[0146] Specifically, for each sub-video, the commentary information corresponding to each sub-video can be generated based on the above embodiments using a large language model. Exemplarily, the prompt structure corresponding to a sub-video can be as follows: "You are a professional football game commentator. This is what happened in one minute of the game, which I have organized into structured data 'Structured prompt information'. Please provide a commentary on the game based on the structured data." Here, the 'Structured prompt information' can specifically refer to the structured data corresponding to this segment.
[0147] By gradually calling the large language model, the pre-match background commentary information corresponding to the start stage 62, the in-game process commentary information corresponding to the process, and the post-match summary commentary information corresponding to the summary stage can be obtained.
[0148] In S440, the editing and synthesis of the game commentary video.
[0149] In the specific implementation of S440, the generated commentary information is synchronized with the game footage in the original video, and the game footage is edited to finally generate a sports event video with commentary content. Specifically, the commentary text can be converted into commentary speech through a text-to-speech (TTS) model, and then a video editing software or an automatic editing algorithm is used to synchronize the commentary speech with the game footage in the original video. Specifically, since the structured data contains timestamp information, the commentary speech is spliced into the corresponding original video file according to the timestamp information.
[0150] In an exemplary embodiment, the commentary speech of a certain sub-video can be personalized, or the commentary speeches corresponding to two adjacent sub-videos can be synthesized through transition effects. Thus, the commentary video can look more natural and coherent.
[0151] The embodiments of the present application adopt mature technologies such as existing object detection capabilities and face recognition capabilities to extract the features of different scenes during the game process, and then structure the features of each scene into prompt information about the scene. Further, the embodiments of the present application make full use of the strong text understanding ability of the large language model and perform targeted optimization based on the highly targeted prompt information, so as to complete the task of generating commentary information with fluent and logical text. In the process of combining the above-mentioned commentary information with the original video, personalized processing can also be added to the commentary information according to the style of the sub-video, enriching the style of the commentary video and the coherence of the commentary video. It can be seen that the commentary video generation scheme provided by the embodiments of the present application can make the commentary information fit the displayed video scene, improve the quality of the commentary video, and at the same time can also enhance the coherence and fluency of the video commentary. Personalized processing is beneficial to meeting the needs of different audiences. In addition, the production cost can be reduced, which is beneficial to saving the manufacturing cost of the commentary video.
[0152] As described above in conjunction with Figures 2 to 7 , the method embodiments for generating the commentary video of the present application have been described in detail. Below in conjunction with Figure 8 , the device embodiments of the present application will be described in detail.
[0153] Figure 8 is a schematic block diagram of a device for generating a commentary video provided by an embodiment of the present application. Refer to Figure 8The generating device 800 of the explanatory video includes: a prompt information determining module 810, an explanatory information determining module 820, and an end video generating module 830;
[0154] Among them, the above-mentioned prompt information determining module 810 is used to determine target prompt information according to the information reflecting the target scene in the original video, where the target scene is any one of the scenes included in the original video; the above-mentioned explanatory information determining module 820 is used to input the above-mentioned target prompt information into a large language model and determine the explanatory information for the above-mentioned target scene according to the output of the above-mentioned large language model; and the above-mentioned explanatory video generating module 830 is used to generate an explanatory video according to the explanatory information for the above-mentioned target scene and the above-mentioned original video.
[0155] In some embodiments, based on the above solution, the above-mentioned prompt information determining module 810 includes: a first determining unit, a second determining unit, and a structuring unit;
[0156] Among them, the above-mentioned first determining unit is used to: determine the target detection object in the above-mentioned original video; the above-mentioned second determining unit is used to: determine the target key event corresponding to the target detection object in the above-mentioned original video, where the information reflecting the target scene includes the above-mentioned target detection object and the target key event corresponding to the target detection object; and the above-mentioned structuring unit is used to: perform structuring processing on the above-mentioned target detection object and the target key event corresponding to the target detection object to obtain the target prompt information for the above-mentioned target scene.
[0157] In some embodiments, based on the above solution, the above-mentioned first determining unit is specifically used to: determine the target detection object in the above-mentioned original video based on the processing of the above-mentioned original video by a trained target detection model;
[0158] The above-mentioned second determining unit is specifically used to: after determining the target detection object, determine the target movement trajectory of the target detection object based on a trained trajectory determination model; and based on a trained behavior analysis model, determine the above-mentioned target key event according to the above-mentioned target detection object and the above-mentioned target movement trajectory.
[0159] In some embodiments, based on the above solution, the above-mentioned structuring unit is specifically used to: determine the target time information of the above-mentioned target scene in the above-mentioned original video; and perform structuring processing on the above-mentioned target time information, the above-mentioned target detection object, and the above-mentioned target key event to obtain the target prompt information for the above-mentioned target scene.
[0160] In some embodiments, based on the above solution, the above-mentioned explanatory video module 830 includes: a sub-video splitting unit and an information adding unit;
[0161] Among them, the above-mentioned sub-video segmentation unit is used to: segment the above-mentioned original video according to the above-mentioned target time information to obtain a plurality of sub-videos; and, the above-mentioned information addition unit is used to: add the process commentary information of the above-mentioned target scene to the sub-video corresponding to the above-mentioned target scene to obtain a commentary video.
[0162] In some embodiments, based on the above solution, the above-mentioned commentary video generation device 800 further includes: a personalized processing module;
[0163] Among them, the above-mentioned personalized processing module is used to: after the above-mentioned sub-video segmentation unit segments the above-mentioned original video according to the above-mentioned target time information to obtain a plurality of sub-videos, perform personalized processing on the commentary information corresponding to the target sub-video; and / or, add special effects processing between the first commentary information and the second commentary information respectively corresponding to two adjacent sub-videos.
[0164] In some embodiments, based on the above solution, the above-mentioned commentary information determination module 820 is specifically used to: input the prompt information of the target scene into the large language model to optimize the above-mentioned large language model; and, determine the commentary information about the above-mentioned target scene according to the output information of the above-mentioned large language model.
[0165] In some embodiments, based on the above solution, the above-mentioned commentary information includes pre-game commentary information about the event; the above-mentioned commentary video generation device 800 further includes: a first information acquisition module;
[0166] Among them, the above-mentioned first information acquisition module is used to: acquire the basic information of the detection object in the event; the above-mentioned prompt information determination module 810 is further used to: determine the first prompt information according to the above-mentioned basic information; the above-mentioned commentary information determination module 820 is further used to: input the above-mentioned first prompt information into the large language model, and determine the pre-game commentary information about the above-mentioned event according to the output of the above-mentioned large language model; the above-mentioned commentary video generation module 830 is further used to: determine the commentary video according to the above-mentioned pre-game commentary information, the commentary information about the above-mentioned target scene, and the above-mentioned original video.
[0167] In some embodiments, based on the above solution, the above-mentioned commentary information includes post-game commentary information about the event; the above-mentioned commentary video generation device 800 further includes: a second information acquisition module;
[0168] Among them, the above-mentioned second information acquisition module is used to: acquire summary information about the event; the above-mentioned prompt information determination module 810 is further used to: determine the second prompt information according to the above-mentioned summary information; the above-mentioned commentary information determination module 820 is further used to: input the above-mentioned second prompt information into a large language model, and determine the post-match commentary information about the above-mentioned event according to the output of the above-mentioned large language model; the above-mentioned commentary video generation module 830 is further used to: determine the commentary video according to the above-mentioned post-match commentary information, the commentary information for the above-mentioned target scenario, and the above-mentioned original video.
[0169] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 8 The shown device for generating the commentary video can execute the embodiments of the above-mentioned method for generating the commentary video, and the foregoing and other operations and / or functions of each module in the device are respectively for implementing the embodiments of the method for generating the commentary video. For the sake of brevity, it will not be elaborated here.
[0170] In the foregoing, the device for generating the commentary video according to the embodiments of the present application has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware of the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above-mentioned method embodiments.
[0171] Figure 9 is a schematic block diagram of an electronic device provided by an embodiment of the present application, Figure 9 The electronic device can be used to execute the above-mentioned method for generating the commentary video.
[0172] As Figure 9 shown, the electronic device 900 may include:
[0173] A memory 910 and a processor 920. The memory 910 is used to store a computer program 930 and transmit the computer program 930 to the processor 920. In other words, the processor 920 can call and run the computer program 930 from the memory 910 to implement the method for generating the commentary video provided by the embodiments of the present application.
[0174] For example, the processor 920 can be used to execute the steps in the method for generating the above-explained video according to the instructions in the computer program 930.
[0175] In some embodiments of the present application, the processor 920 may include but is not limited to:
[0176] a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.
[0177] In some embodiments of the present application, the memory 910 includes but is not limited to:
[0178] a volatile memory and / or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0179] In some embodiments of the present application, the computer program 930 may be divided into one or more modules, which are stored in the memory 910 and executed by the processor 920 to complete the method for generating the explanatory video provided in the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 930 in the electronic device.
[0180] As Figure 9 shown, the electronic device 900 may further include:
[0181] a transceiver 940, which may be connected to the processor 920 or the memory 910.
[0182] Among them, the processor 920 may control the transceiver 940 to communicate with other devices. Specifically, it may send information or data to other devices, or receive information or data sent by other devices. The transceiver 940 may include a transmitter and a receiver. The transceiver 940 may further include an antenna, and the number of antennas may be one or more.
[0183] It should be understood that each component in the electronic device 900 is connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.
[0184] According to one aspect of the present application, there is provided a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method in the above method embodiments. Or rather, the embodiments of the present application also provide a computer program product containing instructions. When the instructions are executed by the computer, the computer executes the method for generating the explanatory video provided in the above method embodiments.
[0185] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for generating the explanatory video provided in the above method embodiments.
[0186] In other words, when implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0187] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0188] In several embodiments provided in this application, it should be understood that the disclosed apparatus and method for generating an explanatory video can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, and the indirect couplings or communication connections of the devices or modules may be in electrical, mechanical, or other forms.
[0189] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in various embodiments of the present application, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0190] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for generating an explanatory video, characterized in that, The method includes: Determining target prompt information according to the information reflecting the target scene in the original video, where the target scene is any one of the scenes included in the original video; Inputting the target prompt information into a large language model, and determining the commentary information for the target scene according to the output of the large language model; Generating a commentary video according to the commentary information for the target scene and the original video.
2. The method according to claim 1, wherein The determining the target prompt information according to the information reflecting the target scene in the original video includes: Determining the target detection object in the original video; Determining the target key event corresponding to the target detection object in the original video, where the information reflecting the target scene includes the target detection object and the target key event corresponding to the target detection object; Structuring the target detection object and the target key event corresponding to the target detection object to obtain the target prompt information for the target scene.
3. The method according to claim 2, wherein The determining the target detection object in the original video includes: Based on the processing of the original video by the trained target detection model, determining the target detection object in the original video; The determining the target key event corresponding to the target detection object in the original video includes: After determining the target detection object, determining the target movement trajectory of the target detection object based on the trained trajectory determination model; Based on the trained behavior analysis model, determining the target key event according to the target detection object and the target movement trajectory.
4. The method according to claim 2, wherein The structuring the target detection object and the target key event corresponding to the target detection object to obtain the target prompt information for the target scene includes: Determining the target time information of the target scene in the original video; Structuring the target time information, the target detection object, and the target key event to obtain the target prompt information for the target scene.
5. The method according to claim 4, wherein The generating a commentary video according to the commentary information for the target scene and the original video includes: Segmenting the original video according to the target time information to obtain a plurality of sub-videos; Adding the commentary information for the target scene to the sub-video corresponding to the target scene to obtain a commentary video.
6. The method according to claim 5, characterized in that After segmenting the original video according to the target time information to obtain a plurality of sub-videos, the method further includes: Performing personalized processing on the commentary information corresponding to the target sub-video; and / or, Adding special effect processing between the first commentary information and the second commentary information corresponding to two adjacent sub-videos respectively.
7. The method according to any one of claims 1 to 6, characterized in that, The inputting the target prompt information into a large language model and determining the commentary information for the target scene according to the output of the large language model includes: Inputting the prompt information for the target scene into the large language model to optimize the large language model; Determining the commentary information for the target scene according to the output information of the large language model.
8. The method according to any one of claims 1 to 6, characterized in that The commentary information includes pre-game commentary information about the event; the method further includes: Obtaining the basic information about the detection object in the event; Determine the first prompt information according to the basic information; Input the first prompt information into a large language model, and determine the pre-game commentary information about the event according to the output of the large language model; Generating a commentary video according to the commentary information for the target scene and the original video includes: Determine the commentary video according to the pre-game commentary information, the commentary information for the target scene, and the original video.
9. The method according to any one of claims 1 to 6, characterized in that, The commentary information includes post-game commentary information about the event; the method further includes: Obtain summary information about the event; Determine the second prompt information according to the summary information; Input the second prompt information into a large language model, and determine the post-game commentary information about the event according to the output of the large language model; Generating a commentary video according to the commentary information for the target scene and the original video includes: Determine the commentary video according to the post-game commentary information, the commentary information for the target scene, and the original video.
10. A generating device for an explanatory video, characterized in that, The device includes: A prompt information determination module, configured to determine target prompt information according to information reflecting a target scene in the original video, where the target scene is any one of the scenes included in the original video; A commentary information determination module, configured to input the target prompt information into a large language model, and determine the commentary information for the target scene according to the output of the large language model; A commentary video generation module, configured to generate a commentary video according to the commentary information for the target scene and the original video.
11. An electronic device, characterized in that, Includes a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the method for generating a commentary video according to any one of claims 1 to 9 above.
12. A computer-readable storage medium, characterized in that, For storing computer programs; The computer program causes the computer to execute the method for generating a commentary video according to any one of claims 1 to 9 above.
Citation Information
Cited By
Football explanation method and device, computer equipment and storage medium
CN121151580A