Content generation method, device, storage medium, and program product

By generating personalized plot summaries and recap videos, the problem of long-form video platforms struggling to provide targeted recap content has been solved, increasing users' continued viewing and completion rates.

CN122160595APending Publication Date: 2026-06-05SHENZHEN SKYWORTH RGB ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN SKYWORTH RGB ELECTRONICS CO LTD
Filing Date
2026-04-01
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing long-form video platforms struggle to generate targeted and personalized replay content, leading users to easily abandon the video and pause for extended periods without resuming playback, resulting in a low completion rate.

Method used

By determining the historical viewing breakpoints of the target video content, semantic information of the plot script and/or keyframes of the viewed video content is extracted to generate personalized plot summaries, and a large language model is used to generate plot review videos.

Benefits of technology

It improves the ease with which users can follow the plot of long videos, increases the possibility of resuming playback after interruption, and improves the completion rate of long videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122160595A_ABST
    Figure CN122160595A_ABST
Patent Text Reader

Abstract

The application discloses a content generation method, device, storage medium and program product. First, in response to a video playing instruction of target video content, a historical viewing breakpoint moment of the target video content is determined, and viewed video content before the historical viewing breakpoint moment in the target video content is determined. Then, semantic information including a plot script and / or a key frame of the viewed video content is extracted. Finally, a plot abstract of the viewed video content is generated according to the semantic information. Thus, the plot abstract of the viewed video content is generated by using the semantic information of the viewed video content before the historical viewing breakpoint moment in the target video content, so that the user can quickly review the historical video content that has been played and viewed according to the plot abstract at a current viewing starting moment, thereby making it easier for the user to connect the plot of a long video, improving the possibility of breakpoint continuation playing of the long video, and finally improving the completion rate of the long video.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202610398547.4, filed on March 27, 2026, entitled “Content Generation Method, Apparatus, Storage Medium and Program Product”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the technical field of human-computer interaction, and in particular to content generation methods, content generation devices, storage media, and computer program products. Background Technology

[0003] Currently, when users watch long videos, if they interrupt the viewing for a period of time, they often forget what happened before when they resume watching. The current solutions typically involve long video platforms providing standardized highlights or previews for the next episode. However, the pre-episode content generated by these features is edited based on fixed video content and time points, failing to provide personalized recaps based on the user's individual breakpoint (i.e., which minute of the episode they were watching). This makes it difficult for users to connect the plot of long videos, making them more likely to abandon resuming playback and ultimately resulting in a low completion rate for long videos.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a content generation method, content generation device, storage medium, and computer program product, aiming to solve the technical problem that the current difficulty in generating targeted and personalized review content for long videos leads to users easily abandoning the video and pausing it for a long time without resuming playback, ultimately resulting in a low completion rate for long videos.

[0006] To achieve the above objectives, this application proposes a content generation method, which includes: In response to a video playback command for the target video content, determine the historical viewing breakpoint of the target video content and the video content that has been viewed before the historical viewing breakpoint in the target video content; Extract semantic information from the viewed video content, including the plot script and / or keyframes; Based on the semantic information, a plot summary of the watched video content is generated.

[0007] In one embodiment, after the step of generating a plot summary of the watched video content based on the semantic information, the method further includes: Based on the semantic information and the plot summary, a plot recap video of the watched video content is generated.

[0008] In one embodiment, after determining the historical viewing breakpoint time of the target video content and the video content already viewed before the historical viewing breakpoint time in the target video content, the method further includes: Determine the current viewing start time when the video playback command is issued, and the playback interval between the current viewing start time and the historical viewing breakpoint time; When the playback interval is greater than a preset upper limit threshold for the playback interval, the step of extracting the semantic information of the viewed video content is executed; When the playback interval is less than a preset lower threshold, the target video content is played.

[0009] In one embodiment, the step of extracting the semantic information of the viewed video content includes: Based on the playback interval, the historical duration for extracting semantic information of the watched video content is determined, and the semantic information of the watched video content for the historical duration is extracted, wherein the playback interval is proportional to the historical duration.

[0010] In one embodiment, the step of extracting the semantic information of the viewed video content further includes: Determine the device performance of the playback device for the target video content and the summary requirements in the video playback command; Based on the device performance and the playback requirements, semantic information of the watched video content is extracted. The higher the device performance and the more detailed the summary requirements, the more semantic content is contained in the semantic information of the watched video content.

[0011] In one embodiment, the method further includes: In response to a query command during the playback of the target video content, the query intent of the query command and the contextual information of the real-time screen during the playback of the target video content are determined. The contextual information includes a story script and / or keyframes of a preset duration prior to the query command. Based on the query intent, the search results are determined within the contextual information. Based on the search results, generate the query results corresponding to the query command.

[0012] In one embodiment, the method further includes: In the pre-built knowledge base of the target video content, the pre-stored knowledge corresponding to the query intent is retrieved, and the pre-stored knowledge is used as part of the information in the search results.

[0013] Furthermore, to achieve the above objectives, this application also proposes a content generation apparatus, which includes: The first module is used to respond to a video playback command for the target video content, and to determine the historical viewing breakpoint time of the target video content, as well as the video content that has been viewed before the historical viewing breakpoint time in the target video content. The second module is used to extract semantic information from the viewed video content, the semantic information including the plot script and / or keyframes; The third module is used to generate a plot summary of the watched video content based on the semantic information.

[0014] In addition, to achieve the above objectives, this application also proposes a content generation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the content generation method as described above.

[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the content generation method described above.

[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the content generation method described above.

[0017] One or more technical solutions proposed in this application have at least the following technical effects: In this application, firstly, in response to a video playback instruction for the target video content, the historical viewing breakpoint of the target video content and the video content already watched before the historical viewing breakpoint are determined. Then, semantic information, including the plot script and / or keyframes, of the watched video content is extracted. Finally, a plot summary of the watched video content is generated based on the semantic information. Thus, by utilizing the semantic information of the video content already watched before the historical viewing breakpoint, a plot summary of the user's watched video content is generated. This allows the user to quickly review their previously played and watched video content based on the plot summary at the current viewing start time, making it easier for the user to connect the plot of long videos, increasing the possibility of resuming playback of long videos from breakpoints, and ultimately improving the completion rate of long videos. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating the first embodiment of the content generation method of this application; Figure 2 A flowchart illustrating the second embodiment of the content generation method of this application; Figure 3 This application provides an application diagram illustrating the use of Method 1 for the content of this application. Figure 4 This is a schematic diagram of the module structure of the content generation device in the embodiments of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the content generation method in the embodiments of this application.

[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0024] Currently, when users watch long videos, if they interrupt the viewing for a period of time, they often forget what happened before when they resume watching. The current solutions typically involve long video platforms providing standardized highlights or previews for the next episode. However, the pre-episode content generated by these features is edited based on fixed video content and time points, failing to provide personalized recaps based on the user's individual breakpoint (i.e., which minute of the episode they were watching). This makes it difficult for users to connect the plot of long videos, making them more likely to abandon resuming playback and ultimately resulting in a low completion rate for long videos.

[0025] The main solution in this application's embodiments is: First, in response to the video playback command of the target video content, the historical viewing breakpoint of the target video content and the video content that has been viewed before the historical viewing breakpoint are determined; then, semantic information, including the plot script and / or keyframes, of the viewed video content is extracted; finally, a plot summary of the viewed video content is generated based on the semantic information.

[0026] Therefore, by utilizing the semantic information of the video content watched before the historical viewing breakpoint in the target video content, a plot summary of the video content watched by the user is generated. This allows the user to quickly review the historical video content they have played and watched based on the plot summary at the moment of starting the current viewing, making it easier for the user to connect the plot of long videos, increasing the possibility of resuming playback of long videos from breakpoints, and ultimately improving the completion rate of long videos.

[0027] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or content generation device capable of performing the above functions. The following description uses a content generation device as an example to illustrate this embodiment and the subsequent embodiments.

[0028] Based on this, embodiments of this application provide a content generation method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the content generation method of this application.

[0029] In this embodiment, the content generation method includes steps S10 to S30: Step S10: In response to the video playback command of the target video content, determine the historical viewing breakpoint time of the target video content and the video content that has been viewed before the historical viewing breakpoint time in the target video content. Understandably, the target video content can be online or offline. It can be videos from online education (such as automatically generated course reviews), live sports events (such as real-time tactical commentary and highlight replays), VR / AR immersive viewing, etc. For example, the target video content can be video content from TV series, movies, documentaries, or competitions.

[0030] The video playback command for the target video content can be a video playback command issued by the user on the playback device of the target video content, or it can be a video playback command issued by a digital device such as a content generation device that is communicatively connected to the playback device of the target video content. That is to say, the playback device and the content generation device of the target video content can be different devices or the same device. In addition, in this embodiment, there is no limitation on the method of issuing the video playback command, which can be a touch screen, voice, gesture, etc.

[0031] The historical viewing breakpoint of the target video content refers to the last moment when a user stopped watching during their historical viewing process. For example, if the target video content is a TV series, the corresponding historical viewing breakpoint could be the moment when the user exited watching a particular episode.

[0032] The viewed video content before the historical viewing breakpoint in the target video content can correspond to different types and formats of target video content. For example, if the target video content is a TV series, the viewed video content is the watched episodes such as episodes 1-5. If the target video content is a movie, the viewed video content is the movie content that has been played, such as the content from 00:00:00 to 1:20:26 (hours:minutes:seconds).

[0033] It should be noted that the historical viewing breakpoint of the target video content and the video playback command corresponding to the target video content can be the same or different. Whether they are the same or different does not affect the implementation and effect of the content generation method in this embodiment. For example, if the historical viewing breakpoint of the target video content is episode 3 of a TV series, the video playback command corresponding to the target video content can be episode 3 or episode 5.

[0034] In one embodiment, breakpoint recording is achieved by recording the user's viewing progress (Timestamp) in real time, thereby saving the user's historical viewing breakpoints for the target video content, as well as the video content that has been watched before the historical viewing breakpoints.

[0035] Step S20: Extract semantic information of the viewed video content, including the plot script and / or keyframes; In this embodiment, semantic information of the viewed video content can be extracted using a large language model. As a preferred embodiment, this can be done by taking into account the user's need for plot review, such as full plot review or simplified plot review, and / or the performance of the playback device and content generation device of the target video content, and extracting all plot scripts and some keyframes, thereby improving the efficiency and accuracy of semantic information extraction.

[0036] In one feasible implementation, steps A1 to A3 may be included after step S10: Step A1: Determine the current viewing start time when the video playback command is issued, and the playback interval between the current viewing start time and the historical viewing breakpoint. Step A2: When the playback interval is greater than the preset upper limit threshold for the playback interval, execute step S20. Step A3: Play the target video content when the playback interval is less than the preset lower threshold of the playback interval.

[0037] Unlike traditional video platforms where all users see the same highlights, this embodiment generates plot recap content in real time based on the user's current viewing start time when they issue the video playback command, the playback interval (Time-Gap) between the current viewing start time and the historical viewing breakpoint, and the viewing progress (Watch-History). When the playback interval is greater than a preset upper threshold (e.g., 12 hours or 1 month), step S20 and subsequent steps are executed to generate plot recap content; when the playback interval is less than a preset lower threshold (e.g., 1 hour or 1 week), the target video content is played directly without generating plot recap content.

[0038] In one embodiment, when the playback interval is greater than the preset upper limit threshold of the playback interval, step S20 is executed to generate detailed plot review content and perform a detailed plot summary; when the playback interval is less than the preset lower limit threshold of the playback interval, step S20 is executed to generate brief plot review content, which may only generate a one-sentence summary.

[0039] Steps A1 to A3 above, through playback intervals, enable the generation of personalized dynamic summaries tailored to each individual, thereby improving the efficiency and accuracy of plot recap generation.

[0040] In one feasible implementation, step S20 may include the following steps: Based on the playback interval, the historical duration for extracting semantic information of the viewed video content is determined, and the semantic information of the viewed video content for the historical duration is extracted. The playback interval is proportional to the historical duration.

[0041] When extracting semantic information from viewed video content, the extraction duration and amount of content are determined based on the playback interval. This means extracting semantic information from viewed video content of historical duration. The playback interval is directly proportional to the historical duration; a larger playback interval results in a longer extracted historical duration, and vice versa. Therefore, for video content with a large playback interval, a longer extraction duration can increase the detail and coverage of the plot recap, while for video content with a small playback interval, a shorter extraction duration can increase the accuracy of the plot recap.

[0042] In another feasible implementation, step S20 may further include the following steps: Determine the device performance of the playback device for the target video content and the summary requirements in the video playback instructions; Based on device performance and playback requirements, semantic information of watched video content is extracted. The higher the device performance and the more detailed the summary requirements, the more semantic content is contained in the semantic information of watched video content.

[0043] The device performance of the playback device for the target video content can be evaluated and determined through the following steps: Collect the hardware parameters and real-time status data of the playback device, including but not limited to CPU computing power, GPU performance, memory capacity, storage read / write speed, network bandwidth, and current load. Using a preset performance evaluation model, quantify the above parameters into a device performance level, such as high, medium, and low, or a continuous numerical score.

[0044] Regarding the summary requirements in video playback instructions, the summary requirements can be parsed and determined through the following steps: Receive the summary requirement instruction input by the user through an interactive interface, voice, gesture, or other human-computer interaction interface. This instruction may include summary type (such as keyframe summary, text summary, structured event line, etc.), level of detail (such as concise, standard, detailed), and key video elements (such as dialogue, scene changes, specific objects, action sequences, etc.). Then, parse the summary requirement into a set of structured parameters.

[0045] In one embodiment, a corresponding semantic extraction strategy can be matched from a predefined strategy matrix based on the device performance level and summary detail parameters determined in the preceding steps. The strategy defines the following adjustable dimensions: 1. Analysis depth: granularity of visual feature extraction (e.g., frame-level, segment-level, scene-level), and complexity of natural language processing (e.g., keyword extraction, entity recognition, sentiment analysis, relational reasoning). 2. Processing breadth: the number of analysis modules activated (e.g., face recognition, object detection, action recognition, speech-to-text, text summarization, etc.). 3. Resource allocation scheme: determining the resource scheduling method of local computing, cloud collaboration, or hierarchical processing.

[0046] If the device has high performance and the summary requirement is detailed, a multi-level, multi-modal deep analysis module is enabled. For example, high-precision scene segmentation is performed on the video, combined with speech recognition and natural language understanding to generate detailed description text, extract key objects and relationships between people, and generate a structured timeline. The processing can fully utilize local high-performance hardware for parallel computing. If the device performance is average and the requirement is standard, a balanced strategy is adopted, such as performing medium-granularity scene detection, combining basic speech recognition to generate summary text, and extracting significant keyframes and main entities. If the device performance is low or the requirement is a concise summary, a lightweight strategy is adopted, such as extracting only representative keyframes, or using a lightweight cloud API for fast speech-to-text conversion and generating a short summary to reduce the local computing load.

[0047] Finally, semantic information is synthesized and output: the extracted semantic elements (such as text descriptions, keyframe timestamps, tags, entity relationship graphs, etc.) are integrated according to the user's required format and output as a unified semantic information report.

[0048] In this way, by dynamically adapting to device performance and user needs, the semantic information extraction becomes flexible and intelligent. Specifically, it achieves resource optimization: avoiding processing bottlenecks on low-performance devices while fully utilizing the computing power potential of high-performance devices. It also enhances the user experience: providing users with just the right amount of semantic information richness to meet the needs of different scenarios, from quick review to in-depth analysis. Furthermore, it is highly scalable: the strategy matrix can be continuously updated with technological advancements, supporting new analysis modules and required parameters.

[0049] Step S30: Generate a plot summary of the watched video content based on semantic information.

[0050] A plot summary of watched video content refers to a summary and overview of the plot of the watched video content. In one embodiment, a plot summary of watched video content can be generated based on a large language model using plot scripts and / or keyframes.

[0051] In one possible implementation, step S30 may be followed by step S40: Step S40: Generate a plot recap video of the watched video content based on semantic information and plot summary.

[0052] A recap video of the viewed video content refers to a summary and overview of the plot of the viewed video content presented in video format. In one embodiment, a recap video of the viewed video content can be generated using multimodal synthesis technology based on the plot script and / or keyframes, and the plot summary determined in step S30. In one embodiment, a large language model is used to summarize the plot of the watched video content into a textual summary, generating a recap script. This can be combined with TTS speech synthesis and key scene editing technology to generate a recap video of the watched video content through multimodal synthesis, which is then automatically played before the main feature.

[0053] In this way, accurate plot summaries can be automatically generated based on the user's personalized viewing history, helping the user quickly recall the plot. These summaries are dynamically generated based on previously watched video content that the user may have forgotten, making them more user-centric than fixed edits and effectively reducing the rate of abandoning shows.

[0054] Currently, during video playback and viewing, users often have questions about plot details, actor information, or background knowledge, such as "Who is this person?" or "What is this foreshadowing?" The current approach to answering these questions often requires users to pause playback and manually search for answers on a search engine, which interrupts the user's immersion in the viewing experience. Furthermore, existing voice assistants typically only handle simple device controls, such as adjusting the volume and brightness of playback devices, and cannot understand the deeper semantic information of video content. In other words, users face obstacles in obtaining information while watching video content.

[0055] Therefore, based on the first embodiment of this application, the content that is the same as or similar to the first embodiment in the second embodiment of this application can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 The content generation method further includes steps T10 to T30: Step T10: In response to a query command during the playback of the target video content, determine the query intent of the query command and the contextual information of the real-time screen during the playback of the target video content. The contextual information includes a story script and / or keyframes of a preset duration prior to the query command. The query command during the playback of the target video content is similar to the video playback command in the first embodiment. The query command can be issued by the user on the playback device of the target video content, or it can be issued by a digital device such as a content generation device that is communicatively connected to the playback device of the target video content. Furthermore, in this embodiment, the method of issuing the query command is not limited; it can be via touchscreen, voice, gestures, etc.

[0056] In one embodiment, the query intent of a query instruction is determined by intent recognition technology. In this embodiment, the method for determining the query intent of a query instruction is not limited.

[0057] The contextual information of the real-time frame during the playback of the target video content refers to the plot script and / or keyframes of the target video content that has been played for a preset duration before the query command. The method for extracting the contextual information of the real-time frame during the playback of the target video content can refer to step S20 in the first embodiment, and is similar to the extraction method in step S20, so it will not be described again here.

[0058] Step T20: Determine the search results based on the contextual information according to the query intent; Once a user submits a query and determines the intent behind it, the context information corresponding to the query intent can be retrieved from the contextual information and used as the search result.

[0059] Step T30: Based on the search results, generate the query results corresponding to the query command.

[0060] In one embodiment, a large language model generates query results corresponding to the query instruction based on the retrieval results, i.e., based on the contextual information corresponding to the query intent.

[0061] In this way, context awareness during video playback becomes possible. That is, the system can understand the current frame and plot of the target video content, providing real-time and accurate answers to user questions and queries. Users can obtain in-depth plot analysis without leaving the playback screen, thus achieving immersive interaction.

[0062] In one feasible implementation, the content generation method further includes the step of: In a pre-built knowledge base of the target video content, retrieve the pre-stored knowledge corresponding to the query intent, and use the pre-stored knowledge as part of the information in the search results.

[0063] In this embodiment, a knowledge base for the target video content is pre-built. This knowledge base may include pre-stored knowledge such as character relationship graphs, episode plots, and behind-the-scenes footage from the target video content. When answering user questions or query commands, this knowledge base can be retrieved to ensure the accuracy of the answer and avoid model illusion.

[0064] In one embodiment, when a user initiates a voice question, such as "Why is the female protagonist crying?", an explanatory answer is generated by combining the contextual information corresponding to the query intent with the plot knowledge base, and the explanatory answer is fed back to the user through picture-in-picture text or voice.

[0065] Compared to traditional speech recognition and voice answering, which can only search for movie titles and achieve simple device control, this embodiment utilizes a pre-built vector database, i.e. a knowledge base, to store pre-stored knowledge such as scripts and metadata, thereby achieving deep semantic understanding and real-time question answering of target video content.

[0066] In the application scenario of Method 1 for generating content in this application, refer to Figure 3When a user clicks play, if the playback interval is less than the preset lower threshold (i.e., a short interval), or when watching the target video content for the first time, the target video content is played directly. If the playback interval is greater than the preset upper threshold (i.e., a long interval), the system responds to the video playback command of the target video content by determining the historical viewing breakpoint and the viewed video content before that breakpoint. The system extracts the plot script and / or keyframes of the viewed video content. Based on semantic information, it uses a large language model to generate a plot summary of the viewed video content, which is a summary text. Then, using TTS speech synthesis and keyframe editing technology, a multimodal synthesis is used to generate a plot recap video of the viewed video content, which is automatically played before the main feature. For example, if a user clicks to play episode 5 of "TV Series A," the system detects that the user last watched up to the end of episode 3, and that more than 7 days have passed since the last viewing. At this point, the plot outlines of episodes 1 to 3 are extracted, along with their keyframes. An LLM (Large Language Model) is then used to generate a 300-word recap. Before episode 5 begins, a one-minute recap video is created, narrated by AI: "Last time we talked about…".

[0067] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method of generating the content of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0068] This application also provides a content generation apparatus, please refer to... Figure 4 The content generation device includes: The first module 10 is used to determine the historical viewing breakpoint time of the target video content and the video content that has been watched before the historical viewing breakpoint time in the target video content in response to the video playback command of the target video content. The second module 20 is used to extract semantic information of the viewed video content, the semantic information including the plot script and / or keyframes; The third module 30 is used to generate a plot summary of the watched video content based on the semantic information.

[0069] In one embodiment, the content generation apparatus further includes a fourth module for: After the step of generating a plot summary of the watched video content based on the semantic information: a plot recap video of the watched video content is generated based on the semantic information and the plot summary.

[0070] In one embodiment, the content generation apparatus further includes a fifth module for: After determining the historical viewing breakpoint time of the target video content and the video content already viewed before the historical viewing breakpoint time in the target video content: determine the current viewing start time when the video playback command is issued, and the playback interval between the current viewing start time and the historical viewing breakpoint time; when the playback interval is greater than a preset upper limit threshold for playback interval, perform the step of extracting the semantic information of the viewed video content; when the playback interval is less than a preset lower limit threshold for playback interval, play the target video content.

[0071] In one embodiment, the fifth module is further configured to: Based on the playback interval, the historical duration for extracting semantic information of the watched video content is determined, and the semantic information of the watched video content for the historical duration is extracted, wherein the playback interval is proportional to the historical duration.

[0072] In one embodiment, the fifth module is further configured to: Determine the device performance of the playback device for the target video content and the summary requirements in the video playback command; Based on the device performance and the playback requirements, semantic information of the watched video content is extracted. The higher the device performance and the more detailed the summary requirements, the more semantic content is contained in the semantic information of the watched video content.

[0073] In one embodiment, the content generation apparatus further includes a sixth module for: In response to a query command during the playback of the target video content, the query intent of the query command and the contextual information of the real-time screen during the playback of the target video content are determined. The contextual information includes a story script and / or keyframes of a preset duration prior to the query command. Based on the query intent, the search results are determined within the contextual information. Based on the search results, generate the query results corresponding to the query command.

[0074] In one embodiment, the sixth module is further configured to: In the pre-built knowledge base of the target video content, the pre-stored knowledge corresponding to the query intent is retrieved, and the pre-stored knowledge is used as part of the information in the search results.

[0075] The content generation apparatus provided in this application, employing the content generation method described in the above embodiments, can solve the technical problem of low completion rates for long videos due to the difficulty in generating targeted and personalized review content, which often leads to users abandoning the video after prolonged pauses and failure to resume playback. Compared with the prior art, the beneficial effects of the content generation apparatus provided in this application are the same as those of the content generation method described in the above embodiments, and other technical features in the content generation apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0076] This application provides a content generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the content generation method in Embodiment 1 above.

[0077] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a content generation device suitable for implementing embodiments of this application. The content generation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The content generation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0078] like Figure 5As shown, the content generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the content generation device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the content generating device to communicate wirelessly or wiredly with other devices to exchange data. While the figures show content generating devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0079] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0080] The content generation device provided in this application, employing the content generation method described in the above embodiments, can solve the technical problem of low completion rates for long videos due to the difficulty in generating targeted and personalized review content, which often leads to users abandoning the video after prolonged pauses and failure to resume playback. Compared with the prior art, the beneficial effects of the content generation device provided in this application are the same as those of the content generation method described in the above embodiments, and other technical features of this content generation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0081] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0082] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0083] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the content generation method in the above embodiments.

[0084] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0085] The aforementioned computer-readable storage medium may be included in the content generating device or may exist independently and not assembled into the content generating device.

[0086] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a content generation device, cause the content generation device to: in response to a video playback instruction for target video content, determine the historical viewing breakpoint of the target video content and the video content already viewed before the historical viewing breakpoint; extract semantic information from the viewed video content, the semantic information including a plot script and / or keyframes; and generate a plot summary of the viewed video content based on the semantic information.

[0087] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0089] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0090] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described content generation method. This solves the technical problem that, due to the difficulty in generating targeted and personalized review content for long videos, users are prone to abandoning long videos that have been paused for extended periods without resuming playback, ultimately leading to a low completion rate for long videos. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the content generation method provided in the above embodiments, and will not be repeated here.

[0091] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the content generation method described above.

[0092] The computer program product provided in this application can solve the technical problem that the difficulty in generating targeted and personalized review content for long videos leads to users easily abandoning long videos that have been paused for a long time without resuming playback, ultimately resulting in a low completion rate for long videos. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the content generation method provided in the above embodiments, and will not be repeated here.

[0093] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A content generation method, characterized in that, The content generation method includes: In response to a video playback command for the target video content, determine the historical viewing breakpoint of the target video content and the video content that has been viewed before the historical viewing breakpoint in the target video content; Extract semantic information from the viewed video content, including the plot script and / or keyframes; Based on the semantic information, a plot summary of the watched video content is generated.

2. The content generation method as described in claim 1, characterized in that, After the step of generating a plot summary of the watched video content based on the semantic information, the method further includes: Based on the semantic information and the plot summary, a plot recap video of the watched video content is generated.

3. The content generation method as described in claim 1, characterized in that, The step of determining the historical viewing breakpoint time of the target video content and the video content already viewed before the historical viewing breakpoint time in the target video content further includes: Determine the current viewing start time when the video playback command is issued, and the playback interval between the current viewing start time and the historical viewing breakpoint time; When the playback interval is greater than a preset upper limit threshold for the playback interval, the step of extracting the semantic information of the viewed video content is executed; When the playback interval is less than a preset lower threshold, the target video content is played.

4. The content generation method as described in claim 3, characterized in that, The step of extracting semantic information from the viewed video content includes: Based on the playback interval, the historical duration for extracting semantic information of the watched video content is determined, and the semantic information of the watched video content for the historical duration is extracted, wherein the playback interval is proportional to the historical duration.

5. The content generation method as described in claim 3, characterized in that, The step of extracting the semantic information of the viewed video content further includes: Determine the device performance of the playback device for the target video content and the summary requirements in the video playback command; Based on the device performance and the playback requirements, semantic information of the watched video content is extracted. The higher the device performance and the more detailed the summary requirements, the more semantic content is contained in the semantic information of the watched video content.

6. The content generation method as described in claim 1, characterized in that, The method further includes: In response to a query command during the playback of the target video content, the query intent of the query command and the contextual information of the real-time screen during the playback of the target video content are determined. The contextual information includes a story script and / or keyframes of a preset duration prior to the query command. Based on the query intent, the search results are determined within the contextual information. Based on the search results, generate the query results corresponding to the query command.

7. The content generation method as described in claim 6, characterized in that, The method further includes: In the pre-built knowledge base of the target video content, the pre-stored knowledge corresponding to the query intent is retrieved, and the pre-stored knowledge is used as part of the information in the search results.

8. A content generation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the content generation method as described in any one of claims 1 to 7.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the content generation method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the content generation method as described in any one of claims 1 to 7.