Method and device for inserting content in video

By analyzing video content and user interaction data, determining the insertion time point and content, the problem of traditional video insertion methods not matching user needs is solved, and user experience and interactivity are improved.

CN120302109APending Publication Date: 2025-07-11SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510505500.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional video content recommendation and insertion methods cannot match user needs, resulting in disrupting user experience.

Method used

By analyzing the video content and user interaction data of the target video, determine the content to be inserted and its insertion time point to ensure that the content is highly matched with the video context and user needs.

Benefits of technology

Improves user experience, enhances personalization and interactivity of content, and ensures that inserting content does not interfere with user viewing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302109A_ABST
    Figure CN120302109A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for inserting content in a video, computer equipment, a medium and a program product. Relates to the technical field of computers. The method comprises the following steps: acquiring video data of a target video, wherein the video data comprises video content and / or interaction data for the target video; determining to-be-inserted content of the target video and an insertion time point of the to-be-inserted content based on the video data; and inserting the to-be-inserted content into the target video based on the insertion time point. According to the technical scheme, it can be ensured that the to-be-inserted content is highly matched with the context of the target video and the user requirement, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product for inserting content into a video. Background Art

[0002] With the explosive growth of video content, users' demands for personalization, interactivity, and real-time performance are increasing day by day. Traditional video content recommendation and video content insertion usually recommend content based on simple rules or based on users' historical behaviors first, and then insert the recommended content into fixed positions. This way of video content recommendation and insertion easily leads to a mismatch between the inserted content and users' demands, and even interferes with the user experience.

[0003] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention

[0004] Embodiments of the present application provide a method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product for inserting content into a video to solve or alleviate one or more of the above technical problems.

[0005] One aspect of the embodiments of the present application provides a method for inserting content into a video, the method including: determining content to be inserted into the target video and an insertion time point of the content to be inserted based on the video content of the target video and / or interaction data for the target video; inserting the content to be inserted into the target video based on the insertion time point.

[0006] Optionally, the determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: detecting shot transition points in the video content of the target video to obtain shot transition points; using the shot transition points as the insertion time points of the content to be inserted; using the content to be inserted of the first type as the content to be inserted into the target video.

[0007] Optionally, the determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: performing emotion detection processing on the video content of the target video to obtain an emotion detection result; determining emotional high points in the target video according to the emotion detection result, and using the emotional high points as the insertion time points of the content to be inserted; Use the content to be inserted of the second type as the content to be inserted into the target video.

[0008] Optionally, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Perform information-intensive detection processing on the video content of the target video to obtain information-intensive points; Use the information-intensive points as the insertion time points of the content to be inserted; Use the content to be inserted of the third type as the content to be inserted into the target video.

[0009] Optionally, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Perform information-relaxed detection processing on the video content of the target video to obtain information-relaxed points; Use the information-relaxed points as the insertion time points of the content to be inserted; Use the content to be inserted of the fourth type as the content to be inserted into the target video.

[0010] Optionally, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Perform plot turning point prediction processing on the video content of the target video to obtain plot turning points; Use the plot turning points as the insertion time points of the content to be inserted; Use the content to be inserted of the fifth type as the content to be inserted into the target video.

[0011] Optionally, the interaction data includes at least one of the following: Fast-forward data for a target video segment, skip data for a target video segment, pause data for a target video segment, playback data for a target video segment, long-time stay data for a target video segment, high-frequency interaction point data.

[0012] Optionally, inserting the content to be inserted into the target video based on the insertion time point includes: Detect the external environment of the target video, and determine the insertion method of the content to be inserted according to the detection result; Insert the content to be inserted into the target video in the determined insertion method based on the insertion time point.

[0013] Optionally, the external environment includes at least one of the following: The device type for playing the target video and the network condition where the device for playing the target video is currently located.

[0014] Another aspect of the embodiments of the present application provides a device for inserting content into a video. The device includes: An acquisition module, configured to acquire video data of a target video, where the video data includes video content and / or interaction data for the target video; A determination module, configured to determine the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data; An insertion module, configured to insert the content to be inserted into the target video based on the insertion time point.

[0015] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.

[0016] Another aspect of the embodiments of the present application provides a computer-readable storage medium, where computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0017] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0018] The embodiments of the present application adopting the above technical solutions may include the following advantages: By analyzing the video content of the target video and / or the interaction data of the user with respect to the target video, it is possible to ensure that the content to be inserted highly matches the context of the target video and the user's needs, thereby improving the user experience. Description of the Drawings

[0019] The drawings exemplarily show embodiments and constitute a part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0020] Figure 1 Schematically shows an operating environment diagram of a method for inserting content into a video according to Embodiment 1 of the present application; Figure 2Schematically shows a flowchart of a method for inserting content into a video according to Embodiment 1 of the present application; Figure 3 Schematically shows a refined flowchart of steps for determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data in one embodiment; Figure 4 Schematically shows a refined flowchart of steps for determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data in another embodiment; Figure 5 Schematically shows a refined flowchart of steps for determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data in another embodiment; Figure 6 Schematically shows a refined flowchart of steps for determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data in another embodiment; Figure 7 Schematically shows a refined flowchart of steps for determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data in another embodiment; Figure 8 Schematically shows a refined flowchart of steps for inserting the content to be inserted into the target video based on the insertion time point; Figure 9 Schematically shows a block diagram of a device for inserting content into a video according to Embodiment 2 of the present application; and Figure 10 Schematically shows a schematic diagram of the hardware architecture of a computer device according to Embodiment 3 of the present application. Detailed implementation manners

[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0022] It should be noted that in the embodiments of the present application, the descriptions involving "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0023] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and to distinguish each step, and thus should not be construed as a limitation to the present application.

[0024] To facilitate the understanding of those skilled in the art of the technical solutions provided by the embodiments of the present application, the related technologies are described below: With the explosive growth of video content, users' demands for personalization, interactivity, and real-time are increasing day by day. Traditional video content recommendation and video content insertion usually first recommend content based on simple rules or based on the user's historical behavior, and then insert the recommended content into a fixed position. This way of video content recommendation and insertion is likely to result in the inserted content not matching the user's needs, and even interfering with the user experience.

[0025] For this reason, the embodiments of the present application provide a technical solution for inserting content into a video. In this technical solution, by analyzing the video content of the target video and / or the interaction data of the user with respect to the target video, it can be ensured that the content to be inserted highly matches the context of the target video and the user's needs, thereby improving the user experience. See the following for details.

[0026] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0027] As Figure 1 shown, the schematic diagram of the environment includes a service platform 2, a network 4, and a client 6, where: The service platform 2 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing devices can load virtual machines based on virtual images and / or other data that define specific software (e.g., operating systems, dedicated applications, servers) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.

[0028] The service platform 2 can be configured to communicate with the client 6 etc. via the network 4. The network 4 includes various network devices such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 can include physical links such as coaxial cable links, twisted pair cable links, fiber optic links, combinations thereof, etc., or wireless links such as cellular links, satellite links, Wi-Fi links, etc.

[0029] The service platform 2 can provide services such as storage, reading, writing, querying, deleting, etc., such as providing a service for inserting content into a video for the client.

[0030] The client 6 can be an electronic device running an operating system such as Windows, Android™, or iOS, such as a smart phone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, an in-vehicle terminal, a smart TV. Based on the above operating systems, various application programs can be run, such as a program for inserting content into a video.

[0031] The client 6 can provide / configure a user access page for manipulating the service platform 2 or uploading an object, etc.

[0032] It should be noted that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices can be adjusted.

[0033] The technical solutions of the present application will be introduced below through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.

[0034] Embodiment 1 Figure 2 A flowchart of a method for inserting content into a video according to Embodiment 1 of the present application is schematically shown.

[0035] As Figure 2 shown, the method for inserting content into a video may include steps S200 to S204, where: Step S200, obtaining video data of a target video, where the video data includes video content and / or interaction data for the target video.

[0036] Step S202, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data.

[0037] Step S204, inserting the content to be inserted into the target video based on the insertion time point.

[0038] The method for inserting content into a video provided in this embodiment analyzes the video content of the target video and / or the interaction data of the user with respect to the target video, so as to ensure that the content to be inserted highly matches the context of the target video and the user needs, thereby improving the user experience.

[0039] The following combines Figure 1 to elaborate in detail each step in steps S200 to S204 and other optional steps.

[0040] Step S200 to obtain the video data of the target video, where the video data includes video content and / or interaction data with respect to the target video.

[0041] The target video is a video into which content needs to be inserted, and the type of the target video can be an on-demand video or a live video.

[0042] The video content may include the picture content, audio content, and subtitle text content of the target video.

[0043] The interaction data may include fast-forward data for a target video segment, skip data for a target video segment, pause data for a target video segment, playback data for a target video segment, long-time stay data for a target video segment, and high-frequency interaction point data.

[0044] Among them, the target video segment can be any scene segment in the target video. The scene segment is obtained by segmenting the target video based on the scene switching points in the target video. For example, if there are 5 scene switching points in the target video, the target video can be divided into 6 scene segments.

[0045] The high-frequency interaction point is determined based on the density of at least one of the bullet screens, likes, and comments in the target video. For example, first, the bullet screens in the target video can be counted to find out the video segments with the top N bullet screen densities, and then these N found video segments are used as high-frequency interaction video segments. Finally, any time point in the high-frequency interaction video segment can be used as the high-frequency interaction point.

[0046] Step S202 to determine the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data.

[0047] The content to be inserted is the content that needs to be inserted into the target video. The format of the content to be inserted can be a picture, text, animation, video, etc.

[0048] The insertion time point is used to characterize the insertion opportunity of the content to be inserted in the target video.

[0049] In actual application, the content to be inserted in the target video and the insertion time point of the content to be inserted can be determined in various ways. An exemplary way is provided below.

[0050] In an alternative embodiment, refer to Figure 3 , determining the content to be inserted in the target video and the insertion time point of the content to be inserted based on the video data includes: Step S300, detecting shot change points in the video content of the target video to obtain shot change points.

[0051] Step S302, using the shot change points as the insertion time points of the content to be inserted.

[0052] Step S304, using the content to be inserted of the first type as the content to be inserted in the target video.

[0053] In this embodiment, the video content is video frame images.

[0054] In some embodiments, the Scale-Invariant Feature Transform (SIFT) or Speeded-Up Robust Features (SURF) algorithm can be used to extract key points from the video frame images of the target video first. Then, the optical flow analysis (OpticalFlow) can be combined to track the visual changes of the key points between frames. Finally, the shot change points are determined according to the visual changes. In a specific application, when the visual change exceeds a certain threshold, it can be considered that a shot change has occurred. Correspondingly, the time point corresponding to the video frame where the shot change occurs is the shot change point, which is also the insertion time point of the content to be inserted in the target video.

[0055] It should be noted that the key points refer to the points with significant features in the image. These points have unique features in the image and can be stably detected under different perspectives, illuminations, and noise conditions.

[0056] In other embodiments, the TransNetV2 model based on deep learning can also be used to analyze the video frames in the target video to identify the shot change points.

[0057] After obtaining the shot change points, the content to be inserted of the first type can be used as the content to be inserted in the target video. Among them, the content to be inserted of the first type can be lightweight prompt information or bullet screen information.

[0058] It should be noted that the specific content of the prompt information or bullet screen information is not limited. For example, it can be "A camera switch has occurred currently."

[0059] In this embodiment, by inserting lightweight prompt information or bullet screen information at the camera switch point, while achieving the reminder for users, it will not interrupt the users from watching the target video, thus improving the user experience.

[0060] The following provides another exemplary method for determining the content to be inserted and the insertion time point of the content to be inserted in the target video.

[0061] In an alternative embodiment, refer to Figure 4 The determining of the content to be inserted and the insertion time point of the content to be inserted in the target video based on the video data includes: Step S400: Perform an emotion detection process on the video content of the target video to obtain an emotion detection result.

[0062] Step S402: Determine the emotional high point in the target video according to the emotion detection result, and use the emotional high point as the insertion time point of the content to be inserted.

[0063] Step S404: Use the content to be inserted of the second type as the content to be inserted in the target video.

[0064] In this embodiment, the video content includes at least one of video frame images and audio data.

[0065] In one implementation manner, an emotion analysis algorithm can be used to perform an emotion detection process on the user emotion in the video frame images of the target video to obtain an emotion detection result. Among them, the emotion detection result can include happy, sad, surprised, etc.

[0066] It should be noted that the above emotion analysis algorithm can be a convolutional neural network model based on facial expressions, a video emotion analysis model based on a pre-trained model, etc.

[0067] In one implementation manner, the Mel-Spectrogram of the audio data in the target video can be extracted first, and then the audio emotion can be classified by combining natural language processing (NLP) technology to obtain an emotion detection result. In another implementation manner, the Mel Frequency Cepstral Coefficients (MFCC) of the audio data in the target video can also be extracted first, and then the extracted Mel Frequency Cepstral Coefficients are input into an LSTM model for training. By the LSTM model learning the relationship between audio features and emotions, the classification of audio emotions can be realized to obtain an emotion detection result.

[0068] After obtaining the sentiment detection result, the change trend of the sentiment in the target video can be analyzed based on the sentiment detection result, so as to identify the emotional high point and emotional low point in the target video. After identifying the emotional high point, the time point corresponding to the video frame of the emotional high point can be used as the insertion time of the content to be inserted in the target video.

[0069] After obtaining the emotional high point, the content to be inserted of the second type can be used as the content to be inserted in the target video. Among them, the content to be inserted of the second type can be interactive content such as voting, plot prediction, etc.

[0070] It should be noted that the interactive content can be determined according to the context of the current target video.

[0071] In this embodiment, by inserting interactive content at the emotional high point, the interactivity can be enhanced, and the user's participation and stickiness can be improved.

[0072] The following provides another exemplary way to determine the content to be inserted in the target video and the insertion time point of the content to be inserted.

[0073] In an alternative embodiment, refer to Figure 5 , determining the content to be inserted in the target video and the insertion time point of the content to be inserted based on the video data includes: Step S500, perform information density detection processing on the video content of the target video to obtain information density points.

[0074] Step S502, use the information density points as the insertion time points of the content to be inserted.

[0075] Step S504, use the content to be inserted of the third type as the content to be inserted in the target video.

[0076] In this embodiment, the video content includes at least one of subtitle text and audio data.

[0077] In one implementation, NLP technology can be used to perform information density analysis on the subtitle text in the target video to obtain the information density. Among them, the information density can be the keyword frequency, sentence complexity, etc.

[0078] In one implementation, the speech rate of the audio data of the target video can also be analyzed to obtain the speech rate analysis result.

[0079] After obtaining the speech rate analysis result, the knowledge-intensive video segment can be determined by combining the information density. After that, any time point in the knowledge-intensive video segment can be used as the information density point.

[0080] After obtaining the information-intensive points, the content to be inserted of the third type can be used as the content to be inserted in the target video. Among them, the content to be inserted of the third type can be explanatory captions, knowledge supplement content, etc.

[0081] In this embodiment, by inserting explanatory captions, knowledge supplement content, etc. at the information-intensive points, the user's ability to understand the content of the target video can be enhanced, and the user experience can be improved.

[0082] The following provides another exemplary method for determining the content to be inserted in the target video and the insertion time point of the content to be inserted.

[0083] In an alternative embodiment, refer to Figure 6 The determining of the content to be inserted in the target video and the insertion time point of the content to be inserted based on the video data includes: Step S600: Perform information moderation detection processing on the video content of the target video to obtain information moderation points.

[0084] Step S602: Use the information moderation points as the insertion time points of the content to be inserted.

[0085] Step S604: Use the content to be inserted of the fourth type as the content to be inserted in the target video.

[0086] In this embodiment, the video content is video frame images.

[0087] In one implementation, visual change detection can be performed on consecutive video frame images to detect whether the pixel difference value between video frame images is less than a set threshold. If it is less than the set threshold, these consecutive video frame images can be used as content moderation video segments. Then, any time point in the content moderation video segments can be used as an information moderation point.

[0088] After obtaining the information moderation points, the content to be inserted of the fourth type can be used as the content to be inserted in the target video. Among them, the content to be inserted of the fourth type can be recommended pictures, recommended videos, etc.

[0089] In this embodiment, by inserting recommended content at the information moderation points, content recommendation can be achieved on the premise of minimizing the impact on the user's viewing of the target video.

[0090] The following provides another exemplary method for determining the content to be inserted in the target video and the insertion time point of the content to be inserted.

[0091] In an alternative embodiment, refer to Figure 7, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Step S700, perform a plot turning point prediction process on the video content of the target video to obtain plot turning points.

[0092] Step S702, use the plot turning points as the insertion time points of the content to be inserted.

[0093] Step S704, use the content to be inserted of the fifth type as the content to be inserted into the target video.

[0094] In this embodiment, the video content is video frame images and subtitle texts.

[0095] In an implementation manner, a long short-term memory network (LSTM) or a Transformer model can be used to perform plot prediction based on video frame images and subtitle texts, so as to output plot turning points.

[0096] After obtaining the plot turning points, the content to be inserted of the fifth type can be used as the content to be inserted into the target video. Among them, the content to be inserted of the fifth type can be plot interaction content, for example, an ending segment selected by the user.

[0097] In this embodiment, by inserting plot interaction content at the plot turning points, the user interaction can be enhanced and the user experience can be improved.

[0098] Step S204 , insert the content to be inserted into the target video based on the insertion time point.

[0099] In this embodiment, when determining the insertion time point and the content to be inserted, the content to be inserted can be directly inserted into the target video at the corresponding insertion time point.

[0100] In an implementation manner, when the video data is fast-forward data for a target video segment, the content to be inserted and the insertion time point of the content to be inserted can be determined in the following manner: Calculate the fast-forward ratio of the user in the target video segment, compare the fast-forward ratio with a set threshold. If it exceeds the set threshold, it is considered that the user is not interested in the target video segment. In this way, at the end of the user's fast-forward operation, a key plot summary of the target video segment or a relevant recommendation related to the target video segment can be inserted. That is to say, the end time point of the fast-forward operation can be used as the insertion time point, and the key plot summary of the target video segment or a relevant recommendation related to the target video segment can be used as the content to be inserted.

[0101] In one embodiment, when the video data is skip data for a target video segment, the content to be inserted and the insertion time point of the content to be inserted can be determined in the following manner: Calculate the skip ratio of the user for the target video segment, compare this skip ratio with a set threshold. If it exceeds the set threshold, it is considered that the user is not interested in the target video segment. In this way, at the end of the user's skip operation, a key plot summary of the target video segment or a relevant recommendation related to the target video segment can be inserted. That is to say, the time point at the end of the skip operation can be used as the insertion time point, and the key plot summary of the target video segment or a relevant recommendation related to the target video segment can be used as the content to be inserted.

[0102] In one embodiment, when the video data is pause data for a target video segment, the content to be inserted and the insertion time point of the content to be inserted can be determined in the following manner: Monitor the user's pause behavior for the target video segment. When the number of times the user's pause behavior for the target video segment occurs reaches a preset number of times, it is considered that the user does not quite understand the target video segment. In this way, when the user's pause behavior reaches the preset number of times, a detailed explanation or plot analysis of the target video segment can be inserted. That is to say, the time point when the user's pause behavior for the target video segment reaches the preset number of times can be used as the insertion time point, and the detailed explanation or plot analysis of the target video segment for the target video segment can be used as the content to be inserted.

[0103] Similarly, in one embodiment, when the video data is replay data for a target video segment, the content to be inserted and the insertion time point of the content to be inserted can be determined in the following manner: Monitor the user's replay behavior for the target video segment. When the number of times the user's replay behavior for the target video segment occurs reaches a preset number of times, it is considered that the user does not quite understand the target video segment. In this way, when the user's replay behavior reaches the preset number of times, a detailed explanation or plot analysis of the target video segment can be inserted. That is to say, the time point when the user's replay behavior for the target video segment reaches the preset number of times can be used as the insertion time point, and the detailed explanation or plot analysis of the target video segment for the target video segment can be used as the content to be inserted.

[0104] In one embodiment, when the video data is long stay data for a target video segment, the content to be inserted and the insertion time point of the content to be inserted can be determined in the following manner: Monitor the user's dwell time on the target video segment. When it is detected that the user's dwell time on the target video segment reaches the preset duration, it is considered that the user is interested in the target video segment. In this way, when the user's dwell time reaches the preset duration, interactive content for the target video segment can be inserted. Among them, the interactive content can be voting, community discussion, etc. That is to say, the time point when the user's dwell time on the target video segment reaches the preset duration can be used as the insertion time point, and the interactive content of the target video segment can be used as the content to be inserted.

[0105] In one embodiment, when the video data is high-frequency interaction point data, the content to be inserted and the insertion time point of the content to be inserted can be determined in the following manner: Determine the high-frequency interaction points based on at least one of the density of bullet screens, likes, and comments in the target video, and use the high-frequency interaction points as the insertion time points of the content to be inserted. At the same time, topics or hot recommendations for guiding discussions can be used as the content to be inserted.

[0106] In the actual application process, the content to be inserted can be inserted into the target video in various ways. The following provides an exemplary way.

[0107] In an alternative embodiment, refer to Figure 8 , inserting the content to be inserted into the target video based on the insertion time point includes: Step S800, detect the external environment of the target video, and determine the insertion method of the content to be inserted according to the detection result.

[0108] Step S802, insert the content to be inserted into the target video in the determined insertion method based on the insertion time point.

[0109] In one embodiment, the external environment may include at least one of the following: The device type for playing the target video, the network status of the device currently playing the target video.

[0110] Among them, the device type for playing the target video can be identified by parsing the User-Agent information of the device. The network status of the device currently playing the target video can be determined by real-time detecting the network bandwidth of the device.

[0111] The insertion methods of the content to be inserted include bullet screen method, pip method, sidebar method, etc.

[0112] In one embodiment, the insertion method of the content to be inserted corresponding to different detection results can be preset. For example, if the detection result shows that the device type for playing the target video is a mobile device, the corresponding insertion method is the bullet screen method or the sidebar method. Another example is that if the detection result shows that the device type for playing the target video is a PC, the corresponding insertion method is the pip (picture-in-picture) method. Still another example is that if the detection result shows that the current network condition of the device for playing the target video is average, the corresponding insertion method is the bullet screen method. If the detection result shows that the current network condition of the device for playing the target video is good, the corresponding insertion method is the sidebar method. If the detection result shows that the current network condition of the device for playing the target video is excellent, the corresponding insertion method is the pip method.

[0113] In this embodiment, based on the detection result of detecting the external environment of the target video, the insertion method of the content to be inserted is determined, so that the content to be inserted can be inserted in a more appropriate way, ensuring a smooth user experience.

[0114] Embodiment Two Figure 9 The block diagram of the device 900 for inserting content into a video according to Embodiment Two of the present application is schematically shown. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 9 shown, the device 900 may include: an acquisition module 910, a determination module 920, and an insertion module 930, where: The acquisition module 910 is configured to acquire video data of the target video, where the video data includes video content and / or interaction data for the target video; The determination module 910 is configured to determine the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data; The insertion module 930 is configured to insert the content to be inserted into the target video based on the insertion time point.

[0115] As an optional embodiment, the determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Detecting shot change points in the video content of the target video to obtain shot change points; Using the shot change points as the insertion time points of the content to be inserted; Use the content to be inserted of the first type as the content to be inserted into the target video.

[0116] As an optional embodiment, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Perform an emotion detection process on the video content of the target video to obtain an emotion detection result; Determine the emotional high point in the target video according to the emotion detection result, and use the emotional high point as the insertion time point of the content to be inserted; Use the content to be inserted of the second type as the content to be inserted into the target video.

[0117] As an optional embodiment, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Perform an information density detection process on the video content of the target video to obtain information density points; Use the information density points as the insertion time points of the content to be inserted; Use the content to be inserted of the third type as the content to be inserted into the target video.

[0118] As an optional embodiment, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Perform an information moderation detection process on the video content of the target video to obtain information moderation points; Use the information moderation points as the insertion time points of the content to be inserted; Use the content to be inserted of the fourth type as the content to be inserted into the target video.

[0119] As an optional embodiment, determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Perform a plot turning point prediction process on the video content of the target video to obtain plot turning points; Use the plot turning points as the insertion time points of the content to be inserted; Use the content to be inserted of the fifth type as the content to be inserted into the target video.

[0120] As an optional embodiment, the interaction data includes at least one of the following: Fast-forward data for a target video segment, skip data for a target video segment, pause data for a target video segment, playback data for a target video segment, long-time stay data for a target video segment, high-frequency interaction point data.

[0121] As an alternative embodiment, inserting the content to be inserted into the target video based on the insertion time point includes: Detecting the external environment of the target video, and determining the insertion method of the content to be inserted according to the detection result; Inserting the content to be inserted into the target video in the determined insertion method based on the insertion time point.

[0122] As an alternative embodiment, the external environment includes at least one of the following: The device type for playing the target video, the network condition where the device for playing the target video is currently located.

[0123] Embodiment III Figure 10 FIG. schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing a method of inserting content into a video according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. As Figure 10 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or DX memories), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the method for inserting content into a video. In addition, the memory 10010 may also be used to temporarily store various types of data that have been output or will be output.

[0124] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0125] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), a 4G network, a 5G network, Bluetooth, Wi-Fi, etc.

[0126] It should be noted that Figure 10 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be implemented alternatively.

[0127] In this embodiment, the method for inserting content into a video stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.

[0128] Embodiment 4 The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for inserting content into a video in the embodiment are implemented.

[0129] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system and various application software installed on the computer device, such as the program code of the method for inserting content into a video in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or are to be output.

[0130] Embodiment 5 The embodiment of the present application further provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.

[0131] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0132] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for inserting content into a video, characterized in that, The method includes: Obtaining video data of a target video, where the video data includes video content and / or interaction data for the target video; Determining content to be inserted into the target video and an insertion time point of the content to be inserted based on the video data; Inserting the content to be inserted into the target video based on the insertion time point.

2. The method according to claim 1, wherein The determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Detecting shot transition points in the video content of the target video to obtain shot transition points; Using the shot transition points as the insertion time points of the content to be inserted; Using the first type of content to be inserted as the content to be inserted into the target video.

3. The method according to claim 1, wherein The determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Performing emotion detection processing on the video content of the target video to obtain an emotion detection result; Determining emotional high points in the target video according to the emotion detection result, and using the emotional high points as the insertion time points of the content to be inserted; Using the second type of content to be inserted as the content to be inserted into the target video.

4. The method according to claim 1, wherein The determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Performing information density detection processing on the video content of the target video to obtain information density points; Using the information density points as the insertion time points of the content to be inserted; Using the third type of content to be inserted as the content to be inserted into the target video.

5. The method according to claim 1, wherein The determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Performing information relaxation detection processing on the video content of the target video to obtain information relaxation points; Using the information relaxation points as the insertion time points of the content to be inserted; Using the fourth type of content to be inserted as the content to be inserted into the target video.

6. The method according to claim 1, wherein The determining the content to be inserted into the target video and the insertion time point of the content to be inserted based on the video data includes: Performing plot turning point prediction processing on the video content of the target video to obtain plot turning points; Using the plot turning points as the insertion time points of the content to be inserted; Using the fifth type of content to be inserted as the content to be inserted into the target video.

7. The method according to any one of claims 1 to 6, characterized in that The interaction data includes at least one of the following: Fast-forward data for a target video segment, skip data for a target video segment, pause data for a target video segment, playback data for a target video segment, long-time stay data for a target video segment, high-frequency interaction point data.

8. The method according to any one of claims 1 to 6, characterized in that, The inserting the content to be inserted into the target video based on the insertion time point includes: Detecting the external environment of the target video, and determining an insertion method of the content to be inserted according to the detection result; Inserting the content to be inserted into the target video using the determined insertion method based on the insertion time point.

9. The method according to claim 8, wherein The external environment includes at least one of the following: The device type for playing the target video and the network status of the device currently playing the target video.

10. An apparatus for inserting content into a video, characterized in that, The apparatus includes: An acquisition module, configured to acquire video data of a target video, where the video data includes video content and / or interactive data for the target video; A determination module, configured to determine, based on the video data, the content to be inserted into the target video and the insertion time point of the content to be inserted; An insertion module, configured to insert the content to be inserted into the target video based on the insertion time point.

11. A computer device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to claims 1 to 9 are implemented.