Information insertion method and device and computer readable storage medium

By detecting designated elements in the video frame, determining the information insertion area and inserting promotional information, the problem of promotional information blocking important visual information in the existing technology is solved, and the user viewing experience is improved.

CN120676197APending Publication Date: 2025-09-19JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510969456.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The conventional method of inserting promotional information into videos easily blocks important visual information, affecting the user's viewing experience.

Method used

By detecting specified elements in the video frame, the information insertion area is determined and promotional information is inserted in this area, including the detection of faces, item logos, item description text and item images, and the information insertion location is determined using neural network models and deep learning algorithms.

Benefits of technology

Accurately determine the information insertion location to reduce the risk of promotional information blocking important information and improve the user's video viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676197A_ABST
    Figure CN120676197A_ABST
Patent Text Reader

Abstract

The invention provides an information insertion method and device and a computer readable storage medium, and relates to the technical field of video processing. The method comprises the following steps: performing specified element detection on a video frame contained in a video to be processed to obtain a probability value that each pixel point in the video frame belongs to the specified element, the specified element comprising at least two of a face, an article identifier, an article description text and an article image; according to the probability value that each pixel point in the video frame belongs to the specified element, determining the score of each region in a plurality of regions contained in the video frame; determining an information insertion region from the plurality of regions according to the score of each region in the plurality of regions; and inserting promotion information in the information insertion area. Through the method, the information insertion position can be determined more intelligently and accurately, the promotion information is effectively prevented from shielding important information in the video, and the video watching experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to an information insertion method, device, and computer-readable storage medium. Background Art

[0002] With the rise of short video platforms, videos have become an essential part of e-commerce shopping. Adding appropriate promotional information (such as advertisements) to videos can more easily attract users' attention and increase purchase conversion rates.

[0003] The technology for inserting promotional information into videos is not mature. Currently, promotional information is typically inserted into fixed locations within the video or through pop-up windows. These methods of inserting information can obscure important visual elements in the video, significantly negatively impacting the viewer's viewing experience. Summary of the Invention

[0004] The present disclosure provides an information insertion method, an apparatus, and a computer-readable storage medium.

[0005] According to a first aspect of the present disclosure, there is provided an information insertion method, comprising: performing designated element detection on a video frame contained in a video to be processed to obtain a probability value that each pixel point in the video frame belongs to the designated element, wherein the designated elements include at least two of a face, an item identification, a description text of an item, and an item image; determining a score for each of a plurality of areas contained in the video frame based on the probability value that each pixel point in the video frame belongs to the designated element; determining an information insertion area from the plurality of areas based on the score of each area in the plurality of areas; and inserting promotional information in the information insertion area.

[0006] In some embodiments, the performing of specified element detection on the video frames contained in the video to be processed to obtain a probability value that each pixel point in the video frame belongs to the specified element includes: performing face, object identification, and object description text detection or segmentation on the video frame to obtain a probability value that each pixel point in the video frame belongs to the face, a probability value of the object identification, and a probability value of the object description text; performing object image detection or segmentation on the video frame to obtain a probability value that each pixel point in the video frame belongs to the object image.

[0007] In some embodiments, the performing object image detection or segmentation on the video frame to obtain a probability value that each pixel in the video frame belongs to the object image includes: using a first neural network model to perform object image segmentation on the video frame to obtain a probability value that each pixel in the video frame belongs to the object image, wherein the first neural network model is pre-trained based on a training sample set, the training sample set includes a sample image and an object segmentation map corresponding to the sample image, the object segmentation map is obtained by performing object image segmentation on the sample image by a second neural network model and binarizing the segmented image, and the first neural network model is more lightweight than the second neural network model.

[0008] In some embodiments, determining the score of each of the multiple areas contained in the video frame based on the probability value of each pixel point in the video frame belonging to the specified element includes: determining the maximum pixel value of each pixel point in the video frame based on the probability value of each pixel point in the video frame belonging to a face, the probability value of an object logo, the probability value of an object description text, and the probability value of an object image; and determining the score of each area based on the maximum pixel value of all pixels in each area.

[0009] In some embodiments, determining the score of each area based on the maximum pixel value of all pixels in each area includes: averaging the maximum pixel value of all pixels in each area to obtain the pixel average value of each area; and determining the score of each area based on the pixel average value of each area.

[0010] In some embodiments, the multiple regions include at least two of the first to fourth regions, and the setting positions of the first to fourth regions are the top, bottom, left, and right of the video frame respectively; and / or, averaging the maximum pixel values ​​of all pixels in each region to obtain the pixel average value of each region includes: averaging the maximum pixel values ​​of all pixels contained in each region at the setting position to obtain the pixel average value corresponding to the setting position; sliding each candidate region according to a set sliding step size; averaging the maximum pixel values ​​of all pixels contained in each region at each sliding position to obtain the pixel average value corresponding to each sliding position; performing a minimum value operation on the pixel average value corresponding to the setting position of each region and the pixel average value corresponding to all sliding positions to obtain the pixel average value of each region.

[0011] In some embodiments, determining the score of each region based on the pixel average value of each region includes: averaging the pixel average value of each region in the video frame and the pixel average value of each region in at least one adjacent video frame of the video frame to obtain the score of each region; determining the information insertion area from the multiple regions based on the score of each region includes: using the region with the smallest score among the multiple regions as the information insertion area of ​​the video frame and at least one adjacent video frame of the video frame.

[0012] In some embodiments, the promotional information is advertising text, and the method further includes: using a deep learning algorithm and attributes of the information insertion area to determine the attributes of the advertising text, wherein the attributes of the information insertion area include the color of the information insertion area, and the attributes of the advertising text include text color.

[0013] According to a second aspect of the present disclosure, an information insertion apparatus is provided, comprising: a device for executing the information insertion method as described above.

[0014] According to a third aspect of the present disclosure, an information insertion apparatus is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the aforementioned information insertion method based on instructions stored in the memory.

[0015] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the information insertion method as described above is implemented.

[0016] According to a fifth aspect of the present disclosure, a computer program product is provided, on which computer program instructions are stored, and when the instructions are executed by a processor, the information insertion method as described above is implemented.

[0017] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0019] Figure 1 is a flowchart of an information insertion method according to some embodiments of the present disclosure;

[0020] Figure 2 is a flowchart of an information insertion method according to other embodiments of the present disclosure;

[0021] Figure 3is a schematic diagram of image fusion according to some embodiments of the present disclosure;

[0022] Figure 4 is a schematic diagram of a delineated area according to some embodiments of the present disclosure;

[0023] Figure 5 is a schematic diagram of information insertion results according to some embodiments of the present disclosure;

[0024] Figure 6 is a structural diagram of an information insertion device according to some embodiments of the present disclosure;

[0025] Figure 7 is a schematic structural diagram of an information insertion device according to other embodiments of the present disclosure;

[0026] Figure 8 Schematic diagram of the structure of a computer system according to some embodiments of the present disclosure.

[0027] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings. DETAILED DESCRIPTION

[0028] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0029] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0030] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0031] Technologies, methods and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods and equipment should be considered part of the authorization specification.

[0032] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0033] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0034] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0035] In order to solve the problems existing in the related art, the present disclosure proposes an information insertion method, device and computer-readable storage medium, so as to more intelligently and accurately determine the information insertion position in the video, effectively prevent promotional information from blocking important information in the video, and improve the user's video viewing experience.

[0036] Figure 1 FIG. 1 is a flow chart of an information insertion method according to some embodiments of the present disclosure. Figure 1 As shown, the information insertion method includes steps S11 to S14.

[0037] Step S11: performing designated element detection on the video frames included in the video to be processed to obtain a probability value of each pixel in the video frame belonging to the designated element.

[0038] A video frame refers to an image extracted from a video to be processed. Before step S11, the video to be processed may be subjected to frame extraction, for example, one frame is extracted per second, thereby obtaining one or more frames of image.

[0039] The designated elements include at least two of a face, an item identifier, a description text of the item, and an item image.

[0040] In some examples, the designated elements include faces and object images. In these examples, step S11 may specifically include: using a face detection model to detect a video frame extracted from the video to be processed to obtain the probability that each pixel in the video frame belongs to a face; using an object image detection model to detect the video frame to obtain the probability that each pixel in the video frame belongs to an object image. For example, the face detection model can be a multi-task cascaded convolutional network (MTCNN) model, or other network models that can be used for face detection. For example, the object image detection model can be a You Only Look Once (YOLO) model, or other network models that can be used for object image detection.

[0041] In other examples, the designated elements include item images, item description text, and item identifiers. In these examples, step S11 may specifically include: performing item image detection on the video frame using an item image detection model to obtain the probability that each pixel in the video frame belongs to the item image; detecting the item description text in the video frame using an optical character recognition model to obtain the probability that each pixel in the video frame belongs to the item description text; and detecting the item identifier in the video frame using an item identifier detection model to obtain the probability that each pixel in the video frame belongs to the item identifier. For example, the item image detection model may be a YOLO (You Only Look Once) model, or other network models that can be used for item image detection. For example, the optical character recognition model may be a Differentiable Binarization Network (DBNet) model, or other network models that can be used for text detection. For example, the item identifier detection model may be an open collection item identifier detection model, or other network models that can be used for item identifier detection.

[0042] In some further examples, the designated elements include faces, item images, item description text, and item identifiers. In these examples, step S11 may specifically include: detecting or segmenting the video frame for faces, item identifiers, and item description texts to obtain the probability value of each pixel in the video frame belonging to a face, the probability value of the item identifier, and the probability value of the item description text; detecting or segmenting the video frame for item images to obtain the probability value of each pixel in the video frame belonging to the item image. By detecting the above four designated elements, it helps to improve the rationality of the subsequently determined information insertion area, further alleviate the problem of promotional information blocking important information in the video, and improve the user's video viewing experience. Moreover, by considering the above four designated elements, it is more suitable for e-commerce scenarios.

[0043] Step S12: determining a score for each of the multiple regions included in the video frame according to the probability value of each pixel point in the video frame belonging to the specified element.

[0044] In some examples, the multiple regions included in the video frame are regions located at fixed positions and have fixed sizes.

[0045] For example, the multiple areas include a first area and a second area, the first area is a rectangular area located at the top of the video frame, and the size of the area meets the following requirements: the length is 80% of the horizontal length of the video frame, and the width is 10% of the vertical length of the video frame; the second area is a rectangular area located at the bottom of the video frame, and the size of the area meets the following requirements: the length is 80% of the horizontal length of the video frame, and the width is 10% of the vertical length of the video frame.

[0046] For another example, the multiple regions include first to fourth regions, the first region is located at the top of the video frame, the second region is located at the bottom of the video frame, the third region is located at the left of the video frame, and the fourth region is located at the right of the video frame.

[0047] In other examples, the video frame includes multiple regions of fixed size that can be slid to change values. For example, the multiple regions include a first region located at the top of the video frame and slidable within a specified range along the horizontal coordinate direction; and a second region located at the bottom of the video frame and slidable within a specified range along the horizontal coordinate direction. By allowing these multiple regions to slidably change values, it is possible to more flexibly and accurately determine information insertion areas for different video content, thereby further improving the user's video viewing experience.

[0048] Step S12 can be implemented in a variety of ways, and two implementations are described below as examples.

[0049] In the first embodiment, step S12 includes: determining the maximum pixel value of each pixel in the video frame based on the probability value of each pixel in the video frame belonging to all specified elements; and determining the score of each area based on the maximum pixel value of all pixels in each area.

[0050] Furthermore, in the above embodiment, the maximum pixel value of each pixel in the video frame can be determined according to the following exemplary method: the maximum value of the probability values ​​of each pixel in the video frame belonging to all designated elements is used as the maximum pixel value of the pixel. For example, if the designated elements include a face and an object image, the maximum value of the probability value of the pixel belonging to the face and the probability value of the pixel belonging to the object image can be used as the maximum pixel value of the pixel. Similarly, the maximum pixel value of each pixel in the video frame can be obtained.

[0051] Furthermore, in the above embodiment, the score of each area can be determined according to the following exemplary method: averaging the maximum pixel values ​​of all pixels in each area to obtain the pixel average value of each area; and determining the score of each area based on the pixel average value of each area.

[0052] In specific implementations, the score of each region can be determined individually for each video frame. For example, video frames A, B, and C are extracted from the video to be processed. A designated element detection is performed on video A to obtain a probability value for each pixel in video frame A belonging to the designated element. Based on the probability value for each pixel in video frame A belonging to the designated element, a score is determined for each of the multiple regions contained in video frame A. Similarly, a score can be determined for each of the multiple regions contained in video frame B, as well as a score for each of the multiple regions contained in video frame C.

[0053] In addition, the score of each region can also be determined based on a video frame and its adjacent video frames. Specifically, the method includes: first determining the pixel average value of each region in the video frame and the pixel average value of each region in at least one adjacent video frame of the video frame; then averaging the pixel average value of each region in the video frame and the pixel average value of each region in at least one adjacent video frame of the video frame to obtain a score for each region. For example, assume that video frames A, B, and C are extracted from the video to be processed, and each video frame contains first to fourth regions. First, the maximum pixel value of all pixels in the first region contained in video frame A can be averaged to obtain the pixel average value of the region contained in video frame A. Then, the pixel average value of the first region contained in video frame A and the pixel average value of the first regions contained in video frames B and C that are temporally adjacent to video frame A are averaged to obtain a score for the first region. Similarly, a score for each of the four regions can be obtained.

[0054] In the disclosed embodiments, the aforementioned implementation enables a more accurate assessment of the likelihood that each region contains a specified element, thereby improving the accuracy of the subsequently determined information insertion zones. Furthermore, by averaging the average pixel values ​​of the regions within a video frame and its adjacent frames to obtain a final score for each region, not only does this facilitate the subsequent, single-step determination of the promotional information's position within multiple frames of video, but it also ensures that the subsequently determined information insertion position takes into account the requirement to display important information across multiple frames without obstructing the display, thereby improving both the accuracy and processing efficiency of promotional information insertion.

[0055] In a second embodiment, step S12 includes: determining the pixel average value of each pixel in the video frame based on the probability value of each pixel in the video frame belonging to all specified elements; and determining the score of each area based on the pixel average value of all pixels in each area.

[0056] For example, if the specified elements include faces and object images, the average of the probability that a pixel belongs to a face and the probability that the pixel belongs to an object image can be used as the pixel average for that pixel. Similarly, the pixel average for each pixel in the video frame can be obtained. Then, the pixel averages of all pixels in a region are averaged to obtain the pixel average for that region, and the score for that region is determined based on the pixel average. Similarly, the scores for each of multiple regions can be obtained.

[0057] Step S13: determining an information insertion area from the plurality of areas according to the score of each area among the plurality of areas.

[0058] Step S13 can be implemented in a variety of ways, and three implementations are described below as examples.

[0059] In a first embodiment, each extracted video frame is processed separately to obtain an information insertion area for each video frame. For example, video frames A, B, and C are extracted from the video to be processed. A designated element detection is performed on video A to obtain a probability value for each pixel in video frame A belonging to the designated element. Based on the probability value for each pixel in video frame A belonging to the designated element, a score is determined for each of the multiple regions contained in video frame A. Then, in step S13, the region with the lowest score among the multiple regions contained in video frame A is determined as the information insertion area for video frame A. Similarly, the information insertion area for video frame B and the information insertion area for video frame C can be determined.

[0060] In a second embodiment, a unified information insertion area is jointly determined based on the extracted video frame and at least one adjacent video frame, including: the area with the lowest score among multiple areas is used as the information insertion area for the video frame and at least one adjacent video frame of the video frame. For example, assuming that the information insertion area is determined to be the first area based on video frame A, video frame B and video frame C adjacent to video frame A, then in step S13, the first area is used as the information insertion area for video frames A, video frame B, and video frame C. Through this processing, not only can the inserted promotional information be prevented from jumping frequently, but the determined information insertion location can also be made more accurate, thereby helping to improve the user's viewing experience.

[0061] In a third embodiment, in step S13, the information insertion area determined for a single video frame is used as the information insertion area for that video frame and at least one adjacent video frame. For example, assuming the information insertion area determined for video frame A is the first area, the first area is used as the information insertion area for video frame A, video frames B and C, which are adjacent to video frame A. This process prevents the inserted promotional information from jumping around frequently, thereby improving the user's viewing experience.

[0062] Step S14: inserting promotional information into the information insertion area.

[0063] In some embodiments, considering the display stability of the promotion information in the video timing, a set duration (e.g., 5 seconds or other duration) or a set number of frames is used as a calculation cycle, and in each calculation cycle, Figure 1 The process shown generates a unified information insertion area and inserts promotional information into the information insertion area. This can prevent the inserted promotional information from jumping around too frequently, thereby helping to improve the user's viewing experience.

[0064] The promotional information may be in the form of advertising text, images, videos, etc.

[0065] In some embodiments, when the promotional information is an advertisement text, Figure 1 The illustrated method may further include: Step S15, determining the attributes of the advertisement text using a deep learning algorithm and the attributes of the information insertion area. The attributes of the information insertion area include the color of the information insertion area, and the attributes of the advertisement text include the text color. Furthermore, in a specific implementation, the attributes of the information insertion area and the advertisement text may also include other common attributes. For example, the advertisement text may include common attributes such as text size, font, and font style.

[0066] In some examples, in step S15, the color of the information insertion area can be first calculated using a deep learning algorithm, and then the color of the advertising text can be determined based on the color of the information insertion area and the color correspondence between the information insertion area and the advertising text. This process helps improve the display effect of the promotional information and enhance the user's viewing experience.

[0067] In the disclosed embodiment, by detecting a variety of specified elements and selecting the information insertion area based on the detection results according to the above process, the insertion position of the promotional information in the video can be determined more intelligently and accurately. This is particularly suitable for e-commerce scenarios and can effectively alleviate the problem of promotional information blocking important information in the video, thereby improving the user's video viewing experience.

[0068] Figure 2 FIG. 1 is a flow chart of an information insertion method according to other embodiments of the present disclosure. Figure 2 As shown, the information insertion method includes steps S21 to S27.

[0069] Step S21: Face detection.

[0070] In step S21 , face detection may be performed on the video frames included in the video to be processed based on a face detection model to obtain the probability that each pixel in the video frame belongs to a face.

[0071] Furthermore, before step S21 , the information insertion method may further include: extracting frames from the video to be processed, for example, extracting one frame per second, thereby obtaining one or more image frames (ie, video frames).

[0072] Step S22: Detecting the item identification.

[0073] In this step, the object identification detection model can be used to detect the object identification in the video frame to obtain the probability that each pixel in the video frame belongs to the object identification.

[0074] Step S23: Detecting the item description text.

[0075] In this step, an optical character recognition model may be used to detect the item description text in the video frame to obtain the probability that each pixel in the video frame belongs to the item description text.

[0076] In specific implementations, steps S21 through S23 can be performed sequentially or in parallel. For example, face detection can be performed on video frame A first, followed by object identification detection on video frame A after face detection, and then object description text detection on video frame A after object identification detection, to determine the probability that each pixel in video frame A belongs to a face, the probability that it belongs to an object identification, and the probability that it belongs to an object description text. Furthermore, the order in which steps S21 through S23 are performed can be flexibly adjusted.

[0077] Step S24: object image detection.

[0078] In this step, an object image detection or segmentation model may be used to perform object image detection or segmentation on the video frame to obtain the probability that each pixel in the video frame belongs to an object image.

[0079] In some examples, step S24 specifically includes: using a first neural network model to segment the video frame into objects to obtain a probability value for each pixel in the video frame belonging to an object. The first neural network model is pre-trained based on a training sample set, which includes a sample image and an object segmentation map corresponding to the sample image. The object segmentation map is obtained by performing object image segmentation on the sample image using a second neural network model and binarizing the segmented image. The first neural network model is more lightweight than the second neural network model.

[0080] For example, a white-background item image from an e-commerce platform can be used as a sample image. A second neural network model (e.g., a large visual segmentation model such as the Segment Anything Model) can be used to segment the item image and binarize the segmentation results to obtain an item segmentation map. Then, a first neural network model (e.g., a YOLO model) is trained using a training sample set containing the sample image and the item segmentation map to obtain a trained first neural network model. When executing step S24, the video frame is segmented into item images based on the trained first neural network model to obtain a probability value for each pixel in the video frame belonging to an item.

[0081] In the disclosed embodiment, by first using a more accurate but time-consuming second neural network model to segment sample images during the training phase to obtain an item segmentation map, and then training the first neural network model based on this map, and then performing product segmentation on video frames based on the more lightweight first neural network model during the model usage phase, the execution efficiency of the information insertion method and the accuracy of information insertion are further improved. Furthermore, by using a white background image as the sample image required for model training, since the white background image only contains products without other background interference, it helps to improve the training effect of the first neural network model, and thus helps to improve the accuracy of the subsequent information insertion area selection.

[0082] Step S25: Determine the maximum pixel value of each pixel in the video frame.

[0083] In this step, the maximum pixel value of each pixel in the video frame is determined based on the probability value of each pixel in the video frame belonging to a face, the probability value of the item identifier, the probability value of the item description text, and the probability value of the item image. For example, the maximum value of the probability values ​​of a pixel belonging to a face, the item identifier, the item description text, and the item image is used as the maximum pixel value within the pixel band.

[0084] In some examples, the above process of determining the maximum pixel value of pixels can be regarded as an image fusion processing process. Figure 3An exemplary image fusion processing process is shown. The process includes: performing face detection, object identification detection, and object description text detection on the video frame to obtain a restricted area map 31, in which the pixel value of each pixel point in the restricted area map is the probability value of the pixel point being a face, an object identification, and an object description text; performing object segmentation on the video frame to obtain Figure 3 The object segmentation map 32 is shown, where the pixel value of each pixel in the object segmentation map is the probability value that the pixel is an object image (or an object). Then, the restricted area map 31 and the object segmentation map 32 are fused to obtain a confidence distribution map 33. The pixel value of each pixel in the confidence distribution map is obtained by retaining the maximum value of the pixel values ​​at the same location in the restricted area map and the object segmentation map.

[0085] Step S26: Determine a score for each of the plurality of regions.

[0086] In some examples, the plurality of regions are first to fourth regions set in the video frame. Figure 4 As shown, the first area 41 is set at the top of the video frame, the second area 42 is set at the bottom of the video frame, the third area 43 is set at the left part of the video frame, and the fourth area 44 is set at the right part of the video frame.

[0087] In step S26, the score of each region can be determined based on various methods, which are exemplified below with reference to two implementations.

[0088] In a first embodiment, the first to fourth regions are fixed in position. In this embodiment, the score of each region is determined as follows: the maximum pixel values ​​of all pixels in each region are averaged to obtain a pixel average value for each region; and the score of each region is determined based on the pixel average value for each region.

[0089] In a second embodiment, the first to fourth regions can all slide within a certain range. For example, the first and second regions can slide along the horizontal coordinate direction of the video frame, and the third and fourth regions can slide along the vertical coordinate direction of the video frame. In this embodiment, the score of each region can be determined as follows: the maximum pixel values ​​of all pixels contained in each region at its initial set position are averaged to obtain a pixel average corresponding to the set position; each candidate region is slid according to a set sliding step size; the maximum pixel values ​​of all pixels contained in each region at each sliding position are averaged to obtain a pixel average corresponding to each sliding position; the pixel average corresponding to each region at the set position and the pixel average corresponding to all sliding positions are minimized to obtain a pixel average for each region, and the position with the minimum pixel average between the initial set position and each sliding position is used as the final position of the region; and the score of each region is determined based on the pixel average of each region. By making multiple regions slidable, more candidate locations for information insertion areas can be covered, thereby facilitating more flexible and accurate determination of information insertion areas for different video scenarios.

[0090] Step S27: Determine the information insertion area.

[0091] In step S27, the area with the lowest score among the multiple areas may be used as the information insertion area. Thereafter, the promotional information is inserted into the information insertion area.

[0092] The promotion information can be an advertisement text. In some embodiments, the promotion information and a video can be used as input. Figure 2 The algorithm flow shown above obtains the video after the promotion information is inserted. Figure 5 As shown in FIG, the advertising text "aerospace-grade aluminum-magnesium alloy fuselage" and a video are used as input, and the algorithm process is used to obtain the video with the advertising text inserted.

[0093] In the disclosed embodiment, the above process can more intelligently and accurately determine the information insertion position in the video, effectively prevent the promotional information from blocking important information in the video, and improve the user's video viewing experience.

[0094] Figure 6 Schematic diagram of the structure of the information insertion device according to some embodiments of the present disclosure. Figure 6 As shown, the information insertion device 60 is used to execute the information insertion method as described above, and includes a detection module 61 , a first determination module 62 , a second determination module 63 , and an insertion module 64 .

[0095] The detection module 61 is configured to perform designated element detection on the video frames included in the video to be processed, so as to obtain a probability value of each pixel in the video frame belonging to the designated element.

[0096] The designated elements include at least two of a face, an item identifier, a description text of the item, and an item image.

[0097] In some examples, the detection module 61 is configured to: detect or segment faces, object identifiers, and descriptive text of objects in video frames to obtain a probability value that each pixel in the video frame belongs to a face, a probability value of an object identifier, and a probability value of a descriptive text of an object; and perform object image detection or segmentation on the video frame to obtain a probability value that each pixel in the video frame belongs to an object image.

[0098] The first determination module 62 is configured to determine a score for each of the multiple regions included in the video frame according to a probability value of each pixel point in the video frame belonging to a specified element.

[0099] In some examples, the first determination module 62 determines the score of the region as follows: based on the probability value of each pixel in the video frame that belongs to a face, the probability value of the object identification, the probability value of the object description text, and the probability value of the object image, the maximum pixel value of each pixel in the video frame is determined; and based on the maximum pixel value of all pixels in each region, the score of each region is determined.

[0100] In the above example, after determining the maximum pixel value of each pixel, the first determination module 62 may determine the region score in the following exemplary manner: averaging the maximum pixel values ​​of all pixels within each region to obtain a pixel average value for each region; and determining the score for each region based on the pixel average value. For example, the pixel average value of a region within a video frame may be used as the score for that region.

[0101] In addition, in the above example, after determining the maximum pixel value of each pixel point, the first determination module 62 can also determine the regional score according to the following exemplary method: averaging the pixel average value of each region in the video frame and the pixel average value of each region in at least one adjacent video frame of the video frame to obtain a score for each region.

[0102] The second determining module 63 is configured to determine an information insertion area from the multiple areas according to the score of each area in the multiple areas.

[0103] In some examples, the second determination module 63 processes each extracted video frame individually to obtain an information insertion region for each video frame. For example, the region with the lowest score among multiple regions included in video frame A is used as the information insertion region for video frame A; and the region with the lowest score among multiple regions included in video frame B is used as the information insertion region for video frame B.

[0104] In other examples, the second determination module 63 uses the information insertion area determined based on a single video frame as the information insertion area for the video frame and at least one adjacent video frame. For example, assuming that the information insertion area is determined to be the first area based on video frame A, the first area is used as the information insertion area for video frame A, video frame B adjacent to video frame A, and video frame C.

[0105] In some further examples, the second determination module 63 jointly determines a unified information insertion area based on the extracted video frame and at least one adjacent video frame, including: after determining a score for each of the multiple areas based on the extracted video frame and at least one adjacent video frame, selecting the area with the lowest score as the information insertion area for the video frame and at least one adjacent video frame of the video frame. For example, assuming that the information insertion area is determined to be the first area based on video frame A, video frame B adjacent to video frame A, and video frame C, the first area is used as the information insertion area for video frames A, video frame B, and video frame C.

[0106] The inserting module 64 is configured to insert promotional information into the information insertion area.

[0107] In some examples, the promotional information is advertising text. In this instance, the information insertion device may further include a third determination module. The third determination module is configured to determine the attributes of the advertising text using a deep learning algorithm and attributes of the information insertion area. The attributes of the information insertion area include the color of the information insertion area, and the attributes of the advertising text include the text color.

[0108] In the disclosed embodiment, the above device can more intelligently and accurately determine the information insertion position in the video, effectively prevent the promotional information from blocking important information in the video, and improve the user's video viewing experience.

[0109] Figure 7 Schematic diagram of the structure of the information insertion device according to other embodiments of the present disclosure. Figure 7 As shown, the information insertion apparatus 70 includes a memory 71 and a processor 72 coupled to the memory 71. The memory 71 is used to store instructions for executing the corresponding embodiments of the information insertion method. The processor 72 is configured to execute the information insertion method in any of the embodiments of the present disclosure based on the instructions stored in the memory 71.

[0110] Figure 8 FIG. 1 is a schematic diagram of the structure of a computer system according to some embodiments of the present disclosure. Figure 8 As shown, the information insertion apparatus can be embodied in the form of a general-purpose computing device. A computer system 80 includes a memory 81, a processor 82, and a bus 83 connecting different system components.

[0111] The memory 81 may include, for example, a system memory, a non-volatile storage medium, and the like. The system memory may store, for example, an operating system, application programs, a boot loader, and other programs. The system memory may include volatile storage media, such as random access memory (RAM) and / or cache memory. The non-volatile storage medium may store, for example, instructions for executing at least one of the corresponding embodiments of the information insertion method. Non-volatile storage media include, but are not limited to, disk storage, optical storage, flash memory, and the like.

[0112] The processor 82 can be implemented as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, or as discrete hardware components such as discrete gates or transistors. Accordingly, each module, such as the detection module and the first determination module, can be implemented by a central processing unit (CPU) executing instructions in a memory that execute corresponding steps, or by dedicated circuits that execute corresponding steps.

[0113] The bus 83 may use any of a variety of bus architectures, including, but not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, and a Peripheral Component Interconnect (PCI) bus.

[0114] These interfaces 84, 85, and 86 in computer system 80, as well as memory 81 and processor 82, can be connected via bus 83. Input / output interface 84 provides a connection interface for input / output devices such as a display, mouse, and keyboard. Network interface 85 provides a connection interface for various networked devices. Storage interface 86 provides a connection interface for external storage devices such as floppy disks, USB flash drives, and SD cards.

[0115] Here, various aspects of the present disclosure are described with reference to flowcharts and / or block diagrams of methods, devices, and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks, can be implemented by computer-readable program instructions.

[0116] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable device to produce a machine, so that the processor executes the instructions to produce means for implementing the functions specified in one or more blocks in the flowcharts and / or block diagrams.

[0117] These computer-readable program instructions may also be stored in a computer-readable memory, which cause the computer to operate in a specific manner to produce an article of manufacture, including instructions for implementing the functions specified in one or more blocks in the flowcharts and / or block diagrams.

[0118] The present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.

[0119] The information insertion method, apparatus, and computer-readable storage medium according to the present disclosure have been described in detail. To avoid obscuring the concepts of the present disclosure, some details known in the art have been omitted. Based on the above description, those skilled in the art will fully understand how to implement the technical solutions disclosed herein.

Claims

1. A method for inserting information, comprising: Performing designated element detection on video frames included in the video to be processed to obtain a probability value of each pixel in the video frame belonging to the designated element, wherein the designated elements include at least two of a face, an item identifier, a description text of the item, and an item image; determining a score for each of a plurality of regions included in the video frame according to a probability value of each pixel in the video frame belonging to the designated element; determining an information insertion area from among the plurality of areas based on the score of each of the plurality of areas; Insert promotional information in the information insertion area.

2. The information insertion method according to claim 1, wherein: The performing designated element detection on the video frame included in the video to be processed to obtain a probability value of each pixel in the video frame belonging to the designated element includes: Detecting or segmenting faces, item identifiers, and item description text on the video frame to obtain a probability value of each pixel in the video frame belonging to a face, a probability value of the item identifier, and a probability value of the item description text; Object image detection or segmentation is performed on the video frame to obtain a probability value of each pixel in the video frame belonging to an object image.

3. The information insertion method according to claim 2, wherein: The performing object image detection or segmentation on the video frame to obtain a probability value of each pixel in the video frame belonging to an object image includes: The video frame is segmented into an object image using a first neural network model to obtain a probability value that each pixel in the video frame belongs to an object image, wherein the first neural network model is pre-trained based on a training sample set, and the training sample set includes a sample image and an object segmentation map corresponding to the sample image. The object segmentation map is obtained by segmenting the sample image into an object image using a second neural network model and binarizing the segmented image. The first neural network model is more lightweight than the second neural network model.

4. The information insertion method according to claim 1, wherein: Determining the score of each of the multiple regions included in the video frame according to the probability value of each pixel point in the video frame belonging to the designated element includes: Determining a maximum pixel value for each pixel in the video frame based on a probability value of each pixel in the video frame belonging to a face, a probability value of an item identifier, a probability value of a description text of the item, and a probability value of an item image; The score of each area is determined according to the maximum pixel value of all pixels in each area.

5. The information insertion method according to claim 4, wherein: Determining the score of each region according to the maximum pixel value of all pixels in each region includes: Performing an averaging operation on the maximum pixel values ​​of all pixels in each area to obtain a pixel average value of each area; A score for each region is determined according to an average value of pixels in each region.

6. The information insertion method according to claim 5, wherein: The plurality of regions include at least two of the first region to the fourth region, and the first region to the fourth region are sequentially arranged at the top, bottom, left, and right of the video frame; and / or, The averaging operation of the maximum pixel values ​​of all pixels in each region to obtain the pixel average value of each region includes: Performing an averaging operation on the maximum pixel values ​​of all pixel points included in each area at the set position to obtain a pixel average value corresponding to the set position; Slide each candidate area according to the set sliding step size; The maximum pixel values ​​of all pixel points contained in each area at each sliding position are averaged to obtain the pixel average value corresponding to each sliding position; the minimum value operation is performed on the pixel average value corresponding to each area at the set position and the pixel average value corresponding to all sliding positions to obtain the pixel average value of each area.

7. The information insertion method according to claim 5, wherein: Determining the score of each region according to the pixel average value of each region includes: performing an averaging operation on the pixel average value of each region in the video frame and the pixel average value of each region in at least one adjacent video frame of the video frame to obtain the score of each region; Determining the information insertion area from the multiple areas according to the score of each area includes: using the area with the smallest score among the multiple areas as the information insertion area of ​​the video frame and at least one adjacent video frame of the video frame.

8. The information insertion method according to claim 1, wherein: The promotion information is an advertisement text, and the method further includes: The attributes of the advertisement text are determined using a deep learning algorithm and the attributes of the information insertion area, wherein the attributes of the information insertion area include the color of the information insertion area, and the attributes of the advertisement text include the text color.

9. An information insertion device, comprising: Used to execute the information insertion method described in any one of claims 1 to 8.

10. An information insertion device, comprising: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the information insertion method according to any one of claims 1 to 8 based on instructions stored in the memory.

11. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the information insertion method as described in any one of claims 1 to 8.

12. A computer program product having computer program instructions stored thereon, wherein when the instructions are executed by a processor, the information insertion method according to any one of claims 1 to 8 is implemented.