A processing method and device based on video data

By generating and rendering personalized video borders based on video content, the problems of horizontal and vertical screen switching and video adaptation in the prior art are solved, and the user experience is improved.

CN115811582BActive Publication Date: 2025-06-24CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211472928.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-06-24
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problems of horizontal and vertical screen switching and video adaptation on devices with different screen proportions, resulting in poor user experience.

Method used

By acquiring video data, determine video clips, keyframes and saliency areas, generate video border images based on video tags, and render color based on color block sequences of saliency areas to achieve personalized video borders.

Benefits of technology

It realizes the generation of personalized video borders based on video content, solves the problem of unchanging background filling in the prior art, and improves the user's visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115811582B_ABST
    Figure CN115811582B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method and apparatus for processing video data. The method includes: obtaining video data, and determining one or more video segments from the video data; for each video segment, determining a key frame; determining one or more salient regions from the key frames; and generating a video border image for each video segment in the video data by combining the video tags of the video data and the one or more salient regions. Through the embodiments of the present invention, generating a video border according to video content is realized, so that the video border presents personalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video technology, and particularly to a method and apparatus for processing video data. Background Art

[0002] With the wide use of video applications by users, especially short video applications, both the number of users and their usage duration have been increasing year by year. During the use of video applications, due to the difference between landscape and portrait orientations when recording videos, when users watch videos, there are often situations where the recorded video is in landscape orientation but the user watches it in portrait orientation. In addition, video applications can usually play videos on multi-platform devices such as mobile phones, tablets, and PCs, and the screen ratios of these platforms are different, resulting in video incompatibility.

[0003] In the prior art, to solve the problems of horizontal and vertical screen switching and adapting to video playback on devices with different screen ratios, black backgrounds are usually filled, or the video edges are blurred and then filled. This filling method is relatively single and the user experience is not good. Summary of the Invention

[0004] In view of the above problems, a method and apparatus for processing video data are proposed to overcome or at least partially solve the above problems, including:

[0005] A method for processing video data, the method includes:

[0006] Obtain video data, and determine one or more video segments from the video data;

[0007] For each video segment, determine a key frame;

[0008] Determine one or more salient regions from the key frames;

[0009] Combine the video tags of the video data and the one or more salient regions to generate a video border image for each video segment in the video data.

[0010] Optionally, it further includes:

[0011] Determine multiple border image regions from the video border image;

[0012] For each border image region, determine the salient region with the closest distance;

[0013] Combine the salient region with the closest distance to render the border image region.

[0014] Optionally, the combining the salient region with the closest distance to render the border image region includes:

[0015] For the nearest significant region, determine color block sequence information; wherein, the color block sequence information includes color block information sorted by color block proportion in the significant region;

[0016] According to the color block sequence information, perform color rendering on the border image region.

[0017] Optionally, the determining color block sequence information for the nearest significant region includes:

[0018] For the nearest significant region, determine multiple color data;

[0019] Cluster the multiple color data to obtain multiple color block information and their color block proportions;

[0020] Generate color block sequence information according to the multiple color block information and their color block proportions.

[0021] Optionally, the generating a video border image for each video segment in the video data by combining the video label of the video data and the one or more significant regions includes:

[0022] According to the video label of the video data, determine a target significant region from the one or more significant regions, and determine the arrangement mode of the target significant region;

[0023] Arrange the target significant region according to the arrangement mode to obtain a video border image for each video segment in the video data.

[0024] Optionally, the target significant region is the significant region with the largest area, including any one of the following:

[0025] The significant region with the largest area, the N significant regions ranked in the front by area size;

[0026] Wherein, N is a positive integer greater than 1.

[0027] Optionally, the arrangement mode of the target significant region includes any one of the following:

[0028] Repeat and arrange in a circle according to the border width;

[0029] Arrange the target significant regions in a ring.

[0030] A processing device based on video data, the device includes:

[0031] A video segment determining module, configured to obtain video data and determine one or more video segments from the video data;

[0032] A key frame determination module, configured to determine a key frame for each video segment.

[0033] A salient region determination module, configured to determine one or more salient regions from the key frames.

[0034] A video border image generation module, configured to generate a video border image for each video segment in the video data by combining the video tags of the video data and the one or more salient regions.

[0035] An electronic device includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the processing method based on video data as described above is implemented.

[0036] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processing method based on video data as described above is implemented.

[0037] The embodiments of the present invention have the following advantages:

[0038] In the embodiments of the present invention, by obtaining video data, determining one or more video segments from the video data, determining a key frame for each video segment, determining one or more salient regions from the key frames, and generating a video border image for each video segment in the video data by combining the video tags of the video data and the one or more salient regions, generating a video border according to the video content is realized, so that the video border presents personalization. Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 is a flowchart of the steps of a processing method based on video data provided by an embodiment of the present invention;

[0041] Figure 2 is a flowchart of the steps of another processing method based on video data provided by an embodiment of the present invention;

[0042] Figure 3a is a schematic diagram of a salient region provided by an embodiment of the present invention;

[0043] Figure 3b It is a schematic diagram of a video border image layout provided by an embodiment of the present invention;

[0044] Figure 3c It is a schematic diagram of a border image area division provided by an embodiment of the present invention;

[0045] Figure 4 It is a flowchart of steps of another video data-based processing method provided by an embodiment of the present invention;

[0046] Figure 5 It is a flowchart of steps of another video data-based processing method provided by an embodiment of the present invention;

[0047] Figure 6 It is a structural block diagram of a video data-based processing device provided by an embodiment of the present invention. Detailed implementation manners

[0048] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] In the embodiments of the present invention, by dividing video data into shots, selecting a key frame in each shot, determining the salient regions of each key frame, then clustering the color data of the salient regions in each key frame to determine a color block sequence, and then determining a video border image according to the video label in combination with the position of the salient region, and then dividing it into multiple regions, and performing color rendering on each region according to the color block sequence of the salient region closest to each region for video display, realizing the generation of a border background according to the video content. Specifically, as Figure 1 shown, it may include the following:

[0050] S1. Divide the video data into shots, and select a video frame in each shot.

[0051] S2. Based on the video features of the selected video frames, determine the salient regions of each video frame.

[0052] S3. Extract the color data of the salient regions of each video key frame, cluster the color data, and determine a color block sequence.

[0053] S4. Determine a video border image according to the video label in combination with the position of the salient region.

[0054] S5. Divide the determined video border image into four blocks, and perform color rendering on each border image according to a color block sequence of a salient region closest to each border image for video display.

[0055] The embodiments of the present invention have the following beneficial effects:

[0056] 1. By determining multiple video border images according to the lens, the lens switching is realized, and the video border image switches accordingly, which solves the defect of the current existing technology that the background filling is unchanging. Users can see different borders in the background filling at the edge of the video, and the visual experience is better.

[0057] 2. In the video border processing of each shot, the video border image is determined by combining the video label with the position distribution of the salient area. Taking into account the characteristics of different video types, the determined video border image conforms to the user's viewing characteristics and has a better viewing effect.

[0058] 3. By overlaying color rendering on the blurred video border image according to the color block sequence of the salient area, it not only enriches the user's visual experience, but also does not produce an abrupt feeling when watching the video, effectively improving the user experience.

[0059] The following is further explained:

[0060] Reference Figure 2 , shows a flowchart of a processing method based on video data provided by an embodiment of the present invention, which may specifically include the following steps:

[0061] Step 201: Obtain video data, and determine one or more video segments from the video data.

[0062] After the user uploads the video data to the server, the server can extract all video frames from the video data, segment the video data according to the continuity of adjacent video frames, obtain all shot information of the video data, and then divide the video data into shots according to the shot information. Each shot can correspond to a video clip, and each video clip can include multiple video frames.

[0063] Step 202: determine a key frame for each video segment.

[0064] Among the video frames constituting each shot, a video frame is selected, and a key frame can be selected from the video frame to achieve selecting a video frame in each shot.

[0065] For example, the video data is divided into shots to obtain 20 shots, and a key frame is selected from a set of video frames of each shot to obtain 20 key frames.

[0066] Step 203: Determine one or more significant regions from the key frames.

[0067] For a key frame, the video features of the frame can be input into a trained saliency detection model, and then the saliency map data of the key frame can be output. The saliency detection model can use models such as LeNet, FCN, VGG-Net, RCNN, fast-RCNN, SPP, etc.

[0068] In a specific implementation, the saliency map data can include the saliency values predicted by the model for each pixel point. The larger the saliency value, the more significant the pixel point. After obtaining the saliency map data, the significant regions in the key frame can be determined based on the saliency values in the saliency map data. For example, Figure 3a determine significant region ①, significant region ②, and significant region ③ from the key frame.

[0069] For example, the video features of 20 key frames are respectively input into a trained saliency detection model to obtain the significant regions of the 20 key frames.

[0070] Step 204: Combine the video labels of the video data and the one or more significant regions to generate a video border image for each video segment in the video data.

[0071] In a specific implementation, a video label recognition module can be used to automatically label the video data with type labels. When labeling the video, multi-modal information such as the video text title, video cover image, and video content can be used for analysis to determine the video label. The video label can include news, funny, food, travel, game, pet, beauty, home, fitness, film and television, etc.

[0072] After obtaining the video label, considering the characteristics of different video types, the video category of the video label can be determined. Then, the video border image can be determined by combining the video category and the position of the significant region. The obtained video border images all come from the corresponding video frames and are part of the video frames, and mainly intercept partial region images through the video label and the position of the significant region.

[0073] As an example, the video labels can be divided into three major categories:

[0074] Category A (less emphasis in video content): News, Funny.

[0075] Category B (line of sight focus in video content): Food, Pet, Fitness, Beauty.

[0076] Category C (multiple focuses in video content): Travel, Game, Film and Television, Home.

[0077] In one embodiment of the present invention, generating a video border image for each video segment in the video data by combining the video tags of the video data and the one or more salient regions may include:

[0078] According to the video tags of the video data, determining a target salient region from the one or more salient regions, and determining the arrangement mode of the target salient region; arranging the target salient region according to the arrangement mode to obtain a video border image for each video segment in the video data.

[0079] In a specific implementation, a target salient region (which may be one or more) may be selected from the one or more salient regions according to different video tags, and the arrangement mode of the target salient region may be determined according to different video tags, and then the target salient region is arranged according to the arrangement mode to obtain a video border image for the current video segment.

[0080] As an example, the target salient region is the salient region with the largest area, and may include any one of the following: the salient region with the largest area, the top N salient regions in terms of area size ranking;

[0081] wherein, N is a positive integer greater than 1.

[0082] As an example, the arrangement mode of the target salient region includes any one of the following:

[0083] Repeating and arranging in a circle according to the border width;

[0084] Arranging the target salient regions in a ring.

[0085] For example, when the video tag belongs to category A, the image of one week of the video frame edge can be directly selected as the video border image.

[0086] For another example, when the video tag belongs to category B, the image of the salient region with the largest area is selected and repeated and arranged in a circle according to the border width as the video border image, such as Figure 3b , the video tag is of the pet category and is classified into category B of the video categories. In each video frame, the image of the salient region with the largest range (cat face) is selected and repeated and arranged in a circle according to the border width as the video border image.

[0087] For still another example, when the video tag belongs to category C, the circular image composed of the top 3 salient regions arranged from large to small in area, such as Figure 3a the circular dotted line in, is used as the video border image.

[0088] In an embodiment of the present invention, the video border image is determined by combining video tags with the position distribution of salient regions, taking into account the characteristics of different video types. For example, for Category A videos, there are usually not many key points, and the border background is filled with edge images, providing a better visual experience for users. In the video content of Category B videos, there are usually areas of visual focus. The video border image is generated by repeating the image of the salient region with the largest range in a loop, highlighting the key points and providing a better experience. For Category C videos, there are usually multiple key points, and a circular image formed by selecting the three most salient regions is used to fill the border background, which conforms to the user's visual perception.

[0089] In an embodiment of the present invention, it may further include:

[0090] Determine multiple border image regions from the video border image; for each border image region, determine the salient region with the closest distance; and render the border image region in combination with the salient region with the closest distance.

[0091] In a specific implementation, the obtained video border image can be divided to obtain multiple border image regions, such as being divided into four regions: upper left, upper right, lower left, and lower right.

[0092] For each border image region, determine the salient region with the closest distance, and then the border image region can be rendered in combination with the salient region with the closest distance.

[0093] In an embodiment of the present invention, the rendering of the border image region in combination with the salient region with the closest distance may include:

[0094] Determine color block sequence information for the salient region with the closest distance; and perform color rendering on the border image region according to the color block sequence information.

[0095] In a specific implementation, for the salient region with the closest distance, its color block sequence information can be determined by analyzing its color. The color block sequence information may include color block information sorted by the proportion of color blocks in the salient region. Then, the border image region can be color-rendered according to the color block sequence information, so that color rendering is superimposed on the blurred video border image, which not only enriches the user's visual experience but also does not produce a sudden feeling during video viewing, effectively improving the user experience.

[0096] In an embodiment of the present invention, the determining of the color block sequence information for the salient region with the closest distance may include:

[0097] For the nearest significant region, determine a plurality of color data; cluster the plurality of color data to obtain a plurality of color block information and their color block ratios; generate color block sequence information according to the plurality of color block information and their color block ratios.

[0098] In a specific implementation, images of a plurality of significant regions in each video frame can be extracted, color data of each significant region image can be obtained, and clustering is performed on the color data of each significant region image, and corresponding color block sequences can be obtained. The color block sequences are color block information sorted according to the ratio after clustering.

[0099] For example, a video frame has three significant regions A1, A2, and A3. After clustering, the clustering result of A1 is: red 30%, yellow 25%, blue 18%, green 12%... If the TOP3 is taken to generate a color block sequence, it is (red 30%, yellow 25%, blue 18%).

[0100] After obtaining the video border image, it can be first blurred, and then divided into four parts: upper left, upper right, lower left, and lower right according to the quadrant positions, that is, the border image regions. Then, the nearest significant regions in the corresponding video frame image are obtained. According to the color block sequence of the nearest significant region, color rendering is performed on this part of the video border image.

[0101] Such as Figure 3c , there are 6 significant regions in the video frame image. It can be determined that the nearest to the upper left border image is significant region ①, the nearest to the lower left border image is significant region ②, the nearest to the upper right border image is significant region ⑤, and the nearest to the lower right border image is significant region ⑥. Taking the upper left border image as an example, according to the color block sequence (red 30%, yellow 25%, blue 18%) of significant region ①, color rendering is performed on the upper left border image. And so on, the rendering of each border is completed.

[0102] After color rendering, each processed video border image can be used as the background filling for the corresponding shot. For example, if the video data has 20 shots, there are 20 rendered video border images. According to the shot switching, the video border images change accordingly to complete the video playback display.

[0103] In the embodiments of the present invention, by obtaining video data, and determining one or more video segments from the video data, for each video segment, determining a key frame, and determining one or more significant regions from the key frame, combining the video tags of the video data and the one or more significant regions to generate video border images for each video segment in the video data, it realizes generating video borders according to video content, making the video borders present personalization.

[0104] Reference Figure 4 , a step flowchart of another video data - based processing method provided by an embodiment of the present invention is shown, which may specifically include the following steps:

[0105] Step 401, obtain video data, and determine one or more video segments from the video data.

[0106] After the user uploads the video data to the server, the server can extract all video frames from the video data, segment the video data according to the continuity of adjacent video frames, obtain all shot information of the video data, and then can divide the video data according to the shot information. Each shot can correspond to a video segment, and each video segment can include multiple video frames.

[0107] Step 402, for each video segment, determine a key frame.

[0108] Among the video frames that make up each shot, select a video frame, and this video frame can select a key frame among them, so as to achieve selecting a video frame in each shot.

[0109] For example, divide the video data into shots, and get 20 shots. Select a key frame from the set of video frames of each shot to obtain 20 key frames.

[0110] Step 403, determine one or more salient regions from the key frames.

[0111] For the key frames, the video features of the frames can be input into a trained saliency detection model, and then the saliency map data of the key frames can be output. The saliency detection model can use the following models: LeNet, FCN, VGG - Net, RCNN, fast - RCNN, SPP, etc.

[0112] In a specific implementation, the saliency map data can include the saliency values predicted by the model for each pixel point. The larger the saliency value, the more salient the pixel point. After obtaining the saliency map data, the salient regions in the key frames can be determined based on the saliency values in the saliency map data, such as Figure 3a , determine salient region ①, salient region ②, and salient region ③ from the key frames.

[0113] For example, input the video features of 20 key frames into the trained saliency detection model respectively to obtain the salient regions of 20 key frames.

[0114] Step 404, combine the video tags of the video data and the one or more salient regions to generate a video border image for each video segment in the video data.

[0115] In a specific implementation, the video tag recognition module can be used to automatically assign type tags to video data. When tagging a video, multimodal information such as the video text title, video cover image, and video content can be utilized for analysis to determine the video tags. The video tags can include news, comedy, food, travel, games, pets, beauty, home, fitness, movies, etc.

[0116] After obtaining the video tags, considering the characteristics of different video types, the video category corresponding to the video tags can be determined. Furthermore, the video border image can be determined by combining the video category and the position of the salient region. The obtained video border images all originate from the corresponding video frames and are part of the video frames. The partial region images are mainly intercepted through the video tags and the position of the salient region.

[0117] As an example, the video tags can be divided into three major categories:

[0118] Category A (less key points in video content): News, Comedy.

[0119] Category B (there are visual key points in video content): Food, Pets, Fitness, Beauty.

[0120] Category C (there are multiple key points in video content): Travel, Games, Movies, Home.

[0121] Step 405: Determine multiple border image regions from the video border image.

[0122] Step 406: For each border image region, determine the salient region with the closest distance.

[0123] Step 407: Render the border image region in combination with the salient region with the closest distance.

[0124] In a specific implementation, the images of multiple salient regions in each video frame can be extracted, the color data of each salient region image can be obtained, and the color data of each salient region image can be clustered to obtain the corresponding color block sequence. The color block sequence is the color block information sorted by proportion after clustering.

[0125] For example, if a video frame has three salient regions A1, A2, and A3, after clustering, the clustering result of A1 is: red 30%, yellow 25%, blue 18%, green 12%..., if the TOP3 is taken to generate the color block sequence, it is (red 30%, yellow 25%, blue 18%).

[0126] After obtaining the video border image, it can be blurred first, and then divided into four parts: upper left, upper right, lower left, and lower right according to the quadrant positions, that is, the border image areas. Then, the most significant areas closest to the four borders in the corresponding video frame image are obtained. According to the color block sequence of the closest significant area, color rendering is performed on this part of the video border image.

[0127] For example Figure 3c , there are 6 significant areas in the video frame image. It can be determined that the one closest to the upper left border image is significant area ①, the one closest to the lower left border image is significant area ②, the one closest to the upper right border image is significant area ⑤, and the one closest to the lower right border image is significant area ⑥. Taking the upper left border image as an example, according to the color block sequence of significant area ① (30% red, 25% yellow, 18% blue), color rendering is performed on the upper left border image. And so on, the rendering of each border is completed.

[0128] After color rendering, each processed video border image can be used as the background filling for the corresponding shot. For example, if the video data has 20 shots, there will be 20 rendered video border images. According to the shot switching, the video border images change accordingly to complete the video playback display.

[0129] Referring to Figure 5 , the flowchart of steps of another processing method based on video data provided by an embodiment of the present invention is shown, which may specifically include the following steps:

[0130] Step 501, obtain video data and determine one or more video segments from the video data.

[0131] After the user uploads the video data to the server, the server can extract all video frames from the video data, divide the video data according to the continuity of adjacent video frames, obtain all shot information of the video data, and then divide the video data according to the shot information. Each shot can correspond to a video segment, and each video segment can include multiple video frames.

[0132] Step 502, determine a key frame for each video segment.

[0133] Among the video frames that make up each shot, select a video frame, and this video frame can be selected as the key frame therein to achieve selecting one video frame in each shot.

[0134] For example, divide the video data into shots, obtain 20 shots, and select a key frame from the video frame set of each shot to obtain 20 key frames.

[0135] Step 503: Determine one or more significant regions from the key frames.

[0136] For key frames, the video features of the frames can be input into a trained saliency detection model, and then the saliency map data of the key frames can be output. The saliency detection model can use models such as LeNet, FCN, VGG-Net, RCNN, fast-RCNN, SPP, etc.

[0137] In a specific implementation, the saliency map data can include the saliency values predicted by the model for each pixel point. The larger the saliency value, the more significant the pixel point. After obtaining the saliency map data, the significant regions in the key frames can be determined based on the saliency values in the saliency map data, such as Figure 3a , determine significant region ①, significant region ②, and significant region ③ from the key frames.

[0138] For example, input the video features of 20 key frames into the trained saliency detection model respectively to obtain the significant regions of the 20 key frames.

[0139] Step 504: According to the video label of the video data, determine the target significant region from the one or more significant regions, and determine the layout mode of the target significant region.

[0140] Step 505: Arrange the target significant region according to the layout mode to obtain a video border image for each video segment in the video data.

[0141] In a specific implementation, according to different video labels, the target significant region (which can be one or more) can be selected from the one or more significant regions, and the layout mode of the target significant region can be determined according to different video labels. Then, arrange the target significant region according to the layout mode to obtain a video border image for the current video segment.

[0142] Step 506: Determine a plurality of border image regions from the video border image.

[0143] Step 507: For each border image region, determine the nearest significant region.

[0144] Step 508: Render the border image region in combination with the nearest significant region.

[0145] In a specific implementation, the images of multiple significant regions in each video frame can be extracted, the color data of each significant region image can be obtained, and the color data of each significant region image can be clustered to obtain the corresponding color block sequence. The color block sequence is the color block information sorted according to the proportion after clustering.

[0146] For example, a video frame has three significant regions A1, A2, and A3. After clustering, the clustering results of A1 are: 30% red, 25% yellow, 18% blue, 12% green..., if we take the TOP3 to generate a color block sequence, it is (30% red, 25% yellow, 18% blue).

[0147] After obtaining the video border image, it can be first blurred, and then divided into four parts: upper left, upper right, lower left, and lower right according to the quadrant positions, that is, the border image regions. Then, the significant regions closest to the four borders in the corresponding video frame image are obtained. According to the color block sequence of the closest significant region, color rendering is performed on this part of the video border image.

[0148] Such as Figure 3c , there are 6 significant regions in the video frame image. It can be determined that the significant region closest to the upper left border image is significant region ①, the significant region closest to the lower left border image is significant region ②, the significant region closest to the upper right border image is significant region ⑤, and the significant region closest to the lower right border image is significant region ⑥. Taking the upper left border image as an example, according to the color block sequence of significant region ① (30% red, 25% yellow, 18% blue), color rendering is performed on the upper left border image. And so on, the rendering of each border is completed.

[0149] After color rendering, each processed video border image can be used as the background filling for the corresponding shot. If the video data has 20 shots, there will be 20 rendered video border images. According to the shot switching, the video border images change accordingly, and the video playback display is completed.

[0150] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0151] Referring to Figure 6 , a schematic structural diagram of a processing device based on video data provided by an embodiment of the present invention is shown, which may specifically include the following modules:

[0152] A video segment determination module 601, configured to obtain video data and determine one or more video segments from the video data.

[0153] A key frame determination module 602, configured to determine a key frame for each video segment.

[0154] A saliency region determination module 603, configured to determine one or more saliency regions from the key frames.

[0155] A video border image generation module 604, configured to generate a video border image for each video segment in the video data by combining the video tags of the video data and the one or more saliency regions.

[0156] In an embodiment of the present invention, it further includes:

[0157] A border image region determination module, configured to determine a plurality of border image regions from the video border image.

[0158] A nearest saliency region determination module, configured to determine the nearest saliency region for each border image region.

[0159] A border image region rendering module, configured to render the border image region by combining the nearest saliency region.

[0160] In an embodiment of the present invention, the border image region rendering module includes:

[0161] A color block sequence information determination sub-module, configured to determine color block sequence information for the nearest saliency region; wherein, the color block sequence information includes color block information sorted by color block proportion in the saliency region.

[0162] A rendering sub-module according to the color block sequence information, configured to perform color rendering on the border image region according to the color block sequence information.

[0163] In an embodiment of the present invention, the color block sequence information determination sub-module includes:

[0164] A color data determination unit, configured to determine a plurality of color data for the nearest saliency region.

[0165] A color block information and its color block proportion obtaining unit, configured to cluster the plurality of color data to obtain a plurality of color block information and their color block proportions.

[0166] A color block sequence information generation unit, configured to generate color block sequence information according to the plurality of color block information and their color block proportions.

[0167] In an embodiment of the present invention, the video border image generation module 604 includes:

[0168] A target salient region and its layout determination module, configured to determine a target salient region from the one or more salient regions according to the video tags of the video data, and determine the layout of the target salient region.

[0169] A target salient region layout module, configured to layout the target salient region according to the layout, and obtain a video border image for each video segment in the video data.

[0170] In an embodiment of the present invention, the target salient region is the salient region with the largest area, including any one of the following:

[0171] The salient region with the largest area, the top N salient regions sorted by area size;

[0172] wherein, N is a positive integer greater than 1.

[0173] In an embodiment of the present invention, the layout of the target salient region includes any one of the following:

[0174] Repeating and arranging in a circle according to the border width;

[0175] Arranging the target salient regions in a ring.

[0176] In an embodiment of the present invention, by acquiring video data, determining one or more video segments from the video data, determining a key frame for each video segment, determining one or more salient regions from the key frame, and combining the video tags of the video data and the one or more salient regions, generating a video border image for each video segment in the video data, realizing generating a video border according to the video content, so that the video border presents personalization.

[0177] An embodiment of the present invention further provides an electronic device, which may include a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the above video data-based processing method is implemented.

[0178] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above video data-based processing method is implemented.

[0179] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment.

[0180] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0181] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0182] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0183] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0184] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0185] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0186] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the above elements.

[0187] The above has introduced in detail a method and device for processing video data. Specific examples are used in this text to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A method for processing video data, characterized in that, The method includes: Obtaining video data, and determining one or more video segments from the video data; For each video segment, determining a key frame; Determining one or more salient regions from the key frames; Combining the video tags of the video data and the one or more salient regions to generate a video border image for each video segment in the video data; Determining a plurality of border image regions from the video border image; For each border image region, determining the salient region with the closest distance; Combining the salient region with the closest distance to render the border image region; The combining the salient region with the closest distance to render the border image region includes: For the salient region with the closest distance, determining color block sequence information; wherein, the color block sequence information includes color block information sorted by the proportion of color blocks in the salient region; Performing color rendering on the border image region according to the color block sequence information; The determining the color block sequence information for the salient region with the closest distance includes: For the salient region with the closest distance, determining a plurality of color data; Clustering the plurality of color data to obtain a plurality of color block information and their color block proportions; Generating color block sequence information according to the plurality of color block information and their color block proportions.

2. The method according to claim 1, wherein The combining the video tags of the video data and the one or more salient regions to generate a video border image for each video segment in the video data includes: According to the video tags of the video data, determining a target salient region from the one or more salient regions, and determining the arrangement mode of the target salient region; Arranging the target salient region according to the arrangement mode to obtain a video border image for each video segment in the video data.

3. The method according to claim 2, characterized in that, The target salient region is the salient region with the largest area, including any one of the following: The salient region with the largest area, the first N salient regions sorted by area size; Wherein, N is a positive integer greater than 1.

4. The method according to claim 2, wherein The arrangement mode of the target salient region includes any one of the following: Repeating and arranging in a circle according to the border width; Arranging the target salient regions in a ring shape.

5. A processing device based on video data, characterized in that, The apparatus includes: A video segment determination module, configured to obtain video data, and determine one or more video segments from the video data; A key frame determination module, configured to determine a key frame for each video segment; A salient region determination module, configured to determine one or more salient regions from the key frames; A video border image generation module, configured to combine the video tags of the video data and the one or more salient regions to generate a video border image for each video segment in the video data; A border image region determination module, configured to determine a plurality of border image regions from the video border image; A closest salient region determination module, configured to determine the salient region with the closest distance for each border image region; A border image area rendering module, configured to render the border image area in combination with the nearest saliency area; The border image area rendering module includes: A color patch sequence information determination sub-module, configured to determine color patch sequence information for the nearest saliency area; wherein, the color patch sequence information includes color patch information sorted according to the proportion of color patches in the saliency area; A rendering sub-module according to the color patch sequence information, configured to perform color rendering on the border image area according to the color patch sequence information; The color patch sequence information determination sub-module includes: A color data determination unit, configured to determine a plurality of color data for the nearest saliency area; A unit for obtaining color patch information and its color patch proportion, configured to cluster the plurality of color data to obtain a plurality of color patch information and their color patch proportions; A color patch sequence information generation unit, configured to generate color patch sequence information according to the plurality of color patch information and their color patch proportions.

6. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the video data-based processing method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the video data-based processing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Video picture black edge processing method and device

    CN111970556A

  • Video frame detection method and device and electronic equipment

    CN113255812A