A video processing method, apparatus, device, and storage medium
By extracting and mapping target objects from video data, the problem of border occlusion is solved, enabling effective information expression and enhanced visual effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI SHANGQU PLAY NETWORK TECH CO LTD
- Filing Date
- 2022-09-27
- Publication Date
- 2026-05-05
AI Technical Summary
When creating borders for video data, existing technologies cause the target object to be obscured, affecting the expression of information.
By extracting the target object from the original video data and adding a border to a portion of the image data, the target object is then mapped back into the original image data, ensuring that the target object is positioned above the border to avoid occlusion, and dynamic effects are created by changing the border size.
It effectively reduces the obstruction of information by borders, ensures the normal expression of target objects, and creates a three-dimensional visual effect through dynamic changes in borders, thereby enhancing the user experience.
Smart Images

Figure CN115661303B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia technology, and more particularly to a video processing method, apparatus, device, and storage medium. Background Technology
[0002] In promoting games, electronic products, and other similar products, video data is often used to introduce these products. Video data presents information about these products through visuals and sound, making it easier for users to read.
[0003] After recording the original video data, art staff often use professional video production tools to process the video data in post-production to improve its quality.
[0004] In a common post-production process, artists use video editing tools to add uniform borders to video data. However, these borders can obscure parts of the video data, affecting the information conveyed by the video data. Summary of the Invention
[0005] This invention provides a video processing method, apparatus, device, and storage medium to address how to reduce the impact on the information conveyed by video data when creating borders for video data.
[0006] According to one aspect of the present invention, a video processing method is provided, comprising:
[0007] Acquire raw video data, which contains multiple frames of raw image data;
[0008] Extract the specified target object from the original image data in each frame;
[0009] Add borders to at least a portion of the original image data;
[0010] If the border is added, the target object is mapped back into the original image data to obtain the target image data;
[0011] In the original video data, the target image data is replaced with the original image data to obtain the target video data.
[0012] According to another aspect of the present invention, a video processing apparatus is provided, comprising:
[0013] The raw video data acquisition module is used to acquire raw video data, which contains multiple frames of raw image data.
[0014] The target object extraction module is used to extract a specified target object from the original image data in each frame;
[0015] A border-adding module is used to add borders to at least a portion of the original image data;
[0016] The target object mapping module is used to map the target object back into the original image data after the border is added, so as to obtain the target image data.
[0017] The target video data generation module is used to replace the original image data with the target image data in the original video data to obtain the target video data.
[0018] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video processing method according to any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program configured to cause a processor to execute and implement the video processing method according to any embodiment of the present invention.
[0023] According to another aspect of the present invention, a computer program product is provided, characterized in that the computer program product includes a computer program that, when executed by a processor, implements the video processing method according to any embodiment of the present invention.
[0024] In this embodiment, raw video data is acquired, which contains multiple frames of raw image data. A specified target object is extracted from each frame of raw image data. A border is added to at least a portion of the raw image data. If the border addition is complete, the target object is mapped back into the raw image data to obtain target image data. The target image data replaces the raw image data in the raw video data to obtain target video data. This embodiment first extracts the target object from the raw image data, adds a border to the raw image data, and then maps the target object back into the raw image data, ensuring that the target object is above the border. This avoids the border obscuring the target object, ensuring that the target video data normally expresses the main information. Furthermore, a visual gap is created between the target object and non-target objects, which can create a three-dimensional visual effect. There will be continuous changes between the target object and the border in multiple consecutive frames of target image data, which can create various dynamic effects of the target object jumping between the inside and outside of the border.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart of a video processing method provided according to Embodiment 1 of the present invention;
[0028] Figures 2A to 2D This is an example diagram of adding a border according to Embodiment 1 of the present invention;
[0029] Figure 3 This is a schematic diagram illustrating how to set the size of a border according to Embodiment 1 of the present invention;
[0030] Figure 4 This is a schematic diagram of the structure of a video processing device according to Embodiment 2 of the present invention;
[0031] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] Example 1
[0035] Figure 1 This is a flowchart of a video processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where target objects are excluded when creating borders for raw video data. The method can be executed by a video processing device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0036] Step 101: Obtain the raw video data.
[0037] In practical applications, users can collect or create video data of different content in their daily lives, studies, and work, which is recorded as raw video data. The raw video data awaits the addition of borders in post-processing.
[0038] If the user is an artist, they can create raw video data for business objects. Business objects refer to objects that have the business characteristics of different business scenarios.
[0039] Furthermore, the business object can be a physical item, such as a mobile phone, tablet, smartwatch, etc., or it can be virtual data, mostly third-party applications, such as games, game distribution applications, short video applications, shopping applications, etc. This embodiment does not limit this.
[0040] To enable those skilled in the art to better understand the present invention, in this embodiment, a game is used as an example of a business object for illustration.
[0041] The types of games can include MOBA (Multiplayer Online Battle Arena), RPG (Role-playing game), SLG (Simulation Game), etc., and this embodiment does not limit them.
[0042] For a given target audience, promotion may be carried out on different channels. Different channels have differences in terms of duration, content, etc. In order to facilitate subsequent promotion, the art staff can pre-create one or more video data that can cover different channels, referred to as raw video data. This raw video data can be edited according to the channel specifications. For example, the raw video data may be longer than the duration limit of all channels, allowing the art staff to edit it for specific channels. The raw video data may not have background music, allowing the art staff to add background music for specific channels, and so on.
[0043] Furthermore, the content of the raw video data (including images and sound) is related to the business object and can be used to introduce and promote the business object.
[0044] Taking games as an example, the content of raw video data can be divided into two main forms: game content and real-life storylines. Game content can include introductions to the user's control of the game, introductions by a spokesperson, or introductions by a spokesperson wearing in-game costumes. Storylines can be further divided into the following categories:
[0045] 1. Pseudo-food sharing
[0046] The original video data included some food-related materials to attract users' attention and incorporate gameplay elements of eating food while playing games.
[0047] 2. Themes closely related to users' daily lives
[0048] The content of the raw video data closely reflects users' current lifestyles, integrating games into various aspects of life, such as playing games while eating or buying snacks. The first half of this type of material primarily features dialogue between two people, while the second half includes segments showcasing the game integration.
[0049] 3. Exaggerated situational drama
[0050] The raw video data contains material from sitcoms, some of which features exaggerated storylines designed to attract users' attention.
[0051] Of course, the above-described raw video data is merely an example. When implementing this embodiment, other raw video data can be set according to actual circumstances, and this embodiment does not impose any limitations on this. Furthermore, in addition to the above-described raw video data, those skilled in the art can use other raw video data as needed, and this embodiment does not impose any limitations on this either.
[0052] In practical applications, the original video data contains multiple frames of image data. In order to promote business targets, information such as icons (logos), banners, and ending cards (ECs) are usually configured in different image data.
[0053] The icon (Logo) is the identifier of the business object itself, and it can be a text icon (containing the name of the business object) or a graphic icon.
[0054] Banner information is generally rectangular and is usually located at the top and / or bottom of image data. It can record information about the business object itself (such as a scene in a game, a character in the game, or a name) or information to attract users to purchase or download the business object (such as a gift code).
[0055] The end segment EC contains the identifier of the download business object, such as the business object's own information (e.g., in-game graphics, characters, and names), and the method of purchasing or downloading the business object (e.g., the icon of the application distribution platform, the name and icon of the application distribution platform, the name and icon of the shopping platform, etc.).
[0056] Step 102: Extract the specified target object from each frame of raw image data.
[0057] Generally, raw video data contains multiple frames of image data, referred to as raw image data. Each frame of raw image data contains different elements, which can be displayed elements or virtual elements, such as human image data, subtitle data, tree data, building data, etc. This embodiment does not impose any restrictions on this.
[0058] In this embodiment, the information provided by each element in the original video data varies for different services. Therefore, one or more elements that provide more information can be selected according to the needs of different services to prevent the prohibition from being blocked by the border. These elements can be recorded as target objects.
[0059] Taking games as an example, when promoting games, the raw video data mainly consists of in-game characters, which are the focus of users. Meanwhile, the subtitle data that matches the voice-over can provide users with accurate information. Therefore, both in-game data and subtitle data can be selected as target objects.
[0060] Therefore, for a given target object, we can traverse each frame of the original video data, perform semantic recognition on each frame of the original video data, and detect and extract the target object in each frame of the original video data.
[0061] In one example, if the target object is human portrait data (i.e., a description of a character in a two-dimensional or three-dimensional form), then in this example, a pre-trained Relevance Vector Machine (RVM) can be loaded into memory and run. This RVM has been trained to segment the human portrait data.
[0062] The training of the relevance vector machine is conducted within a Bayesian framework. Based on the structure of prior parameters, it uses automatic relevance determination (ARD) to remove irrelevant points, thus obtaining a sparse model. During the iterative learning process on the sample data, the posterior distribution of most parameters tends to zero and is independent of the predicted values. The points corresponding to those non-zero parameters are called relevance vectors, reflecting the main features of the data.
[0063] Then, the raw image data is input into the correlation vector machine for semantic recognition in order to extract human image data as the target object.
[0064] Generally, characters and their accessories (such as hats, game props, etc.) are semantically similar. Therefore, when segmenting portrait data, the correlation vector machine may segment the accessories into the portrait data along with the characters.
[0065] For example, such as Figure 2A The original image data shown depicts a circus ringmaster wearing a hat. A correlation vector machine is used to analyze the data... Figure 2A The original image data segmentation input and output shown Figure 2B The image of the person wearing a hat shown in the image data has been segmented out.
[0066] In another example, the target object is caption data, which is usually located in the center of the image data background and is easily obscured when a border is added.
[0067] Considering that image data may contain other text data in addition to caption data, and that distinguishing between caption data and non-caption data requires a lot of computation, the target object can be expanded from caption data to text data.
[0068] In this example, a text recognition network can be loaded, which supports optical character recognition (OCR).
[0069] To reduce the computational cost of optical character recognition, deep learning methods such as TexRNet (deep text matching) and HRNet (high-resolution network) or rules (regions located at the bottom of the image data) can be used to segment regions containing text from the original image data, thus creating text image data.
[0070] The region image data is input into a text recognition network for optical character recognition in order to extract text data as the target object.
[0071] Of course, the target objects and extraction methods described above are merely examples. When implementing the embodiments of the present invention, other target objects and extraction methods can be set according to actual circumstances, and the embodiments of the present invention do not impose any limitations on this. Furthermore, in addition to the target objects and extraction methods described above, those skilled in the art can also employ other target objects and extraction methods according to actual needs, and the embodiments of the present invention do not impose any limitations on this.
[0072] Step 103: Add borders to at least a portion of the original image data.
[0073] In practice, at least some (i.e., some or all) of the original image data can be filtered out based on business requirements (art design), content of the original video data, etc., and borders can be added to the filtered image data.
[0074] When adding borders to all raw image data, the borders of multiple frames of raw video data are neat and uniform, which can reduce the amount of computation and improve the efficiency of post-processing.
[0075] When adding borders to a portion of the original image data, the borders change across multiple frames of the original image data, creating a dynamic effect that better matches the content of the original video data and improves the quality of post-processing.
[0076] In one embodiment of the present invention, step 103 may include the following steps:
[0077] Step 1031: Set the border size for at least a portion of the original image data.
[0078] For the original image data to be filtered for borders, a uniform size can be set for the borders, or a variable size can be set for the borders based on the content of the original image data. This embodiment does not impose any restrictions on this.
[0079] In one embodiment of the present invention, if borders are added to all original image data, then the size of the borders of all original image data can be unified.
[0080] In this embodiment, the attributes of the original image data can be queried to obtain the width and height of the original image data.
[0081] like Figure 2C As shown, the width Width of the original image data is taken by a preset first ratio (the first ratio is an empirical value, such as 10%), that is, the width Width of the original image data is multiplied by the preset first ratio (such as 10%) to obtain the size A of the border 201 in the horizontal direction.
[0082] The height of the original image data is taken by a preset second ratio (the second ratio is an empirical value, such as 10%), that is, the height of the original image data is multiplied by the preset second ratio (such as 10%) to obtain the size B of the border 201 in the vertical direction.
[0083] In another embodiment of the present invention, step 1031 may further include the following steps:
[0084] Step 10311: Query the time range of the portrait data in the original video data.
[0085] Generally, the characters in the original video data are one of the main elements that drive the plot forward. The main characters in the original video data (such as spokespeople, protagonists in games, etc.) have a particularly significant impact on the plot.
[0086] Therefore, in this embodiment, the border can be dynamically set with reference to the distribution of human image data in the original video data, so that the border matches the plot development of the original video data, and the user's attention is focused on the plot development.
[0087] Furthermore, the time range of the portrait data is marked on the timeline of the original video data. The portrait data can be the portrait data of any character or the portrait data of the main character. This embodiment does not limit this.
[0088] Borders are set for the raw image data within these time ranges, but no borders are set for the raw image data outside these time ranges.
[0089] In practice, the time point of the portrait data can be queried in the timeline of the original video data. If the portrait data is the target object, the time point of the portrait data cached when the target object was first identified can be queried.
[0090] Facial recognition is performed on human image data to identify facial data.
[0091] For facial data of the same person, all time points can be traversed. For any two adjacent time points, the difference between the two adjacent time points can be calculated to obtain the time difference between them. This time difference is then compared with a preset distance threshold. If the time difference between two adjacent time points is less than or equal to the preset distance threshold, it means that the two adjacent time points are relatively close. In this case, the two adjacent time points can be connected. After traversing all time points and connecting the closer time points, a range formed by connecting one or more time points is obtained, which is used as the time range.
[0092] In addition, it can be determined whether the time ranges corresponding to different roles overlap. If the time ranges corresponding to different roles overlap, a first role range and a second role range can be set. The first role range is the time range corresponding to the role that is ranked first on the timeline of the original video data, and the second role range is the time range corresponding to the role that is ranked last on the timeline of the original video data.
[0093] On the one hand, the duration of overlap between the first role range and the second role range is calculated, and the ratio between that duration and the first role range is calculated.
[0094] On the other hand, the difference between the range of the first role and the range of the second role is calculated as the difference range.
[0095] If the ratio is less than the preset overlap threshold and the difference range is less than the preset range threshold, it means that the overlap between the first character range and the second character range is small and the duration between the first character range and the second character range is relatively close. That is, the first character range and the second character range are independent of each other in terms of plot. In order to avoid the border from appearing frequently in a short period of time and to ensure focus, the first character range or the second character range can be ignored.
[0096] If the percentage is greater than or equal to the preset overlap threshold, or the difference range is greater than or equal to the preset range threshold, it means that there is a large overlap between the first character range and the second character range, and the first character range and the second character range have a large difference in duration. That is, the first character range and the second character range are related in the plot. In order to ensure the continuity of the plot, the first character range and the second character range can be superimposed to obtain a new time range.
[0097] Step 10312: Divide the time range into three regions, namely the head range, the middle range, and the tail range.
[0098] Step 10313: Set the border size for the original image data within the head area in an incremental manner.
[0099] Step 10314: Keep the border size unchanged for the original image data within the central area.
[0100] Step 10315: Set the border size for the original image data within the head area in a decreasing manner.
[0101] In this embodiment, as Figure 3 As shown, on the time axis T of the original video data, for each time range [T1, T2], three regions can be divided, which are denoted as the head range [T1, T4], the middle range [T4, T6], and the tail range [T6, T2].
[0102] Among them, the head range [T1, T4] is located before the middle range [T4, T6], and the middle range [T4, T6] is located before the tail range [T6, T2]. That is, the time points of the head range [T1, T4] are all less than the time points of the middle range [T4, T6], and the time points of the middle range [T4, T6] are all less than the time points of the tail range [T6, T2].
[0103] For the head range [T1, T4], the size S of the border (including, for example, ...) of the original image data within the head range [T1, T4] can be set in an incremental manner. Figure 2C The dimensions A of the border 201 in the horizontal direction and B of the border 201 in the vertical direction are shown. That is, within the header range [T1, T4], the dimensions S of the border of the original image data of each frame (including the dimensions A and B of the border 201 in the vertical direction) are shown. Figure 2C The dimensions A of the horizontal direction and B of the vertical direction of the border 201 shown are increased from zero in chronological order until they reach their maximum values.
[0104] The increment method is generally uniform. Assuming the number of frames of the original image data within the head range [T1, T4] is F1, and the size of the border (including...) Figure 2C If the maximum value of the horizontal dimension A and the vertical dimension B of the border 201 shown is S1, then the first step length of the increment between every two frames of original image data is F1 / S1.
[0105] Furthermore, if F1 / S1 is not an integer, F1 / S1 can be rounded up or down, and the rounded value can be set as the first step length that increases between every two frames of original image data.
[0106] For the middle range [T4, T6], the original image data within the middle range [T4, T6] can maintain the bounding box size S (including, for example, ...). Figure 2CThe dimensions A in the horizontal direction and B in the vertical direction of the border 201 shown remain unchanged, so that the size of the border is kept at the maximum value S1.
[0107] For the tail range [T6, T2], the border size S of the original image data within the tail range [T6, T2] can be set in a decreasing manner (including, for example, ...). Figure 2C The dimensions A of the border 201 in the horizontal direction and B of the border 201 in the vertical direction are shown. That is, within the tail range [T6, T2], the dimensions S of the border of the original image data of each frame (including the dimensions A and B of the border 201 in the vertical direction) are shown. Figure 2C The dimensions A of the horizontal direction and B of the vertical direction of the border 201 shown decrease from their maximum values in chronological order until they reach zero.
[0108] The decreasing method is generally uniform. Assuming the number of frames of the original image data within the tail range [T6, T2] is F2, and the size of the border (including...) Figure 2C If the maximum value of the horizontal dimension A and the vertical dimension B of the border 201 shown is S1, then the second step size that decreases between every two frames of original image data is F2 / S1.
[0109] Furthermore, if F2 / S1 is not an integer, F2 / S1 can be rounded up or down, and the rounded value can be set as the second step size that decreases between every two frames of original image data.
[0110] Over the entire time span, the size of the borders in each frame of the original image data will change from zero to a maximum value, remain at the maximum value, and then decrease from the maximum value back to zero. This gradual process introduces the borders by increasing their size and eliminates them by decreasing their size, making the borders appear more natural and reducing any abruptness while attracting the user's attention.
[0111] In the specific implementation, considering that the length of the time range is not fixed, the length of the head range and the length of the tail range can be determined first. The middle range is the other areas in the time range excluding the head range and the tail range. This simplifies the operation of dividing the head range, the middle range and the tail range and reduces the processing time.
[0112] Furthermore, the lengths of the head and tail regions can be default empirical values or dynamically set according to the plot development of the original video data. In this case, this embodiment does not impose any restrictions on this.
[0113] In a dynamically configured manner, such as Figure 3As shown, starting from the beginning point T1 of the time range [T1, T2], the first candidate range [T1, T3] is taken sequentially, and starting from the end point T2 of the time range [T1, T2], the second candidate range [T5, T2] is taken in reverse order. The lengths of the first candidate range [T1, T3] and the second candidate range [T5, T2] can be default empirical values.
[0114] On the one hand, a first brilliant value representing the brilliance of each frame of original image data in the first candidate range is calculated using methods such as a flexible detect to summarize network (DSNet) for video summarization. The summarization network can extract the main part of the original video data to generate a segment, and use this segment to summarize the content of the original video data. The summarization network includes two network frameworks: anchor-based method and anchor-free method.
[0115] In the anchor-based method, a multi-scale interval of proposals (candidate boxes) is provided for dense sampling to extract their long-term, time-dependent features for proposal location regression and importance prediction. Here, positive and negative samples are assigned to generate correctness and completeness information for the summary.
[0116] In the anchor-free method, the importance of each frame of image data and segment position in the video data is directly predicted.
[0117] The region between the starting point T1 of the time range [T1, T2] and the time point T4 where the original data with the highest first highlight value is located is set as the head range [T1, T4]. Then, within this head range [T1, T4], the size S of the border starts from zero and increases. When it reaches the maximum value S1, it represents the most exciting scene (i.e., the original image data) within the local area of the original video data. This makes the increase of the size S of the border within the head range [T1, T4] synchronized with the content within the local area of the original video data, thus improving the fit between the size S of the border within the head range [T1, T2] and the content within the local area of the original video data.
[0118] On the other hand, a second brilliance value, representing the brilliance of each frame of original image data within the second candidate range, is calculated using methods such as a summary generation network.
[0119] The region between the end point T2 of the time range [T1, T2] and the time point T6 where the original data with the second highest highlight value is located is set as the tail range [T6, T2]. Then, in this tail range [T6, T2], the size S of the border starts to decrease from the maximum value S1 of the most highlight scene (i.e., the original image data) in the local area of the original video data, until it becomes zero. This makes the decrease of the size S of the border in the tail range [T6, T2] synchronized with the content in the local area of the original video data, thus improving the fit between the size S of the border in the tail range [T6, T2] and the content in the local area of the original video data.
[0120] Since the head range [T1, T4] and the tail range [T6, T2] are independent of each other, the head range [T1, T4] and the tail range [T6, T2] can be divided asynchronously. When determining the head range [T1, T4] and the tail range [T6, T2], the other regions within the time range [T1, T2] other than the head range [T1, T4] and the tail range [T6, T2] can be set as the middle range [T4, T3].
[0121] Step 1032: Set the target range according to the size of the edges in at least part of the original image data.
[0122] In this embodiment, as Figure 2C As shown, for at least a portion of the original image data to which a border is to be added, the target range can be obtained by extending inward from each edge in each frame of the original image data according to the size of the border (including the size A of the border 201 in the horizontal direction and the size B of the border 201 in the vertical direction).
[0123] Step 1033: Fill the target area with the specified color as a border.
[0124] In this embodiment, a specified color can be filled in the target area, that is, the pixels in the target area are uniformly adjusted to the same color. At this time, the target area is denoted as the border.
[0125] In practice, the color can be a default value or a value that can be dynamically set based on the color of the original image data, so that the color of the border matches the color of the original image data, thereby highlighting the border and avoiding the situation where the color of the border is too similar to the color of the original image data, which would make it difficult to distinguish.
[0126] In one method of dynamically setting the border color, for at least a portion of the original image data to which borders are to be added, K-means clustering can be used to cluster the pixels in the at least portion of the original image data according to the color value of the pixels, resulting in multiple candidate clusters.
[0127] The candidate clusters are selected from all candidate clusters, and the top n (n is a positive integer) candidate clusters are selected based on the number of pixels sorted from largest to smallest.
[0128] Using the target cluster as a reference for the color of at least part of the original image data, color values that meet the deviation condition are selected as target colors. The deviation condition is that the distance between the target color and the color value represented by any target cluster (i.e. the color value represented by the center point of the target cluster) is greater than a preset color threshold. That is, the distance between the target color and the color value represented by any target cluster is large, so that the target color and the color value represented by any target cluster are obviously contrasted, and the existence of the border is visually highlighted.
[0129] At this point, the pixels within the target area are filled with the target color to form the border.
[0130] Step 104: If the border addition is completed, the target object is mapped back into the original image data to obtain the target image data.
[0131] like Figure 2D As shown, if the border 201 has been added to at least part of the original image data, the target object 202 can be mapped back to the original image data to obtain the target image data. That is, each pixel in the target object 202 is mapped back to the original image data according to its coordinates in the original image data to obtain the target image data.
[0132] When the target object overlaps with the border, the target object is located at the edge of the original image data. In this case, the target object is segmented. When no border is added to the original image data, the target object is not occluded. When a border is added to the original image data, the target object is occluded in part of the edge area by the border. Therefore, when the segmented target object is mapped back to the original image data to obtain the target image data, the target object in the target image data does not change relative to the target object in the original image data. However, the target object is located on the border and is not occluded by the border, while other elements besides the target object remain occluded by the border.
[0133] Step 105: Replace the original image data with the target image data in the original video data to obtain the target video data.
[0134] In the original video data, each frame of original image data is traversed. For target image data with added borders and remapped target objects, the target image data can replace the corresponding original image data. For original image data without added borders, the original image data can be kept unchanged. At this time, the original video data is recorded as the target video data.
[0135] Visually, users usually perceive the border as superimposed on the target image data, and the target object of the target image data as mapped onto the border. That is, the image of the target image data is located behind the border, and the target object is located in front of the border. When playing the target video data, multiple consecutive frames of target image data can present a dynamic effect of the target object jumping between the inside and outside of the border.
[0136] Furthermore, as the size of the border gradually increases, multiple consecutive frames of target image data can present a dynamic effect of the target object (especially a character (i.e., portrait data)) jumping out of the border from the inside to the outside of the border as it gets closer. When the size of the border remains constant, multiple consecutive frames of target image data can present a dynamic effect of the target object (especially a character (i.e., portrait data)) jumping out of the border and staying there. When the size of the border gradually decreases, multiple consecutive frames of target image data can present a dynamic effect of the target object (especially a character (i.e., portrait data)) jumping back into the border from the outside to the inside of the border as it gets farther away.
[0137] In some cases, if the target video data contains information related to the business object, then the target video data can be published on designated channels (such as news, short videos, novel reading, sports and health, etc.) so that when the client accesses the channel, the target video data is pushed to the client for playback. When users are interested in the business object, they can search for the business object through the information in the target video data, for example, searching for and downloading games from a game distribution platform, and so on.
[0138] In this embodiment, raw video data is acquired, which contains multiple frames of raw image data. A specified target object is extracted from each frame of raw image data. A border is added to at least a portion of the raw image data. If the border addition is complete, the target object is mapped back into the raw image data to obtain target image data. The target image data replaces the raw image data in the raw video data to obtain target video data. This embodiment first extracts the target object from the raw image data, adds a border to the raw image data, and then maps the target object back into the raw image data, ensuring that the target object is above the border. This avoids the border obscuring the target object, ensuring that the target video data normally expresses the main information. Furthermore, a visual gap is created between the target object and non-target objects, which can create a three-dimensional visual effect. There will be continuous changes between the target object and the border in multiple consecutive frames of target image data, which can create various dynamic effects of the target object jumping between the inside and outside of the border.
[0139] Example 2
[0140] Figure 4 This is a schematic diagram of the structure of a video processing device provided in Embodiment 2 of the present invention. Figure 4 As shown, the device includes:
[0141] The raw video data acquisition module 401 is used to acquire raw video data, which contains multiple frames of raw image data.
[0142] The target object extraction module 402 is used to extract a specified target object in each frame of the original image data;
[0143] Border adding module 403 is used to add borders to at least a portion of the original image data;
[0144] The target object mapping module 404 is used to map the target object back into the original image data after the border is added, so as to obtain the target image data.
[0145] The target video data generation module 405 is used to replace the original image data with the target image data in the original video data to obtain the target video data.
[0146] In one embodiment of the present invention, the target object extraction module 402 is further configured to:
[0147] Load the relevant vector machine;
[0148] The original image data is input into the correlation vector machine for semantic recognition in order to extract human image data as the target object.
[0149] In another embodiment of the present invention, the target object extraction module 402 is further configured to:
[0150] Load the text recognition network;
[0151] The regions containing text are segmented from the original image data to obtain text image data;
[0152] The region image data is input into the text recognition network for optical character recognition in order to extract text data as the target object.
[0153] In one embodiment of the present invention, the border adding module 403 is further configured to:
[0154] Set the border size for at least a portion of the original image data;
[0155] The target range is defined according to the dimensions at least in a portion of the original image data at the edges;
[0156] Fill the target area with the specified color as a border.
[0157] In one embodiment of the present invention, the border adding module 403 is further configured to:
[0158] Query the width and height of the original image data;
[0159] A preset first ratio is taken for the width to obtain the horizontal dimension of the border;
[0160] The height is taken as a preset second ratio to obtain the size of the border in the vertical direction.
[0161] In another embodiment of the present invention, the border adding module 403 is further configured to:
[0162] Query the time range of the human image data in the original video data;
[0163] The time range is divided into three regions, namely the head range, the middle range, and the tail range.
[0164] The size of the border is set incrementally for the original image data within the head area;
[0165] The size of the border is maintained for the original image data within the central region;
[0166] The size of the border is set for the original image data within the head area in a decreasing manner.
[0167] In one embodiment of the present invention, the border adding module 403 is further configured to:
[0168] Query the time point of the portrait data in the original video data;
[0169] Recognize facial data from the aforementioned portrait data;
[0170] For the facial data of the same character, if the time difference between two adjacent time points is less than or equal to a preset distance threshold, then the two adjacent time points are connected to obtain a time range.
[0171] In one embodiment of the present invention, the border adding module 403 is further configured to:
[0172] If the time ranges corresponding to different roles overlap, the duration of the overlap between the first role range and the second role range is calculated. The first role range is the time range corresponding to the role that is ranked first, and the second role range is the time range corresponding to the role that is ranked second.
[0173] Calculate the ratio between the duration and the range of the first role;
[0174] Calculate the difference between the first character range and the second character range, and use it as the difference range;
[0175] If the ratio is less than a preset overlap threshold and the difference range is less than a preset range threshold, then the first role range or the second role range is ignored.
[0176] If the percentage is greater than or equal to a preset overlap threshold, or the difference range is greater than or equal to a preset range threshold, then the first role range and the second role range are superimposed to obtain a new time range.
[0177] In one embodiment of the present invention, the border adding module 403 is further configured to:
[0178] The first candidate range is selected sequentially starting from the beginning of the time range, and the second candidate range is selected in reverse order starting from the end of the time range.
[0179] For each frame of the original image data within the first candidate range, a first brilliance value representing the degree of brilliance is calculated;
[0180] The area between the starting point of the time range and the time point where the original data with the highest first brilliance value is located is set as the head range.
[0181] For each frame of the original image data within the second candidate range, a second brilliance value representing the degree of brilliance is calculated;
[0182] The region between the end of the time range and the time point of the original data with the highest second brilliance value is defined as the tail range.
[0183] The regions other than the head and tail regions within the time frame are defined as the middle region.
[0184] In one embodiment of the present invention, the border adding module 403 is further configured to:
[0185] Clustering at least a portion of the original image data according to color values yields multiple candidate clusters;
[0186] The candidate clusters that are sorted from largest to smallest by the number of pixels are selected as the target clusters.
[0187] Color values that meet the deviation condition are selected as target colors, wherein the deviation condition is that the distance between the color value and any color value represented by the target cluster is greater than a preset color threshold.
[0188] The pixels within the target area are filled with the target color to form a border.
[0189] The video processing apparatus provided in the embodiments of the present invention can execute the video processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the video processing method.
[0190] Example 3
[0191] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0192] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0193] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0194] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as video processing methods.
[0195] In some embodiments, the video processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the video processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the video processing method by any other suitable means (e.g., by means of firmware).
[0196] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0197] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0198] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0199] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0200] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0201] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0202] Example 4
[0203] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the video processing method provided in any embodiment of this invention.
[0204] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0205] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0206] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A video processing method, characterized in that, include: Acquire raw video data, which contains multiple frames of raw image data; Extract the specified target object from the original image data in each frame; Adding a border to at least a portion of the original image data includes: setting the size of the border on at least a portion of the original image data; If the border is added, the target object is mapped back into the original image data to obtain the target image data; In the original video data, the target image data is replaced with the original image data to obtain the target video data; The step of setting the border size for at least a portion of the original image data includes: Query the time range of the human image data in the original video data; The time range is divided into three regions, namely the head range, the middle range, and the tail range. The size of the border is set incrementally for the original image data within the head area; The size of the border is maintained for the original image data within the central region; The size of the border is set for the original image data within the tail region in a decreasing manner; If the time ranges corresponding to different roles overlap, the duration of the overlap between the first role range and the second role range is calculated. The first role range is the time range corresponding to the role that is ranked first, and the second role range is the time range corresponding to the role that is ranked second. Calculate the ratio between the duration and the range of the first role; Calculate the difference between the first character range and the second character range, and use it as the difference range; If the ratio is less than a preset overlap threshold and the difference range is less than a preset range threshold, then the first role range or the second role range is ignored. If the ratio is greater than or equal to a preset overlap threshold, or the difference range is greater than or equal to a preset range threshold, then the first role range and the second role range are superimposed to obtain a new time range.
2. The method according to claim 1, characterized in that, Extracting the specified target object from the original image data in each frame includes: Load the relevant vector machine; The original image data is input into the correlation vector machine for semantic recognition in order to extract human image data as the target object.
3. The method according to claim 1, characterized in that, Extracting the specified target object from the original image data in each frame includes: Load the text recognition network; The regions containing text are segmented from the original image data to obtain text image data; The region image data is input into the text recognition network for optical character recognition in order to extract text data as the target object.
4. The method according to any one of claims 1-3, characterized in that, Adding borders to at least a portion of the original image data includes: The target range is defined according to the dimensions at least in a portion of the original image data at the edges; Fill the target area with the specified color as a border.
5. The method according to claim 4, characterized in that, Setting the border size for at least a portion of the original image data includes: Query the width and height of the original image data; A preset first ratio is taken for the width to obtain the horizontal dimension of the border; The height is taken as a preset second ratio to obtain the size of the border in the vertical direction.
6. The method according to claim 1, characterized in that, The step of querying the time range of the human image data in the original video data includes: Query the time point of the portrait data in the original video data; Recognize facial data from the aforementioned portrait data; For the facial data of the same character, if the time difference between two adjacent time points is less than or equal to a preset distance threshold, then the two adjacent time points are connected to obtain a time range.
7. The method according to claim 1, characterized in that, The process of dividing the time range into three regions, designated as the head region, middle region, and tail region, includes: The first candidate range is selected sequentially starting from the beginning of the time range, and the second candidate range is selected in reverse order starting from the end of the time range. For each frame of the original image data within the first candidate range, a first brilliance value representing the degree of brilliance is calculated; The region between the starting point of the time range and the time point of the original image data with the highest first brilliance value is defined as the head range. For each frame of the original image data within the second candidate range, a second brilliance value representing the degree of brilliance is calculated; The region between the end of the time range and the time point of the original image data with the highest second brilliance value is defined as the tail range; The regions other than the head and tail regions within the time frame are defined as the middle region.
8. The method according to claim 4, characterized in that, The step of filling the target area with a specified color as a border includes: Clustering at least a portion of the original image data according to color values yields multiple candidate clusters; The candidate clusters that are sorted from largest to smallest by the number of pixels are selected as the target clusters. Color values that meet the deviation condition are selected as target colors, wherein the deviation condition is that the distance between the color value and any color value represented by the target cluster is greater than a preset color threshold. The pixels within the target area are filled with the target color to form a border.
9. A video processing apparatus, characterized in that, include: The raw video data acquisition module is used to acquire raw video data, which contains multiple frames of raw image data. The target object extraction module is used to extract a specified target object from the original image data in each frame; A border adding module is used to add a border to at least a portion of the original image data, including: setting the size of the border on at least a portion of the original image data; The target object mapping module is used to map the target object back into the original image data after the border is added, so as to obtain the target image data. The target video data generation module is used to replace the original image data with the target image data in the original video data to obtain the target video data; The border adding module is also used for: Query the time range of the human image data in the original video data; The time range is divided into three regions, namely the head range, the middle range, and the tail range. The size of the border is set incrementally for the original image data within the head area; The size of the border is maintained for the original image data within the central region; The size of the border is set for the original image data within the tail region in a decreasing manner; If the time ranges corresponding to different roles overlap, the duration of the overlap between the first role range and the second role range is calculated. The first role range is the time range corresponding to the role that is ranked first, and the second role range is the time range corresponding to the role that is ranked second. Calculate the ratio between the duration and the range of the first role; Calculate the difference between the first character range and the second character range, and use it as the difference range; If the ratio is less than a preset overlap threshold and the difference range is less than a preset range threshold, then the first role range or the second role range is ignored. If the ratio is greater than or equal to a preset overlap threshold, or the difference range is greater than or equal to a preset range threshold, then the first role range and the second role range are superimposed to obtain a new time range.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video processing method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that enables a processor to execute the video processing method according to any one of claims 1-8.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the video processing method according to any one of claims 1-8.
Citation Information
Patent Citations
Video playing method and device, storage medium and electronic equipment
CN113747227A
Multiclass clustering with side information from multiple sources and the application of converting 2d video to 3D
US20120218382A1