Video processing method, electronic device, storage medium and program product
By extracting the area of interest in the video frame and reducing the sharpness of the background area for image fusion, the problem of blurred foreground content after video compression is solved, and the overall clarity after video compression is achieved.
Patent Information
- Application Number
- CN202510606311.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-08
AI Technical Summary
While maintaining efficient compression, it is difficult for existing video compression technology to maintain the clarity of the foreground content of the video, resulting in a decrease in the overall clarity of the video after compression or distortion of the foreground content.
Extract the area of interest of the video frame, reduce the clarity of the background area and perform image fusion, and then compress the video sequence to ensure that the area of interest allocates more code rates.
While maintaining the clarity of the video foreground content, the encoding coding rate requirement of the video sequence is reduced, and the overall clarity balance after video compression is achieved.
Smart Images

Figure CN120455689A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of multimedia technology, and in particular to a video processing method, electronic equipment, storage medium, and program product. Background Art
[0002] Currently, video compression technology is a fundamental method for efficiently transmitting, storing, or applying video data. The information captured in videos captured in daily life and work typically consists of two parts: background information and foreground content. This information plays different roles in practical applications. For example, in the field of video surveillance, foreground content, such as moving objects and salient areas, is of particular interest, as it provides a critical description of the monitored object. Using video compression technology to compress a video can reduce its overall clarity, blurring the foreground content of interest to the user. At higher compression rates, this content can even be distorted. Therefore, maintaining a clear foreground while maintaining efficient compression has become a technical challenge in video compression. Summary of the Invention
[0003] The embodiments of the present application provide a video processing method, electronic device, storage medium, and program product to alleviate or solve one or more technical problems existing in the prior art.
[0004] In a first aspect, an embodiment of the present application provides a video processing method, comprising:
[0005] extracting a region of interest of a first video frame in a first video sequence;
[0006] Performing image processing on a first background region of the first video frame to obtain a second background region; the first background region is a region other than the region of interest in the first video frame; and the clarity of the second background region is lower than that of the first background region;
[0007] Performing image fusion on the region of interest and the second background region to obtain a second video frame;
[0008] The second video sequence is compressed to obtain a compressed target video sequence; the second video sequence includes a plurality of second video frames.
[0009] In a second aspect, an embodiment of the present application provides a video processing device, comprising:
[0010] An extraction module, configured to extract a region of interest from a first video frame in a first video sequence;
[0011] a processing module configured to perform image processing on a first background region of the first video frame to obtain a second background region; the first background region is a region other than the region of interest in the first video frame; and the clarity of the second background region is lower than that of the first background region;
[0012] a fusion module, configured to fuse the image of the region of interest and the second background region to obtain a second video frame;
[0013] The compression module is used to compress the second video sequence to obtain a compressed target video sequence; the second video sequence includes a plurality of second video frames.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements any method of the embodiments of the present application when executing the computer program.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of any one of the embodiments of the present application is implemented.
[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements any method of the embodiments of the present application when executed by a processor.
[0017] According to the technical solution of the embodiment of the present application, a region of interest is extracted from a first video frame in a first video sequence; a first background region of the first video frame is image processed to obtain a second background region, such that the clarity of the second background region is lower than that of the first background region, where the first background region is the region of the first video frame other than the region of interest. The region of interest and the second background region are then image-fused to obtain a second video frame; and the second video sequence is compressed to obtain a compressed target video sequence, which includes multiple second video frames. As can be seen, before the video sequence is compressed, the clarity of the background region in each video frame of the video sequence is reduced, thereby reducing the encoding bit rate required for the video sequence. During compression of the video sequence, more attention is paid to the region of interest with higher clarity, thereby allocating a higher bit rate to the region of interest and a lower bit rate to the background region. This ensures that the clarity of the region of interest is maintained in the target video sequence after compression, avoiding the situation where the video sequence is not clear as a whole after efficient compression, thus achieving a balance between the clarity of the foreground object and the degree of compression.
[0018] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.
[0020] Figure 1 A flowchart of a video processing method provided by an embodiment of the present application is shown;
[0021] Figure 2 A flowchart of a video processing method provided by another embodiment of the present application is shown;
[0022] Figure 3 A schematic diagram of a video processing method provided in an embodiment of the present application is shown;
[0023] Figure 4 A block diagram of a video processing device provided in an embodiment of the present application is shown;
[0024] Figure 5 A block diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present application. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0026] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the scope of protection of the embodiments of the present application.
[0027] The following describes in detail the technical solution of this application and how it solves the aforementioned technical problems using specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following describes the embodiments of this application in detail with reference to the accompanying drawings.
[0028] Figure 1FIG. 1 shows a flow chart of a video processing method provided by an embodiment of the present application. Figure 1 As shown, the method may include step S101, step S102, step S103 and step S104.
[0029] Step S101: extracting a region of interest of a first video frame in a first video sequence.
[0030] The first video sequence includes multiple video frames, and the first video frame can be any video frame in the first video sequence.
[0031] ROI (Regions of Interest) extraction can be performed as follows: identifying a foreground object in the first video frame, and determining the region of interest based on the foreground area corresponding to the foreground object. When identifying the foreground object, any existing object recognition algorithm can be used, such as the YOLO (You Only Look Once) series of object detection algorithms, the DeepLab series of image segmentation algorithms, and DeepSORT (Deep Learning + SORT, deep learning object tracking algorithm analysis).
[0032] After the foreground object is identified, the area where the foreground object is located in the first video frame is the foreground area corresponding to the foreground object. The foreground area can be represented by a rectangular box containing the foreground object. For example, the foreground area is represented by the position information of each vertex of the rectangular box containing the foreground object. The region of interest can be the entire foreground area or a portion of the foreground area.
[0033] Step S102 : performing image processing on the first background area of the first video frame to obtain a second background area. The first background area is the area other than the area of interest in the first video frame, and the clarity of the second background area is lower than that of the first background area.
[0034] Optionally, image processing of the first background region may be performed as follows: blurring the first background region to obtain a second background region. Image blurring refers to low-pass filtering in the spatial domain, including methods such as mean filtering, Gaussian filtering, and median filtering. In this embodiment, any of these low-pass filtering methods can be used to achieve a blurring effect on the first background region.
[0035] Optionally, image processing of the first background region may be performed as follows: performing frequency domain low-pass filtering on the first background region. Frequency domain low-pass filtering converts the image corresponding to the first background region into the frequency domain, filtering out high-frequency portions of the image, such as details and edges, thereby retaining only low-frequency portions of the image, such as smooth regions, to achieve a blurring effect on the first background region.
[0036] Optionally, image processing of the first background region may be performed as follows: performing a morphological operation on the first background region. The morphological operation includes an opening operation or a closing operation. The opening operation is performed by first eroding and then dilating, which can eliminate small noise points in the first background region and smooth the contour. The closing operation is performed by first dilating and then eroding, which can fill small holes in the first background region and expand flat areas.
[0037] Step S103: performing image fusion on the region of interest and the second background region to obtain a second video frame.
[0038] Step S104 : compress the second video sequence to obtain a compressed target video sequence, where the second video sequence includes a plurality of second video frames.
[0039] In this embodiment, each first video frame in the first video sequence is processed according to steps S101 to S103 to obtain second video frames corresponding to each first video frame. The second video frames corresponding to each first video frame are combined to obtain a second video sequence. That is, compared to the first video sequence, the first background region in the second video sequence is processed into the second background region, while the region of interest remains unchanged.
[0040] Optionally, when compressing the second video sequence, the second video sequence can be input into an encoder, which encodes the second video sequence and outputs a video stream. Subsequently, when the video is used, the video stream is input into a decoder for decoding to obtain a reconstructed video sequence, i.e., the compressed target video sequence.
[0041] When encoding the second video sequence using an encoder, the encoder can reduce the quantization parameter value based on information such as the complexity and motion of the region of interest, allocating more bitrate to the region of interest, thereby performing fine encoding and ensuring image quality (including clarity) in the region of interest. For the second background region, the quantization parameter value is correspondingly increased, allocating less bitrate. Thus, given a given encoder bitrate, the lower clarity of the second background region requires less bitrate resources during encoding, allowing more bitrate to be allocated to the region of interest, ensuring image quality in the region of interest after video compression.
[0042] According to the technical solution of the embodiment of the present application, a region of interest is extracted from a first video frame in a first video sequence; a first background region of the first video frame is image processed to obtain a second background region, such that the clarity of the second background region is lower than that of the first background region, where the first background region is the region of the first video frame other than the region of interest. The region of interest and the second background region are then image-fused to obtain a second video frame; and the second video sequence is compressed to obtain a compressed target video sequence, which includes multiple second video frames. As can be seen, before the video sequence is compressed, the clarity of the background region in each video frame of the video sequence is reduced, thereby reducing the encoding bit rate required for the video sequence. During compression of the video sequence, more attention is paid to the region of interest with higher clarity, thereby allocating a higher bit rate to the region of interest and a lower bit rate to the background region. This ensures that the clarity of the region of interest is maintained in the target video sequence after compression, avoiding the situation where the video sequence is not clear as a whole after efficient compression, thus achieving a balance between the clarity of the foreground object and the degree of compression.
[0043] In some embodiments, determining the region of interest based on the foreground region corresponding to the foreground object may be performed by the following steps: screening the foreground region that matches the data transmission parameter from the foreground region as the region of interest.
[0044] The data transmission parameters may include at least one of the following: data transmission bandwidth and data transmission bit rate. The area of the region of interest is positively correlated with the data transmission bandwidth and data transmission bit rate, respectively. A foreground region matching the data transmission parameters refers to a foreground region selected so that its area matches the data transmission parameters.
[0045] If the current data transmission parameters indicate good data transmission conditions, the area of the ROI is relatively large. If the current data transmission parameters indicate poor data transmission conditions, the area of the ROI is relatively small. For example, when the data transmission bit rate is low, part of the foreground area can be selected as the ROI. When the data transmission bit rate is high, if the clarity requirement of all foreground objects is met during video compression, the entire foreground area can be selected as the ROI.
[0046] Optionally, a correlation relationship between a data transmission parameter and an area of the region of interest is pre-configured, and the correlation relationship may include at least one of the following: a correlation relationship between a data transmission bandwidth and an area of the region of interest, and a correlation relationship between a data transmission bit rate and an area of the region of interest. In the correlation relationship, both the data transmission parameter and the area of the region of interest may be in the form of a numerical value or a numerical range. For example, when the data transmission parameter is pre-configured, the area range corresponding to the data transmission parameter is determined based on the correlation relationship between the data transmission parameter and the area of the region of interest, and then, based on the determined area range, the foreground area within the area range is screened out as the region of interest.
[0047] For example, if five foreground objects are identified in the first video frame, each corresponding to a foreground region, the first video frame includes five foreground regions. If the sum of the areas of three of the foreground regions is exactly within the determined area range, then these three foreground regions can be determined as regions of interest.
[0048] In some embodiments, determining the region of interest based on the foreground region corresponding to the foreground object may be performed by the following steps: filtering the region of interest from the foreground region based on preset filtering parameters.
[0049] The preset screening parameters include at least one of the following: target category of the foreground target corresponding to the foreground area, motion information of the foreground target, number of regions in the foreground area, region area, and region position.
[0050] The target category of the foreground target may be, for example, a person, an animal, a plant, a vehicle, or the like. The motion information of the foreground target may include whether the foreground target is in motion. If the foreground target is in motion, the motion information may further include the motion trajectory of the foreground target in the first video sequence. The area of the foreground region refers to the total area of all the foreground regions screened out. The area position of the foreground region may be represented by the position information of key points on the foreground region, for example, the position information of each vertex of the rectangular box corresponding to the foreground region, the position information of the center point of the foreground region, and so on.
[0051] Optionally, the preset screening parameters include a target category of the foreground object. When the target category is determined, the region of interest can be screened from the foreground area based on the target category. For example, if the first video frame includes foreground objects of two categories, a person and a vehicle, and the target category corresponding to the foreground area to be screened is pre-configured to be a person, then based on the target category, the foreground area containing the foreground object "person" is screened from all identified foreground areas as the region of interest.
[0052] Optionally, the preset screening parameters include the motion trajectory of the foreground object. When screening the ROI from the foreground area based on the preset screening parameters, the following steps may be performed: based on the motion trajectory of the foreground object, foreground areas corresponding to foreground objects that meet preset trajectory conditions are screened as the ROI; the preset trajectory conditions include at least one of the following: the trajectory direction matches a preset direction, the trajectory length reaches a preset length threshold, or the trajectory length is the longest.
[0053] The motion trajectory of the foreground object is determined according to position information of the foreground object in a plurality of first video frames.
[0054] For example, the preset trajectory condition is that the trajectory length is the longest. Then, after determining the motion trajectory of each foreground target, the foreground target with the longest trajectory is selected from multiple foreground targets, and the foreground area corresponding to the foreground target is the region of interest.
[0055] For another example, if the preset trajectory condition is that the trajectory length reaches a preset length threshold, then after determining the motion trajectory of each foreground target, foreground targets whose trajectory length reaches the preset length threshold are selected from multiple foreground targets. There may be one or more foreground targets that meet the requirement of a trajectory length reaching the preset length threshold. If only one foreground target's trajectory length reaches the preset length threshold, the foreground region corresponding to that foreground target is determined to be the region of interest. If multiple foreground targets' trajectory lengths reach the preset length threshold, the sum of the foreground regions corresponding to these multiple foreground targets is determined to be the region of interest.
[0056] When determining the motion trajectory of the foreground object, a plurality of consecutive first video frames including the foreground object may be selected from the first video sequence, and position information of the foreground object in each of the selected first video frames may be determined. The motion trajectory of the foreground object may then be determined based on changes in the position information of the foreground object in the plurality of consecutive first video frames.
[0057] Optionally, the preset screening parameters include a number of regions in the foreground area. When screening the ROI from the foreground area according to the preset screening parameters, foreground areas that meet the preset number of regions can be screened as the ROI. For example, if a total of five foreground objects are identified in the first video frame and the preset number of regions is three, then three foreground areas corresponding to the five identified foreground objects can be screened as the ROI.
[0058] When filtering foreground areas that meet the number of regions, you can use any of the following methods to filter until a foreground area that meets the number of regions is found: random filtering, filtering in descending order of area size, filtering in descending order of distance between the foreground area and the center point of the video frame, and so on.
[0059] Optionally, the preset screening parameters include the area of the foreground region. When screening the region of interest from the foreground region according to the preset screening parameters, the screening can be performed in any of the following ways: screening the foreground region with the largest area, or screening the foreground region with an area reaching a preset area threshold.
[0060] Optionally, the preset screening parameters include a location of the foreground area. When screening the region of interest from the foreground area based on the preset screening parameters, the screening can be performed in any of the following ways: screening the foreground area whose location is closest to the center point of the video frame, or screening the foreground area whose distance from the center point of the video frame is less than or equal to a preset distance threshold.
[0061] In this embodiment, by filtering the ROI based on preset filtering parameters, which include at least one of the following: the target category of the foreground target corresponding to the foreground region, the motion information of the foreground target, the number of regions in the foreground region, the region area, and the region location, the ROI can be flexibly selected according to actual needs during video compression, thereby achieving a processing effect for a specific foreground target and meeting the user's clarity requirements for the specific foreground target.
[0062] In some embodiments, blurring the first background area may be performed by performing the following steps A1 and A2:
[0063] Step A1: determining a first fuzzy parameter value corresponding to the first background area according to the correlation between the data transmission parameter and the fuzzy parameter value; the data transmission parameter includes at least one of the following: data transmission bandwidth and data transmission bit rate.
[0064] Step A2: blurring the first background area based on the first blur parameter value.
[0065] Table 1 below takes the kernel size in the fuzzy parameter as an example to illustrate the effects of several fuzzy parameter value ranges.
[0066] Table 1
[0067]
[0068] In the relationship between data transmission parameters and blur parameter values, the data transmission parameters can correspond to different blur parameter value ranges. The magnitude of the blur parameter value represents the degree of blur applied to the first background area; larger blur parameter values correspond to higher blur levels. The data transmission parameter and the degree of blur are inversely proportional. For example, taking the data transmission bitrate as the data transmission parameter, a higher bitrate corresponds to a lower degree of blur and a smaller corresponding blur parameter value. Assuming that, based on the relationship between the data transmission bitrate and the blur parameter value, the blur parameter value corresponding to the current data transmission bitrate is determined to be 1, then when blurring the first background area, a kernel size between (3,3) and (15,15) can be used. Of course, the blur parameter value range in the above example can be adjusted to meet actual needs, for example, to a finer range of blur parameter values, where the difference between the minimum and maximum kernel values within each blur parameter value range is no greater than 5.
[0069] Optionally, a blur parameter value may be pre-specified, so that when blurring is performed on the first background area, the blurring may be performed according to the pre-specified blur parameter value.
[0070] The purpose of blurring the first background area is to reduce its clarity to accommodate current parameters such as the data transmission bandwidth and data transmission bit rate. For example, the blur parameter value for the first background area is adaptively adjusted based on the current data transmission bit rate. When the data transmission bit rate increases, the blur parameter value decreases, corresponding to a lower degree of blur; when the data transmission bit rate decreases, the blur parameter value increases, corresponding to a higher degree of blur.
[0071] In this embodiment, the first fuzzy parameter value corresponding to the first background area is determined by the correlation between the data transmission parameters and the fuzzy parameter value, and then the first background area is blurred based on the first fuzzy parameter value, so that the blurring of the first background area can match the current data transmission parameters, thereby achieving the effect of adaptively determining the degree of blur according to the data transmission parameters.
[0072] Figure 2 The flowchart of the video processing method provided by the embodiment of the present application is shown. Figure 3 The schematic diagram of the video processing method provided by the embodiment of the present application is shown. Assume that the first video sequence is represented by the following set: Video = {F1, F2, ..., F n}. Among them, F1, F2, ..., F n Respectively represent the first frame, the second frame, ... the nth frame in the first video sequence, and each frame can be used as Figure 1 The first video frame in the illustrated embodiment.
[0073] like Figure 2As shown, the video processing method may include the following steps S201 to S208.
[0074] Step S201 : Initialize the number of times T of processing the first video sequence, so that T=1, and the first video frame to be processed is the first frame in the first video sequence.
[0075] Step S202 : identifying a foreground object in the first video frame, representing a foreground area corresponding to the foreground object as a rectangular frame, and obtaining a foreground area set.
[0076] Optionally, the foreground region set is represented as {R (1) ,R (2) ,...,R (m)}, where R (1) ,R (2) ,...,R (m) They respectively represent the position information of the first foreground target, the second foreground target, ... the mth foreground target in the first video frame.
[0077] Reference Figure 3 As shown, the first video frame includes three foreground targets, namely target 1, target 2 and target 3, so the identified foreground region set can be expressed as {R (1) ,R (2) ,R (3)}.
[0078] Step S203: Filter out the region of interest from the foreground region set.
[0079] The method of screening the region of interest from the foreground region set has been described in detail in the above embodiment and will not be repeated here. Figure 3 As shown, it can be seen that the foreground region corresponding to the foreground object 3 is screened out from the foreground region set as the region of interest.
[0080] Step S204 , performing blur processing on the first background region except the region of interest to obtain a second background region.
[0081] After the first background area is blurred, the clarity of the obtained second background area is lower than that of the first background area.
[0082] Step S205 , performing image fusion on the region of interest and the second background region to obtain a second video frame.
[0083] exist Figure 3 In the figure, the image area is blurred by the diagonal filling method. It can be seen that in the second video frame, the region of interest (i.e., the foreground area corresponding to the foreground object 3) remains clear, while the background area other than the region of interest is blurred.
[0084] Step S206 : Set the number of times the first video sequence is processed, T=T+1, and use the next video frame as the first video frame to be processed.
[0085] After setting T=T+1, the process goes to step S202 to continue processing the current first video frame until F1, F2, ..., F in the first video sequence are n All are processed and the second video sequence {F′1, F′2, ..., F′ n}. Among them, F′1, F′2, ..., F′ n Represent F1, F2, ..., F n The corresponding second video frame.
[0086] Step S207: input the second video sequence into the encoder for encoding to obtain a video code stream.
[0087] Among them, since the clarity of the background area in each video frame of the second video sequence is reduced, the second video sequence has a reduced demand for encoding bit rate. Therefore, when compressing the second video sequence, more attention can be paid to the area of interest with higher clarity, and a higher bit rate can be allocated to the area of interest, while a lower bit rate is allocated to the background area, thereby obtaining a video stream that meets the current data transmission parameters.
[0088] Step S208: Input the video code stream into the decoder for decoding to obtain a compressed target video sequence.
[0089] According to the technical solution of the embodiment of the present application, a region of interest is extracted from a first video frame in a first video sequence; a first background region of the first video frame is image processed to obtain a second background region, such that the clarity of the second background region is lower than that of the first background region, where the first background region is the region of the first video frame other than the region of interest. The region of interest and the second background region are then image-fused to obtain a second video frame; and the second video sequence is compressed to obtain a compressed target video sequence, which includes multiple second video frames. As can be seen, before the video sequence is compressed, the clarity of the background region in each video frame of the video sequence is reduced, thereby reducing the encoding bit rate required for the video sequence. During compression of the video sequence, more attention is paid to the region of interest with higher clarity, thereby allocating a higher bit rate to the region of interest and a lower bit rate to the background region. This ensures that the clarity of the region of interest is maintained in the target video sequence after compression, avoiding the situation where the video sequence is not clear as a whole after efficient compression, thus achieving a balance between the clarity of the foreground object and the degree of compression.
[0090] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a video processing device.
[0091] Figure 4 A block diagram of a video processing device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the video processing device includes:
[0092] An extraction module 41 is configured to extract a region of interest from a first video frame in a first video sequence;
[0093] a processing module 42 configured to perform image processing on a first background region of the first video frame to obtain a second background region; the first background region is a region other than the region of interest in the first video frame; and the clarity of the second background region is lower than that of the first background region;
[0094] a fusion module 43, configured to fuse the region of interest and the second background region to obtain a second video frame;
[0095] The compression module 44 is configured to perform compression processing on the second video sequence to obtain a compressed target video sequence; the second video sequence includes a plurality of second video frames.
[0096] In some embodiments, the extraction module 41 performs the following steps when extracting the region of interest of the first video frame in the first video sequence:
[0097] identifying a foreground object in the first video frame;
[0098] The region of interest is determined according to the foreground area corresponding to the foreground object.
[0099] In some embodiments, when the extraction module 41 determines the region of interest based on the foreground area corresponding to the foreground object, it performs the following steps:
[0100] Screening a foreground area that matches a data transmission parameter from the foreground area as the region of interest; the data transmission parameter includes at least one of the following: data transmission bandwidth and data transmission bit rate;
[0101] The area of the region of interest is positively correlated with the data transmission bandwidth and the data transmission code rate.
[0102] In some embodiments, when the extraction module 41 determines the region of interest based on the foreground area corresponding to the foreground object, it performs the following steps:
[0103] The region of interest is filtered from the foreground area according to preset filtering parameters; the preset filtering parameters include at least one of the following: the target category of the foreground target corresponding to the foreground area, the motion information of the foreground target, the number of regions of the foreground area, the area of the region, and the position of the region.
[0104] In some embodiments, the preset screening parameters include motion information of the foreground target; the motion information includes a motion trajectory;
[0105] When the extraction module 41 selects the region of interest from the foreground region according to the preset screening parameters, the extraction module 41 performs the following steps:
[0106] According to the motion trajectory of the foreground target, a foreground area that meets preset trajectory conditions is screened out as the region of interest; the preset trajectory conditions include at least one of the following: the trajectory direction matches the preset direction, the trajectory length reaches a preset length threshold, and the trajectory length is the longest; the motion trajectory is determined based on the position information of the foreground target in multiple first video frames.
[0107] In some embodiments, when the processing module 42 performs image processing on the first background area of the first video frame to obtain the second background area, the processing module 42 executes the following steps:
[0108] The first background area is blurred to obtain the second background area.
[0109] In some embodiments, the processing module 42 performs the following steps when blurring the first background area:
[0110] Determining a first fuzzy parameter value corresponding to the first background area according to a correlation between a data transmission parameter and a fuzzy parameter value; the data transmission parameter includes at least one of the following: a data transmission bandwidth and a data transmission bit rate;
[0111] The first background area is blurred based on the first blur parameter value.
[0112] According to an embodiment of the present application, a device extracts a region of interest from a first video frame in a first video sequence; performs image processing on a first background region of the first video frame to obtain a second background region, such that the second background region has lower clarity than the first background region, where the first background region is the region of the first video frame other than the region of interest; then performs image fusion on the region of interest and the second background region to obtain a second video frame; and compresses the second video sequence to obtain a compressed target video sequence, where the second video sequence includes multiple second video frames. As can be seen, before the video sequence is compressed, the clarity of the background region in each video frame of the video sequence is reduced, thereby reducing the encoding bitrate required for the video sequence. During compression of the video sequence, more attention is paid to the region of interest with higher clarity, thereby allocating a higher bitrate to the region of interest and a lower bitrate to the background region. This ensures that the clarity of the region of interest is maintained in the target video sequence after compression, avoiding the situation where the video sequence is not clear as a whole after efficient compression, thereby achieving a balance between the clarity of foreground objects and the degree of compression.
[0113] The functions of each module in each device in the embodiment of the present application can be referred to the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.
[0114] Figure 5 A block diagram of an electronic device for implementing the embodiments of the present application. Figure 5 As shown, the electronic device includes a memory 501 and a processor 502. The memory 501 stores a computer program that can be executed on the processor 502. When the processor 502 executes the computer program, the method of the above embodiment is implemented. The number of memory 501 and processor 502 can be one or more. In a specific implementation, the electronic device may also include a communication interface 503 for communicating with external devices and performing data exchange.
[0115] In a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are implemented independently, the memory 501, the processor 502, and the communication interface 503 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0116] Optionally, in a specific implementation, if the memory 501 , the processor 502 , and the communication interface 503 are integrated on a chip, the memory 501 , the processor 502 , and the communication interface 503 may communicate with each other through an internal interface.
[0117] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which implements the method provided in the embodiment of the present application when the program is executed by a processor.
[0118] An embodiment of the present application provides a computer program product, including a computer program, which implements the method provided in the embodiment of the present application when executed by a processor.
[0119] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.
[0120] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.
[0121] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.
[0122] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM) and direct memory bus random access memory (DR RAM).
[0123] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0124] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0125] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0126] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes other implementations in which the functions may be performed in a different order than shown or discussed, including performing the functions substantially simultaneously or in reverse order depending on the functions involved.
[0127] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with such instruction execution systems, apparatuses or devices.
[0128] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0129] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.
[0130] The above is merely an exemplary embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope described in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A video processing method, characterized in that: include: extracting a region of interest of a first video frame in a first video sequence; Performing image processing on a first background region of the first video frame to obtain a second background region; the first background region is a region other than the region of interest in the first video frame; and the clarity of the second background region is lower than that of the first background region; Performing image fusion on the region of interest and the second background region to obtain a second video frame; compressing the second video sequence to obtain a compressed target video sequence; The second video sequence includes a plurality of second video frames.
2. The method according to claim 1, characterized in that The extracting the region of interest of the first video frame in the first video sequence includes: identifying a foreground object in the first video frame; The region of interest is determined according to the foreground area corresponding to the foreground object.
3. The method according to claim 2, characterized in that The determining the region of interest according to the foreground area corresponding to the foreground object includes: Screening a foreground area that matches a data transmission parameter from the foreground area as the region of interest; the data transmission parameter includes at least one of the following: data transmission bandwidth and data transmission bit rate; The area of the region of interest is positively correlated with the data transmission bandwidth and the data transmission code rate.
4. The method according to claim 2, characterized in that The determining the region of interest according to the foreground area corresponding to the foreground object includes: The region of interest is filtered from the foreground area according to preset filtering parameters; the preset filtering parameters include at least one of the following: the target category of the foreground target corresponding to the foreground area, the motion information of the foreground target, the number of regions of the foreground area, the area of the region, and the position of the region.
5. The method according to claim 4, characterized in that The preset screening parameters include motion information of the foreground target; the motion information includes a motion trajectory; The step of screening the region of interest from the foreground region according to the preset screening parameters includes: According to the motion trajectory of the foreground target, a foreground area meeting a preset trajectory condition is screened out as the region of interest; The preset trajectory condition includes at least one of the following: the trajectory direction matches the preset direction, the trajectory length reaches a preset length threshold, and the trajectory length is the longest; the motion trajectory is determined based on the position information of the foreground target in multiple first video frames.
6. The method according to claim 1, characterized in that The performing image processing on the first background area of the first video frame to obtain the second background area includes: The first background area is blurred to obtain the second background area.
7. The method according to claim 6, characterized in that The blurring of the first background area includes: Determining a first fuzzy parameter value corresponding to the first background area according to a correlation between a data transmission parameter and a fuzzy parameter value; the data transmission parameter includes at least one of the following: a data transmission bandwidth and a data transmission bit rate; The first background area is blurred based on the first blur parameter value.
8. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer program product, characterized in that A computer program is included which, when executed by a processor, implements the method according to one of claims 1 to 7.