Fine-grained scene knowledge base construction and action range decision method

Through the fine-grained scene knowledge base construction method, smaller video sub-segments and multiple knowledge images are selected as references, the problem that the existing technology cannot effectively remove cross-scenes and long-term redundancy in videos is solved, and the efficiency and effect of video encoding is significantly improved.

CN120111215APending Publication Date: 2025-06-06ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311648901.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing scene knowledge base methods cannot effectively remove cross-scene and long-term redundancy in videos, especially when the scene changes slowly.

Method used

Through a fine-grained scene knowledge base construction method, the random access segment is divided into smaller sub-segments, and knowledge image selection and scope decisions are made based on the sub-segments, allowing a single random access segment to select multiple knowledge images as references for different regions.

Benefits of technology

It significantly improves the effectiveness of the scene knowledge base algorithm in slow-changing videos, can effectively eliminate long-term redundancy, and maintains excellent performance when the I-frame interval of video is large.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111215A_ABST
    Figure CN120111215A_ABST
Patent Text Reader

Abstract

The invention discloses a fine-grained scene knowledge base construction and action range decision-making method, which comprises the following steps of: selecting candidate knowledge images, namely selecting a plurality of candidate knowledge images from a single random access segment; a single random access segment is divided into a plurality of sub-segments, each sub-segment independently refers to any frame of reference knowledge image in the knowledge image set, and earnings and cost are calculated based on the referenced knowledge image and sub-segment content, determining an optimal knowledge image set from the candidate knowledge image sets based on a calculation result to construct a scene knowledge base; and knowledge image action range decision: distributing an optimal long-term reference frame for each sub-segment. According to the method provided by the invention, knowledge image selection and action range decision are carried out by taking the sub-segment as an income cost calculation unit, so that a single random access segment is allowed to select a plurality of knowledge images as references of different regions, and the action effect of a scene knowledge base algorithm on a scene slow change video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image or video compression technology, and more specifically to a fine-grained scene knowledge base construction and scope decision method. Background Art

[0002] In recent years, with the continuous development of online conferencing, live broadcasting, on-demand and other technologies, the proportion of video traffic in network traffic has also continued to increase. In order to reduce the storage and transmission costs of more and more videos without affecting the viewer's viewing experience, researchers have continuously improved the existing video encoding algorithms, from H.264, H.265 to H.266, through better removal of temporal redundancy and spatial redundancy, to achieve clearer video content with a smaller bit rate.

[0003] In the process of removing temporal redundancy in videos, the selection of reference frames has an important impact on the effect of removing temporal redundancy, so researchers are constantly studying better reference frame selection strategies. Based on the scope of application, reference frames in video encoding can be divided into short-term reference frames and long-term reference frames. Short-term reference frames are only used as references for the remaining frames in a small time period before and after their playback timestamps, which are used to reduce short-term temporal redundancy in videos. Long-term reference frames are not restricted by timestamps to a certain extent, and provide references for frames that are far apart in the video, which are used to reduce long-term temporal redundancy.

[0004] Early long-term reference frame selection algorithms usually adopt an interval-based selection strategy, where a long-term reference frame is selected every several frames as a reference for subsequent video sequences, and only one long-term reference frame exists at a time. This strategy can effectively reduce long-term redundancy within the same scene, but it cannot remove long-term redundancy across scenes, such as when multiple scenes appear alternately and repeatedly.

[0005] The video coding based on scene knowledge base proposed by the researchers has the characteristic of allowing caching of multiple long-term reference frames. When the same scene appears again after a period of time, the long-term reference frames of the corresponding scene cached in the scene knowledge base can be used as a reference, thereby avoiding repeated encoding of the scene content.

[0006] Existing scene knowledge base methods, such as the crowdsourcing method proposed in the Chinese patent application "A Crowdsourcing Video Coding Method and Device", mainly focus on removing the long-term correlation of I frames across scenes and random access segments in the video.

[0007] I frames in videos are mainly generated by scene switching and random access requirements. When a scene switch occurs, or the number of consecutive non-I frame encoding frames in a video exceeds a certain value, the encoder will insert an I frame in the video as a random access point. The frames from the random access point to the next random access point are called random access segments. Existing scene knowledge base methods mainly select random access points as candidates for knowledge images in the scene knowledge base, calculate the net benefit based on the benefits of the candidate knowledge images to the random access segment and the cost of their own encoding, and select the optimal knowledge image set based on the net benefit through crowdsourcing algorithms or KMeans clustering algorithms. This method can effectively select a knowledge image set that is approximately optimal for the I frame set composed of random access points.

[0008] However, the existing scene knowledge base method also has its shortcomings, that is, it ignores the reference of non-I frames in the video to the knowledge image. In fact, although there are adjacent frames as reference objects, there are also a considerable number of blocks in non-I frames that can only be encoded intra-frame. For example, for a video in which a lens moves slowly in one direction, the blocks at the edge of the lens movement direction in the picture will always fail to find a reference object, so only intra-frame encoding can be performed. If the lens moves a complete picture width, then these intra-frame encoded blocks will actually form a hidden I frame. If the lens moves back and forth, then I frames with alternating and repeated content will be generated, and long-term redundancy will be generated between these I frames. However, since the movement of the lens is a slowly changing process, the scene switching algorithm usually does not determine that there is a scene switch in this process. Therefore, when the set random access segment length is large, this repeated scene change will be separated into a single random access segment and will not be considered by the current scene knowledge base algorithm. In fact, the repeated gradual changes of this scene are common. The slow translation, rotation, zoom, and back-and-forth movement of the camera, the gradual change of light and dark, the gradual change of the position of the light source, the slow movement, deformation, replacement, and quantity change of the main objects in the scene, etc., will cause significant changes in the video content in a single random access segment, and will not trigger the scene switching detection to lead to the division of new random access segments. If these changes have a certain periodicity, then they will bring long-term redundancy that cannot be removed by the existing scene knowledge base method. Based on this, it is necessary to further improve the effect of scene knowledge base methods on removing long-term redundancy. Summary of the invention

[0009] In order to overcome the shortcomings of the above technologies, the present invention provides a fine-grained scene knowledge base construction and scope decision method, which further divides the random access segment into smaller sub-segments, and uses the sub-segments as benefit-cost calculation units to perform knowledge image selection and scope decision, thereby allowing a single random access segment to select multiple knowledge images as references for its different areas, thereby improving the effect of the scene knowledge base algorithm on videos with slowly changing scenes.

[0010] The technical solution adopted by the present invention to overcome its technical problems is: a fine-grained scene knowledge base construction and scope decision method proposed by the present invention, including: candidate knowledge image selection, selecting multiple candidate knowledge images from a single random access segment; scene knowledge base construction: dividing the single random access segment into a number of sub-segments, each sub-segment independently refers to any frame of reference knowledge image in a knowledge image set, calculating benefits and costs based on the reference knowledge image and the sub-segment content, and selecting the optimal knowledge image set from the candidate knowledge image set based on the calculation results to construct the scene knowledge base, and the reference knowledge image is not limited to the candidate knowledge image; knowledge image scope decision: assigning an optimal long-term reference frame to each sub-segment, and the optimal long-term reference frame is not limited to selecting a knowledge image in the optimal knowledge image set.

[0011] Furthermore, the candidate knowledge image selection includes at least: random access segment segmentation, dividing the random access segment into a number of segmented sub-segments according to preset rules, and candidate knowledge image extraction, extracting frames from the segmented sub-segments as candidate knowledge images according to preset rules.

[0012] Furthermore, the calculation of benefits and costs based on the reference knowledge image and sub-segment content includes at least: segmentation of sub-segments, segmenting the random access segment into several sub-segments according to preset rules; similarity calculation, calculating the similarity of each candidate knowledge image corresponding to each sub-segment; optimal knowledge image selection: calculating the benefits and costs of the candidate knowledge image set based on the similarity, and selecting the optimal knowledge image based on the benefits and costs.

[0013] Furthermore, the propagation information calculated based on the improved cutree algorithm is used to calculate the similarity of each candidate knowledge image corresponding to each sub-segment.

[0014] Furthermore, the improved cutree algorithm adds a long-term reference frame hypothesis to the cutree algorithm of X256, and calculates the propagation information flowing to the long-term reference frame based on the prediction loss of the existing short-term reference frame hypothesis and the added long-term reference frame hypothesis, thereby calculating the reference value of a given long-term reference frame.

[0015] Furthermore, the improved cutree algorithm adds a reference relationship of a long-term reference frame to the cutree algorithm of X256, and calculates the propagation information flowing to the long-term reference frame based on the existing short-term reference frame hypothesis and the prediction loss of the added long-term reference frame hypothesis, specifically including: making a reference relationship hypothesis between the long-term reference frame group and the short-term reference frame group for each frame in the sub-segment; inter-frame prediction of the long-term reference frame group and inter-frame prediction of the short-term reference frame group; calculating the propagation information of the reference block based on the intra-frame prediction result and the inter-frame prediction result and each block of each frame traversed in the sub-segment; and taking the average value after normalizing the propagation information calculation of all sub-blocks with the intra-frame prediction loss to obtain the similarity of the candidate knowledge image for each video sub-segment.

[0016] Furthermore, the calculation of propagation information of the reference block based on the intra-frame prediction result and the inter-frame prediction result and traversing each block of each frame of the first video sub-segment specifically includes: calculating the total amount of long-term reference frame group information based on the intra-frame prediction loss and the inter-frame prediction loss of the long-term reference frame group; calculating the total amount of short-term reference frame group information based on the intra-frame prediction loss and the inter-frame prediction loss of the short-term reference frame group; calculating the reference probability of each long-term reference frame group and short-term reference frame group. Correcting the total amount of information based on the reference probability; accumulating the corrected total amount of information and passing it to the reference frame as propagation information.

[0017] Furthermore, the knowledge image scope decision allocates the optimal long-term reference frame to each sub-segment, specifically including: dividing a single random access segment into several video sub-segments, traversing the video sub-segments, calculating the net benefits of several long-term reference frame selection modes for each video sub-segment, and selecting the optimal long-term reference frame mode for each video sub-segment based on the net benefit comparison.

[0018] Furthermore, the long-term reference frame selection mode includes at least one of a first mode, a second mode, a third mode and a fourth mode. In the first mode, the long-term reference frame of the previous video sub-segment is used as the long-term reference frame of the current video sub-segment; in the second mode, the long-term reference frame of the first frame of the video sub-segment is used as the long-term reference frame of the first frame of the video sub-segment. n long-term reference frame; in the third mode, the knowledge image with the greatest similarity to the video sub-segment is used as the long-term reference frame of the video sub-segment; in the fourth mode, no long-term reference frame is used.

[0019] Furthermore, the optimal long-term reference frame mode for each video sub-segment is selected based on the net gain comparison. If the net gain value of the fourth mode is none >MAX(gain last , gain newlt , gain newlp ), then select the fourth mode; otherwise if Then select the first mode; otherwise if Then select the third mode; otherwise select the second mode; wherein the first threshold TH_REL_USE_NEW, the second threshold TH_REL_USE_LP, gain last , gain newlt , gain newlp , gain none They are the net profit values ​​of the first to fourth modes respectively.

[0020] The beneficial effects of the present invention are:

[0021] 1. The knowledge image selection and scope decision algorithm proposed in this method can significantly improve the effectiveness of the scene knowledge base algorithm in video sequences with slowly changing scenes. At the same time, the performance of the scene knowledge base algorithm is no longer significantly related to the I-frame interval affected by the random access segment or scene switching interval. Even when the video I-frame interval is large, the scene knowledge base algorithm using this method can still maintain excellent performance, and the gain will not decrease significantly compared to a smaller I-frame interval.

[0022] 2. The similarity calculation algorithm based on the improved cutree proposed in this method has good scalability. It can not only be used in the selection of knowledge images and the scope of action decision, but also in the pre-processing links such as time domain filtering. Therefore, this method can become the pre-basis for more subsequent optimizations;

[0023] 3. The scene knowledge base algorithm optimized based on this method can adapt to more complex scene changes and random access partitioning methods, and has a wider range of applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A schematic diagram of a process of constructing a fine-grained scenario knowledge base and determining a scope of action in an embodiment of the present invention;

[0025] Figure 2 This is an example of image content changes caused by video camera movement or slow scene changes;

[0026] Figure 3 It is a flow chart of using propagation information calculated based on the improved cutree algorithm as similarity in an embodiment of the present invention;

[0027] Figure 4 A schematic diagram of a process of allocating reference frames to long-term reference frames and short-term reference frame groups based on an improved cutree algorithm according to an embodiment of the present invention;

[0028] Figure 5 A schematic diagram of a process for calculating propagation information of each sub-block of a candidate knowledge image based on an improved cutree algorithm according to an embodiment of the present invention;

[0029] Figure 6A schematic diagram of an algorithm flow for constructing a scene knowledge base according to an embodiment of the present invention;

[0030] Figure 7 A schematic diagram of a process for determining the scope of action of a knowledge image library according to an embodiment of the present invention;

[0031] Figure 8 The figure is a flow chart of selecting the optimal long-term reference frame mode for each video sub-segment based on net benefit comparison according to an embodiment of the present invention. DETAILED DESCRIPTION

[0032] First, some abbreviations and key terms mentioned in the present invention are explained.

[0033] Knowledge image: Knowledge images are some special reference frames selected in the reference frame selection step of video encoding. They are special long-term reference frames used in the scene knowledge base algorithm. The difference between them and ordinary long-term reference frames is that multiple frames can exist at the same time and can exist across random access segments to achieve mutual reference of video content over a large time span.

[0034] Scene knowledge base: The collection of selected knowledge images is called scene knowledge base.

[0035] Gain: It is divided into subjective gain and objective gain. Subjective gain is measured by the subjective evaluation of the subject on the video quality, and objective gain is measured by the objective indicators psnr bdrate or ssim bdrate, which can be understood as the average percentage decrease in the video bit rate of the video encoded based on the optimized algorithm under the same PSNR or SSIM indicators compared with the video encoded based on the algorithm before optimization, that is, the percentage of video volume reduction when the quality remains unchanged;

[0036] SSIM: Structural Similarity, a measure of the similarity between two images.

[0037] PSNR: Peak Signal-to-Noise Ratio, peak signal-to-noise ratio, is one of the indicators to measure image quality.

[0038] Bd-rate: Bjontegaard delta rate, a gain measurement method, which fits the RD curve through the data of multiple quality points, and calculates the percentage of bit rate reduction under the average distortion index of the encoder RD curve before and after optimization based on the integral method, which is used as the gain indicator.

[0039] X265: An open source video encoder based on the H265 standard.

[0040] cutree: The time-domain adaptive quantization algorithm in x265, its main function is to calculate the propagation information of the reference frame based on the time-domain distortion transfer algorithm, which is used as the basis for the adaptive quantization adjustment of the reference frame. Since the propagation information represents the impact of the distortion of the reference frame on the distortion of subsequent frames, it can also be used as a reference value for other optimization decisions.

[0041] POC: video playback order index, for example, the first frame poc is 0, the second frame poc is 1, and so on

[0042] miniGOP: mini group of pictures. MiniGOP refers to the frames from the next B frame after the P frame to the next P frame. For example, if the video frame structure is BBPBBBP, then the BBP and BBBP in it are each a miniGOP.

[0043] SATD: A measure of the size of the video residual signal. The sum of the absolute values ​​of the video residual after the Hadamard transform can be used as a prediction of the video residual coding bit rate.

[0044] In order to facilitate those skilled in the art to better understand the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The following is only exemplary and does not limit the protection scope of the present invention.

[0045] like Figure 1 As shown, the present embodiment describes a method for constructing a fine-grained scene knowledge base and determining a scope of action. The following describes a method for constructing a fine-grained scene knowledge base and determining a scope of action according to a specific embodiment.

[0046] S1, candidate knowledge image selection, selects multiple candidate knowledge images from a single random access segment.

[0047] The random access end is obtained from the video to be encoded. In order to further improve the effect of removing long-term redundancy by the scene knowledge base method, the video to be encoded is divided into sub-segments smaller than the random access segment. Specifically, the steps are as follows:

[0048] S11, random access segment segmentation, dividing the random access segment into a plurality of segmented sub-segments according to a preset rule.

[0049] The segmented sub-segments may be segmented from the random access segment at fixed intervals, or may be adaptively segmented based on motion conditions, such as starting a new sub-segment when the change in the content within the segmented sub-segment is greater than a given value.

[0050] It should be noted that the length of random access segments varies depending on the application, usually from a few seconds to tens of seconds. For example, if a random access segment is 5 seconds and the video frame rate is 30fps, then the random access segment interval is 150 frames. Random access segment segmentation is performed after random access segment division is completed.

[0051] S12, candidate knowledge image extraction, extracting frames from the segmented sub-segments as candidate knowledge images according to preset rules.

[0052] The candidate knowledge image is selected by selecting a specific frame of each segmented sub-segment, such as the first frame or the middle frame.

[0053] In another embodiment, on this basis, a knowledge image outside the video to be encoded can be obtained based on some prior knowledge as a supplement. For example, when the usage scenario of the current video input device to be encoded is relatively fixed, the scene knowledge base of the previous encoding or previous multiple encodings can be used as a candidate supplement.

[0054] For example, if the camera position is fixed and the room background does not change, the knowledge image of the previous meeting can be used as the candidate knowledge image of the current meeting. It should be noted that the candidate knowledge image is not limited to being selected from the video to be encoded, but can also be selected from videos of similar scenes based on prior knowledge.

[0055] In another embodiment, when the background remains unchanged for a long time, a knowledge image can be synthesized as a candidate supplement through a background modeling method.

[0056] In yet another embodiment, a hybrid method of multiple modes may be adopted, such as adding equally spaced extracted frames and synthetic background frames to the candidate image set at the same time, so that the image set contains as much information as possible.

[0057] It should be noted that the advantage of this candidate knowledge image extraction method is that for videos with slowly changing scenes rather than suddenly changing scenes, the original scene knowledge base technical solution will only select one knowledge image candidate, because at this time there is only a single I frame in this scene (usually the encoding algorithm will detect the scene switch and insert an I frame, but when the scene changes slowly, the scene switch cannot be detected). However, using the method of this application, multiple candidates can be selected to cover all scenes. Figure 2 In the BasketballDrive video shown, there are at least three different background switches in the random access segment ranging from 0 to 150 frames, namely the center-right, right side, and center-left of the basketball court. Originally, the candidate knowledge image selection could only select one frame from 0 to 150, so there must be two scenes that were not covered. Now, one candidate is selected every 16 frames, so all scenes can be covered.

[0058] S2, scene knowledge base construction: a single random access segment is divided into several sub-segments, each of which independently refers to any frame of reference knowledge image in the knowledge image set, and benefits and costs are calculated based on the reference knowledge image and the sub-segment content. Based on the calculation results, the scene knowledge base is constructed from the candidate knowledge images and the optimal knowledge image set is selected, wherein the reference knowledge images are not limited to the candidate knowledge images, and therefore the knowledge image set is not limited to the candidate knowledge image set composed of the candidate knowledge images.

[0059] S21, segmentation of sub-segments, segmenting the random access segment into a number of sub-segments according to a preset rule.

[0060] The random access segment of the video to be encoded is usually divided into sub-segments in a fixed interval manner. Each sub-segment can independently refer to any frame of reference knowledge image in the knowledge image set, and the knowledge image set here is not limited to the candidate knowledge image set composed of candidate knowledge images.

[0061] The sub-segment segmentation can be performed at fixed intervals or adaptively based on the motion situation. For example, when the change in the content within a sub-segment is greater than a given value, a new sub-segment is started.

[0062] S22, similarity calculation, calculating the similarity of each candidate knowledge image corresponding to each sub-segment;

[0063] First, the similarity between each candidate knowledge image in the candidate image set and the divided video sub-segment is calculated. The similarity calculation result will be used as the measure of the subsequent knowledge image set benefits and costs. The similarity calculation can be done in a variety of ways, including but not limited to image similarity calculation, intra-frame and inter-frame prediction result calculation, and deep learning feature calculation.

[0064] In some embodiments, the histogram information entropy is used based on the image similarity calculation. The similarity is calculated based on the histogram information entropy of the video to be encoded, and the similarity between the candidate knowledge image and the video sub-segment is the cumulative difference between the histogram self-information entropy of each frame in the video sub-segment and the conditional entropy of the candidate knowledge image.

[0065] In some embodiments, the similarity is calculated based on the intra-frame and inter-frame prediction results using a modified X265 program cutree algorithm. The difference between the intra-frame and inter-frame prediction losses is used as the amount of information brought by the reference block to the current block, and the propagation information accumulation value of the reference blocks on all reference frames is calculated using the information propagation in the reverse direction of the reference chain. This accumulation value can represent the reference value of the reference block to all coding blocks that directly or indirectly reference it, and the average propagation information of all sub-blocks of the candidate knowledge image is used as the similarity between the candidate frame and the sub-segment involved in the calculation.

[0066] In some other embodiments, a deep learning model is used to extract features from candidate frames and sub-segments respectively, and then the distance between the features is calculated as the similarity. For sub-segments, the image feature extraction method can be used to extract the features of each frame respectively, and then compared with the features of the candidate frame to calculate the average value. Alternatively, 3D convolution can be used to output the entire video sub-segment as a feature, and its similarity result with the candidate frame is calculated through a specially trained comparison network.

[0067] In one embodiment of the present invention, an improved cutree algorithm is used to calculate similarity. This algorithm can conveniently calculate the reference value of a knowledge image for a certain sub-segment as the degree of similarity, and has high scalability. The advantages of using the cutree algorithm over other similarity calculation algorithms are as follows.

[0068] (1) Indirect reference can be achieved. For example, frame A refers to frame B, and frame B refers to frame C. Although frame A does not directly refer to frame C, frame A will have propagation information transmitted to frame C through the reference chain.

[0069] (2) Supports flexible reference structures. The calculation framework of cutree does not have any rigid restrictions on the frame type and reference structure of video frames. For example, the video frame type and reference structure of IPPPP or IBBBPBBBP can be used to calculate the reference value of cutree without remodeling each structure. It is easier to add new reference structures such as the reference structure of long-term reference frames, so the calculation framework of cutree is selected. It is more versatile and scalable and can be used in encoders in the industry.

[0070] (3) The cutree algorithm can obtain block-level reference values ​​and subsequently support some block-level optimization decisions. In the future, other optimizations can also be achieved with the help of the cutree similarity calculation system.

[0071] Other similarity calculation algorithms, such as the knowledge image benefit calculation algorithm in crowdsourcing, take the coding loss of the random access segment without referring to the knowledge image, minus the coding loss of the random access segment referring to the knowledge image, and perform benefit calculation. They cannot achieve the three advantages (1)-(3) mentioned above.

[0072] In this implementation example, the cutree algorithm is improved to some extent, and the improvements include the following two points.

[0073] (1) Based on the original x265 cutree algorithm, long-term reference relationships are considered when performing inter-frame prediction. Since the knowledge image is used as a long-term reference frame, cutree will select reference frames for inter-frame prediction based on both short-term reference relationship assumptions and long-term reference relationship assumptions during calculation;

[0074] (2) Considering the long-term reference relationship when propagating information, the total amount of propagated information under the two reference relationship assumptions is calculated respectively, the total amount of information is corrected based on the reference probability and back propagated to obtain the propagated information of the knowledge image as a long-term reference frame

[0075] The flow chart of the propagation information calculated based on the improved cutree algorithm as the similarity is as follows Figure 3 As shown, the specific steps are as follows: Assume that the candidate knowledge image whose similarity needs to be calculated is LP.

[0076] S221, dividing the video to be encoded into a plurality of video sub-segments based on a preset segmentation interval by using the random access segment.

[0077] In one embodiment of the present invention, the preset segmentation interval is K, and in one implementation, K = 10. The random access segment is divided into a number of video segments {s1, s2, ..., sN} at intervals of K, wherein segment sn contains frames with poc greater than or equal to (n-1)×K and less than n×K.

[0078] S222, traverse each video sub-segment and each candidate knowledge image, calculate the propagation information of each sub-block of the candidate knowledge image based on the improved cutree algorithm, and take the average value of the normalized propagation information of all blocks of the candidate knowledge image as the similarity.

[0079] S2221, downsampling the first candidate knowledge image, dividing the candidate knowledge image into sub-blocks, and calculating the satd loss value of the candidate knowledge image by performing intra-frame prediction based on the divided sub-blocks;

[0080] In one embodiment of the present invention, Figure 4 Taking the flowchart shown in FIG. 1 as an example, the propagation information of each sub-block of the candidate knowledge image calculated based on the improved cutree algorithm is explained.

[0081] Downsample the candidate knowledge image LP by 2 times and divide it into sub-blocks of 16×16 size. For the downsampled version, divide it into sub-blocks of 8×8 size. Traverse all sub-blocks for intra-frame prediction. Intra-frame prediction is performed on the downsampled frame. After the prediction is completed, the intra-frame prediction satd loss needs to be saved to the intracost array of the frame, where intracost[culdx] stores the intra-frame prediction satd loss of the culdx-th 16×16 block. The specific process of intra-frame prediction for each block is as follows.

[0082] (1) Calculation of DC mode satd loss

[0083] (2) Calculate the Planar mode satd loss and compare it with the DC mode loss to obtain the optimal non-angle mode.

[0084] (3) 35 angle modes are sampled by selecting one out of every five modes for testing, and the optimal angle mode is roughly selected through the satd loss.

[0085] (4) Take the smaller of the optimal angle mode satd and the optimal non-angle mode satd to obtain the intra-frame prediction satd value of the block.

[0086] Repeat the above process (1)-(4) until the intra-frame prediction of sub-blocks of all candidate knowledge images in all candidate image sets is completed.

[0087] S2222, downsampling each frame of the first video sub-segment, dividing the frame into sub-blocks, and calculating the satd loss value of the video sub-segment by intra-frame prediction based on the divided frame sub-blocks, and deciding the frame type of each frame of the video sub-segment.

[0088] (1) Take the middle frame of the video sub-segment and the candidate knowledge image to perform image-level correlation calculation. The correlation calculation can use histogram entropy or inter-frame prediction. If the correlation is lower than the threshold T, it means that the video sub-segment is completely irrelevant to the content of the candidate knowledge image. The subsequent steps are skipped directly and the correlation calculation of the next video sub-segment is performed. When the correlation is greater than or equal to T, the subsequent steps are executed.

[0089] (2) Each frame in the video sub-segment is downsampled by 2 times and divided into 16X16 frame sub-blocks, and a coarse intra-frame prediction is performed. The intra-frame prediction method is the same as the sub-block calculation method of the candidate knowledge image in steps (1)-(4).

[0090] (3) Assign a suitable frame type to each frame. The frame type can be fixed or determined using the X256 frame type decision algorithm, thereby determining the frame type of all frames in the video subsegment.

[0091] For example, three B frames and one P frame are allocated for every 4 frames to form a BBBPBBBP structure. When there are multiple consecutive B frames at the same time, the middle B frame is used as a Bref frame, which can be used as a reference frame, while the remaining B frames cannot be used as reference frames.

[0092] S2223 , allocating a reference frame to each frame of the video subsegment based on the long-term reference frame allocation scheme and the short-term reference frame group reference frame allocation scheme, and based on the frame type decision result.

[0093] (1) Preset the reference frame for each frame in the video sub-segment.

[0094] Two groups of reference relationship assumptions are preset. The first group is the short-term reference frame group, which assumes that only short-term reference relationships exist. This is the original reference relationship assumption of the cutree algorithm. For P frames, it refers to the previous P frame, and for B frames, it refers to the two most recent reference frames. The second group is the long-term reference frame group, which assumes that only long-term reference relationships exist. This is a new modification. For P frames and B frames, it is assumed that they can only refer to the most recent candidate knowledge images.

[0095] (2) performing inter-frame prediction of a long-term reference frame group and a short-term reference frame group;

[0096] Based on the assumed relationship between the short-term reference frame group and the long-term reference frame group, inter-frame prediction is performed on each frame of the video sub-segment for each reference relationship, including inter-frame prediction of the long-term reference frame group and the short-term reference frame group.

[0097] When performing inter-frame prediction, each block of the frame to be predicted in the current video subsegment is traversed, and the inter-frame prediction loss intercost using the short-term reference frame group reference frame is calculated separately, that is, it does not refer to the long-term reference frame group, but only refers to the short-term reference frame. ST , and the inter-frame pre-satd loss intercost using only a short-term reference frame group (i.e., it only refers to the knowledge image and does not refer to other frames) LP And the corresponding motion vectors MVST and MVLP.

[0098] Inter-frame prediction uses a relatively simple HEX template to perform on the downsampled version of the frame to be predicted and its reference frame. The inter-frame prediction results are saved to the intercost array, where intercost ST [cuIdx] represents the prediction loss of the block with the serial number culdx using the reference frame of the short-term reference frame group, intercost LP [cuIdx] represents the prediction loss of the reference frame of the knowledge image group used by the block number culdx.

[0099] S2224, calculating propagation information of the reference block based on the intra-frame prediction result and the inter-frame prediction result and traversing each block of each frame of the video sub-segment.

[0100] The propagation information is calculated from the back to the front in the reverse reference chain. For example, for the BBPBBBP structure, the propagation information of the last mini-image group, i.e., BBBP, is calculated first, and then the propagation information of the previous mini-image group, i.e., BBP, is calculated. For a single mini-image group miniGOP, the propagation information of the B frame that is not a reference frame to its reference frame (BREF and the nearest P frame) is calculated first, and then the propagation information of the BREF frame that is a reference frame to its reference frame (the two P frames before and after) is calculated. Finally, the propagation information of the P frame to the reference frame of the previous mini-image group miniGOP (the P frame of the previous mini-image group miniGOP) is calculated. The calculation of the propagation information requires traversing all blocks of the frame to be calculated. For the block to be calculated with the index culdx, the propagation information passed to the reference block is calculated through the following steps.

[0101] (1) Define the propagation information input of this block as p in , p in is the sum of the propagation information transmitted to the block by all blocks that refer to the block. If the block is not used as a reference block of any block (such as a block of a B frame that is not a reference frame), then p in =0.

[0102] (2) Obtain the intra-frame and inter-frame losses of the block.

[0103] a. Intra prediction loss block_intracost = intracost[cuIdx];

[0104] b. Inter-frame prediction loss block_intercost of the reference frame of the long-term reference frame group LP =intercost LP [cuIdx];

[0105] c. Inter-frame prediction loss block_intercost of the reference frame of the short-term reference frame group ST =intercost ST [cuIdx].

[0106] (3) The total amount of information pa propagated from the current block to the reference block is calculated according to the following formula. Two cases need to be calculated: the total amount of information pa only referring to the long-term reference frame LP , and the total amount of information pa referring only to the short-term reference frame ST The flow chart is as follows Figure 5 shown.

[0107] The total amount of information referring only to the long-term reference frame

[0108] The total amount of information referring only to the short-term reference frame

[0109] (4) Based on the reference probability, pa LP and pa ST Correction by pa LP and pa ST The ratio of is normalized by the sigmoid function to obtain the probability of the actual current block referring to the knowledge image. Before that, two different reference frames are used for inter-frame prediction to obtain pa LP and pa ST However, the actual current block can only select one of the two reference frame groups as its actual reference frame. Therefore, it is necessary to calculate the probability of the current block selecting the long-term reference frame group and the short-term reference frame group. The actual propagation information transmitted to the reference block is the result of multiplying the total amount of original information by the reference probability.

[0110] Estimation of the probability of referring to a long-term reference frame group

[0111] Defining the total information ratio

[0112] Corrected pa LP :pa LP =sigmoid(r)×pa LP

[0113] Corrected pa ST :pa ST =(1-sigmoid(r))×pa ST

[0114] (5) Propagate the total amount of information to the reference block of the current block, respectively LP and pa ST To propagate, for the total amount of information of each reference frame group, perform the following steps.

[0115] a. Determine whether the current block uses intra prediction. If pa LP or pa ST If it is less than or equal to 0, it means that the current block selects the intra-frame mode when using this set of reference frame groups for encoding, and the subsequent process of propagating the total amount of information is skipped.

[0116] It should be noted that the propagation ratio propagate_fraction of the cutree algorithm is (intra_cost - intercost) / intracost. Therefore, if intracost < intercost, it means that the intra-frame prediction loss is less than the inter-frame prediction loss. This block is intra-frame encoded and does not refer to the reference frame. The pa passed to the reference frame is 0. Therefore, it can be determined whether this block is intra-frame encoded under the reference relationship assumption of the short-term reference frame group or the knowledge image group through pa. If it is intra-frame encoded, there is no propagation information, so the subsequent propagation process of the total amount of information is skipped.

[0117] b. Determine the reference block of the current block on this group of reference frames based on the MVST or MVLP obtained when the current block performs inter-frame prediction using this group of reference frames.

[0118] c. Perform the propagation of the total amount of information. If the reference block is unique, directly accumulate the total amount of information to the p of the corresponding reference block. in In the case where there are four reference blocks and only part of each block is referenced by the current block, the total amount of information is divided into four parts according to the reference area and accumulated to the p of the four reference blocks respectively. in In the case where the current block uses the bidirectional prediction mode and references the reference blocks on two reference frames at the same time, first divide the total amount of information equally, and then propagate it to the reference blocks of their respective reference frames according to the aforementioned rules respectively.

[0119] S2225. Normalize the propagation information of all sub-blocks in the candidate knowledge image by the intra-frame prediction loss and then take the average value to obtain the similarity of the candidate knowledge image to the current video segment.

[0120] S23. Optimal knowledge image selection: Calculate the benefits and costs of the candidate knowledge image set based on the similarity, and select the optimal knowledge image set based on the benefits and costs.

[0121] Select the knowledge images included in the final scene knowledge base based on the similarity calculation result to obtain the optimal knowledge image set, so that the net benefit of the selected set to the entire video sequence is maximized.

[0122] The net benefit is represented by the benefit and the cost. For each given knowledge image set, the benefit represents the reference value of the set to the entire video, and the cost represents the additional encoding cost of the set. The benefit and the cost are calculated from the candidate similarity calculated in the previous step.

[0123] In one embodiment, given a set of knowledge images, the benefit can be calculated by assigning the best knowledge image to each video sub-segment, that is, the knowledge image with the greatest similarity, and then accumulating the similarity of the best knowledge image to the corresponding sub-segment. The cost is calculated by the degree of content overlap between the knowledge images. Since the newly appearing content in the video requires at least one intra-frame encoding, if the content of the knowledge image only appears once, its encoding cost can be basically ignored. Only the content that appears multiple times in the knowledge image set needs to calculate the cost. Therefore, the encoding cost can be calculated by the degree of content overlap between the knowledge images. For each knowledge image, its overlap is the similarity of the most similar knowledge image to it. The overlap of all knowledge images is accumulated to obtain the overlap of the entire knowledge image set as its encoding cost.

[0124] After determining the benefits and costs of a given knowledge image set, the net benefit of the set can be calculated and the set with the greatest benefit can be selected. On this basis, different optimal knowledge image set search methods can be used. When the number of candidate knowledge images is small, all optional knowledge image combinations can be searched through a full search. When the number of knowledge images is large, some local optimal algorithms can be used to obtain an approximate optimal combination.

[0125] In some implementations, the candidate knowledge images may be clustered in advance by KMeans clustering, and the cluster centers may be selected as knowledge images to obtain knowledge image sets with different numbers of cluster centers, and the optimal set may be selected from these knowledge image sets by net gain.

[0126] In some implementations, knowledge images may be alternately added and deleted through a crowdsourcing algorithm, so that the net benefit of the existing set increases continuously until it reaches a maximum value, at which point the knowledge image set is the final selected set.

[0127] In one embodiment of the present invention, the scene knowledge base is constructed as a whole using a crowdsourcing algorithm. The difference from the traditional crowdsourcing algorithm is that a new method for calculating the benefits and costs of the selection set is used. In order to better obtain the best knowledge image set for the entire random access segment, the candidate similarity is used to calculate the benefits, and the degree of overlap between the content of the knowledge images is used as the cost.

[0128] In one embodiment of the present invention, the method for calculating the knowledge image set benefit is as follows.

[0129] S231, obtaining the similarity of the candidate knowledge image to each video sub-segment.

[0130] The random access segment is divided into several video sub-segments according to a preset rule.

[0131] In some implementations, the similarity between the candidate knowledge image and the divided sub-segments is obtained based on the S22 method.

[0132] S232, traverse the video sub-segments sn (1≤n≤N), and for each video sub-segment, sort the candidate similarities between all candidate knowledge images in the candidate knowledge image set and the corresponding video sub-segment, and set the largest candidate similarity value to similarity_first n The second largest correlation value is similarity_second n , and so on.

[0133] S233, the size of the benefit of the knowledge image set is calculated as follows:

[0134] The calculation of knowledge image benefits is illustrated by the following specific example.

[0135] The video to be encoded selects 3 frames of candidate knowledge images into A frame, B frame and C frame. The video to be encoded has a total of 50 frames divided into 10 video sub-segments, and the video sub-segments are s0-s4 respectively. Video sub-segment s0 refers to frame A, and the similarity is calculated as a1 through the improved cutree algorithm. Video sub-segment s1 refers to frame A, and the similarity is calculated as b1 through the improved cutree algorithm. Video sub-segment s2 refers to frame C, and the similarity is calculated as C1 through the improved cutree algorithm. Video sub-segment s3 refers to frame B, and the similarity is calculated as b2 through the improved cutree algorithm. S4 refers to frame A, and the similarity is calculated as a2 through the improved cutree algorithm. Then the gain calculation gain = a1+b1+c1+b2+a2.

[0136] In one embodiment of the present invention, the method for calculating the cost of the knowledge image set is as follows.

[0137] (1) Treat the knowledge image as a sub-segment of length 1, and calculate the similarity between all knowledge images in the knowledge image set based on the improved cutree algorithm;

[0138] (2) Traverse each candidate knowledge image LP n (1≤n≤M), is the number of candidate knowledge images, sort the similarities between each candidate knowledge image and the other candidate knowledge images, and let the largest similarity value be similarity_first n The second largest correlation value is similarity_second n ;

[0139] (3) Traverse each candidate knowledge image LP n(1≤n≤M), calculate the average intra-frame prediction loss of all blocks of its downsampled version frame_intracost n ;

[0140] (4) The cost of the knowledge image set is calculated as follows:

[0141] Taking ABCDE as an example, we calculate the similarity of AB, AC, AD, and AE based on the improved cutree algorithm, and take the maximum value as similarity_first. n As the content overlap of A, repeat this process for each knowledge image and multiply the obtained overlap by the intra-frame coding cost frame_intracost n Add up to get the cost.

[0142] In one embodiment of the present invention, the calculation method of the net income of the knowledge image set is as follows:

[0143] (1) Define the weight w of the knowledge image cost, w = 0.02 in one implementation;

[0144] (2) The net gain net_gain = gain - w*cost is obtained by subtracting the cost multiplied by the weight from the gain of the knowledge image.

[0145] The algorithm flow chart for constructing the scene knowledge base is as follows: Figure 6 shown.

[0146] (1) Initialize the knowledge image set to an empty set;

[0147] (2) Record the net gain of the current knowledge image set origin ;

[0148] (3) Traversing all candidate knowledge images that have not been added to the knowledge image set, calculating the benefits after adding the candidate knowledge image to the knowledge image set, and determining the candidate knowledge image that can maximize the net benefit of the knowledge image set by adding it;

[0149] (4) adding the candidate knowledge image with the largest benefit to the knowledge image set. If the net benefit of the knowledge image set is improved after the addition, repeat steps (3)-(4); otherwise, execute step (5);

[0150] (5) Traversing all knowledge images in the knowledge image set, calculating the benefits of deleting the knowledge image from the knowledge image set, and determining the knowledge image that can maximize the net benefit of the knowledge image set by deleting it;

[0151] (6) Delete the knowledge image with the largest deletion benefit. If the net benefit of the knowledge image set is improved after deletion, repeat steps (5)-(6). Otherwise, execute step (7).

[0152] (7) Record the net gain of the current knowledge image set update If gain update > gain origin , then jump to step (2), otherwise, it is considered that the optimal knowledge image set has been found.

[0153] In another embodiment, the calculation method of the set cost and net benefit when constructing the scene knowledge base takes into account the influence of time-domain adaptive quantization when selecting knowledge images. Since the quantization intensity will significantly affect the reference effect of the knowledge image, when adding a new knowledge image, the increased cost needs to consider not only the degree of overlap of the knowledge image, but also the reduction in the time-domain adaptive quantization adjustment value of the remaining knowledge images caused by the new knowledge image. This reduction is essentially due to the fact that the new knowledge image has taken away the frames in the video that originally referenced the knowledge image in the set before the addition, resulting in a decrease in the scope of the knowledge image in the original set, and a decrease in the reference value calculated by the time-domain adaptive quantization algorithm. Therefore, in this implementation example, the overlap of the scope of knowledge images is added as an additional cost of the knowledge image set, so as to consider the influence of time-domain adaptive quantization when selecting knowledge images. Among them, the calculation method of the overlap of the scope of action of candidate knowledge images is as follows.

[0154] (1) Obtain the candidate similarity between each candidate knowledge image and each video sub-segment;

[0155] (2) Traverse the video sub-segments sn (1≤n≤N), and for each sub-segment, sort the relevance of all candidate knowledge images in the candidate knowledge image set to the video sub-segment, and set the largest candidate similarity value to similarity_first n The second largest candidate similarity value is similarity_second n , and so on.

[0156] (3) The calculation method of the overlap degree of the scope of the candidate knowledge image set is: The sum of the relevance sizes of the second ranked candidate knowledge images is used as the scope overlap.

[0157] The calculation method of the net benefit of the candidate knowledge image set based on the overlap of the candidate knowledge image ranges is as follows.

[0158] (1) Define the benefit of knowledge image as gain and the cost based on content overlap as cost c , the weight of the cost based on content overlap is w c, based on the overlap of the scope of action, the cost is cost r , the weight of the cost based on the overlap of the scope is w r

[0159] In an embodiment of the present invention, w c =0.01, w r =0.25.

[0160] (2) The net benefit of the knowledge image set is gain_net = gain-cost c w c -cost r w r .

[0161] Take the calculation of the overlap of the scope of knowledge images as an example.

[0162] The video to be encoded selects 3 frames of candidate knowledge images and divides them into A frame, B frame and C frame. The video to be encoded has a total of 50 frames divided into 10 video sub-segments, and the video sub-segments are s0-s4 respectively. The most similar knowledge image of video sub-segment s0 is A, and the second similar knowledge image is B. The similarity of the second similar knowledge image calculated by the improved cutree algorithm is b1. The second similar knowledge image of video sub-segment s1 is C, and the similarity of the second similar knowledge image calculated by the improved cutree algorithm is c1. The second similar knowledge image of video sub-segment s2 is A, and the similarity of the second similar knowledge image calculated by the improved cutree algorithm is a2. The second similar knowledge image of video sub-segment s3 is B, and the similarity of the second similar knowledge image calculated by the improved cutree algorithm is b2. The second similar knowledge image of video sub-segment s4 is C, and the similarity of the second similar knowledge image calculated by the improved cutree algorithm is c2. Then the overlap of the scope of action is b1+c1+a2+b2+c2.

[0163] S3, knowledge image scope decision: assign the optimal long-term reference frame to each sub-segment.

[0164] It should be noted that the optimal long-term reference frame is not limited to selecting knowledge images from the optimal knowledge image set. When the reference value of the knowledge image is not optimal, a video frame of a non-knowledge image can be selected as a long-term reference frame. Since knowledge images are mainly used as references for content patterns that appear repeatedly at intervals, non-repeated content is somewhat neglected when selecting knowledge images. Therefore, knowledge images may not be the optimal reference frames for non-repeated content. Based on this, when deciding on the scope of knowledge images, a mixture of knowledge images and ordinary long-term reference frames can provide references for both repeated and non-repeated content, further improving the gain effect of the scene knowledge base algorithm on the video. The decision on the scope of knowledge images is achieved by comparing the net benefits of different decision-making methods and selecting the best decision based on a set of predefined thresholds. The specific steps are as follows. The flow chart is shown below. Figure 7 shown.

[0165] S31, segmenting the video into sub-segments as in S21.

[0166] S32, traversing the video sub-segments, and calculating the net benefits of several long-term reference frame selection modes for each video sub-segment.

[0167] In one embodiment of the present invention, the four selection modes of the long-term reference frame are as follows. Assume that the current decision is the sth n The long-term reference frame of the segment (1≤n≤N), N is the total number of video segments.

[0168] Select the first mode: Use the s n-1 The long-term reference frame of the segment is used as the long-term reference frame of the current video sub-segment. When n=1, the loss of this selection is set to infinity;

[0169] Select the second mode: Use video subsegments n The first frame at the beginning of , that is, the frame of poc = (n-1) × K, K is the segmentation interval of , as the current video sub-segment s n The long-term reference frame of the segment;

[0170] Select the third mode: use the knowledge image with the greatest similarity to the current video sub-segment as the long-term reference frame of the current video sub-segment;

[0171] Select the fourth mode: do not use long-term reference frames.

[0172] Based on the above four long-term reference frame selection modes, the benefits of the long-term reference frame selection mode are calculated. The specific steps are as follows.

[0173] S321, if the current long-term reference frame is not a knowledge image, the similarity of the long-term reference frame selected in the current mode to the video segment consisting of all frames from the start frame of the current video sub-segment to the end of the random access segment to which the current video sub-segment belongs is calculated as a gain.

[0174] Since a scene may be blocked and then reappear, the reference value of the long-term reference frame does not always decrease with distance. The current long-term reference frame may not have a significant gain for subsequent closer videos, but may have a significant gain for future farther videos. Therefore, it is necessary to calculate the correlation from the starting frame of the current video sub-segment to the end of the random access segment.

[0175] S322, if the current long-term reference frame is a knowledge image, refer to the above-mentioned knowledge image similarity calculation to calculate the similarity of the knowledge image to the current video sub-segment and the L video sub-segments after the current video sub-segment, and scale the sum of the similarities as the gain. Let the sum of the relevance of L+1 sub-segments be similarity L , the total length of the area where the knowledge image calculates the correlation is Len L , the total length of the area from the current segment to the end of the random access segment is Len R , then the benefit of the knowledge image as a long-term reference frame is

[0176] S323, if there is no long-term reference frame currently, the gain is 0.

[0177] In one embodiment of the present invention, the loss of the long-term reference frame is calculated as follows. Since the long-term reference frame itself may replace an original short-term reference frame that is closer, the reference value of the replaced short-term reference frame can be considered to be the loss of the long-term reference frame. Currently, this loss is represented by a smaller fixed threshold.

[0178] When there is a long-term reference frame, the loss is a fixed threshold CC and a smaller value can be selected, such as 0.1; when there is no long-term reference frame, the loss is 0.

[0179] Then net profit = profit - cost.

[0180] S33, selecting the optimal long-term reference frame mode for each video sub-segment based on the net benefit comparison.

[0181] The net benefits of the four selection modes are compared to determine the final long-term reference frame selection. The specific steps are as follows.

[0182] (1) defining a threshold TH_REL_USE_NEW, which is used to determine whether to select a new non-knowledge image frame as a long-term reference frame;

[0183] In one embodiment of the present invention, the threshold TH_REL_USE_NEW is defined as 2.5.

[0184] (2) defining a threshold TH_REL_USE_LP, which is used to determine whether to select the knowledge image as a long-term reference frame;

[0185] In one embodiment of the present invention, the threshold TH_REL_USE_LP is defined as 1.5.

[0186] (3) Define gain last , gain newlt ,gai newtp , gain none They are the net benefits of the aforementioned selection modes 1 to 4 respectively;

[0187] (4) Make a selection decision, as shown in the flowchart Figure 8 shown.

[0188] If gain none >MAX(gain last , gain newlt , gain newlp ), then select the fourth mode, that is, do not select the long-term reference frame;

[0189] Otherwise, if Then the first mode is selected, that is, the long-term reference frame of the previous sub-segment is selected as the current long-term reference frame;

[0190] Otherwise, if Then choose the third mode, i.e., knowledge image;

[0191] Otherwise, the second mode is selected, i.e., a new long-term reference frame of a non-knowledge image is selected.

[0192] The video encoding method based on the above-mentioned knowledge image scope decision method uses the scene knowledge base constructed in the fine-grained knowledge image selection method and the long-term reference frame selected in the knowledge image scope decision method for video encoding.

[0193] It should be noted that reference frame selection belongs to the category of video encoding. Reference frame selection can be performed before or during the main video encoding process. This type of algorithm is part of the codec core algorithm.

[0194] It should be noted that similarity calculation can be used in knowledge image selection and knowledge image scope decision steps. When used in knowledge image selection, the improved cutree algorithm is called on the candidate knowledge images in the candidate knowledge image set.

[0195] When used in the knowledge image scope decision step, the improved cutree algorithm is called on the selected knowledge image or the ordinary long-term reference frame of the non-knowledge image. At this time, the "knowledge image" here is the selected knowledge image or the ordinary long-term reference frame.

[0196] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.

Claims

1. A fine-grained scenario knowledge base construction and scope decision method, It is characterized in that include: Candidate knowledge image selection, selecting multiple candidate knowledge images from a single random access segment; Construction of scene knowledge base: a single random access segment is divided into several sub-segments, each of which independently refers to any frame of reference knowledge image in the knowledge image set, and benefits and costs are calculated based on the reference knowledge image and the sub-segment content. Based on the calculation results, the optimal knowledge image set is selected from the candidate knowledge image set to construct the scene knowledge base, and the reference knowledge image is not limited to the candidate knowledge image; Knowledge image scope decision: assign an optimal long-term reference frame to each sub-segment, and the optimal long-term reference frame is not limited to selecting knowledge images from the optimal knowledge image set.

2. According to the method for constructing a fine-grained scene knowledge base and determining the scope of action as described in claim 1, It is characterized in that The candidate knowledge image selection at least includes: Random access segment segmentation: split the random access segment into several sub-segments according to preset rules. Candidate knowledge image extraction, extracts frames from the segmented sub-segments as candidate knowledge images according to preset rules.

3. According to the method for constructing a fine-grained scene knowledge base and determining the scope of action as described in claim 1, It is characterized in that The construction of the scenario knowledge base at least includes: Sub-segment segmentation: dividing the random access segment into several sub-segments according to a preset rule; Similarity calculation: calculate the similarity of each candidate knowledge image corresponding to each sub-segment; Optimal knowledge image selection: The benefits and costs of the candidate knowledge image set are calculated based on similarity, and the optimal knowledge image is selected based on the benefits and costs.

4. According to claim 3, a fine-grained scenario knowledge base construction and scope decision method, It is characterized in that The propagation information calculated based on the improved cutree algorithm is used to calculate the similarity of each candidate knowledge image corresponding to each sub-segment.

5. According to claim 4, a fine-grained scenario knowledge base construction and scope decision method, It is characterized in that The improved cutree algorithm adds a long-term reference frame hypothesis on the basis of the cutree algorithm of X256, and calculates the propagation information flowing to the long-term reference frame based on the prediction loss of the existing short-term reference frame hypothesis and the added long-term reference frame hypothesis, thereby calculating the reference value of a given long-term reference frame.

6. According to claim 5, a fine-grained scenario knowledge base construction and scope decision method, It is characterized in that The improved cutree algorithm adds a reference relationship of a long-term reference frame on the basis of the cutree algorithm of X256, and calculates the propagation information flowing to the long-term reference frame based on the prediction loss of the existing short-term reference frame hypothesis and the added long-term reference frame hypothesis, specifically including: For each frame in the sub-segment, a reference relationship assumption between the long-term reference frame group and the short-term reference frame group is made; Inter-frame prediction of long-term reference frame groups and inter-frame prediction of short-term reference frame groups; Calculate the propagation information of the reference block based on the intra-frame prediction result and the inter-frame prediction result and each block of each frame of the traversal subsegment; The propagation information of all sub-blocks is calculated by normalizing the intra-frame prediction loss and taking the average value to obtain the similarity of the candidate knowledge image for each video sub-segment.

7. A fine-grained scenario knowledge base construction and scope decision method according to claim 6, It is characterized in that The calculating of propagation information of the reference block based on the intra-frame prediction result and the inter-frame prediction result and traversing each block of each frame of the first video sub-segment specifically includes: Calculate the total amount of long-term reference frame group information based on intra-frame prediction loss and inter-frame prediction loss of the long-term reference frame group; Calculate the total amount of short-term reference frame group information based on intra-frame prediction loss and inter-frame prediction loss of the short-term reference frame group; The reference probability of each long-term reference frame group and short-term reference frame group is calculated. Correct the total amount of information based on the reference probability; The total amount of corrected information is accumulated and passed to the reference frame as propagation information.

8. According to claim 1, a fine-grained scenario knowledge base construction and scope decision method, It is characterized in that The knowledge image scope decision allocates the best long-term reference frame to each sub-segment, specifically including: Divide a single random access segment into several video sub-segments, Traversing the video sub-segments, and calculating the net benefits of several long-term reference frame selection modes for each video sub-segment; The optimal long-term reference frame mode for each video sub-segment is selected based on the net benefit comparison.

9. A fine-grained scenario knowledge base construction and scope decision method according to claim 8, It is characterized in that The long-term reference frame selection mode includes at least one of a first mode, a second mode, a third mode and a fourth mode. In the first mode, the long-term reference frame of the previous video sub-segment is used as the long-term reference frame of the current video sub-segment; In the second mode, the long-term reference frame of the first frame of the video sub-segment is used as the video sub-segment s n long-term reference frame; In the third mode, the knowledge image with the greatest similarity to the video sub-segment is used as the long-term reference frame of the video sub-segment; The fourth mode does not use long-term reference frames.

10. A fine-grained scenario knowledge base construction and scope decision method according to claim 9, It is characterized in that The optimal long-term reference frame mode for each video sub-segment is selected based on the net benefit comparison, If the net gain value of the fourth mode is none >MAX(gain last , gain newlt , gain newlp ), then select the fourth mode; Otherwise if Then select the first mode; Otherwise if Then select the third mode; Otherwise, select the second mode; Among them, the first threshold TH_REL_USE_NEW, the second threshold TH_REL_USE_LP, gain last , gain newlt , gain newlp , gain none They are the net profit values ​​of the first to fourth modes respectively.