Video coding method and related device

By obtaining historical detection information from a set of reference video frames for prediction, the encoding parameters of the video frames are determined, which solves the problem of poor encoding flexibility due to fixed quantization parameters in existing technologies, and achieves high efficiency and flexibility in video encoding.

CN121125993APending Publication Date: 2025-12-12ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511190149.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

In existing video coding methods, the fixed quantization parameters result in poor coding flexibility and make it difficult to balance image quality and compression ratio.

Method used

By acquiring the target video frame in the current detection round, a set of reference video frames is obtained. Based on historical detection information, prediction is performed, and the encoding parameter information is determined to encode the video frame.

Benefits of technology

It improves the efficiency and flexibility of video encoding, enabling higher compression rates and quality while preserving key details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125993A_ABST
    Figure CN121125993A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding method and a related device, and the method comprises the steps: obtaining an initial video frame sequence obtained through collection, and determining a target video frame matched with a current detection round from the initial video frame sequence; obtaining a reference video frame set based on the target video frame; wherein the reference video frame set comprises historical video frames matched with a plurality of historical detection rounds, and the historical video frames are matched with historical detection information; obtaining target prediction information matched with the target video frame based on the historical detection information matched with the reference video frame set; and based on the target prediction information, determining coding parameter information matched with the target video frame, and coding the target video frame by using the coding parameter information. In this way, the video coding efficiency and flexibility can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a video encoding method and related device. BACKGROUND

[0002] In the video encoding process, quantization plays a crucial role, which is the core means to achieve high compression rate and is also the key regulating valve between image quality and code rate control. The strength of this lossy operation is directly controlled by the quantization parameter (QP). The quantization parameter essentially defines the size of the quantization step: the larger the quantization parameter value, the larger the quantization step, resulting in more significant compression effect, but also introducing more serious image quality loss; on the contrary, the smaller the quantization parameter value, the finer the quantization step, the higher the accuracy of coefficient preservation, and the better the reconstructed image quality, but the required code rate also rises. The current traditional video encoding method mainly relies on fixed quantization parameter for encoding, resulting in poor encoding flexibility.

[0003] Therefore, how to propose a video encoding method with high efficiency and flexibility has become a problem to be solved. SUMMARY

[0004] The technical problem solved by the present application is to provide a video encoding method and related device, which can improve the efficiency and flexibility of video encoding.

[0005] To solve the above technical problem, one technical solution adopted by the present application is to provide a video encoding method, comprising: acquiring an initial video frame sequence collected, determining a target video frame matched with a current detection round from the initial video frame sequence; acquiring a reference video frame set based on the target video frame; wherein the reference video frame set includes historical video frames matched with a plurality of historical detection rounds, and the historical video frames are matched with historical detection information; acquiring target prediction information matched with the target video frame based on the historical detection information matched with the reference video frame set; determining encoding parameter information matched with the target video frame based on the target prediction information, and encoding the target video frame by using the encoding parameter information.

[0006] To solve the above technical problem, another technical solution adopted by the present application is to provide an electronic device, comprising: a memory and a processor coupled with each other, the memory has program instructions stored therein, and the processor is configured to execute the program instructions to implement the method mentioned in the above technical solution.

[0007] To solve the above technical problems, the application adopts another technical solution: providing a computer readable storage medium, which stores program instructions, and the program instructions are executed by a processor to implement the method mentioned in the above technical solution.

[0008] The application has the following advantages: Different from the prior art, the video encoding method provided by the application obtains a corresponding reference video frame set according to a target video frame matched with a current detection round. The position information of a target object in each target video frame is predicted according to historical detection information matched with each historical video frame in the reference video frame set, to obtain target prediction information matched with each target video frame. The quantization parameters matched with each region in the target video frame are determined according to the target prediction information, so that the key detail information related to the target object in the target video frame can be retained and the video compression rate can be improved in the encoding process, greatly improving the quality and efficiency of video compression. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort. Among them:

[0010] Figure 1 is a flowchart of an embodiment of the video encoding method of the application;

[0011] Figure 2 is Figure 1 corresponds to another embodiment of step S102 in

[0012] Figure 3 is a schematic diagram of an embodiment of the initial video frame sequence of the application;

[0013] Figure 4 is a schematic diagram of an embodiment of the reference video frame sequence of the application

[0014] Figure 5 is Figure 1 corresponds to another embodiment of step S103 in

[0015] Figure 6 is a flowchart of an embodiment of the reference confidence acquisition method of the application;

[0016] Figure 7 is Figure 1 corresponds to another embodiment of step S104 in

[0017] Figure 8 is a schematic diagram of an embodiment of a target video frame corresponding to the present application;

[0018] Figure 9 is a structural schematic diagram of an embodiment of an electronic device of the present application;

[0019] Figure 10 is a structural schematic diagram of an embodiment of a computer readable storage medium of the present application. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application, and the embodiments can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0021] Please refer to Figure 1 , Figure 1 is a flow schematic diagram of an embodiment of a video encoding method of the present application, and the method comprises:

[0022] S101: acquiring an initial video frame sequence obtained by collection, and determining a target video frame matched with a current detection round from the initial video frame sequence.

[0023] In an embodiment, a video stream collected by a collection device in real time is acquired, and a corresponding initial video frame sequence is obtained according to the video stream. A target video frame matched with a current detection round is acquired from the initial video frame sequence.

[0024] In some implementation scenarios, the video stream collected by the collection device is composed of multiple initial video frames, each initial video frame is matched with a timestamp, and all initial videos are arranged in turn in an ascending order of the timestamps to obtain an initial video frame sequence. The initial video frame sequence corresponds to multiple detection rounds, and at least one target video frame corresponding to a current detection round is determined. The above-mentioned collection device is a monocular camera, and the detection rounds before the current detection round are historical detection rounds.

[0025] In a specific application scenario, in order to perform subsequent coding tasks and / or video processing tasks, in the process of acquiring the initial video frame sequence constructed by the initial video frames collected in real time, a fixed number of continuous initial video frames are taken as a detection round, so that target detection is performed on a certain initial video frame in each detection round, that is, each detection round matches a fixed number of initial video frames. For example, when each detection round matches ten initial video frames, the ten video frames after the last historical detection round are taken as the target video frames matched by the current detection round. The above target detection can be realized by using an existing image detection model, and the specific process will not be described in detail here.

[0026] Of course, it should be noted that the number of video frames matched by each detection round can be set according to actual conditions, and in order to save the consumption of computing resources, target detection is performed on the last initial video frame in each detection round. In addition, when the acquisition device collects video streams matched with multiple different channels in real time, the initial video frame sequence corresponding to each channel is constructed for each channel video stream, and each initial video frame sequence is labeled with corresponding channel information.

[0027] In another embodiment, a complete video stream that has been collected is acquired, and an initial video frame sequence is obtained according to the time stamps corresponding to each initial video frame in the complete video stream. According to the initial video frame sequence, a plurality of detection rounds are divided so that. For the current detection round, a plurality of target video frames matched from the initial video frame sequence are determined.

[0028] S102: Based on the target video frame, a reference video frame set is acquired; wherein the reference video frame set includes historical video frames matched with a plurality of historical detection rounds, and each historical video frame is matched with historical detection information.

[0029] In an embodiment, according to the target video frame matched by the current detection round, historical video frames matched by a plurality of historical detection rounds are acquired to construct a reference video frame set. Each historical video frame in the reference video frame set is matched with historical detection information, and the historical detection information includes position information of the target object in the corresponding historical video frame.

[0030] In some implementation scenarios, in response to the fact that each historical detection round matches a plurality of initial video frames, for each historical detection round, the initial video frame after target detection is taken as a historical video frame, that is, the historical video frame matched by each historical detection round corresponds to historical detection information.

[0031] S103: Based on the historical detection information matched with the reference video frame set, target prediction information matched with the target video frame is acquired.

[0032] In an embodiment, motion data of the target object is determined according to historical detection information matched by each historical video frame in the set of reference video frames. Each target video frame matched in the current detection round is predicted using the determined motion data to determine target prediction information matched by each target video frame in the current detection round.

[0033] In some implementation scenarios, the motion data of the target object includes at least one of a displacement vector, a speed, and an acceleration of the target object. The position information of the target object in each target video frame is predicted according to the motion data to obtain corresponding target prediction information.

[0034] S104: Based on the target prediction information, determine encoding parameter information matched by the target video frame, and encode the target video frame using the encoding parameter information.

[0035] In an embodiment, the encoding parameter information matched by the target video frame is set according to the target prediction information matched by the target video frame, and the encoding parameter information includes a quantization parameter (QP) matched by different regions in the target video frame. The target video frame is encoded using the corresponding encoding parameter information to obtain an encoded video stream file for subsequent processing.

[0036] In some implementation scenarios, the quantization parameter is used to adjust the quantization intensity in the encoding process during video encoding, and directly affects the compression rate and image quality of the video. The smaller the value of the quantization parameter, the smaller the quantization step, the more details retained after encoding, and the higher the image quality, but the lower the compression rate. The larger the value of the quantization parameter, the larger the quantization step, the more details lost after encoding, and the lower the image quality, but the higher the compression rate. In response to the determined target prediction information including the position information of the target object in the corresponding target video frame, a lower quantization parameter is set for the region in the target video frame where the target object is located, and a higher quantization parameter is set for other regions in the target video frame except the region where the target object is located, so that the key detail information is retained while the video compression rate is improved.

[0037] The video encoding method proposed in the present application obtains a set of reference video frames according to the target video frames matched in the current detection round. The position information of the target object in each target video frame is predicted according to the historical detection information matched by each historical video frame in the set of reference video frames to obtain target prediction information matched by each target video frame. The quantization parameter matched by each region in the target video frame is determined according to the target prediction information, so that both the key detail information related to the target object in the target video frame and the video compression rate can be retained during the encoding process, greatly improving the quality and efficiency of video compression.

[0038] Please refer to Figure 2 , Figure 2 is Figure 1 corresponding to another embodiment of the flowchart of step S102. Specifically, the implementation process of step S102 includes:

[0039] S201: Based on the target video frame, obtain historical video frames matched with multiple historical detection rounds, and determine a reference video frame sequence composed of all historical video frames.

[0040] In an embodiment, please refer to Figure 3 and Figure 4 , Figure 3 is a schematic diagram of the initial video frame sequence of the present application corresponding to an embodiment, Figure 4 is a schematic diagram of the reference video frame sequence of the present application corresponding to an embodiment. According to the target video frame matched with the current detection round, historical video frames matched with multiple historical detection rounds are obtained from the initial video frame sequence collected, and the historical video frames are composed into a reference video frame sequence.

[0041] Specifically, as shown in Figure 3 and Figure 4 , the initial video frame sequence collected by the collection device includes video frames matched with multiple detection rounds, i.e. video frames matched with the current detection round, historical detection round a, historical detection round b, historical detection round c, etc. In response to each historical detection round matching a historical video frame subjected to target detection, historical video frames matched with multiple historical detection rounds are extracted from the initial video frame sequence to compose a reference video frame sequence.

[0042] In some implementation scenarios, historical video frames matched with all historical detection rounds before the current detection round are obtained, and all historical video frame sequences are sorted in chronological order according to the corresponding time stamps to obtain the reference video frame sequence.

[0043] Alternatively, historical video frames within a preset time range are extracted from the initial video frame sequence to compose the reference video frame sequence. For example, in response to the first time stamp in the current detection round matching a first time stamp, historical video frames with a time difference between the corresponding time stamp and the above-mentioned first time stamp within a preset time range are obtained to compose the reference video frame sequence.

[0044] Or, in chronological order, a preset number of historical video frames before the current detection round are sequentially selected to compose the reference video frame sequence. The above-mentioned preset number can be set according to actual scenarios.

[0045] In addition, it should be noted that, since the video stream is real-time and continuously captured, the number of historical video frames in each reference video frame set is constantly updated or increased as the number of initial video frames in the initial video frame sequence increases. For example, as the iteration of the detection round, the latest acquired historical video frame is added to the corresponding reference video frame set by the above-mentioned implementation; or, in response to the addition of the latest acquired historical video frame to the corresponding reference video frame set, the number of historical video frames in the reference video frame set is greater than the preset number threshold, the historical video frame corresponding to the minimum timestamp is deleted from the reference video frame set.

[0046] S202: Obtain a plurality of preset interval steps, and obtain historical video frames from the reference video frame sequence to form a reference video frame set based on the preset interval steps; wherein the number of historical video frames in each reference video frame is consistent.

[0047] In an implementation, a plurality of preset interval steps are obtained in advance, and corresponding historical video frames are selected from the reference video frame sequence to form a reference video frame set according to the preset interval steps. It should be noted that the number of historical video frames in each obtained reference video frame set is consistent.

[0048] In a specific application scenario, the plurality of preset interval steps include 0, 1, 2, 3, etc. When the preset interval step is 0, a plurality of continuous historical video frames are obtained from the reference video sequence to obtain a reference video frame set; when the preset interval step is 1, one historical video frame is extracted from the reference video frame sequence each time, and the extracted historical video frame is added to the reference video frame set; when the preset interval step is 2, two historical video frames are extracted from the reference video frame sequence each time, and the extracted historical video frame is added to the reference video frame set. When the preset interval step is other, the process of constructing the reference video frame set is similar to the above process, which will not be described in detail here.

[0049] It should be noted that the specific value of the preset interval step and the number of preset interval steps can be set according to the actual scene. The more the number of preset interval steps is set, the more the number of reference video frame sets obtained.

[0050] The above scheme, after determining the reference video frame sequence composed of a plurality of historical video frames, determines a plurality of reference video frame sets according to different preset interval steps, so that subsequent target prediction information matched with the target video frame is obtained according to different reference video frame sets. Unlike the implementation of determining the target prediction information from a single reference video frame set, the probability of false prediction is reduced, which helps to improve the quality of video coding.

[0051] Please refer to Figure 5 ,Figure 5 is Figure 1 The step S103 corresponds to a flowchart of another embodiment. Specifically, the implementation process of the step S103 includes:

[0052] S301: For each set of reference video frames, obtain motion data of the target object based on the historical detection information matched by each historical video frame.

[0053] In an embodiment, for each set of reference video frames, obtain motion data of the target object according to the historical detection information matched by each historical video frame in the corresponding set of reference video frames.

[0054] In some implementation scenarios, the historical detection information includes a bounding box labeled in the corresponding historical video frame, and the bounding box is used to represent the position information of the target object in the historical video frame. For each historical video frame in the set of reference video frames, determine the motion data of the target object according to the bounding boxes in adjacent historical video frames.

[0055] Specifically, the historical video frames in the set of reference video frames are arranged in ascending order of time stamp, and for adjacent historical video frames, the center point coordinates of the corresponding bounding boxes are determined. According to the center point coordinates corresponding to the adjacent historical video frames and the time difference corresponding to the adjacent historical video frames, the motion data of the target object in the adjacent historical video frames is calculated, which includes at least one of displacement vector, velocity and acceleration.

[0056] In a specific application scenario, in response to the motion data including displacement vector, velocity and acceleration, the calculation formula of the motion data of the target object in any adjacent historical video frame is as follows:

[0057] Δp = (x i -x i , y i+1 -y i+1 )

[0058]

[0059] wherein (x i+1 , y ) represents the center point coordinates of the first historical video frame, (x

[0001] , y ) represents the center point coordinates of the second historical video frame, the first historical video frame and the second historical video frame are adjacent in the corresponding set of reference video frames, and the time stamp corresponding to the first historical video frame is earlier than the time stamp corresponding to the second historical video frame; Δp represents the displacement vector of the target object obtained according to the first historical video frame and the second historical video frame; v represents the velocity of the target object; and a represents the acceleration of the target object.

[0060] In another embodiment, the historical detection information of the historical video frame matching includes a bounding box matched with a plurality of different target objects, and each bounding box is matched with a corresponding object identifier, which is used to distinguish different target objects in the same historical video frame. For each reference video frame set, motion data of each target object is obtained according to the historical detection information of each historical video frame matching; that is, for each object identifier, the corresponding motion data of each target object is obtained through the above embodiment, so that the position information of each target object in the target video frame is predicted according to the motion data subsequently.

[0061] It should be noted that since the number of reference video frame sets is multiple, the motion data corresponding to each reference video frame set can be calculated through the above embodiment.

[0062] S302: input the motion data obtained based on the reference video frame set into the trained prediction model, and obtain the candidate prediction result matched with the target video frame output by the prediction model.

[0063] In an embodiment, a pre-trained prediction model is obtained. The motion data corresponding to each reference video frame set is input into the prediction model, and the motion trend of the target object is predicted by using the prediction model to obtain the output candidate prediction result. The candidate prediction result includes candidate prediction information matched with each target video frame in the current detection round.

[0064] In some implementation scenarios, the prediction model is obtained by training a plurality of training samples, and the model structure of the prediction model can refer to a neural network model structure, such as an R-CNN model or a YOLO model. By inputting the motion data corresponding to each reference video frame set into the prediction model, the prediction model outputs the candidate prediction information matched with each target video frame in the current detection round, and each candidate prediction information includes the position information of at least one predicted target object.

[0065] In another embodiment, the prediction model can also be a large language model with better data analysis capability. The motion data corresponding to each reference video frame set and the task text are input into the large language model, the motion data and the task text are analyzed by using the large language model, and the candidate prediction result is generated. The task text is used to prompt the large language model to predict the motion trend of the target object in the target video frame according to the input motion data.

[0066] In a specific application scenario, the large language model can include, but is not limited to, deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTM), and generative pre-training Transformer models, and the like. The specific construction and specific deployment of the large language model are not limited here.

[0067] S303: determining a target prediction result from all candidate prediction results; wherein the target prediction result includes target prediction information matched by all target video frames in the current detection round.

[0068] In an embodiment, the candidate prediction result predicted by the prediction model according to each reference video frame set includes candidate prediction information matched by each target video frame in the current detection round. In response to the number of reference video frame sets being multiple, each target video frame is matched with multiple candidate prediction information, and each candidate prediction information includes a bounding box matched with the target object. For multiple candidate prediction information matched with the same target video frame, an overlapping region corresponding to all bounding boxes is obtained, and the overlapping region is taken as the target prediction information matched with the corresponding target video frame.

[0069] In another embodiment, each reference video frame set is matched with a reference confidence, and the reference confidence is determined based on the matching degree between the historical detection information and the historical prediction information corresponding to the corresponding historical detection round. Based on this, after obtaining the candidate prediction result matched by the target video frame, the maximum confidence is determined from the reference confidence matched by all reference video frame sets, and the candidate prediction result corresponding to the maximum confidence is taken as the target prediction result.

[0070] In a specific application scenario, the reference video frame set A, the reference video frame set B and the reference video frame set C are determined by the above-mentioned corresponding embodiments, and the corresponding candidate prediction results are obtained by the prediction model according to the reference video frame set A, the reference video frame set B and the reference video frame set C respectively. In response to the reference confidence corresponding to the reference video frame set A being 0.7, the reference confidence corresponding to the reference video frame set B being 0.6, and the reference confidence corresponding to the reference video frame set C being 0.8, it is determined that 0.8 is the maximum confidence, and the candidate prediction result obtained according to the reference video frame set C is taken as the target prediction result, so as to determine the target prediction information matched by each target video frame in the current detection round.

[0071] The scheme is to obtain a candidate prediction result matched with the target video frame according to the reference video frame set, and select a target prediction result with higher accuracy from the candidate prediction result, so as to help subsequent division of each target video frame according to the target prediction result, thereby improving the effect and efficiency of video coding.

[0072] Please refer to Figure 6 , Figure 6 is a flowchart of an embodiment of the confidence reference acquisition method. Specifically, the process of obtaining the reference confidence in the above-mentioned embodiment includes:

[0073] S401: For each reference video frame set, obtain the matching degree between the historical detection information and the historical prediction information corresponding to each historical detection round; wherein the matching degree is determined based on the overlapping area of the target object in the historical detection information and the historical prediction information.

[0074] In an embodiment, each reference video frame set includes historical video frames matched with multiple historical detection rounds, and the historical video frames are matched with historical prediction information obtained by using a prediction model and historical detection information obtained by target detection. For each historical video frame, the matching degree between the corresponding historical detection information and the historical prediction information is obtained, which is determined according to the overlapping area between the target object in the historical detection information and the target object in the historical prediction information.

[0075] Specifically, for each historical video frame, the overlapping area of the target object in the corresponding historical detection information and the target object in the historical prediction information is obtained. According to the above-mentioned overlapping area, the matching degree is determined. The larger the above-mentioned overlapping area, the higher the matching degree.

[0076] In another embodiment, for each historical video frame, the overlapping area of the target object in the corresponding historical detection information and the historical prediction information is obtained. The actual area of the target object in the historical detection information is obtained, and the ratio between the above-mentioned overlapping area and the above-mentioned actual area is determined. According to the above-mentioned ratio, the matching degree is determined. The larger the above-mentioned ratio, the higher the matching.

[0077] S402: Based on the matching degree, determine the successful detection round that meets the preset condition from all historical detection rounds.

[0078] In an embodiment, for each historical video frame, it is judged whether the corresponding matching degree is greater than a preset matching threshold. If it is greater than or equal to, the historical detection round matched with the corresponding historical video frame is regarded as a successful detection round. If it is less than, it is considered that the accuracy of the historical detection information matched with the historical video frame is low, and the corresponding historical detection round is regarded as a failed detection round.

[0079] S403: determining a reference confidence of the matched reference video frame set based on the number of successful detection rounds.

[0080] In an embodiment, for each reference video frame set, the number of all successful detection rounds of the matched is determined, and the reference confidence of the matched reference video frame set is determined according to the number.

[0081] Specifically, for each reference video frame set, the first number corresponding to all successful detection rounds of the matched is obtained, and the second number corresponding to all historical detection rounds of the matched is obtained. The ratio of the first number and the second number is taken as the reference confidence threshold of the matched reference video frame set.

[0082] The above scheme determines the reference confidence according to the historical detection information and the historical prediction information corresponding to each historical detection round, so that the target prediction result is determined from the multiple candidate prediction results according to the reference confidence in the actual application process, and the accuracy of obtaining the target prediction result is improved.

[0083] In an embodiment, the reference confidence of each reference video frame set is dynamically changed, and after determining the target prediction result matched in the current detection round, the following steps are included: in response to the current detection round including multiple first target video frames and a second target video frame for detection, obtaining the current detection information matched by the second target video frame in the current detection round. For the second target video frame, the reference confidence of each reference video frame set is updated based on the current detection information and the target prediction information matched by each reference video frame set.

[0084] In some implementation scenarios, in response to obtaining all target video frames matched in the current detection round, for all target video frames matched in the current detection round, the target video frame corresponding to the maximum timestamp is taken as the second target video frame; and the multiple target video frames located before the second target video frame are taken as the first target video frames. The target video frame is detected to obtain the matched current detection information. The reference confidence of each reference video frame set is updated according to the current detection information corresponding to the second target video frame and the target prediction information matched by the second target video frame according to each reference video frame set.

[0085] Please refer to Figure 7 , Figure 7 is Figure 1 the flowchart of another embodiment corresponding to step S104 in

[0086] S501: dividing the corresponding target video frame into multiple coding regions based on the target prediction information.

[0087] In an embodiment, in response to the target prediction information comprising position information of the target object in the corresponding target video frame, the target video frame is divided into a plurality of coding regions according to the position information of the target object in the target video frame.

[0088] In some implementation scenarios, based on the target prediction information, a first coding region in which the target object is located in the corresponding target video frame is obtained. A second coding region around the first coding region in the target video frame is obtained. Other regions outside the first coding region and the second coding region in the target video frame are taken as a third coding region.

[0089] Specifically, referring to Figure 8 , Figure 8 is a schematic diagram of an embodiment corresponding to the target video frame of the present application. In response to the target prediction information comprising a labeling box for labeling the target object in the target video frame, a region corresponding to the labeling box in the target prediction information is taken as a first coding region. In addition, a part of the region around the first coding region is taken as a second coding region, and the remaining region in the target video frame is taken as a third coding region. Among them, the first coding region includes detailed information related to the target object.

[0090] In another embodiment, in response to the target prediction information comprising a labeling box for labeling the target object in the target video frame, the region corresponding to the labeling box in the target prediction information is appropriately expanded to obtain a first coding region. In addition, a part of the region around the first coding region is taken as a second coding region, and the remaining region in the target video frame is taken as a third coding region.

[0091] In another embodiment, according to the target prediction information and the motion data of the target object, the corresponding target video frame is divided into a plurality of coding regions.

[0092] Specifically, according to the motion data of the target object, the motion intensity of the target object is determined, and the motion intensity is determined based on at least one of the speed, displacement vector and acceleration of the target object in a preset time period. According to the motion intensity of the target object, the region corresponding to the labeling box in the target prediction information is expanded to obtain a first coding region. A part of the region around the first coding region is taken as a second coding region, and the remaining region in the target video frame is taken as a third coding region. Among them, the greater the motion intensity of the target object, the greater the area of the region corresponding to the labeling box after the expansion, so as to retain more detailed information in the subsequent encoding process.

[0093] In some implementation scenarios, the faster the speed of the target object in a preset time period, the greater the corresponding motion intensity; the greater the acceleration of the target object in a preset time period, the greater the corresponding motion intensity; and the greater the displacement vector of the target object in a preset time period, the greater the corresponding motion intensity.

[0094] S502: Obtain the encoding parameter corresponding to each encoding region based on the position information of the encoding region.

[0095] In an embodiment, the encoding parameter corresponding to each encoding region is obtained. The encoding parameter corresponding to the encoding region where the target object is located is smaller than the encoding parameter corresponding to other encoding regions. The smaller the distance between the encoding region and the target object, the smaller the encoding parameter corresponding to the encoding region.

[0096] In some implementation scenarios, in response to the target video frame being divided into a first encoding region, a second encoding region and a third encoding region, a first quantization parameter matched with the first encoding region, a second quantization parameter matched with the second encoding region and a third quantization parameter matched with the third encoding region are obtained. The first quantization parameter is smaller than the second quantization parameter, and the second quantization parameter is smaller than the third quantization parameter.

[0097] In a specific application scenario, a default quantization parameter is obtained. A first adjustment parameter matched with the first encoding region is obtained, and the difference between the default quantization parameter and the first adjustment parameter is taken as the first quantization parameter. A second adjustment parameter matched with the second encoding region is obtained, and the difference between the default quantization parameter and the second adjustment parameter is taken as the second quantization parameter. A third adjustment parameter matched with the third encoding region is obtained, and the difference between the default quantization parameter and the third adjustment parameter is taken as the third quantization parameter. The first adjustment parameter is greater than the second adjustment parameter, and the second adjustment parameter is greater than the third adjustment parameter.

[0098] The above scheme divides the target video frame into multiple different encoding regions according to the predicted position information of the target object, determines the matched quantization parameter according to different encoding regions, and encodes the target video frame according to the quantization parameter, so as to improve the compression rate while retaining key detail information.

[0099] In an embodiment, after the target video frame matched with the current detection round is encoded by using the determined encoding parameter information according to the above embodiment, the current detection round is updated to a historical detection round, and the step S101 is returned to determine the target video frame matched with the current detection round from the initial video frame sequence and encode the target video frame matched with the current detection round by using the corresponding embodiment.

[0100] Please refer to Figure 9 , Figure 9is a structural schematic diagram of an embodiment of an electronic device of the present application. The electronic device comprises a memory 10 and a processor 20 coupled with each other. The memory 10 stores program instructions, and the processor 20 is configured to execute the program instructions to implement the method mentioned in any of the above embodiments. Specifically, the electronic device includes but is not limited to a desktop computer, a notebook computer, a tablet computer, a server, etc., which are not limited herein. In addition, the processor 20 can also be referred to as a CPU (Central Processing Unit). The processor 20 can be an integrated circuit chip with signal processing capability. The processor 20 can also be a general processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 20 can be implemented by an integrated circuit chip together.

[0101] Please refer to Figure 10 , Figure 10 is a structural schematic diagram of an embodiment of a computer readable storage medium of the present application. The storage medium 30 stores program instructions 40 capable of being executed by a processor, and the program instructions 40 are executed by the processor to implement the method mentioned in any of the above embodiments.

[0102] In several embodiments provided in the present application, it should be understood that the disclosed method and device can be implemented in other ways. For example, the above-described device embodiment is only schematic, for example, the division of the module or unit is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed units can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0103] The unit described as a separate component can or can not be physically separated, and the component shown as a unit can or can not be a physical unit, that is, it can be located in one place, or it can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the present embodiment scheme.

[0104] In addition, each of the functional units in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0105] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (such as a personal computer, a server, or a network device) or a processor to perform all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: various types of U disks, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks, and various other media that can store program codes.

[0106] The above merely describes the embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation based on the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A video encoding method, characterized in that, include: Acquire the initial video frame sequence obtained from the acquisition, and determine the target video frame that matches the current detection round from the initial video frame sequence; Based on the target video frame, a reference video frame set is obtained; wherein, the reference video frame set includes historical video frames that match multiple historical detection rounds, and the historical video frames are matched with historical detection information; Based on the historical detection information that matches the reference video frame set, target prediction information that matches the target video frame is obtained; Based on the target prediction information, encoding parameter information matching the target video frame is determined, and the target video frame is encoded using the encoding parameter information.

2. The method according to claim 1, characterized in that, The step of obtaining a reference video frame set based on the target video frame includes: Based on the target video frame, obtain the historical video frames that match multiple historical detection rounds, and determine a reference video frame sequence composed of all the historical video frames; Multiple preset interval steps are obtained, and based on the preset interval steps, the historical video frames are obtained from the reference video frame sequence to form the reference video frame set; wherein, the number of historical video frames in each reference video frame is the same.

3. The method according to claim 1, characterized in that, The historical detection information includes the position information of the target object in the corresponding historical video frame. The step of obtaining target prediction information matching the target video frame based on the historical detection information matched with the reference video frame set includes: For each set of reference video frames, motion data of the target object is obtained based on the historical detection information matched by each historical video frame; The motion data obtained based on the reference video frame set is input into the trained prediction model to obtain the candidate prediction results output by the prediction model that match the target video frame; The target prediction result is determined from all the candidate prediction results; wherein the target prediction result includes the target prediction information matched by all the target video frames in the current detection round.

4. The method according to claim 3, characterized in that, Each of the reference video frame sets is matched with a reference confidence score, which is determined based on the matching degree between historical detection information and historical prediction information corresponding to the corresponding historical detection round. Determining the target prediction result from all the candidate prediction results includes: The maximum confidence score is determined from the reference confidence scores matched across all the reference video frame sets, and the candidate prediction result corresponding to the maximum confidence score is used as the target prediction result.

5. The method according to claim 4, characterized in that, After determining the target prediction result from all the candidate prediction results, the method further includes: In response to the current detection round including multiple first target video frames and second target video frames for detection, the current detection information matching the second target video frame in the current detection round is obtained; For the second target video frame, the reference confidence score matched by each of the reference video frames is updated based on the current detection information and the target prediction information matched by each of the reference video frame sets.

6. The method according to claim 4, characterized in that, The steps for obtaining the reference confidence level include: For each set of reference video frames, the matching degree between historical detection information and historical prediction information corresponding to each historical detection round is obtained; wherein, the matching degree is determined based on the overlapping area of ​​the target object region in the historical detection information and the historical prediction information. Based on the matching degree, successful detection rounds that meet the preset conditions are determined from all the historical detection rounds; Based on the number of successful detection rounds, the reference confidence level corresponding to the set of reference video frames is determined.

7. The method according to claim 1, characterized in that, The step of determining the encoding parameter information matching the target video frame based on the target prediction information includes: Based on the target prediction information, the corresponding target video frame is divided into multiple coding regions; Based on the location information of the encoded regions, the encoding parameters corresponding to each encoded region are obtained.

8. The method according to claim 7, characterized in that, The step of dividing the target video frame into multiple coding regions based on the target prediction information includes: Based on the target prediction information, obtain the first coding region where the target object is located in the corresponding target video frame; and Obtain the second coding region surrounding the first coding region in the target video frame; and, The regions in the target video frame other than the first and second coding regions are used as the third coding region; The step of obtaining the encoding parameters corresponding to each encoding region based on the location information of the encoding region includes: Obtain a first quantization parameter that matches the first coding region, a second quantization parameter that matches the second coding region, and a third quantization parameter that matches the third coding region; wherein the first quantization parameter is less than the second quantization parameter, and the second quantization parameter is less than the third quantization parameter.

9. An electronic device, characterized in that, include: A memory and a processor are coupled to each other, the memory storing program instructions, and the processor executing the program instructions to implement the method as described in any one of claims 1-8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the method as described in any one of claims 1-8.