Video compression method and device
By acquiring the positional relationship between the sensor and the video acquisition device and real-time motion data, the rotation motion vectors of each original image in the video sequence are calculated and motion compensation and image compensation are performed, the problem of high complexity in the calculation of motion estimation and motion compensation in the prior art is solved, and the efficiency and quality of video compression are improved.
Patent Information
- Application Number
- CN202111586946.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-23
AI Technical Summary
In the existing video compression technology, the calculation complexity of motion estimation and motion compensation is high, and is not suitable for deployment on devices with limited computing capabilities, and will lead to matching errors when the image features are not significant.
By acquiring the positional relationship between the sensor and the video acquisition device and real-time motion data, the rotation motion vector of each original image in the video sequence is calculated, and motion compensation and image compensation are performed to form a target video sequence for compression processing.
It improves the effectiveness and accuracy of image stabilization during video compression, reduces the calculation complexity of motion compensation, and improves the efficiency and quality of video compression.
Smart Images

Figure CN114339259B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video compression technology, and in particular to a video compression method and device. Background Art
[0002] Video compression technology can effectively reduce the storage capacity and transmission bandwidth required for video by reducing the spatial and temporal redundancy between images while ensuring picture quality. Electronic image stabilization technology is a modern scientific technology that uses digital image processing and computer application techniques to process jittery video sequences, eliminate irregular random motion in video sequences, retain subjective motion, and thus reduce or completely eliminate video jitter to a certain extent. Electronic image stabilization technology is mainly divided into three parts: motion estimation, motion compensation, and image compensation. The key is to obtain the real-time motion vector between images during motion estimation, and use the motion vector to correct the image during motion compensation to achieve the purpose of reducing the difference between video frames.
[0003] At present, the processing methods of motion estimation and motion compensation in the video compression process are usually as follows: in the motion estimation stage, the feature points of two adjacent frames are identified through FAST feature point detection, and the feature points are homogenized to reduce the calculation time and mismatching phenomenon of the next step of feature point matching; the feature points after rough matching are finely matched to improve the matching accuracy; finally, the optimized feature points are substituted into the motion model to calculate the motion vectors of the two frames. In the motion compensation stage, the current frame is reversely compensated by the filtered motion vector.
[0004] However, the existing motion estimation and motion compensation processing methods in the video compression process use algorithms such as feature point detection to estimate the motion vector of the frame, which has high computational complexity and is not suitable for deployment on devices with limited computing power. Moreover, when the features of the image are not significant enough, matching errors will also occur. Summary of the invention
[0005] In view of this, embodiments of the present application provide a video compression method and apparatus to eliminate or improve one or more defects existing in the prior art.
[0006] One aspect of the present application provides a video compression method, comprising:
[0007] According to the positional relationship between the sensor and the video acquisition device and the real-time motion data acquired by the sensor, the rotation motion vector corresponding to each original image in the original video sequence acquired by the video acquisition device is acquired;
[0008] Performing motion compensation on each of the original images based on the rotational motion vectors corresponding to each of the original images to obtain a target image corresponding to each of the original images;
[0009] Image compensation is performed on each of the target images to obtain a target video sequence composed of the target images that have undergone image compensation, so as to perform compression processing on the target video sequence.
[0010] In some embodiments of the present application, performing image compensation on each of the target images to obtain a target video sequence composed of the image-compensated target images, so as to compress the target video sequence, includes:
[0011] Performing image compensation on each of the target images based on a maximum rectangle interception algorithm to obtain a target video sequence composed of the target images after image compensation;
[0012] The target video sequence is compressed.
[0013] In some embodiments of the present application, obtaining the rotational motion vector corresponding to each original image in the original video sequence captured by the video capture device according to the positional relationship between the sensor and the video capture device and the real-time motion data captured by the sensor includes:
[0014] Receiving real-time motion data collected by the sensor, and acquiring a current initial rotation angle value of the video acquisition device according to a positional relationship between the sensor and the video acquisition device and the real-time motion data collected by the sensor;
[0015] An original video sequence captured by the video capture device is received, and according to the initial rotation angle value, the sampling frequency of the sensor and the sampling frequency of the video capture device, a rotation motion vector corresponding to each original image in the original video sequence is obtained by linear interpolation.
[0016] In some embodiments of the present application, the receiving of real-time motion data collected by the sensor and obtaining the current initial rotation angle value of the video acquisition device according to the positional relationship between the sensor and the video acquisition device and the real-time motion data collected by the sensor include:
[0017] Receive real-time attitude data collected by the sensor, and read angular velocity and time data from the real-time attitude data;
[0018] Performing smoothing filtering on the angular velocity and time data based on a Kalman filter method to obtain a current angle value;
[0019] According to the angle value and the positional relationship between the sensor and the video acquisition device, a current initial rotation angle value of the video acquisition device is determined.
[0020] In some embodiments of the present application, performing motion compensation on each of the original images based on the rotational motion vectors corresponding to each of the original images to obtain a target image corresponding to each of the original images includes:
[0021] Based on the rotation motion vectors corresponding to the original images, lossless rotation correction processing is performed on the original images to obtain target images corresponding to the original images.
[0022] In some embodiments of the present application, performing image compensation on each of the target images based on the maximum rectangle interception algorithm to obtain a target video sequence composed of the target images after image compensation includes:
[0023] From each of the target images, based on a maximum rectangle interception algorithm, an image with the same proportion as the corresponding original image and with the largest area of the defined area is obtained;
[0024] Obtaining a target width based on an average value of widths of each of the target images and the corresponding image with the largest area, and obtaining a target height based on an average value of heights of each of the target images and the corresponding image with the largest area;
[0025] respectively intercepting the target width and the target height from each of the target images as intercepted images;
[0026] The undefined areas in each of the captured images are filled with pixels and restored by interpolation to a result image with the same size as the original image, so as to form a target video sequence composed of each of the result images.
[0027] In some embodiments of the present application, the compressing the target video sequence includes:
[0028] The target video sequence is subjected to H.264 compression encoding processing, and a code stream subjected to H.264 compression numbering processing is output.
[0029] Another aspect of the present application provides a video compression device, comprising:
[0030] A motion estimation module, used to obtain the rotation motion vector corresponding to each original image in the original video sequence collected by the video acquisition device according to the positional relationship between the sensor and the video acquisition device and the real-time motion data collected by the sensor;
[0031] A motion compensation module, used to perform motion compensation on each of the original images based on the rotation motion vectors corresponding to each of the original images, so as to obtain a target image corresponding to each of the original images;
[0032] The image compensation module is used to perform image compensation on each of the target images to obtain a target video sequence composed of the target images after image compensation, so as to perform compression processing on the target video sequence.
[0033] Another aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the video compression method when executing the computer program;
[0034] The electronic device is communicatively connected with a wireless multimedia sensor network and a video acquisition device to receive real-time motion data from sensors in the wireless multimedia sensor network and receive original video sequences from the video acquisition device.
[0035] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the video compression method when executed by a processor.
[0036] The video compression method of the present application obtains the rotational motion vector corresponding to each original image in the original video sequence captured by the video capture device according to the positional relationship between the sensor and the video capture device and the real-time motion data captured by the sensor, and then performs motion compensation on each original image according to the rotational motion vector corresponding to each original image to obtain each target image corresponding to the original video sequence. This method can effectively improve the effectiveness and accuracy of stabilizing each image in the video during the video compression process, effectively reduce the computational complexity of motion compensation, improve the efficiency of motion compensation, and thus effectively improve the effectiveness, accuracy and efficiency of video compression.
[0037] Additional advantages, purposes, and features of the present application will be partially described in the following description, and will become partially apparent to those skilled in the art after studying the following, or may be learned from the practice of the present application. The purposes and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0038] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present application are not limited to the above specific description, and the above and other purposes that can be achieved by the present application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings described herein are used to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. The components in the drawings are not drawn to scale, but are only for the purpose of illustrating the principles of the present application. In order to facilitate the illustration and description of some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:
[0040] Figure 1 This is a schematic diagram of a first flow chart of a video compression method in an embodiment of the present application.
[0041] Figure 2 Schematic diagram of a second flow chart of a video compression method in an embodiment of the present application.
[0042] Figure 3 Schematic diagram of a specific process of a video compression method in one embodiment of the present application.
[0043] Figure 4 A flowchart of a video compression method based on sensor information provided for an application example of this application.
[0044] Figure 5 Schematic diagram of the sensor and camera locations provided for the application example of this application.
[0045] Figure 6 Schematic diagram of the geometric relationship between images with width w<height h provided for the application example of this application.
[0046] Figure 7 Schematic diagram of the geometric relationship between images with width w>=height h provided for the application example of this application.
[0047] Figure 8 Schematic diagram of the geometric relationship between images a, b, and c provided for the application example of this application.
[0048] Fig. 9 It is a structural schematic diagram of a video compression device in another embodiment of the present application.
[0049] Fig.10 It is a schematic diagram of the structure of an electronic device in another embodiment of the present application. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the implementation modes and the accompanying drawings. Here, the illustrative implementation modes and descriptions of the present application are used to explain the present application, but are not intended to limit the present application.
[0051] It should also be noted here that in order to avoid obscuring the present application due to unnecessary details, only the structures and / or processing steps closely related to the scheme according to the present application are shown in the accompanying drawings, while other details that are not very relevant to the present application are omitted.
[0052] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0053] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection with an intermediate.
[0054] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0055] Video compression technology can effectively reduce the storage capacity and transmission bandwidth required for video by reducing the spatial and temporal redundancy between images while ensuring the picture quality. H.264 / AVC is a new generation of video coding standards jointly proposed by ITU-T and ISO / IEC working groups in 2003. Compared with the previously proposed standards, H.264 has the advantages of low bit rate, wide application target range, strong fault tolerance, good image transmission quality and strong network adaptability. The H.264 / AVC coding standard adopts a hybrid coding framework, which is mainly divided into intra-frame, inter-frame prediction, transformation, quantization and entropy coding. In particular, inter-frame prediction coding uses the temporal redundancy of the video to calculate the residual data and related parameters between video frames, and converts them into corresponding codewords in the entropy coding link. The greater the difference between video frames, the more codewords are generated by the coding, and the lower the video compression ratio.
[0056] Electronic image stabilization (EIS) is a modern scientific technology that uses digital image processing and computer application techniques to process jittery video sequences, eliminate irregular random motion in video sequences, retain subjective motion, and thus reduce or completely eliminate video jitter to a certain extent. Electronic image stabilization technology is mainly divided into three links: motion estimation, motion compensation, and image compensation. The key is to obtain the real-time motion vector between images in the motion estimation module, and use the motion vector to correct the image to achieve the purpose of reducing the difference between video frames. The acquisition of motion vectors is mainly through hardware and software. The hardware method is to detect the offset of the camera through sensors and other devices and convert it into the corresponding motion vector. Its operation speed is fast, and the image stabilization accuracy depends on the accuracy of the sensor, and it is not easily disturbed by the motion foreground. The software method is to directly match images through various motion estimation algorithms to obtain the motion offset between frames, such as block matching method, bit plane matching method, feature matching method, etc. The image stabilization result is greatly affected by the image quality, and the complexity of the motion estimation algorithm is relatively high.
[0057] The existing processing methods of motion estimation and motion compensation in the video compression process introduce software-based electronic image stabilization technology into H.264 inter-frame prediction coding, and smooth the video sequence through image stabilization. Specifically, in the motion estimation stage, the feature points of two adjacent frames are identified through FAST feature point detection, and the feature points are homogenized, which reduces the calculation time and mismatching phenomenon of the next feature point matching; the feature points after rough matching are finely matched to improve the matching accuracy; finally, the optimized feature points are substituted into the motion model to calculate the motion vectors of the two frames. In the motion compensation stage, the current frame is reversely compensated by the filtered motion vector. In the image compensation stage, the undefined area around the compensated image is cropped, and the interpolation method is used to transform the size and restore it to the initial size. Then the video after image stabilization is compressed by H.264, which reduces the difference between video frames, thereby reducing the residual data generated by the inter-frame prediction stage of encoding and improving the compression ratio.
[0058] In other words, the existing motion estimation and motion compensation methods use algorithms such as feature point detection to estimate the motion vector of the frame, which has a high computational complexity and is not suitable for deployment on devices with limited computing power. When the features of the image are not significant enough, matching errors will also occur.
[0059] Based on this, an embodiment of the present application provides a video compression method, which obtains the rotational motion vector corresponding to each original image in the original video sequence captured by the video capture device according to the positional relationship between the sensor and the video capture device and the real-time motion data captured by the sensor; performs motion compensation on each original image based on the rotational motion vector corresponding to each original image to obtain a target image corresponding to each original image; performs image compensation on each target image to obtain a target video sequence composed of each image-compensated target image, so as to compress the target video sequence, which can effectively improve the effectiveness and accuracy of stabilizing each image in the video during the video compression process, and can effectively reduce the computational complexity of motion compensation, improve the efficiency of motion compensation, and thus effectively improve the effectiveness, accuracy and efficiency of video compression.
[0060] In one or more embodiments of the present application, the video acquisition device may be a video acquisition device such as an independent camera or a client device with a video acquisition function.
[0061] In one or more embodiments of the present application, the sensor may be any sensor node in a wireless multimedia sensor network WMSNs.
[0062] In one or more embodiments of the present application, a wireless multimedia sensor network WMSNs (Wireless multimedia sensor networks) includes a large number of sensor nodes, which have limited battery capacity, computing power and storage capacity. The nodes communicate with each other through wireless networks to complete real-time or non-real-time detection and transmission of scalar data and multimedia data. Deploying WMSNs systems in environments that need to be monitored, such as oceans and vehicles, can achieve device interconnection and information integration. Video information is the main body of multimedia data services. In the actual process of video data collection, the carrier device of the collected data will be affected by the real environment. For example, the fluctuation of sea waves will cause the buoy device to shake, resulting in large differences between the collected video images. The compression rate of the video is reduced, but it will still generate a large amount of data, which will put a certain pressure on the data transmission.
[0063] Based on the above content, the present application also provides a video compression device for implementing the video compression method provided in one or more embodiments of the present application. The video compression device can be a processor, a server or a controller, etc. The video compression device can communicate and connect with each sensor node in the wireless multimedia sensor network WMSNs in sequence by itself or through a third-party server, etc. to receive real-time motion data collected by the sensor nodes; the video compression device can also communicate and connect with a video acquisition device by itself or through a third-party server, etc. to receive the original video sequence collected by the video acquisition device; the video compression device can also communicate and connect with a client device by itself or through a third-party server, etc. to send the compressed code stream to the client device for storage and / or playback, etc.
[0064] The part of the video compression device that performs video compression can be executed in a server, controller or processor as described above, and in another practical application scenario, all operations can also be completed in the client device. The specific selection can be based on the processing capability of the client device and the limitations of the user's usage scenario. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor for specific processing of video compression.
[0065] It is understandable that the client device may include any mobile device capable of loading applications, such as a smart phone, a tablet electronic device, a network set-top box, a portable computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0066] The client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.
[0067] The server and the client device may communicate with each other using any suitable network protocol, including network protocols that have not yet been developed on the date of filing this application. The network protocols may include, for example, TCP / IP, UDP / IP, HTTP, HTTPS, etc. Of course, the network protocols may also include, for example, RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer) protocols used on top of the above protocols.
[0068] The details are described in detail through the following embodiments and application examples.
[0069] In order to reduce the computational complexity of motion estimation during video compression and improve the accuracy of motion compensation, the present application provides an embodiment of a video compression method, see Figure 1 The video compression method executed by the video compression device specifically includes the following contents:
[0070] Step 100: According to the positional relationship between the sensor and the video acquisition device and the real-time motion data acquired by the sensor, the rotation motion vector corresponding to each original image in the original video sequence acquired by the video acquisition device is acquired.
[0071] It is understandable that the positional relationship between the sensor and the video acquisition device may be pre-acquired and stored locally in the video compression device, and the sensor may be a sensor node in a wireless multimedia sensor network WMSNs.
[0072] Step 200: performing motion compensation on each of the original images based on the rotational motion vectors corresponding to each of the original images to obtain a target image corresponding to each of the original images.
[0073] It can be understood that each of the original images refers to an original image frame in an original video sequence, and the target image refers to an image obtained after motion compensation of each of the original images.
[0074] Step 300: Perform image compensation on each of the target images to obtain a target video sequence composed of each of the image-compensated target images, so as to compress the target video sequence. From the above description, it can be seen that the video compression method provided in the embodiment of the present application obtains the rotational motion vector corresponding to each of the original images in the original video sequence captured by the video acquisition device according to the positional relationship between the sensor and the video acquisition device and the real-time motion data captured by the sensor; performs motion compensation on each of the original images according to the rotational motion vector corresponding to each of the original images to obtain each of the target images corresponding to the original video sequence, which can effectively improve the effectiveness and accuracy of stabilizing each image in the video during the video compression process, and can effectively reduce the computational complexity of motion compensation, improve the efficiency of motion compensation, and thus effectively improve the effectiveness, accuracy and efficiency of video compression.
[0075] Based on the above content, it can be known that in the existing video compression method, the undefined area around the compensated image is cropped during the image compensation process, and the size is transformed and restored to the initial size using the interpolation method. However, when the jitter is more severe, the undefined area of the image after motion compensation increases. If the direct cropping method is adopted, a large amount of image edge information will be lost, affecting the image quality, which is not conducive to the subsequent video compression encoding and real-time monitoring of the environment.
[0076] Based on this, on the basis of reducing the computational complexity of motion estimation and improving the accuracy of motion compensation during video compression, in order to further solve the technical problem that edge information is easily lost in image compensation during video compression, so as to avoid the loss of edge information in image compensation and improve the accuracy of image compensation during video compression, in an embodiment of a video compression method provided by the present application, see Figure 2 , the step 300 specifically includes the following contents:
[0077] Step 310: performing image compensation on each of the target images based on a maximum rectangle interception algorithm to obtain a target video sequence composed of the target images after image compensation;
[0078] Step 320: compress the target video sequence.
[0079] It can be understood that the specific method of compressing the target video sequence in step 320 can be to compress the target video sequence using a video compression coding standard, wherein the video compression coding standard can be H.264, H.265 and AVS, etc., which can be selected according to the actual application scenario.
[0080] From the above description, it can be seen that the video compression method provided in the embodiment of the present application, by performing image compensation on each of the target images based on the maximum rectangle interception algorithm, can effectively avoid the loss of edge information in the image compensation during the video compression process, can effectively improve the accuracy of image compensation, and thus can further improve the accuracy and reliability of video compression, especially for situations where the jitter is more severe during video acquisition, can effectively guarantee the image quality, and is more conducive to the accuracy of subsequent video compression encoding and real-time monitoring of the environment.
[0081] In order to further improve the efficiency and accuracy of obtaining the rotation motion vector corresponding to each original image in the original video sequence, in one embodiment of the video compression method provided in the present application, step 100 of the video compression method specifically includes the following contents:
[0082] Step 110: receiving real-time motion data collected by the sensor, and acquiring a current initial rotation angle value of the video acquisition device according to a positional relationship between the sensor and the video acquisition device and the real-time motion data collected by the sensor;
[0083] Step 120: Receive the original video sequence captured by the video acquisition device, and obtain the rotation motion vector corresponding to each original image in the original video sequence by linear interpolation according to the initial rotation angle value, the sampling frequency of the sensor and the sampling frequency of the video acquisition device.
[0084] Specifically, the rotation angle value of the camera in the video frame time can be estimated using linear interpolation technology based on the sampling frequency of the sensor and the camera, and then the rotation motion vector of the original image can be obtained through the rotation angle of the camera in the previous step.
[0085] From the above description, it can be seen that the video compression method provided in the embodiment of the present application uses a linear interpolation method to obtain the rotational motion vector corresponding to each original image in the original video sequence according to the initial rotation angle value, the sampling frequency of the sensor, and the sampling frequency of the video acquisition device, which can further improve the efficiency and accuracy of obtaining the rotational motion vector corresponding to each original image in the original video sequence.
[0086] In order to improve the accuracy and effectiveness of determining the current initial rotation angle value of the video acquisition device, in one embodiment of the video compression method provided by the present application, in one embodiment of the video compression method provided by the present application, see Figure 3 , the step 110 specifically includes the following contents:
[0087] Step 111: Receive real-time posture data collected by the sensor, and read angular velocity and time data from the real-time posture data.
[0088] Step 112: Perform smoothing filtering on the angular velocity and time data based on a Kalman filter method to obtain a current angle value.
[0089] Step 113: Determine the current initial rotation angle value of the video acquisition device according to the angle value and the positional relationship between the sensor and the video acquisition device.
[0090] Specifically, the sensor is initialized first, the port number and sampling frequency of the built-in attitude sensor of the device are set, and when the sensor outputs a frame of data, the angular velocity and time information are read from the data. Then, the collected data is smoothed and filtered using Kalman filtering technology to obtain the accurate angle value of the current data time, and then the rotation angle value of the camera at the current data time is derived based on the angle of the previous step and the position relationship between the camera and the sensor.
[0091] From the above description, it can be seen that the video compression method provided in the embodiment of the present application, by performing smoothing filtering on the angular velocity and time data based on the Kalman filtering method to obtain the current angle value, can effectively improve the accuracy and effectiveness of determining the current initial rotation angle value of the video acquisition device based on the angle value and the positional relationship between the sensor and the video acquisition device, and thus can further improve the effectiveness and accuracy of stabilizing each image in the video during the video compression process.
[0092] In order to improve the accuracy and reliability of motion compensation in the video compression process, in one embodiment of the video compression method provided by the present application, in one embodiment of the video compression method provided by the present application, see Figure 3 , the step 200 specifically includes the following contents:
[0093] Step 210: performing lossless rotation correction processing on each of the original images based on the rotation motion vectors corresponding to each of the original images, so as to obtain a target image corresponding to each of the original images.
[0094] From the above description, it can be seen that the video compression method provided in the embodiment of the present application can effectively improve the accuracy and reliability of motion compensation in the video compression process by performing lossless rotation correction processing on each of the original images to obtain a target image corresponding to each of the original images, so as to further improve the accuracy and reliability of video compression.
[0095] In order to improve the accuracy and effectiveness of image compensation for each of the target images, in one embodiment of the video compression method provided by the present application, see Figure 3, the step 310 specifically includes the following contents:
[0096] Step 311: From each of the target images, based on a maximum rectangle interception algorithm, obtain a maximum area image that has the same proportion as the corresponding original image and is a defined area.
[0097] Step 312: Obtain the target width based on the average width of each of the target images and the corresponding maximum area image, and obtain the target height based on the average height of each of the target images and the corresponding maximum area image.
[0098] Step 313: respectively cut out the target width and the target height from each of the target images as cut-out images.
[0099] Step 314: Fill the undefined areas in each of the captured images with pixels, and interpolate and restore them into result images with the same size as the original images, so as to form a target video sequence composed of each of the result images.
[0100] Specifically, for the motion compensated image b, the maximum rectangle interception algorithm is used to calculate the image with the same proportion as image a and the largest area that does not contain the undefined area, which is recorded as image c; a part of the image is intercepted again in image b, and its width is the average width of image b and image c, and its height is the average height of image b and image c, which is image d; the undefined area around image d is filled with pixels, and interpolated and restored to the size of image a to obtain the result image.
[0101] From the above description, it can be seen that the video compression method provided in the embodiment of the present application obtains the maximum area image with the same proportion as the corresponding original image and the largest area of the defined area from each of the target images based on the maximum rectangle interception algorithm, and intercepts the target width and the target height from each of the target images as the intercepted image. This can effectively improve the accuracy and effectiveness of image compensation for each of the target images, especially for situations where the jitter is more severe during video acquisition, and can further ensure the image quality, which is more conducive to the accuracy of subsequent video compression encoding and real-time monitoring of the environment.
[0102] In order to reduce the residual data generated by the inter-frame prediction link of the encoding, in one embodiment of the video compression method provided by the present application, in one embodiment of the video compression method provided by the present application, see Figure 3 , the step 320 specifically includes the following contents:
[0103] Step 321: Perform H.264 compression encoding processing on the target video sequence, and output a code stream after H.264 compression numbering processing.
[0104] From the above description, it can be seen that the video compression method provided in the embodiment of the present application can effectively reduce the difference between video frames by performing H.264 compression encoding processing on the target video sequence, thereby reducing the residual data generated by the inter-frame prediction link of the encoding and improving the compression ratio.
[0105] Because the carrier device will shake under the interference of the monitoring environment, the collected video images will have large differences, which is not conducive to the subsequent video compression and data transmission. The jitter of the carrier device will cause a certain translation and rotation of the image, among which the rotation factor has a greater impact on the subsequent environmental image monitoring.
[0106] Based on this, for the embodiment of the above-mentioned video compression method, the present application also provides a specific application example of the video compression method for further explanation. The application example of the present application proposes a video compression method based on sensor information.
[0107] A video compression optimization method based on sensor information is proposed. The real-time motion data of the device is obtained through the sensors in WMSNs, and the motion data is processed and converted into the rotation motion vector of each frame image in the camera video. The current image is reversely compensated according to the motion vector, and more image information is retained as much as possible to obtain a relatively stable inter-frame sequence. The processed video is then compressed with H.264 to reduce the coding block residual generated during the video inter-frame prediction coding process and reduce the amount of coding data generated to adapt to the transmission bandwidth.
[0108] See also Figure 4 , the video compression method based on sensor information specifically includes the following contents:
[0109] 1. Sensor data acquisition:
[0110] 1) Initialize the sensor and set the port number and sampling frequency of the built-in attitude sensor of the device.
[0111] 2) When the sensor outputs a frame of data, the angular velocity and time information are read from the data.
[0112] 2. Sensor data preprocessing:
[0113] 1) By using Kalman filtering technology, the collected data is smoothed and filtered to obtain the accurate angle value of the current data time.
[0114] 2) Based on the angle in the previous step and the positional relationship between the camera and the sensor, derive the rotation angle value of the camera at the current data time.
[0115] 3. Motion Estimation:
[0116] 1) According to the sampling frequency of the sensor and camera, use linear interpolation technology to estimate the rotation angle value of the camera in the video frame time.
[0117] 2) The rotation motion vector of the original image is obtained through the rotation angle of the camera in the previous step, and the original image is recorded as image a.
[0118] 4. Motion compensation: Perform lossless rotation correction on the image according to the rotation motion vector, and the new image is recorded as image b.
[0119] 5. Image compensation:
[0120] 1) For the motion compensated image b, a maximum rectangle interception algorithm is used to calculate an image with the same proportion as image a and the largest area that does not contain undefined areas, which is recorded as image c.
[0121] 2) Cut out a portion of the image from image b again, whose width is the average of the widths of image b and image c, and whose height is the average of the heights of image b and image c, as image d.
[0122] 3) Fill the undefined area around image d with pixels and interpolate to restore it to the size of image a to obtain the result image.
[0123] 6. H.264 compression: Perform H.264 compression encoding on the video sequence after image stabilization and output the encoded bit stream.
[0124] Based on this, taking the buoy equipment of WMSNs deployed on the ocean as an example, the video compression method based on sensor information provided by the application example of this application is specifically embodied in the following example:
[0125] In offshore buoy equipment, there is a certain positional correlation between digital attitude sensors and cameras:
[0126] The coordinate system of the camera includes x, y, and z axes, where the camera shooting direction is defined as the z axis, the direction perpendicular to the camera is defined as the x axis, and the direction parallel to the camera is defined as the y axis.
[0127] The sensor's coordinate system also includes the x, y, and z axes, where the heading angle direction is defined as the x-axis, the roll angle direction is defined as the z-axis, and the pitch angle direction is defined as the y-axis.
[0128] The sensor equipment and camera on the buoy are installed in parallel. The position relationship between the two is as follows: Figure 5 shown.
[0129] At the same time, when the sensor generates an angle around its own z-axis When the roll angle is , the angle generated by the camera rotating around the z-axis is The image captured by the camera is rotated by an angle of θ, and the relationship between the three angles is given by Figure 5 It can be seen that the following formula (1) is:
[0130]
[0131] Based on the above background, the process of the video compression method based on sensor information is further described in detail:
[0132] S1: Sensor data acquisition:
[0133] <1> Before collecting motion data, the attitude sensor needs to set some parameters, including the port number of the connection and the sampling frequency, i.e., the baud rate. The higher the baud rate, the faster the sampling frequency and the higher the output data accuracy.
[0134] <2> The sensor outputs data in real time at a baud rate. One frame of data contains two data packets (acceleration packet and angle packet) and system time. According to the background information, the current time and angular velocity in the roll angle direction need to be obtained from the data packet.
[0135] S2: Sensor data preprocessing:
[0136] <1> Since the marine environment is unstable and changeable, the z-axis angular velocity obtained from the data packet may have noise errors, and the angle obtained by direct integration calculation is not accurate enough. The Kalman filter technology is introduced to optimally estimate the collected angular velocity data, remove the influence of noise, obtain a relatively smooth data value, and then obtain the rotation angle of the sensor around the z-axis through integration calculation.
[0137] <2> According to the rotation angle of the sensor and formula (1), the rotation angle α of the camera around the z-axis at the sensor data time is obtained.
[0138] S3: Motion Estimation:
[0139] <1> Since the sampling frequency of the attitude sensor is different from that of the camera, the video frame at a certain moment may not have a corresponding angle value. In order to obtain a more accurate effect, linear interpolation technology is required to estimate the rotation angle of the camera around the z-axis at the current video frame time. angle The conversion of and α is shown in formula (2), where (t1, α1), (t2, α2) represent the angle values α1 and α2 obtained by the sensor at time t1 and t2. Indicates the rotation angle of the video image at time t
[0140]
[0141] <2> According to the rotation angle of the camera and formula (1), the rotation motion vector of the current image, that is, the rotation angle θ, is obtained.
[0142] S4: Motion Compensation:
[0143] After the current image is losslessly rotated according to formula (3), the width and height of the new image are calculated, and then the original image is losslessly rotated according to formula (4). Where w is the width of the original image a, w New is the width of the new image, h is the height of the original image a, h New is the height of the new image, (c x ,c y ) are the coordinates of the center point of the new image, (x1, y1) represent the coordinates of a pixel in the original image, (x2, y2) represent the coordinates of the corresponding pixel in the rotated image, and θ is the rotation angle. The new image after lossless rotation is recorded as image b.
[0144] w New =h*sinθ+w*cosθ
[0145] h New =w*sinθ+h*cosθ Formula (3)
[0146]
[0147] S5: Image compensation:
[0148] <1> The motion compensated image b has some undefined areas around it, i.e., black borders. The first step of image compensation uses the maximum rectangle interception algorithm to intercept the largest area image in image b that is the same proportion as the original image a and does not contain black borders.
[0149] The maximum rectangle interception algorithm is analyzed from two cases. One case is that the width w of the original image a is less than the height h. The geometric relationship between the images is as follows: Figure 6 As shown, w′ is the width value of the maximum area image, and h′ is the height value of the maximum area image.
[0150] Similarly, for the case where the width value w of the original image a is greater than or equal to the height value h, the image geometric relationship is as follows: Figure 7 As shown, formula (5) calculates the length OA = OB + BA in two cases, where
[0151]
[0152]
[0153] Formula (5) is derived to obtain the calculation method of the ratio, as shown in formula (6), and the image with the largest area is recorded as image c.
[0154] k=r*sinθ+cosθ
[0155]
[0156]
[0157] <2> When the offshore buoy equipment shakes slightly, the generated image rotation vector is relatively small, so the image c obtained in the previous step does not contain the undefined area generated by the rotation correction, and the information loss of the image edge is small. However, when the offshore buoy equipment is disturbed by bad weather and the generated rotation vector is large, although the black border is removed in image c, the image information around it is lost more.
[0158] The geometric relationship between images a, b, and c is as follows Figure 8 As shown, here we take the case where the width value w of the original image a is smaller than the height value h.
[0159] Depend on Figure 8 It can be seen that image b contains all the information of the original image a, and there is a black border with the largest area around it, while image c contains part of the information of the original image a, but there is no black border around it. In order to retain more image information as much as possible, the width value of image b and image c is taken as the average Average height The image is intercepted and the obtained image is recorded as image d. The undefined area around image d is filled with pixels and interpolated to restore it to the size of image a to obtain the result image.
[0160] S6: H.264 compression: Use the x264 encoder to perform video compression on the rotation-corrected video image and output an encoded bit stream.
[0161] In summary, the application example of this application applies the information collected by WMSNs sensors to the compression and coding optimization of video data. Specifically, the real-time motion data of the device is obtained through sensor information, and the motion data is processed and converted into the rotation motion vector of each frame image in the camera video; the current image is reversely compensated according to the motion vector, and more image information is retained as much as possible to obtain a relatively stable video sequence; the processed video is then compressed with H.264 video to reduce the coding block residual generated during the inter-frame prediction coding process and reduce the amount of coding data generated to adapt to the transmission bandwidth.
[0162] The application example of this application applies the information collected by the WMSNs sensor to the compression coding optimization of the video data. In order to better adapt to the unstable monitoring environment, the present invention processes and applies the acquired sensor information to assist in the compression coding of the video data, thereby improving the compression rate of the video while ensuring the quality of the video as much as possible to adapt to the bandwidth of data transmission.
[0163] Based on the above content, the present application also provides a video compression device for implementing the video compression method provided in one or more embodiments of the present application. The specific implementation of the video compression device can be a processor, a controller or a server. In a specific example, see Fig. 9 , the video compression device specifically includes the following contents:
[0164] The motion estimation module 10 is used to obtain the rotation motion vector corresponding to each original image in the original video sequence captured by the video capture device according to the positional relationship between the sensor and the video capture device and the real-time motion data captured by the sensor.
[0165] The motion compensation module 20 is used to perform motion compensation on each of the original images based on the rotation motion vectors corresponding to each of the original images, so as to obtain a target image corresponding to each of the original images.
[0166] The image compensation module 30 is used to perform image compensation on each of the target images to obtain a target video sequence composed of the image-compensated target images, so as to perform compression processing on the target video sequence.
[0167] The embodiment of the video compression device provided in the present application can be specifically used to execute the processing flow of the embodiment of the video compression method in the above-mentioned embodiment. Its functions will not be repeated here, and reference can be made to the detailed description of the above-mentioned video compression method embodiment.
[0168] From the above description, it can be seen that the video compression device provided in the embodiment of the present application obtains the rotational motion vector corresponding to each original image in the original video sequence captured by the video acquisition device according to the positional relationship between the sensor and the video acquisition device and the real-time motion data captured by the sensor; and performs motion compensation on each original image according to the rotational motion vector corresponding to each original image to obtain each target image corresponding to the original video sequence. This can effectively improve the effectiveness and accuracy of stabilizing each image in the video during the video compression process, and can effectively reduce the computational complexity of motion compensation, improve the efficiency of motion compensation, and thus effectively improve the effectiveness, accuracy and efficiency of video compression.
[0169] The embodiment of the present invention further provides a computer device (ie, electronic device), such as Fig.10 As shown, the computer device may include a processor 810, a memory 820 and an image acquisition device 830, wherein the processor 810 and the memory 820 may be connected via a bus or other means. Fig.10 The bus connection is taken as an example. The image acquisition device 830 can be connected to the processor 810 and the memory 820 by wire or wirelessly. The computer device is connected to the wireless multimedia sensor network and the video acquisition device in communication to receive real-time motion data from the sensors in the wireless multimedia sensor network and receive the original video sequence from the video acquisition device.
[0170] The processor 810 may be a central processing unit (CPU). The processor 810 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0171] The memory 820 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the vehicle display device key shielding method in the embodiment of the present invention. The processor 810 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory 820, that is, implementing the method in the above method embodiment.
[0172] The memory 820 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created by the processor 810, etc. In addition, the memory 820 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 820 may optionally include a memory remotely arranged relative to the processor 810, and these remote memories may be connected to the processor 810 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0173] The one or more modules are stored in the memory 820, and when executed by the processor 810, the video compression method in the embodiment is performed.
[0174] In some embodiments of the present disclosure, the user equipment may include a processor, a memory and a transceiver unit, which may include a receiver and a transmitter. The processor 810, the memory 820, the receiver 830 and the transmitter 840 may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.
[0175] As an implementation method, the functions of the receiver and the transmitter in the present invention may be implemented by a transceiver circuit or a dedicated chip for transceiver, and the processor may be implemented by a dedicated processing chip, a processing circuit or a general-purpose chip.
[0176] As another implementation, it is possible to use a general-purpose computer to implement the server provided in the embodiment of the present invention, that is, the program code for implementing the functions of the processor, receiver, and transmitter is stored in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.
[0177] The embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned video compression method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0178] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0179] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0180] In the present application, features described and / or illustrated for one embodiment may be used in the same manner or in a similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.
[0181] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A video compression method, characterized in that: include: According to the positional relationship between the sensor and the video acquisition device and the real-time motion data acquired by the sensor, the rotation motion vector corresponding to each original image in the original video sequence acquired by the video acquisition device is acquired; Performing motion compensation on each of the original images based on the rotation motion vectors corresponding to each of the original images to obtain a target image corresponding to each of the original images; Performing image compensation on each of the target images based on a maximum rectangle interception algorithm to obtain a target video sequence composed of the target images after image compensation; Performing compression processing on the target video sequence; The method of performing image compensation on each of the target images based on the maximum rectangle interception algorithm to obtain a target video sequence composed of the target images after image compensation includes: From each of the target images, based on a maximum rectangle interception algorithm, an image with the same proportion as the corresponding original image and with the largest area of the defined area is obtained; Obtaining a target width based on an average value of widths of each of the target images and the corresponding image with the largest area, and obtaining a target height based on an average value of heights of each of the target images and the corresponding image with the largest area; respectively intercepting the target width and the target height from each of the target images as intercepted images; The undefined areas in each of the captured images are filled with pixels and restored by interpolation to a result image with the same size as the original image, so as to form a target video sequence composed of each of the result images.
2. The video compression method according to claim 1, characterized in that: The step of obtaining the rotational motion vectors corresponding to the respective original images in the original video sequence collected by the video collection device according to the positional relationship between the sensor and the video collection device and the real-time motion data collected by the sensor includes: Receiving real-time motion data collected by the sensor, and acquiring a current initial rotation angle value of the video acquisition device according to a positional relationship between the sensor and the video acquisition device and the real-time motion data collected by the sensor; An original video sequence captured by the video capture device is received, and according to the initial rotation angle value, the sampling frequency of the sensor and the sampling frequency of the video capture device, a rotation motion vector corresponding to each original image in the original video sequence is obtained by linear interpolation.
3. The video compression method according to claim 2, characterized in that: The receiving of real-time motion data collected by the sensor and obtaining the current initial rotation angle value of the video acquisition device according to the positional relationship between the sensor and the video acquisition device and the real-time motion data collected by the sensor include: Receive real-time attitude data collected by the sensor, and read angular velocity and time data from the real-time attitude data; Performing smoothing filtering on the angular velocity and time data based on a Kalman filter method to obtain a current angle value; According to the angle value and the positional relationship between the sensor and the video acquisition device, a current initial rotation angle value of the video acquisition device is determined.
4. The video compression method according to claim 1, characterized in that: The performing motion compensation on each of the original images based on the rotational motion vectors corresponding to each of the original images to obtain the target images corresponding to each of the original images includes: Based on the rotation motion vectors corresponding to the original images, lossless rotation correction processing is performed on the original images to obtain target images corresponding to the original images.
5. The video compression method according to claim 1, characterized in that: The compressing the target video sequence comprises: The target video sequence is subjected to H.264 compression encoding processing, and a code stream subjected to H.264 compression numbering processing is output.
6. A video compression device, characterized in that: include: A motion estimation module, used to obtain the rotation motion vector corresponding to each original image in the original video sequence collected by the video acquisition device according to the positional relationship between the sensor and the video acquisition device and the real-time motion data collected by the sensor; A motion compensation module, used to perform motion compensation on each of the original images based on the rotation motion vectors corresponding to each of the original images, so as to obtain a target image corresponding to each of the original images; An image compensation module is used to perform image compensation on each of the target images based on a maximum rectangle interception algorithm to obtain a target video sequence composed of the target images after image compensation; and to perform compression processing on the target video sequence; The method of performing image compensation on each of the target images based on the maximum rectangle interception algorithm to obtain a target video sequence composed of the target images after image compensation includes: From each of the target images, based on a maximum rectangle interception algorithm, an image with the same proportion as the corresponding original image and with the largest area of the defined area is obtained; Obtaining a target width based on an average value of widths of each of the target images and the corresponding image with the largest area, and obtaining a target height based on an average value of heights of each of the target images and the corresponding image with the largest area; respectively intercepting the target width and the target height from each of the target images as intercepted images; The undefined areas in each of the captured images are filled with pixels and restored by interpolation to a result image with the same size as the original image, so as to form a target video sequence composed of each of the result images.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the video compression method according to any one of claims 1 to 5 is implemented; The electronic device is communicatively connected with a wireless multimedia sensor network and a video acquisition device to receive real-time motion data from sensors in the wireless multimedia sensor network and receive original video sequences from the video acquisition device.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video compression method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method for electronic image stabilization of video on mobile terminal
CN104902142A