Image coding method and device
By obtaining the foreground position information of the image to be encoded and the difference information from the reference image, and generating a difference map and a segmented map for encoding, the problem of high-resolution image storage cost in the monitoring scene is solved, and the storage space saving and image quality guarantee is achieved.
Patent Information
- Application Number
- CN202410073240.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
In security protection monitoring scenarios, the storage of high-resolution images leads to high storage costs and waste of code rates, and the prior art has not effectively solved it.
By obtaining the foreground position information of the image to be encoded and the difference information from the reference image, a difference value diagram and a segmentation diagram are generated, and encoding and storage are performed to reduce the duplicate encoding of the background area and save the code rate.
It reduces storage costs, reduces storage space requirements, and ensures lossless quality of prospective goals, improving image encoding efficiency and quality.
Smart Images

Figure CN120343276A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image encoding method and device. Background Art
[0002] In the monitoring scenarios of security protection, cameras deployed can instantaneously capture passing vehicles and pedestrians, and upload these captured images to a storage server for storage. When the platform side needs to trace back key events and perform big data analysis, the required images can be obtained from the storage server. With the continuous progress of information technology, the resolution of cameras is getting higher and higher, which results in a large number of high-quality and high-resolution images that need to be stored, posing challenges to storage technologies. For example, in the monitoring scenarios of security protection, if it is required to count the number of passing vehicles in a place in a day, then each camera in the monitoring scenario needs to capture at least 5000 images per day. Taking an image with a resolution of 4K and a size of 2MB as an example, then at least about 10G of storage space is required per day to store the images captured by each camera. In addition, there is a requirement for a 90-day cache of the captured images in the monitoring scenarios of security protection. Thus, more storage space is required, resulting in a relatively high storage cost.
[0003] To reduce the storage cost, in the prior art, each image captured by a camera can be separately encoded before storage. However, the prior art cannot well solve the problem of large storage cost, and there is also a problem of waste of bitrate. Summary of the Invention
[0004] The embodiments of this application provide an image encoding method and device, which solve the problems of large storage cost and waste of bitrate caused by separately encoding each image in a continuous scenario.
[0005] To achieve the above object, the embodiments of this application provide the following technical solutions:
[0006] In a first aspect, an image encoding method is provided. The method may include:
[0007] First, obtain a first image to be encoded; based on the first image, obtain a second image, where the second image is used to indicate the position information of the foreground included in the first image; based on the first image, a reference image, and the second image, obtain a third image, where the third image includes: the foreground of the first image and the difference information between the first image and the reference image; encode the second image and the third image to obtain a first bitstream, and store the first bitstream as the encoding information of the first image.
[0008] By adopting the above technical solution, the difference information between the image to be encoded and the reference image and the segmentation map of the image to be encoded are encoded and stored as the encoding information of the image to be encoded, which not only reduces the overhead of repeated encoding in the background area, saves the bit rate, but also reduces the storage space required for storing the image and the storage cost. In addition, the background of the reference image, the difference information between the image to be encoded and the reference image, and the segmentation map of the image to be encoded can be directly used to reconstruct the image to be encoded, which can ensure the quality of the foreground object is lossless to a certain extent, and thus ensure the quality of the restored image to be encoded.
[0009] In a possible implementation manner in combination with the first aspect, obtaining the third image based on the first image, the reference image and the second image specifically includes: obtaining the difference information between the first image and the reference image based on the first image and the reference image; obtaining the third image based on the difference information and the second image. The difference information may include the different information in the foreground and the background between the first image and the reference image. In this way, the changed parts in the image can be effectively extracted, and the efficiency and quality of image encoding are greatly improved.
[0010] In a possible implementation manner in combination with the first aspect, obtaining the difference information between the first image and the reference image based on the first image and the reference image specifically includes: determining the pixel difference between the first image and the reference image to obtain a difference result; performing normalization processing on the difference result to obtain the difference information. Exemplarily, the normalization processing of the difference result can be performed according to a predetermined range, and the pixel difference of each pixel point obtained is normalized to the predetermined range. In this way, it helps to improve the efficiency of the algorithm, reduce the storage requirements, ensure the accuracy of color display, and be compatible with various standards and tools.
[0011] In a possible implementation manner in combination with the first aspect, obtaining the second image based on the first image specifically includes: inputting the first image into a segmentation model to obtain the second image; the segmentation model has the function of segmenting the foreground and background of the input image. Exemplarily, after the first image is input into the segmentation model, the segmentation model can output a segmentation map for indicating the position information of the foreground and / or the background in the first image. In this way, the foreground object and the background in the image can be accurately separated through the segmentation model, and the algorithm efficiency and accuracy can be improved.
[0012] In a possible implementation manner in combination with the first aspect, it further includes: decoding the first bitstream to obtain the second image and the third image; decoding the second bitstream of the reference image to obtain the reference image, where the second bitstream is the bitstream obtained after encoding the reference image; restoring the first image based on the reference image, the second image and the third image. In this way, the quality of the foreground object can be ensured to be lossless, and the quality of the image is guaranteed.
[0013] In combination with the first aspect, in a possible implementation, the second image is further used to indicate the position information of the background included in the first image. Based on the reference image, the second image, and the third image, the first image is restored, which specifically includes: according to the position information of the background included in the first image indicated by the second image, adding the pixels at the corresponding positions in the reference image to the pixels at the corresponding positions in the third image to restore the first image. In this way, the overhead of redundant coding in the background area can be reduced, the quality of the foreground object can be ensured without loss, and the quality of the image can be guaranteed.
[0014] In combination with the first aspect, in a possible implementation, the above reference image and the first image are images obtained by photographing the same scene.
[0015] Exemplarily, the technical solution provided by the first aspect can be applied to a device with image encoding and decoding capabilities, such as a storage device.
[0016] In a second aspect, an image encoding method is provided. The method is applied to a first device and may include:
[0017] First, obtain a first image to be encoded; based on the first image, obtain a second image, where the second image is used to indicate the position information of the foreground included in the first image; based on the first image, the reference image, and the second image, obtain a third image, where the third image includes: the foreground of the first image and the difference information between the first image and the reference image; encode the second image and the third image to obtain a first bitstream, and send the first bitstream to a storage device. Among them, the above first device may be a photographing device.
[0018] Among them, the above reference image and the first image are images obtained by photographing the same scene.
[0019] In combination with the second aspect, in a possible implementation, the above obtaining the third image based on the first image, the reference image, and the second image specifically includes: obtaining the difference information between the first image and the reference image based on the first image and the reference image; obtaining the third image based on the difference information and the second image.
[0020] In combination with the second aspect, in a possible implementation, the above obtaining the difference information between the first image and the reference image based on the first image and the reference image specifically includes: determining the pixel difference between the first image and the reference image to obtain a difference result; performing normalization processing on the difference result to obtain the difference information.
[0021] In combination with the second aspect, in a possible implementation, the above obtaining the second image based on the first image specifically includes: inputting the first image into a segmentation model to obtain the second image; the segmentation model has the function of segmenting the foreground and background of the input image.
[0022] In a third aspect, an image decoding method is provided, which is applied to a second device. The method may include: receiving and decoding a first bitstream to obtain a second image and a third image; receiving and decoding a second bitstream of a reference image to obtain the reference image, where the second bitstream is a bitstream obtained by encoding the reference image; and restoring and displaying a first image based on the reference image, the second image, and the third image. The second device may be a display device.
[0023] In a possible implementation manner in combination with the third aspect, the second image is further used to indicate the position information of the background included in the first image. Restoring the first image based on the reference image, the second image, and the third image specifically includes: adding the pixels at the corresponding positions in the reference image and the pixels at the corresponding positions in the third image according to the position information of the background included in the first image indicated by the second image to restore the first image.
[0024] In a possible implementation manner in combination with the third aspect, the above reference image and the first image are images obtained by photographing the same scene.
[0025] In a fourth aspect, a device is provided, and the device has the function of implementing the behavior of the electronic device in the method described in the first aspect above. The function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, for example, a communication unit or module, a processing unit or module, and a storage unit or module.
[0026] In a fifth aspect, an electronic device is provided. The electronic device includes: a processor; a memory; and a computer program, where the computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device is caused to execute the method described in any one of the first aspect, the second aspect, or the third aspect above.
[0027] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the electronic device can be caused to execute the method described in any one of the first aspect, the second aspect, or the third aspect above.
[0028] In a seventh aspect, a computer program product containing instructions is provided, and when it runs on an electronic device, the electronic device can be caused to execute the method described in any one of the first aspect, the second aspect, or the third aspect above.
[0029] In an eighth aspect, an embodiment of the present application provides a chip. The chip includes a processor, and the processor is used to call a computer program in a memory to execute the method described in any one of the first aspect, the second aspect, or the third aspect.
[0030] In a ninth aspect, a parallel configuration system for operators is provided. When it runs on an electronic device, the electronic device can execute the method described in any one of the above first aspect, second aspect, or third aspect.
[0031] It can be understood that for the beneficial effects that can be achieved by the methods described in the above-provided second aspect and third aspect, the device described in the fourth aspect, the electronic device described in the fifth aspect, the computer-readable storage medium described in the sixth aspect, the computer program product described in the seventh aspect, the chip described in the eighth aspect, and the parallel configuration system for operators described in the ninth aspect, reference can be made to the beneficial effects in the first aspect and any possible implementation manner thereof, which will not be elaborated here. Description of the Drawings
[0032] Figure 1 A captured image in a monitoring scenario provided by the related art;
[0033] Figure 2 A schematic diagram of an image processing system provided by an embodiment of the present application;
[0034] Figure 3 A flowchart of a method for encoding an image provided by an embodiment of the present application;
[0035] Figure 4 A schematic diagram of a method for encoding an image provided by an embodiment of the present application;
[0036] Figure 5 A schematic diagram of another method for encoding an image provided by an embodiment of the present application;
[0037] Figure 6 A schematic diagram of another method for encoding an image provided by an embodiment of the present application;
[0038] Figure 7 A schematic diagram of another method for encoding an image provided by an embodiment of the present application;
[0039] Figure 8 A schematic diagram of another image processing system provided by an embodiment of the present application;
[0040] Figure 9 A flowchart of another method for encoding an image provided by an embodiment of the present application;
[0041] Figure 10 A schematic diagram of the composition of an image encoding device provided by an embodiment of the present application;
[0042] Figure 11 A schematic diagram of the hardware structure of an image encoding device provided by an embodiment of the present application. Detailed Embodiments
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0044] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0045] In the following description, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, unless otherwise specified, "a plurality" means two or more.
[0046] It can be understood that some optional features in the embodiments of this application can, in some scenarios, be implemented independently without relying on other features, such as the current solution they are based on, to solve the corresponding technical problems and achieve the corresponding effects. In some scenarios, they can also be combined with other features according to requirements. Correspondingly, the devices given in the embodiments of this application can also implement these features or functions accordingly, which will not be elaborated here.
[0047] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the embodiments of this application are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0048] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained, and the nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0049] 1) Joint Photographic Experts Group (JPEG) encoding: JPEG encoding is a standard method for compressing continuous-tone still images. By using techniques such as predictive coding, discrete cosine transform, and entropy coding, redundant data in the image can be removed to achieve lossy compression, thereby reducing the storage space required for storing the image. JPEG is a common picture file format.
[0050] 2) Resolution: Resolution determines the fineness of image details. Usually, the higher the resolution of an image, the more pixels it contains and the clearer the image. The higher the resolution of an image, the more storage space it will occupy when stored. 4K resolution can refer to the number of pixels per row in the horizontal direction of the image reaching or approaching 4096.
[0051] 3) Encoding quality factor Q: In JPEG encoding, the encoding quality factor Q is an important parameter affecting encoding quality and file size, used to characterize the quantization degree. The value range of the encoding quality factor Q is usually from 1 to 100. Generally speaking, the larger the value of the encoding quality factor Q, the more image details are retained, so the quality loss of the encoded image is smaller, but the bit rate is also larger, which also means a larger storage space is required.
[0052] In recent years, with the continuous progress of information technology, the resolution of cameras has become higher and higher, and the quality and resolution of the obtained images have also become higher. When a large number of high-quality and high-resolution images need to be stored, traditional storage methods and capacities may not be able to meet the requirements. Thus, it poses a challenge to the existing storage technology. In related technologies, generally, images are encoded before storage. Additionally, limited by the image encoding standard protocol, currently, whether it is the camera side for collecting images, the storage side for storing images, or the platform side for using images, standard JPEG encoding is used, and each image to be stored needs to be encoded separately. However, since the positions of cameras in the monitoring scenario are basically fixed, the background areas of different images captured at different times change little, that is, for consecutive images in the monitoring scenario, the background part generally changes little. Therefore, if the encoding method of encoding each image separately is still adopted in the monitoring scenario, there will be a problem of repeatedly encoding the background area, wasting the bit rate. For example, as Figure 1 shown, Figure 1 in (a) and Figure 1 in (b) are images obtained by the camera shooting the same scene at different times. It can be seen that in the two images shown in Figure 1 in (a) and Figure 1 in (b), the backgrounds such as flower beds, roads, trees, and lane lines are extremely similar (or repetitive), that is, the background changes very little. If each image is encoded separately, the similar background areas will be encoded repeatedly, wasting the bit rate and storage space.
[0053] If the bit rate is simply reduced by decreasing the encoding quality factor Q of the image, although the bit rate is reduced and the required storage space is also decreased, this method often leads to a significant decline in image quality. Especially in a monitoring scenario, if the image quality of foreground objects containing important information is lost, it may have a major impact on subsequent key event retrospection or data analysis on the platform side (such as: illegal evidence collection, secondary intelligent analysis, big data analysis, etc.).
[0054] Therefore, it is necessary to seek an encoding method that can take into account such images with less background change, so as to reduce the waste of bit rate and further reduce the demand for storage capacity while ensuring the image quality, and reduce the storage cost.
[0055] To solve the above problems, an embodiment of the present application provides an encoding method for an image, which can be applied to scenarios with continuous shooting requirements (or continuous scenarios), such as a monitoring scenario for security protection (or a security system). Specifically, a second image (or a segmentation map) of the first image to be encoded can be obtained, and based on the first image to be encoded, a reference image, and the second image, a third image (or a back-projected map) including the foreground of the first image and the difference information between the first image and the reference image can be obtained. Then, the second image and the third image can be encoded to obtain the encoding information of the first image and stored. In this way, by encoding the difference information between the image to be encoded and the reference image and the segmentation map of the image to be encoded as the encoding information of the image to be encoded for storage, not only the overhead of repeated encoding of the background area is reduced, the bit rate is saved, but also the storage space required to store the image is reduced, and the storage cost is reduced. In addition, the background of the reference image, the difference information between the image to be encoded and the reference image, and the segmentation map of the image to be encoded can be directly used to reconstruct the image to be encoded, which can ensure the quality of the foreground object is lossless to a certain extent, and thus ensure the quality of the restored image to be encoded.
[0056] In some embodiments of the present application, the encoding method for an image provided by the embodiments of the present application can be applied to an image processing system as Figure 2 shown. Among them, the image processing system may include a shooting device 210, a storage server 220, and a display device 230.
[0057] As a possible implementation, the above shooting device 210 can be used to shoot a shooting scene (such as the above monitoring scenario for security protection) to obtain a corresponding image. Then, the shooting device 210 can send the captured image to the storage server 220. Exemplarily, the shooting device 210 can shoot the shooting scene at a predetermined time interval to obtain multiple images and send them to the storage server 220. For example, the shooting device 210 can encode the captured image (such as encoding it using a standard JPEG encoding method) and send it to the storage server 220 in the form of a bitstream.
[0058] The storage server 220 can receive images from the photographing device 210. The images received by the storage server 220 can be referred to as images to be encoded. For example, the storage server 220 can receive a bitstream from the photographing device 210, decode it (for example, decode it using the standard JPEG decoding method), and then obtain the image to be encoded. After that, the storage server 220 can encode the image to be encoded based on the solution provided in this application and store it in the storage server 220.
[0059] The display device 230 can be an electronic device deployed on the platform side. The user can view the images captured by the photographing device 210 through the display device 230. For example, in the above monitoring scenario, when the user has requirements such as key event backtracking and big data analysis, the display device 230 can be triggered to obtain the corresponding images from the storage server 220.
[0060] After receiving the acquisition request, the storage server 220 can reconstruct the corresponding image based on the stored content and send it to the display device 230 for the display device 230 to display. For example, after reconstructing the corresponding image based on the stored content, the storage server 220 can encode it (for example, encode it using the standard JPEG encoding method) and send it to the display device 230 in the form of a bitstream. After receiving the bitstream, the display device 230 can decode it (for example, decode it using the standard JPEG decoding method) to obtain the image and display it.
[0061] As an example, the above-mentioned photographing device 210 can be an electronic device with a photographing function such as a camera, a mobile phone, a video camera, a drone, a camera (or called a camera), etc. Among them, the photographing device 210 can be fixed in a certain photographing scene, such as a camera with a fixed photographing angle in the above-mentioned monitoring scene. The above-mentioned storage server 220 can be implemented by an independent server or a server cluster composed of multiple servers. The above-mentioned display device 230 can be an electronic device with a display function such as a mobile phone, a tablet computer, a notebook computer, an Ultra-mobile Personal Computer (UMPC), a handheld computer, a netbook, a Personal Digital Assistant (PDA), a wearable electronic device (for example, a smart watch, smart glasses, a smart helmet, a smart bracelet), a television, etc.
[0062] Next, in combination with Figure 2 the image processing system and the accompanying drawings shown, the image encoding method provided in the embodiments of this application will be described.
[0063] Figure 3Schematic flowchart of an image encoding method provided by an embodiment of this application. As Figure 3 shown, the method may include:
[0064] S301. The storage server acquires the image to be encoded.
[0065] Among them, the image to be encoded may include multiple images with basically unchanged backgrounds. These multiple images with basically unchanged backgrounds may include a reference image used as a reference and other images. Generally, the reference image may be one image, which may be any one of the multiple images with basically unchanged backgrounds. For example, it may be the first image among the multiple images with basically unchanged backgrounds. The other images may be one image or multiple images, which are the images other than the reference image among the multiple images with basically unchanged backgrounds. In an embodiment of this application, the other images are referred to as first images. That is to say, the reference image usually refers to a reference image used to compare with the first image to be encoded. It can be understood that the above multiple images with basically unchanged backgrounds may be images obtained by shooting a continuous scene, such as shooting the same shooting scene. For example, the multiple images with basically unchanged backgrounds are images obtained by shooting the same crossroads. Among the multiple images obtained by shooting the same crossroads, the first image may be the above reference image, and the other images except the first image are the above first images.
[0066] In some embodiments of this application, the above image to be encoded may be an image obtained by a shooting device (such as a camera, a mobile phone, an unmanned aerial vehicle aerial photography, etc.) shooting the same shooting scene. The storage server may acquire the image to be encoded from the shooting device. For example, the shooting device may shoot the shooting scene at regular intervals. After the shooting device obtains an image, it may encode the shot image (such as encoding it in a standard JPEG encoding method) and send it to the storage device in the form of a bitstream. After the storage device receives the bitstream, it may decode the bitstream using a corresponding decoding method to obtain the image to be encoded.
[0067] Among them, the above image to be encoded is obtained by the same shooting device. That is to say, the shooting device for shooting the reference image is the same as the shooting device for shooting the first image. Exemplarily, the above image to be encoded may be an image obtained by the same shooting device shooting a certain monitoring scene. For example, the image to be encoded is an image obtained by a camera A shooting a certain crossroads at different time points.
[0068] The above is introduced by taking the storage server obtaining the image to be encoded from the shooting device as an example. The storage server may also obtain the image to be encoded by other means such as screen capture and network download. This application does not specifically limit the method for obtaining the image to be encoded.
[0069] In the embodiment of the present application, after the storage server obtains the image to be encoded, for the reference image, the storage server can encode it using the standard JEPG encoding method to obtain the bitstream of the reference image. The storage server can store the obtained bitstream of the reference image. The bitstream of the reference image can be the second bitstream in the embodiment of the present application.
[0070] In addition, the storage server also needs to directly cache the reference image. Optionally, the storage server can store the image data of the reference image in memory or on disk. When the reference image is needed, the reference image can be directly read from memory or disk.
[0071] Exemplarily, the process of encoding the reference image using the standard JEPG encoding method in this embodiment can be carried out according to the following process: First, perform color space conversion on the reference image to convert the reference image to the Luminance Chrominance - Blue Chrominance - Red (YCbCr) color space to obtain Luminance - Chrominance (YUV) data. Then, perform image segmentation on the reference image to divide the YUV data into 8x8 small blocks. Next, perform Discrete Cosine Transform (DCT) on each 8x8 small block to obtain DCT coefficients. After the DCT transformation, the low - frequency components are concentrated in the upper left corner, and the high - frequency components are concentrated in the lower right corner, thereby separating the high - frequency details (such as noise and edges) and low - frequency information (such as color and brightness) of the reference image. Then, quantize the DCT coefficients using a set quantization matrix to obtain the quantized DCT coefficients. Finally, encode the quantized DCT coefficients through entropy encoding technology to obtain the encoded data, thereby generating a JPEG file, and this JPEG file is the bitstream of the above - mentioned reference image.
[0072] For the first image in the image to be encoded, its encoding can be achieved through the following steps S302 - S305.
[0073] S302. The storage server obtains a second image based on the first image to be encoded.
[0074] Wherein, the second image is used to indicate one or more of the following information of the first image: the position information of the foreground, the position information of the background. In the present application, the second image can also be called a segmentation map.
[0075] After the storage server obtains the first image to be encoded, the storage server can obtain a segmentation map, that is, the second image, which is used to indicate the position information of its foreground and / or background based on the first image.
[0076] In some embodiments, the storage server may input the first image into a segmentation model. The segmentation model has the function of segmenting the foreground and background of the input image. That is to say, after the storage server inputs the first image into the segmentation model, the segmentation model can output a segmentation map for indicating the position information of the foreground and / or the background in the first image. Among them, the segmentation model may be pre-trained based on sample data and stored in the storage server. Or, the segmentation model is stored in another device different from the storage server. Then, the storage server may send the first image to the device so that the device can obtain the segmentation map of the first image based on the segmentation model and return it to the storage server. In this way, the storage server can obtain the segmentation map of the first image.
[0077] Exemplarily, as Figure 4 shown, taking the example of inputting the first image into a segmentation model and obtaining a segmentation map, where the segmentation model is stored in the storage server. The first image is a photo taken at an intersection. The first image includes backgrounds such as zebra crossings, traffic lights, and flower beds, and also includes foregrounds such as vehicles. The storage server may input the first image into the segmentation model, and the segmentation model can output a segmentation map. Among them, there are two pixel values in the segmentation map. One pixel value (such as Figure 4 the white in Figure 4 ) represents the area where the foreground is located in the first image (i.e., the position information of the location where the vehicle is located in the first image), and the other pixel value (such as
[0078] the black in
[0079] ) represents the area where the background is located in the first image (i.e., the position information of the location where the zebra crossing, traffic lights, flower beds, etc. are located in the first image).
[0080] It should be noted that the above embodiments are described by taking the position information of the foreground and / or the background in the image being reflected in the form of a segmentation map as an example, but the present application is not limited to this. For example, the position information of the foreground and / or the background in the image can also be reflected in the form of coordinates, etc.
[0081] The storage server may obtain a third image including the foreground of the first image and the difference information between the first image and the reference image based on the first image, the reference image, and the segmentation map (i.e., the second image). For example, it includes the following S303 and S304.
[0082] In some embodiments, after the storage server obtains the first image to be encoded, the storage server can obtain a reference image from the cache and, based on the reference image, obtain the difference information between the first image and the reference image.
[0083] Exemplarily, the storage server can determine the pixel difference between the first image and the reference image to obtain a difference result. This difference result can be used as the above-mentioned difference information. For example, the storage server can first unify the sizes and color modes of the first image and the reference image. Then, the storage server can traverse each pixel point of the reference image and the first image, respectively obtain the pixel values of the corresponding pixel points, and for each pixel point, calculate the difference between its pixel values in the first image and the reference image, thereby obtaining the difference result, that is, obtaining the above-mentioned difference information. This difference information includes the pixel differences of each pixel point between the first image and the reference image.
[0084] Another exemplarily, the storage server can first determine the pixel difference between the first image and the reference image to obtain a difference result. Then, perform normalization processing on the difference result to obtain the difference information. For example, after determining the difference result by the above method, the storage server can perform normalization processing on the difference result and use the normalized difference result as the above-mentioned difference information. Among them, the difference result can be normalized according to a predetermined range, that is, the pixel difference of each obtained pixel point is normalized to the predetermined range. For example, the predetermined range can be 0-255.
[0085] Among them, the following formula can be used to perform normalization processing on the obtained difference result:
[0086] X' = (X + 255) / 2
[0087] Among them, X' is the normalized pixel value, and X is the pixel difference of the pixel point. For example: when the pixel difference of a certain pixel point is -255, the pixel value of this pixel point can be obtained as 0 through normalization processing. In this way, after normalizing the pixel differences of each pixel point, the above-mentioned difference information can be obtained. It can be considered that after calculating the pixel differences pixel by pixel and performing normalization processing, a difference map representing the difference information between the first image and the reference image is generated. That is, the difference information between the first image and the reference image is represented in the form of an image.
[0088] The above normalization processing can convert negative pixel differences into a predetermined range, such as within the range of 0-255, which helps to improve the efficiency of the algorithm, reduce storage requirements, ensure the accuracy of color display, and be compatible with various standards and tools.
[0089] For example, continuing the above example, such as Figure 5As shown, taking the first image 502 to be encoded, which includes vehicle A1, vehicle B1, vehicle C1, vehicle D1, and the same background as the reference image 501 as an example. By calculating the difference between the first image 502 and the reference image 501 pixel by pixel and then normalizing the calculation result, a difference map 503 can be obtained. It can be seen that the difference map 503 includes the different information in the foregrounds of the first image 502 and the reference image 501, such as vehicle A1, vehicle C1, vehicle D1 in the first image 502 and vehicle A, vehicle C in the reference image 501. The difference map 503 also includes the different information in the backgrounds of the first image 502 and the reference image 501, such as road sign 504.
[0090] It should be noted that the embodiments of the present application do not limit the execution order of the above S302 and S303 here. For example, the above S302 and S303 can be executed simultaneously, or S302 can be executed first and then S303, or S303 can be executed first and then S302.
[0091] S304. The storage server obtains a third image based on the difference information and the second image.
[0092] Among them, the third image includes: the foreground of the first image to be encoded and the difference information between the first image to be encoded and the reference image.
[0093] In the embodiments of the present application, after the storage server obtains the difference information (such as the above difference map) between the first image to be encoded and the reference image and the second image (such as the above segmentation map), the storage server can paste the foreground of the first image onto the difference map according to the segmentation map to obtain the above third image. For example, the storage server can fuse the segmentation map and the difference map to obtain the third image. In the present application, the third image can also be called a back-projected map. It can be understood that since the difference map only contains the different foregrounds and backgrounds between the first image and the reference image, the same foregrounds in the first image and the reference image will not be retained. When restoring the first image, all the foregrounds in the first image need to be restored. Therefore, by fusing the segmentation map and the difference map, the same foregrounds in the first image and the reference image can be restored to retain all the foregrounds in the first image. That is to say, the back-projected map contains all the foregrounds of the first image, the different foregrounds between the first image and the reference image, and the different backgrounds between the first image and the reference image.
[0094] For example, continuing the above example, such as Figure 6As shown, taking the segmentation map 601 including vehicle A1, vehicle B1, vehicle C1, and vehicle D1 and the difference map 602 including vehicle A1, vehicle C1, vehicle D1, vehicle A, vehicle C, and road sign 604 as an example. By fusing the segmentation map 601 with the difference map 602, the same foreground in the first image and the reference image, such as vehicle B1, can be restored, thereby obtaining the back-projected map 603. It can be seen that the back-projected map 603 not only includes all the foregrounds of the first image, such as vehicle A1, vehicle B1, vehicle C1, and vehicle D1. The back-projected map 603 also includes the foregrounds that are different between the first image and the reference image, such as vehicle A and vehicle C, and the backgrounds that are different between the first image and the reference image, such as road sign 604.
[0095] S305. The storage server encodes the second image and the third image to obtain a first bitstream, and stores the first bitstream as the encoding information of the first image to be encoded.
[0096] Among them, the first bitstream is used to indicate the encoding information of the first image to be encoded.
[0097] In the embodiment of the present application, after the storage server obtains the second image (such as the above-mentioned segmentation map) and the third image (such as the above-mentioned back-projected map), it can encode the segmentation map and the back-projected map, obtain the encoded bitstream and store it. Among them, the encoding method can be the same as the method of performing JEPG encoding on the reference image in S301, which will not be elaborated here in the present application.
[0098] S306. The storage server decodes the bitstream to obtain the first image.
[0099] In some embodiments, when the display device needs to retrieve and view the first image, it can send a request to obtain the first image to the storage server. After the storage server obtains the request, it can restore the first image based on the stored bitstream and send it to the display device.
[0100] Exemplarily, the storage server can decode the first bitstream and the second bitstream. Among them, the second image and the third image can be obtained by decoding the first bitstream, and the reference image can be obtained by decoding the second bitstream. Then, the storage server can restore the first image according to the decoded reference image, the second image, and the third image.
[0101] For example, the storage server can add the pixels at the corresponding positions in the reference image to the pixels at the corresponding positions in the third image according to the position information of the background in the first image indicated by the second image to restore the first image. For example, the area where the foreground (such as vehicles, pedestrians, etc.) is located in the second image (i.e., the segmentation map) is called the target area, and the area where the background (such as flower beds, road signs, landmarks, etc.) is located is called the non-target area. The storage server can add the pixels at the corresponding positions in the reference image to the pixels at the corresponding positions in the reply image according to the position information of the non-target area indicated by the second image, so as to restore the background of the first image in the reply image and then reconstruct the first image.
[0102] In some embodiments, after restoring the first image, the storage server can perform JEPG encoding on the restored first image. Send the encoded first image to the display device. After receiving the encoded first image, the display device decodes it (for example, corresponding to using the standard JPEG decoding method for decoding) to obtain the first image and displays it.
[0103] The above process will be described below with specific examples.
[0104] Combined with Figure 7 , taking the example of capturing an image of a passing vehicle at an intersection. The storage server can use the image (a) obtained from the camera as the reference image and cache the reference image. And perform standard JEPG encoding on the reference image, and store the bitstream of the encoded reference image. Then, after the storage server obtains the image (b) to be encoded, it can calculate the pixel difference between the image (a) and the image (b) pixel by pixel, and then perform normalization processing on the obtained pixel difference to obtain the difference map between the first image and the reference image, as Figure 7 the image (d) shown in. In addition, the storage server can also input the image (b) into the segmentation model to obtain the segmentation map, as Figure 7 the image (c) shown in. Then, fuse the image (c) and the image (d) to obtain a reply image including all the foregrounds in the image (b) and the background different between the image (a) and the image (b), as Figure 7 the image (e) shown in. Finally, encode the image (c) and the image (e) to obtain the encoded first bitstream and store the first bitstream. When it is necessary to display the image (b), decode the bitstream of the reference image to obtain the reference image (i.e., the image (a)), decode the first bitstream to obtain the segmentation map (i.e., the image (c)) and the difference map (i.e., the image (e)), and then perform fusion splicing to restore and display the image (b).
[0105] It should be noted that in the above embodiments, the operations of encoding and decoding the image by using the method implemented in this application are all realized in the storage server. That is, the storage server is taken as the execution subject for illustration. In some other embodiments, the above implementation process can also be realized by the cooperation of the shooting device, the storage server, and the display device.
[0106] In some other embodiments of this application, the image encoding method provided in the embodiments of this application can also be applied to, for example, Figure 8 the image processing system shown in the figure. Among them, the image processing system may include a shooting device 810, a storage server 820, and a display device 830.
[0107] The above shooting device 810 can be used to shoot a shooting scene to obtain a corresponding image. At this time, the captured image can be referred to as an image to be encoded. After that, the shooting device 810 can encode the image to be encoded based on the solution provided in this application and send it to the storage server 820. Exemplarily, the shooting device 810 can use the first captured image as a reference image, and use the images captured at regular intervals subsequently as the first images to be encoded, and then encode the obtained first images to be encoded by using the solution provided in the embodiments of this application, and send the obtained code stream to the storage server 820. In addition, the shooting device 810 can also encode the reference image separately and send the encoded code stream to the storage server 820.
[0108] The storage server 820 can receive the code streams of the first images to be encoded and the code stream of the reference image sent by the shooting device 810 and store them.
[0109] The display device 830 can be an electronic device deployed on the platform side. The user can view the images captured by the shooting device 810 through the display device 830. The display device 830 can respond to the instruction to view the first image, and obtain the code stream information of the first image from the storage server 820. The display device 830 can also first decode the code stream of the first image, and then reconstruct and restore the first image according to the solution provided in this application for display. For example, after the storage server 820 receives the acquisition request, it sends the stored code stream of the first image to the display device 230. The display device 830 decodes the acquired code stream (for example, decodes it by using the standard JPEG decoding method), and reconstructs the corresponding image based on this, and finally performs display.
[0110] Next, in combination with Figure 8 the image processing system shown in the figure and the accompanying drawings, the image encoding method provided in the embodiments of this application will be described.
[0111] S901. The shooting device acquires an image to be encoded, where the image to be encoded includes a reference image and a first image.
[0112] S902. The shooting device obtains a second image based on the first image to be encoded.
[0113] S903. The shooting device obtains the difference information between the first image to be encoded and the reference image based on the first image to be encoded and the reference image.
[0114] S904. The shooting device obtains a third image based on the difference information and the second image.
[0115] S905. The shooting device encodes the second image and the third image to obtain a first bitstream, and sends the first bitstream to the storage server.
[0116] In the embodiment of the present application, after the shooting device obtains the second image and the third image, the shooting device may encode the second image and the third image using a standard JEPG encoding method to obtain the encoded first bitstream. Among them, the first bitstream is used to indicate the encoding information of the image to be encoded. After obtaining the first bitstream, the shooting device may send the obtained first bitstream to the storage server.
[0117] S906. The storage server stores the first bitstream.
[0118] In the embodiment of the present application, the storage server may receive the first bitstream sent by the shooting device and store it.
[0119] S907. The display device restores the first image according to the first bitstream.
[0120] In the embodiment of the present application, after receiving the instruction to view the first image, the display device may obtain the first bitstream from the storage server, decode it to obtain the segmentation map (i.e., the second image) and the back projection map (i.e., the third image), and may also obtain the bitstream of the reference image and decode it to obtain the reference image. After that, the display device may also add the pixels at the corresponding positions of the reference image and the pixels at the corresponding positions of the back projection map according to the position information of the background of the first image indicated by the segmentation map, so as to obtain the restored first image. Finally, the display device may display the restored first image.
[0121] It should be noted that the specific implementation of S901 - S907 in this embodiment is similar to the corresponding content in S301 - S306, and the only difference is that the devices performing the corresponding operations are different. The specific implementation can refer to the description of the corresponding content in S301 - S306, and this embodiment will not be elaborated in detail here.
[0122] The technical solution provided by the embodiments of the present application encodes the difference information between the image to be encoded and the reference image and the segmentation map of the image to be encoded as the encoded information of the image to be encoded for storage. This not only reduces the overhead of repeated encoding in the background area, saves the bit rate, but also reduces the storage space required for storing images and lowers the storage cost. In addition, the background of the reference image, the difference information between the image to be encoded and the reference image, and the segmentation map of the image to be encoded can be directly used to reconstruct the image to be encoded, which can ensure the quality of the foreground object is lossless to a certain extent, and thus ensure the quality of the restored image to be encoded.
[0123] The above mainly introduces the solution of the embodiments of the present application from the method perspective. It can be understood that in order for the image encoding device to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0124] The embodiments of the present application can divide the functional units of the image encoding device according to the above method examples. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0125] The embodiments of the present application provide an image encoding device (denoted as image encoding device 100), which can be applied to the above-mentioned electronic device. As Figure 10 shown, it includes: an acquisition unit 1001 and a processing unit 1002. Optionally, it further includes a storage unit 1003. The storage unit 1003 is used to store the data of the image encoding device 100.
[0126] The acquisition unit 1001 is used to acquire a first image to be encoded.
[0127] The processing unit 1002 is used to obtain a second image based on the first image, and the second image is used to indicate the position information of the foreground included in the first image.
[0128] The processing unit 1002 is further configured to obtain a third image based on the first image, the reference image, and the second image, where the third image includes: the foreground of the first image and the difference information between the first image and the reference image.
[0129] The processing unit 1002 is further configured to encode the second image and the third image to obtain a first bitstream, and store the first bitstream as the encoded information of the first image.
[0130] In an implementable manner, the processing unit 1002 is further configured to obtain the difference information between the first image and the reference image based on the first image and the reference image; and obtain the third image based on the difference information and the second image.
[0131] In an implementable manner, the processing unit 1002 is further configured to determine the pixel difference between the first image and the reference image to obtain a difference result; and perform normalization processing on the difference result to obtain the difference information.
[0132] In an implementable manner, the processing unit 1002 is further configured to input the first image into a segmentation model to obtain the second image; where the segmentation model has the function of segmenting the foreground and background of the input image.
[0133] In an implementable manner, the processing unit 1002 is further configured to decode the first bitstream to obtain the second image and the third image; decode the second bitstream of the reference image to obtain the reference image, where the second bitstream is the bitstream obtained after encoding the reference image; and restore the first image based on the reference image, the second image, and the third image.
[0134] In an implementable manner, the second image is further used to indicate the position information of the background included in the first image, and the processing unit 1002 is further configured to add the pixels at the corresponding positions in the reference image and the pixels at the corresponding positions in the third image according to the position information of the background included in the first image indicated by the second image to restore the first image.
[0135] In an implementable manner, the reference image and the first image are images obtained by photographing the same scene.
[0136] Figure 10 The units in can also be referred to as modules. For example, the obtaining unit can be referred to as an obtaining module, and the processing unit can be referred to as a processing module. Additionally, in Figure 10 In the embodiments shown, the names of the respective units may also not be the names shown in the figure. For example, the obtaining unit may also be referred to as a communication unit.
[0137] Figure 10When each unit in [the application] is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of this application. The storage media storing the computer software product include: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0138] The embodiments of this application also provide a schematic diagram of the hardware structure of an image encoding device. Refer to Figure 11 , the image encoding device includes a processor 1101 and a transceiver 1102. Optionally, it further includes a memory 1103 connected to the processor 1101.
[0139] The processor 1101 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the programs of the solutions of this application. The processor 1101 can also include multiple CPUs, and the processor 1101 can be a single-CPU processor or a multi-CPU processor. Here, the processor can refer to one or more devices, circuits, or processing cores for processing data (such as computer program instructions).
[0140] The processor 1101, the memory 1103, and the transceiver 1102 are connected through a bus. The transceiver 1102 is used for parallel configuration devices with other operators. Optionally, the transceiver 1102 can include a transmitter and a receiver. The device in the transceiver 1102 for implementing the receiving function can be regarded as a receiver, and the receiver is used to execute the receiving steps in the embodiments of this application. The device in the transceiver 1102 for implementing the sending function can be regarded as a transmitter, and the transmitter is used to execute the sending steps in the embodiments of this application.
[0141] In the first possible implementation manner, refer to Figure 11, the parallel configuration device of the operator further includes a transceiver 1102. The memory 1103 can be a ROM or other types of static storage devices that can store static information and instructions, a RAM, or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. The embodiments of the present application do not impose any restrictions on this. The memory 1103 can exist independently or be integrated with the processor 1101. Among them, the memory 1103 can contain computer program code. The processor 1101 is used to execute the computer program code stored in the memory 1103, thereby implementing the method provided by the embodiments of the present application.
[0142] The embodiments of the present application also provide a computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute any of the above methods.
[0143] The embodiments of the present application also provide a computer program product containing instructions, which when running on a computer, cause the computer to execute any of the above methods.
[0144] The embodiments of the present application also provide a chip, including: a processor and an interface. The processor is coupled to the memory through the interface. When the processor executes the computer program or instructions in the memory, any of the methods provided by the above embodiments is executed.
[0145] The embodiments of the present application also provide a parallel configuration system of an operator, including: the terminal device and the access network device in the above embodiments.
[0146] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0147] Although the present application has been described in conjunction with various embodiments, however, in the process of implementing the claimed present application, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0148] Although the present application has been described in conjunction with the features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are merely exemplary descriptions of the present application defined by the appended claims and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these changes and modifications.
[0149] As described above, it is only the implementation mode of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the said claims.
Claims
1. A method for encoding an image, characterized in that, Including: Obtain a first image to be encoded; Based on the first image, obtain a second image, where the second image is used to indicate the position information of the foreground included in the first image; Based on the first image, a reference image, and the second image, obtain a third image, where the third image includes: the foreground of the first image and the difference information between the first image and the reference image; Encode the second image and the third image to obtain a first bitstream, and store the first bitstream as the encoding information of the first image.
2. The method according to claim 1, characterized in that, The obtaining the third image based on the first image, the reference image, and the second image includes: Based on the first image and the reference image, obtain the difference information between the first image and the reference image; Based on the difference information and the second image, obtain the third image.
3. The method according to claim 2, wherein The obtaining the difference information between the first image and the reference image based on the first image and the reference image includes: Determine the pixel difference between the first image and the reference image to obtain a difference result; Perform normalization processing on the difference result to obtain the difference information.
4. The method according to any one of claims 1 to 3, characterized in that The obtaining the second image based on the first image includes: Input the first image into a segmentation model to obtain the second image; The segmentation model has the function of segmenting the foreground and background of the input image.
5. The method according to any one of claims 1 to 4, characterized in that The method further includes: Decode the first bitstream to obtain the second image and the third image; Decode the second bitstream of the reference image to obtain the reference image, where the second bitstream is the bitstream obtained after encoding the reference image; Based on the reference image, the second image, and the third image, restore the first image.
6. The method according to claim 5, wherein The second image is further used to indicate the position information of the background included in the first image; The restoring the first image based on the reference image, the second image, and the third image includes: According to the position information of the background included in the first image indicated by the second image, add the pixels at the corresponding positions in the reference image to the pixels at the corresponding positions in the third image to restore the first image.
7. The method according to any one of claims 1-6, characterized in that, The reference image and the first image are images obtained by photographing the same scene.
8. An image encoding device, characterized in that, Including: Functional units for performing the method according to any one of claims 1-7; wherein, the actions performed by the functional units are implemented by hardware or by hardware executing corresponding software.
9. An electronic device, characterized in that, Including: A processor; The processor is connected to a memory, and the memory is used to store computer execution instructions. The processor executes the computer execution instructions stored in the memory so that the electronic device implements the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, Including instructions, when the instructions run on a computer, the computer is caused to execute the method according to any one of claims 1-7.