Image transmission method and image coding device

By acquiring and encoding the form and hollow area data of the current frame image in the image encoding, and using a preset database for decoding in the image decoding device, the problem of not being able to use similar parts of multiple reference frames at the same time in the prior art is solved, and efficient image encoding is achieved.

CN120011592APending Publication Date: 2025-05-16XIAN WANXIANG ELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510205533.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2020-06-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing multi-reference frame image prediction technology cannot simultaneously use similar parts of each frame of multiple reference frames for reference, resulting in ineffective encoding.

Method used

By obtaining the form information in the current frame image, and looking for the corresponding target form identifier in the preset database, encode the form data and hollow area data, and sending it to the image decoding device. This method allows reference to forms in multiple historical frames simultaneously, improving encoding efficiency.

Benefits of technology

It is realized that references to similar parts of multiple reference frames simultaneously in computer synthetic image encoding, reducing the code stream size and improving encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011592A_ABST
    Figure CN120011592A_ABST
Patent Text Reader

Abstract

The invention provides an image transmission method and an image coding device, relates to the technical field of image transmission, and can solve the problem that similar parts in various frames of multiple reference frames cannot be used for reference at the same time in existing multi-reference-frame image prediction. According to the specific technical scheme, window information in a current frame image is obtained; according to each window identifier, searching a preset database for a target window identifier the same as the window identifier; encoding the target window identifier and window data corresponding to the window identifier not included in the preset database to obtain window encoded data; obtaining hollow area data in the current frame image; coding the hollow area data to obtain hollow area coded data; and sending the window coded data and the hollow area coded data to an image decoding device. According to the method and the device, if a plurality of windows which are the same as those in the historical image frames exist in the windows of the current frame image, the windows can be used for reference, namely, a plurality of reference frames are referred at the same time at the moment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention is a divisional application with application number 2020105642217, application date June 18, 2020, and invention name “Image Transmission Method and System”. Technical Field

[0002] The present disclosure relates to the field of image processing, and in particular to an image transmission method and an image encoding device. Background Art

[0003] Computer-generated images often switch back and forth. For example, when a user is surfing the Internet, he may suddenly receive a message from an instant messaging software. The user pops up the communication window, sends and receives the message, and then pops it back up. At the same time, he switches to the text editing window to edit, and then switches to the browser window to check information. The scene of the last surfing scene is very similar to the previous surfing scene. When encoding, the frame image in the earlier surfing scene can be regarded as a reference frame image. Therefore, the encoding scene of computer-generated images is more suitable for encoding and decoding using multiple reference frame images.

[0004] At present, based on the encoding and decoding of multiple reference frame images, only the frame that is "most similar" to the current frame can be selected for reference. However, if multiple reference frames each have a part that is similar to the current frame, the related technology cannot use the similar parts of each frame of the multiple reference frames for reference at the same time. Summary of the invention

[0005] The disclosed embodiments provide an image transmission method and an image coding device, which can solve the problem that the existing multi-reference frame image prediction cannot use similar parts of multiple reference frames for reference at the same time. The technical solution is as follows:

[0006] According to a first aspect of an embodiment of the present disclosure, there is provided an image transmission method, the method being applied to an image encoding device, the method comprising:

[0007] Acquire window information in the current frame image, the window information including: each window identifier and window data corresponding to each window identifier; wherein the window data includes: window pixel data, window position and window size;

[0008] According to each of the window identifiers, searching the preset database for a target window identifier that is identical to the window identifier; the preset database includes: a correspondence between the window identifiers and the window data in the historical frame image; the window identifiers stored in the preset database are all different;

[0009] Encoding the target window identifier and the window data corresponding to the window identifier not included in the preset database to obtain window encoding data;

[0010] Acquire hollow area data in the current frame image, where the hollow area is a display area in the current frame image excluding the window;

[0011] Encoding the hollow area data to obtain hollow area coded data;

[0012] Sending a coded code stream to an image decoding device, the coded code stream comprising: the window coded data and the hollow area coded data.

[0013] The image transmission method provided by the embodiment of the present disclosure includes: obtaining the window information in the current frame image, the window information includes: each window identification and the window data corresponding to each window identification; according to each window identification, searching the target window identification that is the same as the window identification from the preset database; the preset database includes: the correspondence between the window identification and the window data in the historical frame image; the window identifications stored in the preset database are all different; encoding the window data corresponding to the target window identification and the window identification not included in the preset database to obtain the window encoding data; obtaining the hollow area data in the current frame image, the hollow area is the display area in the current frame image except the window; encoding the hollow area data to obtain the hollow area encoding data; sending the window encoding data and the hollow area encoding data to the image decoding device. In the present disclosure, each window in the historical image frame is saved as a unit. If there are multiple windows in the window of the current frame image that are the same as the window in the historical image frame, then these windows can be used for reference. Since these windows come from different historical frames, it is equivalent to referring to multiple reference frames at the same time.

[0014] In one embodiment, encoding the target window identifier and the window data corresponding to the window identifier not included in the preset database to obtain the window encoding data includes:

[0015] Dividing the window corresponding to the target window identifier in the preset database into a plurality of first sub-windows according to a first preset rule;

[0016] Dividing the window corresponding to the target window identifier in the current frame image into a plurality of second sub-windows according to the first preset rule;

[0017] If the window data in the first subwindow at the same position is the same as the window data in the second subwindow, the table entry at the position corresponding to the second subwindow having the same window data in the first preset position table is marked as a first mark; the position of each table entry in the first preset position table corresponds to the position of each second subwindow in the current frame image;

[0018] If the form data in the first sub-window and the form data in the second sub-window at the same position are different, marking the entry at the position corresponding to the second sub-window with different form data in the first preset position table as a second mark;

[0019] Dividing the window in the current frame image corresponding to the window identifier not included in the preset database into a plurality of third sub-windows according to the first preset rule;

[0020] Marking each table item in the second preset position table as the second mark, wherein the position of each table item in the second preset position table corresponds to the position of each third subwindow in the current frame image;

[0021] The target window identifier, the marked first preset position table, the window data in the current frame image corresponding to the second marked table item, the marked second preset position table, and the window data in the current frame image corresponding to each table item in the second preset position table are encoded to obtain window encoding data.

[0022] In one embodiment, the method further comprises:

[0023] Adding the form identifier and the corresponding form data not included in the preset database to the preset database;

[0024] Replacing the window data corresponding to the window identifier in the preset database with the window data corresponding to the window identifier in the current frame image;

[0025] The first update information is carried in the encoded bitstream and sent to the image decoding device, wherein the first update information includes: all the updated data in the preset database.

[0026] In one embodiment, encoding the hollow area data to obtain hollow area coded data includes:

[0027] Detect whether the table items of the preset mapping table include a target table item whose similarity with the window identifier of the current frame image meets the preset conditions, each of the table items of the preset mapping table includes: a mapping relationship between the current table item identifier, historical hollow area data and all window identifiers in the full-frame image corresponding to the hollow area data;

[0028] If included, encoding the current entry identifier included in the target entry to obtain the hollow area encoding data;

[0029] If not included, the hollow area data is encoded to obtain hollow area encoded data.

[0030] In one embodiment, encoding the current entry identifier included in the target entry to obtain the hollowed-out area encoding data includes:

[0031] Dividing the hollowed-out area of ​​the current frame image into a plurality of first sub-hollowed-out areas according to a second preset rule;

[0032] Dividing the hollow area corresponding to the target table item into a plurality of second sub-hollow areas according to the second preset rule;

[0033] If the hollow area data in the first sub-hollow area and the hollow area data in the second sub-hollow area at the same position are the same, the table entry at the position corresponding to the second sub-hollow area with the same hollow area data in the third preset position table is marked as a first mark; the position of each table entry in the third preset position table corresponds to the position of each second sub-hollow area in the current frame image;

[0034] If the hollow area data in the first sub-hollow area and the hollow area data in the second sub-hollow area at the same position are different, marking the entry corresponding to the position of the second sub-hollow area with different hollow area data in the third preset position table as a second mark;

[0035] The hollow area coding data is obtained by encoding the hollow area data in the current frame image corresponding to the current entry identifier, the marked third preset position table, and each entry in the third preset position table marked as the second mark.

[0036] In one embodiment, the method further comprises:

[0037] Detect whether the current frame image is a frame image after scene switching;

[0038] If yes, then add a new table entry in the preset mapping table, and save the hollow area data corresponding to the current frame image and all the window identifiers of the current frame image into the new table entry;

[0039] The second update information is carried in the encoded bitstream and sent to the image decoding device, wherein the second update information includes: the table entry information newly added in the preset mapping table.

[0040] In one embodiment, the detecting whether the current frame image is a frame image after scene switching includes:

[0041] Dividing the current frame image into macroblocks according to a third preset rule;

[0042] Dividing the previous frame of image into macroblocks according to the third preset rule;

[0043] Detecting the similarity between the macroblock corresponding to the current frame image and the macroblock corresponding to the previous frame image;

[0044] When the similarity is greater than a preset value, it is determined that a scene switch occurs.

[0045] According to a second aspect of an embodiment of the present disclosure, there is provided an image transmission method, the method being applied to an image decoding device, the method comprising:

[0046] Receive a coded code stream, the coded code stream includes: window coding data and hollow area coding data; the window coding data is obtained by encoding according to the target window identifier and the window data corresponding to the window identifier not included in the preset database; the preset database includes: the correspondence between the window identifier and the window data in the historical frame image; the window identifiers stored in the preset database are all different; the window data includes: window pixel data, window position and window size; the hollow area is the display area in the current frame image except the window;

[0047] Restoring the window corresponding to the target window identifier according to the preset database and the window encoding data;

[0048] Restoring the form corresponding to the form identifier not included in the preset database according to the form data corresponding to the form identifier not included in the preset database;

[0049] Restoring the hollow area according to the hollow area coding data;

[0050] The current frame image is acquired according to the restored window corresponding to the target window identifier, the window corresponding to the window identifier not included in the preset database, and the hollow area.

[0051] In one embodiment, the encoded code stream further includes: first update information, the first update information includes: all data updated in the preset database, and the method further includes:

[0052] The data in the preset database is updated according to the first update information.

[0053] In one embodiment, the encoded code stream further includes: second update information, the second update information includes: newly added table item information in the preset mapping table, each of the table items in the preset mapping table includes: a mapping relationship between a current table item identifier, historical hollow area data, and all window identifiers in the full-frame image corresponding to the hollow area data; the method further includes:

[0054] The table entries in the preset mapping table are updated according to the second update information.

[0055] In one embodiment, the encoded code stream includes: the target window identifier, the marked first preset position table, the window data in the current frame image corresponding to the second marked table item, the marked second preset position table, and the window data in the current frame image corresponding to each table item in the second preset position table;

[0056] The step of restoring the window corresponding to the target window identifier according to the preset database and the window encoding data includes:

[0057] Restoring the window corresponding to the target window identifier according to the preset database, the target window identifier, the marked first preset position table, and the window data in the current frame image corresponding to the table entry of the second mark;

[0058] The step of restoring the window corresponding to the window identifier not included in the preset database according to the window data corresponding to the window identifier not included in the preset database comprises:

[0059] The window corresponding to the window identifier not included in the preset database is restored according to the preset database, the marked second preset position table and the window data in the current frame image corresponding to each table entry in the second preset position table.

[0060] In one embodiment, the encoded code stream includes: a current table entry identifier, a marked third preset position table, and the hollow area data in the current frame image corresponding to each table entry marked as a second mark in the third preset position table;

[0061] The restoring the hollowed-out area according to the hollowed-out area encoding data comprises:

[0062] The hollow area is restored according to the preset mapping table, the current table item identifier, the marked third preset position table, and the hollow area data in the current frame image corresponding to each table item marked as the second mark in the third preset position table. Each of the table items in the preset mapping table includes: the mapping relationship between the current table item identifier, the historical hollow area data and all the window identifiers in the full-frame image corresponding to the hollow area data.

[0063] According to a third aspect of an embodiment of the present disclosure, there is provided an image transmission system, including: an image encoding device and an image decoding device;

[0064] The image encoding device is used to perform the method steps corresponding to the image encoding device described in any one of the above embodiments;

[0065] The image decoding device is used to execute the method steps corresponding to the image decoding device described in any one of the above embodiments.

[0066] According to a fourth aspect of an embodiment of the present disclosure, there is provided an image transmission device, the device being applied to an image encoding device, the device comprising:

[0067] The first acquisition module is used to acquire the window information in the current frame image, wherein the window information includes: each window identifier and the window data corresponding to each window identifier; wherein the window data includes: window pixel data, window position and window size;

[0068] A search module, used for searching the target window identifier that is the same as the window identifier from the preset database according to the window identifiers; the preset database includes: the correspondence between the window identifiers and the window data in the historical frame images; the window identifiers stored in the preset database are all different;

[0069] A first encoding module, used for encoding the target window identifier and the window data corresponding to the window identifier not included in the preset database to obtain window encoding data;

[0070] A second acquisition module is used to acquire hollow area data in the current frame image, where the hollow area is a display area in the current frame image excluding the window;

[0071] A second encoding module, used for encoding the hollow area data to obtain hollow area encoding data;

[0072] The first sending module is used to send a coded code stream to the image decoding device, wherein the coded code stream includes: the window coded data and the hollow area coded data.

[0073] In one embodiment, the first encoding module includes:

[0074] A first division submodule, used for dividing the window corresponding to the target window identifier in the preset database into a plurality of first subwindows according to a first preset rule;

[0075] A second division submodule, used for dividing the window corresponding to the target window identifier in the current frame image into a plurality of second sub-windows according to the first preset rule;

[0076] A first marking submodule is used for marking the table item at the position corresponding to the second subwindow with the same window data in the first preset position table as a first mark if the window data in the first subwindow at the same position is the same as the window data in the second subwindow; the position of each table item in the first preset position table corresponds to the position of each second subwindow in the current frame image;

[0077] A second marking submodule, for marking the entry at the position corresponding to the second subwindow which is different from the form data in the first preset position table as a second mark if the form data in the first subwindow and the form data in the second subwindow at the same position are different;

[0078] A third division submodule, configured to divide the window in the current frame image corresponding to the window identifier not included in the preset database into a plurality of third sub-windows according to the first preset rule;

[0079] A third marking submodule, used for marking each table item in the second preset position table as the second mark, wherein the position of each table item in the second preset position table corresponds to the position of each third subwindow in the current frame image;

[0080] The first encoding submodule is used to encode the target window identifier, the marked first preset position table, the window data in the current frame image corresponding to the second marked table item, the marked second preset position table, and the window data in the current frame image corresponding to each table item in the second preset position table to obtain window encoding data.

[0081] In one embodiment, the apparatus further comprises:

[0082] A first adding submodule, used for adding the form identifier and the corresponding form data not included in the preset database to the preset database;

[0083] A replacement submodule, used to replace the window data corresponding to the window identifier in the preset database with the window data corresponding to the window identifier in the current frame image;

[0084] The second sending module is used to carry first update information in the encoded code stream and send it to the image decoding device, where the first update information includes: all data updated in the preset database.

[0085] In one embodiment, the second encoding module includes:

[0086] A first detection submodule is used to detect whether the table items in the preset mapping table include a target table item whose similarity with the window identifier of the current frame image meets a preset condition, and each of the table items in the preset mapping table includes: a mapping relationship between the current table item identifier, historical hollow area data and all window identifiers in the full-frame image corresponding to the hollow area data;

[0087] A second encoding submodule, for, if included, encoding the current entry identifier included in the target entry to obtain the hollow area encoding data;

[0088] The third encoding submodule is used to encode the hollow area data to obtain hollow area encoded data if it is not included.

[0089] In one embodiment, the second encoding submodule includes:

[0090] A fourth division submodule, configured to divide the hollowed-out area of ​​the current frame image into a plurality of first sub-hollowed-out areas according to a second preset rule;

[0091] A fifth division submodule, configured to divide the hollow area corresponding to the target table item into a plurality of second sub-hollow areas according to the second preset rule;

[0092] a fourth marking submodule, for marking the table entry at the position corresponding to the second sub-hollowed area having the same hollowed area data in the third preset position table as a first mark if the hollowed area data in the first sub-hollowed area at the same position is the same as the hollowed area data in the second sub-hollowed area; the position of each table entry in the third preset position table corresponds to the position of each second sub-hollowed area in the current frame image;

[0093] a fifth marking submodule, configured to mark the entry corresponding to the position of the second sub-hollow area with different hollow area data in the third preset position table as a second mark if the hollow area data in the first sub-hollow area and the hollow area data in the same position are different;

[0094] The fourth encoding submodule is used to encode the hollow area data in the current frame image corresponding to the third preset position table marked with the current table entry identifier and the second mark in the third preset position table to obtain the hollow area encoded data.

[0095] In one embodiment, the apparatus further comprises:

[0096] A detection module, used to detect whether the current frame image is a frame image after scene switching;

[0097] An adding module, used for adding a new table entry in the preset mapping table when the detection module detects that the current frame image is a frame image after scene switching, and saving the hollow area data corresponding to the current frame image and all the window identifiers of the current frame image to the new table entry;

[0098] The third sending module is used to carry the second update information in the encoded code stream and send it to the image decoding device, where the second update information includes: the newly added table item information in the preset mapping table.

[0099] In one embodiment, the detection module includes:

[0100] A sixth division submodule, used for dividing the current frame image into macroblocks according to a third preset rule;

[0101] A seventh division submodule, used for dividing the previous frame image into macroblocks according to the third preset rule;

[0102] A second detection submodule, used to detect the similarity between the macroblock corresponding to the current frame image and the macroblock corresponding to the previous frame image;

[0103] The determination submodule is used to determine that a scene switch occurs when the similarity is greater than a preset value.

[0104] Based on the above Figure 6 The image transmission method described in the corresponding embodiment is the following embodiment of the device disclosed herein, which can be used to execute the embodiment of the device disclosed herein.

[0105] According to a fifth aspect of an embodiment of the present disclosure, there is provided an image transmission device, the device being applied to an image decoding device, the device comprising:

[0106] A receiving module, used for receiving a coding stream, wherein the coding stream includes: window coding data and hollow area coding data; the window coding data is obtained by coding according to the target window identifier and the window data corresponding to the window identifier not included in the preset database; the preset database includes: the correspondence between the window identifier and the window data in the historical frame image; the window identifiers stored in the preset database are all different; the window data includes: window pixel data, window position and window size; the hollow area is the display area of ​​the current frame image except the window;

[0107] A first recovery module, used for recovering the window corresponding to the target window identifier according to a preset database and the window encoding data;

[0108] A second restoring module, used for restoring the window corresponding to the window identifier not included in the preset database according to the window data corresponding to the window identifier not included in the preset database;

[0109] A third recovery module, used for recovering the hollow area according to the hollow area coding data;

[0110] The acquisition module is used to acquire the current frame image according to the window corresponding to the restored target window identifier, the window corresponding to the window identifier not included in the preset database, and the hollow area.

[0111] In one embodiment, the encoded code stream further includes: first update information, the first update information includes: all data updated in the preset database, and the device further includes:

[0112] The first updating module is used to update the data in the preset database according to the first updating information.

[0113] In one embodiment, the encoded code stream further includes: second update information, the second update information includes: newly added table item information in the preset mapping table, each of the table items in the preset mapping table includes: a mapping relationship between a current table item identifier, historical hollow area data and all window identifiers in the full-frame image corresponding to the hollow area data; the device also includes:

[0114] The second updating module is used to update the table entries in the preset mapping table according to the second updating information.

[0115] In one embodiment, the encoded code stream includes: the target window identifier, the marked first preset position table, the window data in the current frame image corresponding to the second marked table item, the marked second preset position table, and the window data in the current frame image corresponding to each table item in the second preset position table;

[0116] A first recovery module is further used to recover the window corresponding to the target window identifier according to the preset database, the target window identifier, the marked first preset position table and the window data in the current frame image corresponding to the table entry of the second mark;

[0117] The second recovery module is further used to recover the window corresponding to the window identifier not included in the preset database according to the preset database, the marked second preset position table and the window data in the current frame image corresponding to each table item in the second preset position table.

[0118] In one embodiment, the encoded code stream includes: a current table entry identifier, a marked third preset position table, and the hollow area data in the current frame image corresponding to each table entry marked as a second mark in the third preset position table;

[0119] The third recovery module is also used to restore the hollow area according to the preset mapping table, the current table item identifier, the marked third preset position table, and the hollow area data in the current frame image corresponding to each table item marked as the second mark in the third preset position table. Each of the table items in the preset mapping table includes: the mapping relationship between the current table item identifier, the historical hollow area data and all the window identifiers in the full-frame image corresponding to the hollow area data.

[0120] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0122] Figure 1 It is a schematic diagram of a natural video image provided by an embodiment of the present disclosure.

[0123] Figure 2 It is a schematic diagram of a computer synthesized image scene switching provided by an embodiment of the present disclosure.

[0124] Figure 3 is a flow chart of an image transmission method provided by an embodiment of the present disclosure;

[0125] Figure 4 is a window diagram provided by an embodiment of the present disclosure;

[0126] Figure 5 is a window diagram provided by an embodiment of the present disclosure;

[0127] Figure 6 is a flow chart of an image transmission method provided by an embodiment of the present disclosure;

[0128] Figure 7 is a schematic diagram of a logical layer structure of an image encoding device provided by an embodiment of the present disclosure;

[0129] Figure 8 is a schematic diagram of a logical layer structure of an image decoding device provided by an embodiment of the present disclosure;

[0130] Fig. 9 is a schematic diagram of an image transmission system provided by an embodiment of the present disclosure;

[0131] Fig.10 is a structural diagram of an image transmission device provided by an embodiment of the present disclosure;

[0132] Fig.11 It is a structural diagram of an image transmission device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0133] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0134] There are two types of images: natural images and computer-generated images. Natural images refer to the real scenery in nature. The movies and TV programs we see in our daily lives are all natural images. Computer-generated images are artificial images calculated by the computer graphics card using computer graphics technology, such as the interface of the office software Word, game screens, web page text, vector graphics and renderings of CAD software, etc.

[0135] The existing technical solutions are mainly similar to video encoding solutions such as H.264 and H.265, which have high compression rates and good compression effects for natural videos and are widely used in the industry. A common feature of this type of video encoding and decoding solutions is that they basically use the inter-frame prediction method to reduce the bit rate. The so-called inter-frame prediction means that when encoding the current frame, the current frame will be compared with another frame called the reference frame. If a certain area of ​​the current frame is the same as the reference frame, the information of the same area, such as the position, is directly recorded, and the data of the area of ​​the reference frame is directly used during decoding. This solution actually omits the encoding and decoding process of the above-mentioned area of ​​the current frame, which not only reduces the encoding and decoding calculation time, but also reduces the bit rate size. It is a widely used solution at present. However, this solution has shortcomings for computer-synthesized image sequences.

[0136] (1) First, from the perspective of the particularity of the scene, for natural video sequences, scene switching is usually descriptive, narrative, and natural, and there are relatively few cases of repeated switching between different scenes; while for computer-generated images, such switching back and forth is very common. For example, when a user is surfing the Internet, he may suddenly receive a message from an instant messaging software. The user pops up the communication window, sends and receives the message, and then pops it back up. At the same time, he switches to the text editing window to edit, and then switches to the browser window to search for information. Then, the scene of the last Internet search is extremely similar to the previous Internet scene. When encoding, the frame in the earlier Internet scene can be regarded as a reference frame. Therefore, compared with natural videos, the coding scene of computer-generated images has a much more urgent demand for multiple reference frames.

[0137] (2) From the current codec technology, most codecs only support one reference frame. Some codec technology standards support multiple reference frames, such as H.264, but very few products actually implement multiple reference frames. The reason is that it is necessary to save multiple complete reference frames, which puts a huge burden on the storage space of the codec. In addition, when encoding, one reference frame needs to be selected from multiple reference frames, and the optimal algorithm will cause a large amount of computational burden. There are a considerable number of technical documents and patents trying to improve and optimize the selection method of multiple reference frames, but they cannot fundamentally solve the problem of time and efficiency. Therefore, as an industry consensus, multi-reference frame prediction similar to H264 is a very low cost-effective video coding technology. In most cases, it cannot improve the coding efficiency, but will introduce heavy additional computational overhead. It is generally recommended not to use this technology as much as possible.

[0138] (3) Taking a step back, even if the problem of computational complexity is solved, the current multi-reference frame solution can only select the frame that is "most similar" to the current frame for reference. Therefore, if a portion of each of the five reference frames is similar to the current frame, the similar portions of the five reference frames cannot be used for reference at the same time, thereby failing to achieve the best effect and achieving the effect of reducing the bit rate.

[0139] The present disclosure is a solution to the above three problems.

[0140] The original intention of designing this scheme is the particularity of scene switching in computer-synthesized images. Therefore, we must first describe in detail the differences in scene switching between natural video and computer-synthesized image sequences, and make it clear what the particularity of computer-synthesized images is.

[0141] First, let’s take a look at a set of natural video image screenshots, which also include scene changes. Figure 1 is a schematic diagram of a natural video image provided by an embodiment of the present disclosure. Figure 1 As can be seen in the figure, there are large differences between (b) and (c), and between (c) and (d), so these two times can be regarded as scene switching. Although (d) switches back to the green coast, the difference between it and (b) is still very large. If (b) is used as the reference frame when editing (d), good results will not be achieved.

[0142] Below is a sequence of computer-generated images. Figure 2 FIG. 1 is a schematic diagram of a computer-generated image scene switching provided by an embodiment of the present disclosure. Figure 2 As shown, Figure 2This is a simulation of a user operating a computer to write a paper. The user is writing a paper, searching for information online, and drawing and editing. He uses the "task view" that comes with the Windows 10 operating system to switch between multiple windows. Figure 2 As shown, the large picture above is the "Task View" window. The user clicked it 4 times, and 4 desktop images were generated respectively, as shown in the 4 small pictures below. The 4 small pictures below form a sequence. The differences between them are quite large, and it can be considered that 4 scene switches occurred. If the sequence is encoded, it can be found that (d) in the sequence is very similar to (a), and (a) can be used as a reference frame for inter-frame prediction encoding. Similar operations of switching between different windows and repeatedly minimizing and restoring certain windows are very common when users operate computers.

[0143] This sequence is just an example to illustrate the abruptness of the scene switching of computer-generated images and their similarity to a certain frame in history. From this example, you can perceive the difference between it and natural video images.

[0144] Based on the above scenario, the present disclosure provides an image transmission method, such as Figure 3 As shown, the method is applied to an image encoding device, and the image transmission method comprises the following steps:

[0145] 101. Obtain window information in the current frame image, where the window information includes: each window identifier and window data corresponding to each window identifier; the window data includes: window pixel data, window position and window size.

[0146] The image encoding device can receive data to be encoded from the collector, that is, receive a frame of image, usually RGB or YUV original pixel data.

[0147] After receiving a frame of image, the image encoding device can intercept the operating system instruction or interface in the bottom layer software to obtain the window information on the current screen in real time.

[0148] For example, the above-mentioned window identification includes: a window handle.

[0149] The handle is the window identifier (English: ID). Each window corresponds to a unique ID. Therefore, different windows can be identified by this ID. The handle can also be used to obtain the window's position, size and other information from the bottom layer.

[0150] Different operating systems may have different implementations for obtaining the window handle on the current screen. For example, in Windows operating system, you can call GetDesktopWindow and GetNextWindow function interfaces to find all the window handles on the screen. Through the window handle, you can also query the size and position information of each window.

[0151] 102. According to each window identifier, a target window identifier identical to the window identifier is searched from a preset database; the preset database includes: a correspondence between the window identifier and the window data in the historical frame image; the window identifiers stored in the preset database are all different.

[0152] A preset database is maintained in the image encoding device, and the preset database stores the correspondence between the window identifiers and window data in the historical frame images. For example, the preset database may include: a window identifier area and a window data area. The window identifier area is used to store the window identifier, and the window data area is used to store the window data. Moreover, the items in the window identifier area and the window data area are one-to-one corresponding, that is, one window identifier corresponds to one window data. For example, a maximum of N entries can be stored, that is, data of N windows (the data of N windows include: window identifiers of N windows and window data of N windows). The N entries stored may come from different frames or from the same frame, which is not important because the present scheme takes each window as the prediction object, rather than each frame as the prediction object.

[0153] It is worth noting that the form identifiers stored in the preset database in the present disclosure are all different, that is, the same form identifier is only stored once, and thus the form data corresponding to the same form identifier is also only stored once.

[0154] After obtaining the window identifiers in the current frame image, it is necessary to search the preset database to see if the window identifiers that have been edited historically are the same as the window identifiers detected in the current frame. If so, it means that the window to be edited in the current frame has appeared in history, so these two windows are likely to be the same, or partially the same.

[0155] Specifically, several window logos in the current frame image are scanned one by one in the preset database. Figure 4 For example, the current frame image has 3 windows. If these 3 windows have been edited before, their window identifiers will be stored in the preset database. At this time, these 3 window identifiers will be searched in the preset database, so that the window data of these 3 windows can be obtained from the preset database, that is, window 1, window 2 and window 3 in the figure, and the pixel data of the complete window.

[0156] 103. Encode the target form identifier and the form data corresponding to the form identifier not included in the preset database to obtain form encoding data.

[0157] After obtaining the same target form identifier in the preset database, the target form identifier is encoded to obtain encoded data.

[0158] Since the same preset database is also maintained at the image decoding end, the image decoding end can also find the corresponding window data from the preset database maintained by the image decoding end based on the target window identifier in the encoded code stream to restore the window.

[0159] Furthermore, if the window identifier of the current frame image is not found in the preset database, it means that the current window is a new window that has never appeared in the previous historical frames, and the window data of the window is directly encoded, for example, using JPEG encoding.

[0160] Since the same preset database is also maintained at the image decoding end, since the window identifier is not saved in the preset database in the image encoding device, the preset database maintained at the image decoding end will also not have the window identifier, and there will be no window data corresponding to the window identifier. Therefore, the image encoding device will directly encode the window data of the window so that the image decoding end can directly obtain the window data based on the encoded code stream to restore the window.

[0161] The above two types of coded data together constitute the form coded data.

[0162] 104. Obtain hollow area data in the current frame image, where the hollow area is a display area excluding the window in the current frame image.

[0163] Since the computer screen displays not only the window area but also the area outside the window, the display area outside the window in the current frame image is called the hollow area in the present disclosure, and therefore, the hollow area data will also be obtained.

[0164] 105. Encode the hollow area data to obtain hollow area coded data.

[0165] The hollow area data is encoded, for example, using JPEG encoding, to obtain hollow area encoded data.

[0166] 106. Send a coded code stream to an image decoding device, where the coded code stream includes: window coded data and hollow area coded data.

[0167] After obtaining the window encoding data and the hollow area encoding data, the window encoding data and the hollow area encoding data can be packaged and sent to the image decoding device. Since the same preset database is also maintained in the image decoding device, when the image decoding device receives the window encoding data, the window can be restored from the preset database stored locally by combining the window encoding data and the preset database. The hollow area is restored based on the hollow area encoding data, and since the hollow area and the window area are restored at the same time, a complete image of the current frame is obtained.

[0168] The preset database maintained by the image encoding device in the present disclosure is the data of each window when several scene switches occurred in the history, but it does not adopt the solution of "saving multiple reference frames in their entirety and selecting the best one when encoding". It is saved in units of windows, so that when selecting a reference frame, the data of the entire key frame in history is not used as the reference object, but the window is used as the reference object, thereby refining the granularity of the reference object, and since only the data of the window is saved, it will not take up too much storage space, effectively saving the storage space of the encoder.

[0169] If a text editor keeps switching between the word interface and the IE interface, this situation may not be able to use conventional motion vectors and inter-frame prediction to solve the encoding problem, but it is obvious that there are a large number of similar elements between different frames, such as the title menu bar of the window, etc., which are all redundant. If the multi-reference frame of the prior art is used, it is necessary to save multiple complete frames in history. Taking 1920×1080 resolution as an example, the size of one frame is 1920×1080×3≈6MB, and 10 frames are 60MB, which will require a lot of storage space. Moreover, each of these 10 reference frames may have a word interface, which is equivalent to storing only this word interface 10 times, causing great waste. In the present disclosure, it is saved in units of windows, and the same window data will not be saved multiple times, thereby greatly saving storage space. For example, the word interface saved 10 times will only be saved once. Further, in the related art, when encoding the current frame, how to select one of the 10 reference frames that is most similar to the current frame for reference is also a difficult problem. In the present disclosure, however, data is not saved in units of frames, thereby reducing the amount of data to be stored. Since the amount of data to be stored is reduced, the amount of data used in calculations is reduced, thereby reducing the amount of calculations. Most importantly, if 10 window identifiers and corresponding window data are stored, then these 10 windows are likely to come from different frames. When encoding the current frame, if the current frame has several of these 10 windows, then these several windows can be directly compared for reference, that is, several windows can be used for reference at the same time. Since these several windows come from different frames, it is equivalent to referring to multiple reference frames at the same time. This is essentially different from the prior art that can only refer to one reference frame in the end.

[0170] Since several windows can be used for reference at the same time when encoding the current frame image, the window identifier can be directly encoded without encoding the window data. Since the image decoding device also stores 10 window identifiers and corresponding window data, the image decoding end can obtain the corresponding window data from the 10 window identifiers and the corresponding window data based on the window identifier in the encoded code stream to restore the current frame image, thereby reducing the encoded code stream.

[0171] The image transmission method provided by the embodiment of the present disclosure includes: obtaining the window information in the current frame image, the window information includes: each window identification and the window data corresponding to each window identification; according to each window identification, searching the target window identification that is the same as the window identification from the preset database; the preset database includes: the correspondence between the window identification and the window data in the historical frame image; the window identifications stored in the preset database are all different; encoding the window data corresponding to the target window identification and the window identification not included in the preset database to obtain the window encoding data; obtaining the hollow area data in the current frame image, the hollow area is the display area in the current frame image except the window; encoding the hollow area data to obtain the hollow area encoding data; sending the window encoding data and the hollow area encoding data to the image decoding device. In the present disclosure, each window in the historical image frame is saved as a unit. If there are multiple windows in the window of the current frame image that are the same as the window in the historical image frame, then these windows can be used for reference. Since these windows come from different historical frames, it is equivalent to referring to multiple reference frames at the same time.

[0172] Furthermore, the above step 103 includes the following sub-steps:

[0173] A1. Divide a window corresponding to a target window identifier in a preset database into a plurality of first sub-windows according to a first preset rule.

[0174] The sub-window is taken as a macro block for explanation.

[0175] At this time, similar to most encoding and decoding solutions, this solution is also based on macroblocks, and the window corresponding to the target window identifier is divided into multiple first macroblocks according to a preset rule.

[0176] A2. Divide the window corresponding to the target window identifier in the current frame image into a plurality of second sub-windows according to a first preset rule.

[0177] Continuing with the above example, the window corresponding to the target window identifier of the current frame image is divided into a plurality of second macroblocks according to the same first preset rule.

[0178] A3. If the window data in the first subwindow at the same position is the same as the window data in the second subwindow, the table entry at the position corresponding to the second subwindow with the same window data in the first preset position table is marked as a first mark; the position of each table entry in the first preset position table corresponds to the position of each second subwindow in the current frame image.

[0179] Continuing with the above example, the first preset position table at this time can also be called a macroblock mark table, which records whether each first macroblock has the same content as the second macroblock at the same position of the target window, wherein the first mark is, for example, 1. Continuing with the above example:

[0180] If the window data in the first macroblock at the same position is the same as the window data in the second macroblock, the entry at the position corresponding to the second macroblock having the same window data in the first macroblock mark table is marked as 1.

[0181] For example: if the window in the current frame is exactly the same as the content of the target window stored in the preset database, it is similar to the effect of minimizing and then restoring. In this case, the table entry corresponding to each second macroblock in the first macroblock mark table is marked as 1.

[0182] A4. If the form data in the first sub-form and the form data in the second sub-form at the same position are different, the table entry at the position corresponding to the second sub-form with different form data in the first preset position table is marked as a second mark;

[0183] For example: the second marker is 0.

[0184] Continuing with the above example, if the window data in the first macroblock at the same position is different from the window data in the second macroblock, the entry corresponding to the position of the second macroblock with different window data in the first macroblock mark table is marked as 0.

[0185] For example, if the window in the current frame has records in the preset database (the original window is there, but may be temporarily blocked or minimized), but the window data has changed, then there will be a situation where both mark 1 and mark 0 are present.

[0186] A5. Divide the window in the current frame image corresponding to the window identifier not included in the preset database into a plurality of third sub-windows according to a first preset rule.

[0187] A6. Mark each entry in the second preset position table as a second mark, wherein the position of each entry in the second preset position table corresponds to the position of each third subwindow in the current frame image;

[0188] For example, if there is a newly running program in the current frame (the newly running program corresponds to a newly generated window, not a window that is restored after being minimized before), at this time, the window is a window that is not included in the preset database. Continuing with the above example, at this time, all entries in the second macroblock mark table are marked as 0.

[0189] A7. Encode the target window identifier, the marked first preset position table, the window data in the current frame image corresponding to the second marked table item, the marked second preset position table, and the window data in the current frame image corresponding to each table item in the second preset position table to obtain window encoding data.

[0190] Among them, other direct encoding schemes can be adopted for the window data in the current frame image corresponding to the second sub-window with different window data, such as JPEG encoding, or entropy encoding and other schemes. There are many similar direct encoding schemes in the prior art, which is not the content proposed by this patent, so it is not described in detail. Here, JPEG encoding is taken as an example. The data stream after JPEG encoding will also be merged into the encoded code stream finally sent to the image receiving device.

[0191] For example, the editor window in the current frame is not exactly the same as the content saved in the preset database. Figure 5 As shown in the figure, the left picture is the pixel effect of the editor window recorded in the preset database, and the right picture is the pixel effect of the editor window in the current frame (the shadow effect is added later). Figure 5 It can be seen that the menu bar, toolbar, icons, etc. of the form have not changed, but the text in the form has changed. Figure 5 The unchanged area is marked with a shadow. For this situation, in the first macroblock mark table, the entry corresponding to the shaded part (also referred to as the macroblock in the first macroblock mark table) is marked as the first mark, and the macroblock in the first macroblock mark table corresponding to the non-shaded part is marked as the second mark; the image decoding device can know how to decode the macroblock at each position from the first macroblock mark table, and the macroblock marked as the first mark in the first macroblock mark table (shaded part) will be directly copied from the preset database. For the non-shaded part, since it is different from the target window stored in the preset database, it is necessary to adopt other direct encoding schemes, such as JPEG encoding, or entropy encoding. There are many similar direct encoding schemes in the prior art, which is not the content proposed by this patent, so it is not elaborated here. Here, JPEG encoding is taken as an example. The data stream encoded by JPEG will also be merged into the encoded code stream finally sent to the image receiving device, and the remaining window information, such as the position, size and handle of the window, can be directly encoded into the code stream.

[0192] Since the first preset position table and the second preset position table are encoded together during encoding, and since the image decoding device synchronously stores a preset database, when the image decoding device receives the window encoding data, the corresponding window data can be found from the preset database based on the target window identifier in the encoding data in combination with the window encoding data and the preset database, and the table item corresponding to the first mark can be restored based on the first preset position table, and then the table item corresponding to the second mark in the first preset position table can be restored based on the window data in the current frame image corresponding to the second subwindow with different window data, and the new window can be restored based on the marked second preset position table and the encoding data of the window data corresponding to each third subwindow. Since it is not necessary to encode all the window data during encoding, only different window data need to be encoded, and only the window identifier needs to be encoded for the same window data, thereby greatly improving the encoding efficiency and reducing the amount of data sent.

[0193] In one embodiment, the method further comprises the following sub-steps:

[0194] B1. Add the form identifiers and corresponding form data not included in the preset database to the preset database.

[0195] B2. Replace the form data corresponding to the form identifier in the preset database with the form data corresponding to the form identifier in the current frame image.

[0196] In order to make

[0197] B3. Sending first update information in the encoded bitstream to the image decoding device, wherein the first update information includes: all updated data in the preset database.

[0198] After the current frame is edited, the window data in the current frame needs to be updated in the preset database.

[0199] If there is a newly running program in the current frame (the newly running program corresponds to a newly generated window, not one that was previously minimized and now restored), its window identifier and corresponding window data are added to the preset database; if the window in the current frame has a record in the preset database (the original window is there, it may be temporarily blocked or minimized), but the content of the window has changed, the window data corresponding to each window is updated to the preset database.

[0200] Specifically, if the window identifier in the current frame already exists in the preset database, and the window data is exactly the same as that recorded in the preset database, it means that the window has not changed between the two frames, and there is no need to update; if there is a difference with that recorded in the preset database, the window data in the current frame needs to be updated to the preset database so that subsequent frames refer to the latest window data. At this time, the first update information may be, for example: generating a record of "updating the old window", and the record carries the window identifier of the old window.

[0201] Furthermore, if the window identifier of a window in the current frame does not exist in the preset database, the window identifier is added to the preset database. If the preset database is full, the oldest window identifier is removed and a new window identifier is added. Accordingly, the window data of the window is stored in the preset database at the same time.

[0202] The first update information at this time may be, for example: generating a record of "generating a new form", in which the form identifier of the new form is carried.

[0203] Through the above-mentioned updating operation, it can be ensured that when encoding the current frame, the image encoding device stores the latest window identification and window data encountered in history.

[0204] Furthermore, all updated data in the preset database and the previous encoded data are packaged together into an encoded code stream and sent to the image decoding device, in order to notify the image decoding device that a new window has appeared and the window identifier and window data need to be updated synchronously.

[0205] In one embodiment, in order to further improve the coding efficiency, corresponding processing is also performed on the hollowed-out area data. Specifically, the above step 105 includes the following sub-steps:

[0206] C1. Detect whether the table items of the preset mapping table include a target table item whose similarity with the window identifier of the current frame image meets the preset conditions, and each table item of the preset mapping table includes: the mapping relationship between the current table item identifier, the historical hollow area data and all the window identifiers in the full-frame image corresponding to the hollow area data.

[0207] The image coding device also maintains a preset mapping table, which stores the hollow area data of a total of M scene switching frames (also known as key frames) in history, so as to provide a reference for the predictive coding of the hollow area data part and save storage space.

[0208] Each entry in the preset mapping table includes: a mapping relationship between a current entry identifier, historical hollowing data, and all window identifiers in the full-frame image corresponding to the hollowing area data.

[0209] C2. If included, encode the current entry identifier included in the target entry to obtain hollow area encoding data.

[0210] C3. If not included, encode the hollow area data to obtain hollow area coded data.

[0211] It should be noted that the window identifier and the window data in the preset database are one-to-one corresponding, that is, one window identifier corresponds to one window data, and a maximum of N entries can be saved, that is, the data of N windows. The preset mapping table is used to save the data of the hollow area, and a maximum of M hollow area data can be saved according to the preset rules. The number of table entries in the hollow area is not necessarily the same as N, because although the window identifier and the window are one-to-one corresponding, these N windows may come from different frames or from the same frame. This is not important because this scheme is based on each window as a prediction object, not each frame as a prediction object. The hollow area data is only a prediction object for the part outside the window, and its number M is unrelated to N. The M value can be selected according to actual conditions such as the storage capacity of the device and the size of the computing power. The larger the value, the greater the amount of calculation, the larger the storage capacity required, the better the prediction effect, and the lower the bit rate.

[0212] The preset mapping table saves the hollow area data. Specifically, the M data in the preset mapping table not only contains M hollow area data, but also stores the corresponding window identifiers contained in the full-frame picture to which each hollow area data originally belongs. The significance of this design is to quickly determine: when editing the hollow area of ​​the current frame, which hollow area data entry to refer to. In theory, the easiest way is to compare the hollow area data of the current frame with the M hollow area data to see which one is the most similar (there are multiple evaluation criteria, such as the smallest sum of pixel differences), and use the hollow area data as the reference standard. However, this method will result in a large amount of calculation. The new solution proposed in this scheme is as follows:

[0213] Given the window identifiers in the current frame image, search one by one in the preset mapping table to see which entry's window identifier records overlap the most with the window identifier in the current frame image. The hollow area data of the entry is used as the reference data. If there are multiple entries whose window identifiers overlap with the window identifier in the current frame, the hollow area data of the latest entry is used as the reference data. For example, there are 3 windows in the current frame image, and the second entry in the preset mapping table stores 4 windows in the original frame to which it belongs, 3 of which are exactly the same as the 3 window identifiers of the current frame image. That means that the original frame where the hollow area data corresponding to the second entry is located has 3 windows, which are the 3 windows of the current frame. Then the two frames should be extremely similar, and it is also the most appropriate to use the hollow area data corresponding to the second entry as a reference. This selection process completely avoids the process of comparing and selecting the best pixel by pixel. Among them, the hollow area data corresponding to the second entry can also be described as the hollow area data corresponding to the second table item.

[0214] In one embodiment, encoding the current entry identifier included in the target entry to obtain hollow area encoding data includes the following sub-steps:

[0215] D1, dividing the hollowed-out area of ​​the current frame image into a plurality of first sub-hollowed-out areas according to a second preset rule;

[0216] D2. Divide the hollow area corresponding to the target table item into a plurality of second sub-hollow areas according to a second preset rule;

[0217] D3, if the hollow area data in the first sub-hollow area at the same position is the same as the hollow area data in the second sub-hollow area, then mark the table entry at the position corresponding to the second sub-hollow area with the same hollow area data in the third preset position table as the first mark; the position of each table entry in the third preset position table corresponds to the position of each second sub-hollow area in the current frame image;

[0218] D4. If the hollow area data in the first sub-hollow area and the hollow area data in the second sub-hollow area at the same position are different, mark the entry corresponding to the position of the second sub-hollow area with different hollow area data in the third preset position table as a second mark;

[0219] D5. Encode the hollow area data in the current frame image corresponding to the current entry identifier, the marked third preset position table, and each entry marked as the second mark in the third preset position table to obtain hollow area coded data.

[0220] For example, the sub-hollowed area may be a macroblock, and the third preset position table may be a third macroblock mark table.

[0221] The hollow area is also composed of macroblocks. Each macroblock has a flag set to record whether the macroblock is the same as the hollow area data in the target table. If they are not the same, they will be directly encoded, such as JPEG encoding. Finally, the same parts of the hollow area and the target table do not need to be encoded, and the different parts record the results of JPEG encoding.

[0222] In one embodiment, in order to ensure that when encoding the current frame, the image encoding device stores the latest hollowed-out area data encountered in history, the method further includes the following sub-steps:

[0223] E1. Detect whether the current frame image is a frame image after scene switching.

[0224] E2. If yes, then add a new entry to the preset mapping table, and save the hollow area data corresponding to the current frame image and all the window identifiers of the current frame image into the new entry.

[0225] E3. Send the second update information in the encoded bitstream to the image decoding device, where the second update information includes: the newly added entry information in the preset mapping table.

[0226] In addition, if a scene switch occurs, the data outside each window, which is called the "hollow area" in this scheme, needs to be updated to the hollow data area. If the current frame is not a scene switch frame, that is, the current frame is very similar to the previous frame, then the update process of this step is not performed.

[0227] This ensures that the hollow area data of the M entries saved in the preset mapping table come from the latest M scene switching frames. That is, in actual implementation, whenever a scene switch occurs, the hollow area data of the frame is saved in the preset mapping table. If the saved hollow area entry data is greater than M, the oldest entry data is deleted.

[0228] In addition, the update of the preset mapping table needs to be performed only when a scene switch occurs, that is, the difference between the previous frame and the current frame is "huge". If the current frame is a scene switch frame, the hollow area data of the current frame and the window identifier of the current frame are stored in a new entry to complete the update. For example, the second update information mentioned above can include generating a record of "new hollow data".

[0229] In one embodiment, the above step E1 includes the following sub-steps:

[0230] F1. Divide the current frame image into macroblocks according to a third preset rule.

[0231] F2. Divide the previous frame image into macroblocks according to a third preset rule.

[0232] F3. Detect the similarity between the macroblock corresponding to the current frame image and the macroblock corresponding to the previous frame image.

[0233] F4. When the similarity is greater than a preset value, it is determined that a scene switch occurs.

[0234] For example, there are multiple ways to determine whether a scene switch has occurred. One optional implementation method is to divide the current frame image and the previous frame image into macroblocks according to the same rules. After dividing the macroblocks, compare each macroblock to determine whether there is a change. If the number of macroblocks that have not changed is greater than 50% of the number of macroblocks in the full frame, it is considered that a scene switch has occurred. The above 50% judgment standard can be adjusted according to actual needs. Alternatively, if the number of macroblocks that have changed is less than 50% of the number of macroblocks in the full frame, it is considered that a scene switch has occurred. The above 50% judgment standard can be adjusted according to actual needs.

[0235] The present disclosure provides an image transmission method. Figure 6 As shown, the method is applied to an image decoding device, and the image transmission method comprises the following steps:

[0236] 201. Receive a coded code stream, the coded code stream includes: window coding data and hollow area coding data; the window coding data is obtained by encoding the window data corresponding to the target window identifier and the window identifier not included in the preset database, the preset database includes: the correspondence between the window identifier and the window data in the historical frame image; the window identifiers saved in the preset database are all different; the window data includes: window pixel data, window position and window size; the hollow area is the display area excluding the window in the current frame image.

[0237] 202. Restore the form corresponding to the target form identifier according to the preset database and the form coding data.

[0238] 203. Restore the form corresponding to the form identifier not included in the preset database according to the form data corresponding to the form identifier not included in the preset database.

[0239] 204. Restore the hollow area according to the hollow area coding data.

[0240] 205. Acquire a current frame image according to the window corresponding to the restored target window identifier, the window corresponding to the window identifier not included in the preset database, and the hollow area.

[0241] After receiving the coded code stream, the image coding device performs decoding, which is similar to the process at the coding end. It can be mainly divided into two parts, one is the window, and the other is the hollow area outside the window.

[0242] Since the image encoding device and the image decoding device maintain the same preset database, the window data corresponding to the target window identifier can be obtained from the preset database based on the target window identifier in the window encoding data in the encoded code stream, that is, the window pixel data, window position and window size of the window corresponding to the target window identifier are obtained, so that the window corresponding to the target window identifier can be restored based on the window data.

[0243] The form corresponding to the form identifier not included in the preset database is restored according to the form data corresponding to the form identifier not included in the preset database in the encoded code stream.

[0244] The hollow area is restored according to the hollow area coding data in the coded code stream.

[0245] Finally, the current frame image is acquired according to the window corresponding to the restored target window identifier, the window corresponding to the window identifier not included in the preset database, and the hollow area.

[0246] Since the coded code stream contains not the window data of all the windows of the current frame image but the window identifiers of some windows, the coded code stream is reduced. Moreover, since the present disclosure saves the windows in the historical image frames as units, if there are multiple windows in the current frame image that are the same as the windows in the historical image frames, then these windows can be used for reference. Since these windows come from different historical frames, it is equivalent to referring to multiple reference frames at the same time.

[0247] The image transmission method provided by the embodiment of the present disclosure includes: receiving a coded code stream, the coded code stream includes: window coding data and hollow area coding data; the window coding data is obtained by coding according to the target window identifier and the window data corresponding to the window identifier not included in the preset database; the preset database includes: the correspondence between the window identifier and the window data in the historical frame image; the window identifiers stored in the preset database are all different; according to the preset database and the window coding data, the window corresponding to the target window identifier is restored; according to the window data corresponding to the window identifier not included in the preset database, the window corresponding to the window identifier not included in the preset database is restored; according to the hollow area coding data, the hollow area is restored; according to the window corresponding to the restored target window identifier, the window corresponding to the window identifier not included in the preset database, and the hollow area, the current frame image is obtained. Since the present disclosure saves in units of each window in the historical image frame, if there are multiple windows in the window of the current frame image that are the same as the window in the historical image frame, then these windows can be used for reference. Since these windows come from different historical frames, it is equivalent to referring to multiple reference frames at the same time. Furthermore, since the coded code stream contains not the window data of all the windows of the current frame image but the window identifiers of some windows, the coded code stream is reduced.

[0248] In one embodiment, the encoded code stream further includes: first update information, the first update information includes: all data updated in a preset database, each entry in the preset mapping table includes: a mapping relationship between a current entry identifier, historical hollow area data, and all window identifiers in a full-frame image corresponding to the hollow area data; the method further includes:

[0249] The data in the preset database is updated according to the first update information.

[0250] Through the above-mentioned updating operation, it can be ensured that when decoding the coded code stream, the image decoding device stores the latest window identification and window data encountered in the history.

[0251] In one embodiment, the encoded bitstream further includes: second update information, the second update information includes: newly added entry information in the preset mapping table, and the method further includes:

[0252] The table entries in the preset mapping table are updated according to the second update information.

[0253] Through the above-mentioned updating operation, it can be ensured that when decoding the coded code stream, the image decoding device stores the latest hollow area data encountered in history.

[0254] In one embodiment, the encoded code stream includes: a target window identifier, a marked first preset position table, window data in the current frame image corresponding to the second marked table item, a marked second preset position table, and window data in the current frame image corresponding to each table item in the second preset position table;

[0255] Restore the form corresponding to the target form identifier according to the preset database and form coding data, including:

[0256] Restore the window corresponding to the target window identifier according to the preset database, the target window identifier, the marked first preset position table and the window data in the current frame image corresponding to the second marked table item;

[0257] Restoring a form corresponding to a form identifier not included in a preset database according to form data corresponding to the form identifier not included in a preset database includes:

[0258] The window corresponding to the window identifier not included in the preset database is restored according to the preset database, the marked second preset position table and the window data in the current frame image corresponding to each table item in the second preset position table.

[0259] In one embodiment, the encoded code stream includes: the current table entry identifier, the marked third preset position table, and the hollow area data in the current frame image corresponding to each table entry marked as the second mark in the third preset position table;

[0260] Restoring the hollow area according to the hollow area coding data includes:

[0261] The hollow area is restored according to the preset mapping table, the current table item identifier, the marked third preset position table, and the hollow area data in the current frame image corresponding to each table item marked as the second mark in the third preset position table. Each table item in the preset mapping table includes: the mapping relationship between the current table item identifier, the historical hollow area data and all the window identifiers in the full-frame image corresponding to the hollow area data.

[0262] Another embodiment of the present disclosure provides a module diagram of an image encoding device, such as Figure 7 As shown, the image encoding device includes: a window detection module 301, an image device interface virtual layer 302, a window prediction module 303, a window prediction result buffer 304, a code stream generation module 305, a window identification area 306, a window data area 307, a hollow data area 308 and a hollow area prediction encoding module 309.

[0263] The window identifier in the above embodiment is the window handle (Handle), the preset database includes: a window identification area 306 and a window data area 307; the preset mapping table includes: a hollow data area 308; the sub-window is a macroblock, the preset position table is a macroblock mark table, the first update information includes: "update old window" records and "generate new window" records, and the second update information includes: "new hollow data" records.

[0264] This embodiment combines Figure 7 The method in the present disclosure is described in detail.

[0265] Figure 7 The image encoding device receives the data to be encoded from the collector, which is usually RGB or YUV original pixel data.

[0266] After receiving a frame of image, the window detection module 301 will query the image device interface virtual layer 302 to find out how many windows are contained in the current frame. The function of the image device interface virtual layer 302 is to call the operating system related interface to obtain the handles of each window on the current screen in real time (for the specific implementation method of obtaining each window descriptor on the current screen, different operating systems may have different implementation schemes. Taking the Windows operating system as an example, the GetDesktopWindow and GetNextWindow function interfaces can be called to find all the window handles on the screen). Through the handle, the size and position information of each window can be queried.

[0267] The main functions of the window detection module 301 are: first, obtaining the handles of each window in the current frame; second, after the encoding is completed, updating the data in the window data area according to the following method:

[0268] If there is a newly running program in the current frame (the newly running program corresponds to a newly generated window, not a window that has been restored after being minimized before), its handle is inserted into the window identification area 306, and the window pixel data corresponding to the handle is updated to the window data area 307; if the windows in the current frame are all recorded in the window identification area 306 (the original window is there, but may be temporarily blocked or minimized), but the content of the window has changed, the pixel data corresponding to each window is updated to the window data area 307; in addition, if it is determined that a scene switch occurs at present, the data outside each window, that is, the data of the "hollow area" in this scheme needs to be updated to the hollow data area. If the current frame is not a scene switch frame, that is, the current frame is very similar to the previous frame, then the update process of this step is not performed. Specifically, there are multiple implementation methods for determining whether a scene switch occurs currently, one of which is to divide the current frame image and the previous frame image into macroblocks according to the same rule, and after dividing the macroblocks, compare each macroblock to determine whether there is a change, and if the number of changed macroblocks is less than 50% of the number of macroblocks in the full frame, it is considered that a scene switch occurs. The above 50% judgment standard can be adjusted according to actual needs.

[0269] Through the modules of the window detection module 301, the window identification area 306, the window data area 307, the hollow data area 308 and the above operations, it can be ensured that when encoding the current frame, the encoding end saves the latest window mark and data encountered in history.

[0270] It should be noted that the items in the window identification area 306 and the window data area 307 are one-to-one corresponding, that is, one handle corresponds to the data of one window, and a maximum of N items can be stored, that is, the data of N windows. 308 is used to store the data of the hollow area, and according to the preset rules, a maximum of M items of hollow area data can be stored. The number of items in the hollow area is not necessarily the same as N, because although the handle and the window are one-to-one corresponding, these N windows may come from different frames or from the same frame, which is not important, because this scheme is based on each window as the prediction object, rather than each frame as the prediction object. The hollow data area only prepares a prediction object for the part outside the window, and its number M is irrelevant to N. The M value can be selected according to the actual situation, such as the storage capacity of the device and the size of the computing power. The larger the value, the greater the amount of calculation, the larger the storage capacity required, the better the prediction effect, and the lower the bit rate. The hollow area data of the M entries saved in 308 come from the latest M scene switching frames, that is, in actual implementation, whenever a scene switch occurs, the hollow area data of the frame is saved in 308. If the saved hollow area entry data is greater than M, the oldest entry data is deleted.

[0271] The image device interface virtual layer 302 is called by the window detection module to provide it with the number of windows in the current screen and their respective handle information.

[0272] Window prediction module 303: After knowing the window information on the current frame by querying the image device interface virtual layer 302, it is necessary to search in the window identification area 306 to see if there is a window that has been edited in the past with the same handle as the window detected in the current frame. If so, it means that the window to be edited in this frame has appeared in the past, and the two windows are likely to be the same, or partially the same. First, the search method is to scan the handles of several windows in the current picture one by one in the window identification area 306. If there is a window that has been edited in the past, it is necessary to scan the handles of several windows in the current picture one by one in the window identification area 306. Figure 5 For example, the current screen has 3 windows. If these 3 windows have been edited before, their handles will be stored in the window identification area 306. At this time, the window prediction module 303 will search for the handles of these 3 windows in the window identification area 306, so that the data of these 3 windows, that is, the pixel data of the complete window, can be obtained from the window data area 307. After obtaining the predicted data, the window prediction module generates encoding information about these 3 windows. For the convenience of explanation, only one text editor window is taken as an example here, and the processing of other windows is similar. The editor window handle of this frame is searched in the window identification area. If the same one is not found, it means that the current editor window is a new window that has never appeared before. Then each macroblock in the editor window is directly encoded, for example, using JPEG encoding. After the whole frame is edited, a new entry is inserted into the window identification area 306, that is, the current editor window handle, and a new record is inserted into the window data area 307, that is, the pixel data of the current window. And generate an update record, which is passed to the code stream generation module 305 through the window prediction result buffer 304 module as part of the window encoding. The purpose is to notify the decoding end that a new window has appeared and the window identification area and the window data area need to be updated synchronously.

[0273] If the same handle as the editor window is found in the window identification area 306, there will be two situations:

[0274] (1) The editor window in this frame is exactly the same as the content of the editor window saved in the window data area. This is similar to the effect of minimizing and then restoring. In this case, the encoder only needs to record the position, size, and identification number in the title identification area of ​​the current editor window, or directly save the window handle, such as Handle1. This information will be put into the window prediction result buffer 304, and finally put into the code stream generation module 305, that is, the encoded code stream of the area occupied by this editor window in this frame. Since the decoder synchronously saves the three buffers (window identification area 306, window data area 307, and hollow data area 308) with exactly the same content as the encoder, the pixel data of the window can be directly restored at the decoder based on this information. It should be noted that, similar to most encoding and decoding schemes, this scheme is also based on macroblocks, so the window will also be divided into small macroblocks, and each window corresponds to a macroblock mark table, which records whether each macroblock has the same content as the macroblock at the same position of the reference window data. For example, if they are the same, the mark is 1, and if they are not the same, the mark is 0. Obviously, in this case, the flags of each macroblock in the window will be marked as the same. Accordingly, the macroblock flag table will also be encoded into the bitstream and transmitted to the decoding end.

[0275] (2) The editor window in this frame is not exactly the same as the content saved in the window data area. For example, as shown in the figure below, the left picture is the pixel effect of the editor window recorded in the window data area, and the right picture is the pixel effect of the editor window in the current frame (the shadow effect is added later). From the figure, it can be seen that the menu bar, toolbar, icons, etc. of the window have not changed, but the text in the window has changed. The unchanged area is marked with a shadow in the figure. For this situation, it is slightly different from situation (1), which is reflected in:

[0276] a. Macroblock mark table. In this example, the macroblocks in the shaded part are marked as the same, and the macroblocks in the non-shaded part are marked as different; the decoding end can know how to decode the macroblocks in each position from this table, and the macroblocks marked as the same will be directly copied from the form data area.

[0277] b. For the non-shadow part, since it is different from the reference window, other direct encoding schemes need to be adopted, such as JPEG encoding, or entropy encoding. There are many similar direct encoding schemes in the prior art, which is not the content proposed by this patent, so it is not described in detail. Here, JPEG encoding is taken as an example. This part of the work can also be completed by the 303 window prediction module. The data stream after JPEG encoding will also be merged into the final code stream.

[0278] The rest of the information, such as the position, size and handle of the window, can be directly encoded into the bitstream as in case (1).

[0279] (3) After each window of the current frame has been predicted and encoded as described above, the data corresponding to the window will be stored in the window prediction result buffer 304. In the end, it will store the encoded data of all windows. The decoding end should be able to restore the original pixel values ​​of all these windows based on these data.

[0280] The window prediction result buffer 304 already has the coded data of all windows. In that frame, the part outside all windows, that is, the hollow area, has not been coded. In this scheme, the hollow data of a total of M scene switching frames (also understood as key frames) in history will be saved. This can provide a reference for the prediction coding of the hollow data part and save storage space. The hollow data area 308 stores these hollow data. Specifically, the M data in the hollow data area 308 not only contain M hollow data, but also store the handles of each window contained in the full frame picture to which each hollow data originally belongs. The significance of this design is to quickly determine: when compiling the hollow data area of ​​the current frame, which hollow data entry is specifically referenced. In theory, the simplest way is to compare the hollow data of the current frame with the M hollow data pixels to see which one is the most similar (there are multiple evaluation criteria, such as the smallest sum of pixel differences), and use the hollow data as the reference standard. However, this method will result in a large amount of calculation. The new solution proposed in this scheme is as follows:

[0281] The handles of each window in the current frame are known. In the hollow data area 308, search one by one to see which entry has the largest number of overlaps with the window handle records of the window handle in the current frame. The hollow data of the entry is used as the reference data. If there are multiple entries whose window handles overlap with the window handle in the current frame, the hollow data of the latest entry is used as the reference data. For example, the current frame has 3 windows, and the ~Data2 entry in the hollow data area 308 stores that there are 4 windows in the original frame to which it belongs, and 3 of the windows are exactly the same as the 3 window handles of the current frame. That means that the original frame where the ~Data2 hollow data is located has 3 windows, which are the 3 windows of the current frame. Then these two frames should be extremely similar, and it is also most appropriate to use the ~Data2 hollow data as a reference. This selection process completely avoids the process of comparing and selecting the best pixel by pixel, and this work is completed by the hollow area prediction coding module 309.

[0282] Similar to the window prediction, the hollow area is also composed of macroblocks. A mark is set for each macroblock to record whether the macroblock is the same as the reference hollow area data. If they are not the same, they will be directly encoded, such as JPEG encoding. Finally, the parts of the hollow area that are the same as the reference data do not need to be encoded, and the different parts record the results of JPEG encoding. Since the window area has been recorded in 304, the position of the hollow area can be calculated through the window area, so there is no need to save the position and size information of the hollow area additionally. Only the hollow area macroblock is marked, and the entry number in the hollow data area 308 is referenced, and the result after JPEG encoding is passed to the code stream generation module 305.

[0283] After the current frame is edited, the window data in the current frame needs to be updated to the window data area 307, and the steps are:

[0284] (1) If the handle of the window in the current frame already exists in the window identification area 306, and the window data is exactly the same as that recorded in 307, it means that the window has not changed between the two frames and does not need to be updated; if there is a difference with the record in the window data area 307, the data of the window in the current frame needs to be updated to the window data area 307 so that the subsequent frames refer to the latest window data. A record of "updating the old window" is generated, which carries the handle number of the old window.

[0285] (2) If the handle of a window in the current frame does not exist in the window identification area 306, the handle of the window is added to the queue of the window identification area 306. If the queue is full, the oldest window handle is removed and the new window handle is added. Correspondingly, the same process is performed in the window data area 307 to store the pixel data of the window. A record of "generating a new window" is generated, which carries the handle number of the new window.

[0286] In addition, the update of the hollow data area 308 needs to be performed only when a scene switch occurs, that is, the difference between the previous frame and the current frame is "huge". If the current frame is a scene switch frame, the hollow data of the current frame and the handle of the current frame window are stored in a new entry to complete the update. A record of "new hollow data" is generated.

[0287] At the same time, the updated information, that is, the "record" described above, will also be packaged into the coded bitstream through the bitstream generation module 305. The decoding end needs the updated information to synchronize the three tables of the window identification area, the window data area, and the hollow data area of ​​the decoding end.

[0288] The bitstream generation module 305 will eventually integrate the coded data provided by the window prediction result buffer 304, i.e., the coded data of all windows, and the coded data provided by the bitstream generation module 305, i.e., the coded data of the hollow areas outside all windows, and the update information of the three buffer tables (not necessarily present in every frame) to generate the final coded bitstream.

[0289] The above-mentioned updating action of the window identification area 306 , the window data area 307 and the hollow data area 308 can be completed by the window detection module 301 after the encoding is completed.

[0290] Another embodiment of the present invention provides a module diagram of an image encoding device, such as Figure 8 As shown, the image decoding device includes: a prediction decoding module 310, a frame decoding data module 311, a prediction data updating module 312, a window identification area, a window data area and a hollow data area.

[0291] The above is the process of the decoding end, which is mostly the reverse operation of the encoding end. The following describes the functions of its various components.

[0292] The prediction decoding module 310 receives the coded bit stream and performs decoding. Similar to the process at the encoding end, it can be mainly divided into two parts, one is the window part, and the other is the hollow part outside the window. Decoding is performed by referring to the data in the three tables of the window identification area, the window data area, and the hollow data area. After decoding, the pixel data of a complete frame can be obtained, as shown in the frame decoding data module S311.

[0293] When the prediction data update module 312 was encoded into the bitstream at the encoding end, it brought in the update records of the three tables. The prediction data update module 312 obtains these update records, and after the decoding end finishes decoding the current frame, it updates the entries of the corresponding three tables according to the instructions of the update data. For example, a new window is displayed in the current frame, and its handle is x. Its window data can be obtained in the frame decoding data module 311, and the current frame is a scene switching frame, and the hollow data area can also be obtained in the frame decoding data module 311. The following records should be in the encoded bitstream:

[0294] (1) A new form record with handle x.

[0295] (2) New hollowing data record.

[0296] After receiving these two records, the prediction data update module 312 performs the following operations:

[0297] (1) Insert handle number x into the form identification area queue. If the queue is full, remove the oldest inserted handle;

[0298] (2) The pixel data of the window marked by the handle number x is taken out from the frame decoding data module 311, stored in a data block, and the data block is inserted into the queue of the window data area to correspond to the newly inserted entry in the window identification area.

[0299] (3) After the prediction decoding module 310 completes decoding the hollowed-out data area, the decoding end obtains the hollowed-out data area of ​​the current frame. Since the decoding end receives the "new hollowed-out data record" from the encoding end, the decoding end needs to insert the hollowed-out data of the frame into the hollowed-out data area.

[0300] After the above steps, the update of the three tables is completed, and these three tables are synchronized with the encoding end.

[0301] The general idea of ​​this solution has been described above. The main idea is to use the characteristics of "obtaining the window handle" and "obtaining the window size and position information based on the window handle" to save the historical window pixel data, rather than the entire key frame data in history as a reference object. This is the most important technical point of this solution. The benefit of this technical point is that it makes the granularity of the reference object refined. Finally, let's take the simplest example. If a text editor keeps switching between the word interface and the IE interface, this situation may not be able to use conventional motion vectors and inter-frame prediction to solve the encoding problem, but it is obvious that there are a large number of similar elements between different frames, such as the window's Title menu bar, etc., these are all redundant. If the existing technology of multi-reference frame is used, it is necessary to save multiple complete frames in history. Taking 1920x1080 resolution as an example, the size of one frame is 1920x1080x3≈6MB. 10 frames is 60MB, which requires a lot of storage space. Moreover, in these 10 reference frames, each frame may have a word interface, which is equivalent to saving this word interface 10 times, causing great waste. Moreover, for the current encoded frame, how to select one of the 10 reference frames that is most similar to the current frame for reference is also a difficult problem. At present, there is no efficient solution that can be used in industrialization.

[0302] The proposal of this solution solves the problem of multiple historical references.

[0303] In summary, the main advantage of the scheme disclosed in the present invention is that it fully considers the particularity of scene switching of computer-generated images, utilizes the characteristics of being able to obtain window handles and window information through the bottom layer of the operating system, and realizes a content-based multi-reference prediction scheme. The characteristics are that it occupies a small device storage space, has a small amount of calculation, and can reduce the code stream, ensuring that the code stream generated by the encoding is suitable for transmission in a wireless environment.

[0304] Combination Figure 7 and Figure 8The present invention relates to the field of image processing. The main feature of the scheme is that elements in multiple frames can be referenced simultaneously during encoding, and the scheme is completely implemented by software.

[0305] This solution is specifically for computer-generated images, so it fully considers the user switching between different windows, which is usually called "scene switching". A new multi-reference prediction solution is proposed to address this phenomenon. The key invention points are:

[0306] The codec side maintains multiple window data, but it does not adopt the solution of "saving multiple reference frames intact and selecting the best ones during encoding". Instead, it intercepts the operating system instructions or interfaces in the underlying software to obtain the descriptors of each window on the current screen image, or handles (for the sake of uniformity, this article will use the most frequently used term in the industry, namely handle). The handle is the ID of the window. Each window corresponds to a unique ID. Therefore, different windows can be identified by this ID. The handle can also be used to obtain the window's position, size and other information. The codec side mainly maintains the data of each window when several scene switches occurred in history.

[0307] This solution first reduces the amount of data storage, secondly reduces the amount of calculation, and most importantly, if, for example, pixel information of 10 windows is recently stored, then these 10 windows are likely to come from different reference frames. When encoding the current frame, if the current frame has several of these 10 windows, then these several windows can be directly compared for reference, which is equivalent to referring to multiple reference frames at the same time. This is fundamentally different from the existing technology that can only refer to one reference frame in the end.

[0308] Based on the above Figure 3 and Figure 6 Corresponding to the image transmission method provided in the embodiment, the present disclosure also provides an image transmission system, such as Fig. 9 As shown, it includes: an image encoding device 41 and an image decoding device 42;

[0309] The image encoding device is used to perform the method steps corresponding to the image encoding device described in any one of the above embodiments;

[0310] The image decoding device is used to execute the method steps corresponding to the image decoding device described in any one of the above embodiments.

[0311] Based on the above Figure 3 The image transmission method described in the corresponding embodiment is as follows: an embodiment of the device disclosed herein, which can be used to execute the embodiment of the method disclosed herein.

[0312] The present disclosure provides an image transmission device, such as Fig.10As shown, the device is applied to an image encoding device, and the device includes:

[0313] The first acquisition module 51 is used to acquire the window information in the current frame image, wherein the window information includes: each window identifier and the window data corresponding to each window identifier; wherein the window data includes: window pixel data, window position and window size;

[0314] A search module 52 is used to search for a target window identifier that is the same as the window identifier from the preset database according to each window identifier; the preset database includes: a correspondence between the window identifier and the window data in the historical frame image; the window identifiers stored in the preset database are all different;

[0315] A first encoding module 53, used for encoding the target window identifier and the window data corresponding to the window identifier not included in the preset database to obtain window encoding data;

[0316] A second acquisition module 54 is used to acquire hollow area data in the current frame image, where the hollow area is a display area in the current frame image excluding the window;

[0317] A second encoding module 55, used for encoding the hollow area data to obtain hollow area encoding data;

[0318] The first sending module 56 is used to send a coded code stream to the image decoding device, wherein the coded code stream includes: the window coded data and the hollow area coded data.

[0319] In one embodiment, the first encoding module includes:

[0320] A first division submodule, used for dividing the window corresponding to the target window identifier in the preset database into a plurality of first subwindows according to a first preset rule;

[0321] A second division submodule, used for dividing the window corresponding to the target window identifier in the current frame image into a plurality of second sub-windows according to the first preset rule;

[0322] A first marking submodule is used for marking the table item at the position corresponding to the second subwindow with the same window data in the first preset position table as a first mark if the window data in the first subwindow at the same position is the same as the window data in the second subwindow; the position of each table item in the first preset position table corresponds to the position of each second subwindow in the current frame image;

[0323] A second marking submodule, for marking the entry at the position corresponding to the second subwindow which is different from the form data in the first preset position table as a second mark if the form data in the first subwindow and the form data in the second subwindow at the same position are different;

[0324] A third division submodule, configured to divide the window in the current frame image corresponding to the window identifier not included in the preset database into a plurality of third sub-windows according to the first preset rule;

[0325] A third marking submodule, used for marking each table item in the second preset position table as the second mark, wherein the position of each table item in the second preset position table corresponds to the position of each third subwindow in the current frame image;

[0326] The first encoding submodule is used to encode the target window identifier, the marked first preset position table, the window data in the current frame image corresponding to the second marked table item, the marked second preset position table, and the window data in the current frame image corresponding to each table item in the second preset position table to obtain window encoding data.

[0327] In one embodiment, the apparatus further comprises:

[0328] A first adding submodule, used for adding the form identifier and the corresponding form data not included in the preset database to the preset database;

[0329] A replacement submodule, used to replace the window data corresponding to the window identifier in the preset database with the window data corresponding to the window identifier in the current frame image;

[0330] The second sending module is used to carry first update information in the encoded code stream and send it to the image decoding device, where the first update information includes: all data updated in the preset database.

[0331] In one embodiment, the second encoding module includes:

[0332] A first detection submodule is used to detect whether the table items in the preset mapping table include a target table item whose similarity with the window identifier of the current frame image meets a preset condition, and each of the table items in the preset mapping table includes: a mapping relationship between the current table item identifier, historical hollow area data and all window identifiers in the full-frame image corresponding to the hollow area data;

[0333] A second encoding submodule, for, if included, encoding the current entry identifier included in the target entry to obtain the hollow area encoding data;

[0334] The third encoding submodule is used to encode the hollow area data to obtain hollow area encoded data if it is not included.

[0335] In one embodiment, the second encoding submodule includes:

[0336] A fourth division submodule, configured to divide the hollowed-out area of ​​the current frame image into a plurality of first sub-hollowed-out areas according to a second preset rule;

[0337] A fifth division submodule, configured to divide the hollow area corresponding to the target table item into a plurality of second sub-hollow areas according to the second preset rule;

[0338] a fourth marking submodule, for marking the table entry at the position corresponding to the second sub-hollowed area having the same hollowed area data in the third preset position table as a first mark if the hollowed area data in the first sub-hollowed area at the same position is the same as the hollowed area data in the second sub-hollowed area; the position of each table entry in the third preset position table corresponds to the position of each second sub-hollowed area in the current frame image;

[0339] a fifth marking submodule, configured to mark the entry corresponding to the position of the second sub-hollow area with different hollow area data in the third preset position table as a second mark if the hollow area data in the first sub-hollow area and the hollow area data in the same position are different;

[0340] The fourth encoding submodule is used to encode the hollow area data in the current frame image corresponding to the third preset position table marked with the current table entry identifier and the second mark in the third preset position table to obtain the hollow area encoded data.

[0341] In one embodiment, the apparatus further comprises:

[0342] A detection module, used to detect whether the current frame image is a frame image after scene switching;

[0343] An adding module, used for adding a new table entry in the preset mapping table when the detection module detects that the current frame image is a frame image after scene switching, and saving the hollow area data corresponding to the current frame image and all the window identifiers of the current frame image to the new table entry;

[0344] The third sending module is used to carry the second update information in the encoded code stream and send it to the image decoding device, where the second update information includes: the newly added table item information in the preset mapping table.

[0345] In one embodiment, the detection module includes:

[0346] A sixth division submodule, used for dividing the current frame image into macroblocks according to a third preset rule;

[0347] A seventh division submodule, used for dividing the previous frame image into macroblocks according to the third preset rule;

[0348] A second detection submodule, used to detect the similarity between the macroblock corresponding to the current frame image and the macroblock corresponding to the previous frame image;

[0349] The determination submodule is used to determine that a scene switch occurs when the similarity is greater than a preset value.

[0350] Based on the above Figure 6 The image transmission method described in the corresponding embodiment is the following embodiment of the device disclosed herein, which can be used to execute the embodiment of the device disclosed herein.

[0351] The present disclosure provides an image transmission device, such as Fig.11 As shown, the device is applied to an image decoding device, and the device includes:

[0352] The receiving module 61 is used to receive a coded code stream, wherein the coded code stream includes: window coding data and hollow area coding data; the window coding data is obtained by encoding the window data corresponding to the target window identifier and the window identifier not included in the preset database; the preset database includes: the correspondence between the window identifier and the window data in the historical frame image; the window identifiers stored in the preset database are all different; the window data includes: window pixel data, window position and window size; the hollow area is the display area of ​​the current frame image except the window;

[0353] A first recovery module 62, used to recover the window corresponding to the target window identifier according to a preset database and the window encoding data;

[0354] A second restoring module 63, configured to restore a window corresponding to the window identifier not included in the preset database according to the window data corresponding to the window identifier not included in the preset database;

[0355] A third recovery module 64, configured to recover the hollowed-out area according to the hollowed-out area coding data;

[0356] The acquisition module 65 is used to acquire the current frame image according to the restored window corresponding to the target window identifier, the window corresponding to the window identifier not included in the preset database, and the hollow area.

[0357] In one embodiment, the encoded code stream further includes: first update information, the first update information includes: all data updated in the preset database, and the device further includes:

[0358] The first updating module is used to update the data in the preset database according to the first updating information.

[0359] In one embodiment, the encoded code stream further includes: second update information, the second update information includes: newly added table item information in the preset mapping table, each of the table items in the preset mapping table includes: a mapping relationship between a current table item identifier, historical hollow area data and all window identifiers in the full-frame image corresponding to the hollow area data; the device also includes:

[0360] The second updating module is used to update the table entries in the preset mapping table according to the second updating information.

[0361] In one embodiment, the encoded code stream includes: the target window identifier, the marked first preset position table, the window data in the current frame image corresponding to the second marked table item, the marked second preset position table, and the window data in the current frame image corresponding to each table item in the second preset position table;

[0362] A first recovery module is further used to recover the window corresponding to the target window identifier according to the preset database, the target window identifier, the marked first preset position table and the window data in the current frame image corresponding to the table entry of the second mark;

[0363] The second recovery module is further used to recover the window corresponding to the window identifier not included in the preset database according to the preset database, the marked second preset position table and the window data in the current frame image corresponding to each table item in the second preset position table.

[0364] In one embodiment, the encoded code stream includes: a current table entry identifier, a marked third preset position table, and the hollow area data in the current frame image corresponding to each table entry marked as a second mark in the third preset position table;

[0365] The third recovery module is also used to restore the hollow area according to the preset mapping table, the current table item identifier, the marked third preset position table, and the hollow area data in the current frame image corresponding to each table item marked as the second mark in the third preset position table. Each of the table items in the preset mapping table includes: the mapping relationship between the current table item identifier, the historical hollow area data and all the window identifiers in the full-frame image corresponding to the hollow area data.

[0366] Based on the above Figure 3In accordance with the image transmission method described in the embodiment, the embodiment of the present disclosure further provides a computer-readable storage medium. For example, the non-temporary computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, or an optical data storage device. The storage medium stores computer instructions for executing the above Figure 3 The data transmission method described in the corresponding embodiment will not be repeated here.

[0367] Based on the above Figure 6 In accordance with the image transmission method described in the embodiment, the embodiment of the present disclosure further provides a computer-readable storage medium. For example, the non-temporary computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, or an optical data storage device. The storage medium stores computer instructions for executing the above Figure 6 The data transmission method described in the corresponding embodiment will not be repeated here.

[0368] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0369] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image transmission method, characterized in that: The method is applied to an image encoding device, and the method comprises: When it is detected that the current frame image is a frame image after scene switching, the window information in the current frame image is obtained, the window information includes: each window identifier and the window data corresponding to each window identifier; wherein the window data includes: window pixel data, window position and window size; According to each of the window identifiers, a target window identifier identical to the window identifier is searched from a preset database; the preset database includes: a correspondence between the window identifiers and the window data in the historical frame image; the window identifiers stored in the preset database are all different; The target window identifier and the window data corresponding to the window identifier not included in the preset database are encoded to obtain window encoding data, and after obtaining the window encoding data, the preset database is updated according to the window encoding data.

2. The method according to claim 1, characterized in that The step of encoding the target window identifier and the window data corresponding to the window identifier not included in the preset database to obtain window encoding data includes: Dividing the window corresponding to the target window identifier in the preset database into a plurality of first sub-windows according to a first preset rule; Dividing the window corresponding to the target window identifier in the current frame image into a plurality of second sub-windows according to the first preset rule; If the window data in the first subwindow and the window data in the second subwindow at the same position are the same, the table entry at the position corresponding to the second subwindow having the same window data as the first subwindow in the first preset position table is marked as a first mark; the position of each table entry in the first preset position table corresponds to the position of each second subwindow in the current frame image; If the form data in the first sub-window and the form data in the second sub-window at the same position are different, marking the entry at the position corresponding to the second sub-window that is different from the form data of the first sub-window in the first preset position table as a second mark; Dividing the window in the current frame image corresponding to the window identifier not included in the preset database into a plurality of third sub-windows according to the first preset rule; Marking each table item in the second preset position table as the second mark, wherein the position of each table item in the second preset position table corresponds to the position of each third subwindow in the current frame image; The target window identifier, the marked first preset position table, the window data in the current frame image corresponding to the second marked table item, the marked second preset position table, and the window data in the current frame image corresponding to each table item in the second preset position table are encoded to obtain window encoding data.

3. The method according to claim 1, characterized in that After the form coding data is obtained, updating the preset database according to the form coding data includes: Adding the form identifier and the corresponding form data not included in the preset database to the preset database; The window data corresponding to the window identifier in the preset database is replaced by the window data corresponding to the window identifier in the current frame image.

4. The method according to claim 5, characterized in that The method further comprises: Detect whether the current frame image is a frame image after scene switching; If yes, then add a new table entry in the preset mapping table, and save the hollow area data corresponding to the current frame image and all the window identifiers of the current frame image into the new table entry; The second update information is carried in the encoded bitstream and sent to the image decoding device, wherein the second update information includes: the table entry information newly added in the preset mapping table.

5. The method according to claim 1, characterized in that The detecting whether the current frame image is a frame image after scene switching includes: Dividing the current frame image into macroblocks according to a third preset rule; Dividing the previous frame of image into macroblocks according to the third preset rule; Detecting the similarity between the macroblock corresponding to the current frame image and the macroblock corresponding to the previous frame image; When the similarity is less than a preset value, it is determined that a scene switch occurs.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Acquire hollow area data in the current frame image, where the hollow area is a display area in the current frame image excluding the window; Encoding the hollow area data to obtain hollow area coded data; Generate a coding stream according to the window coding data and the hollow area coding data, and send the coding stream to the image decoding device.

7. The method according to claim 6, characterized in that The step of encoding the hollow area data to obtain hollow area coded data includes: Detect whether the table items of the preset mapping table include a target table item whose similarity with the window identifier of the current frame image meets a preset condition, wherein each of the table items of the preset mapping table includes: a mapping relationship between the current table item identifier, historical hollow area data and all window identifiers in the full-frame image corresponding to the hollow area data; If included, encoding the current entry identifier included in the target entry to obtain the hollow area encoding data; If not included, the hollow area data is encoded to obtain hollow area encoded data.

8. The method according to claim 7, characterized in that The step of encoding the current entry identifier included in the target entry to obtain the hollowed-out area encoding data includes: Dividing the hollowed-out area of ​​the current frame image into a plurality of second sub-hollowed-out areas according to a second preset rule; Dividing the hollow area corresponding to the target table item into a plurality of first sub-hollow areas according to the second preset rule; If the hollow area data in the first sub-hollow area and the hollow area data in the second sub-hollow area at the same position are the same, the table entry corresponding to the position of the second sub-hollow area having the same hollow area data as the first sub-hollow area in the third preset position table is marked as a first mark; the position of each table entry in the third preset position table corresponds to the position of each second sub-hollow area in the current frame image; If the hollow area data in the first sub-hollow area and the hollow area data in the second sub-hollow area at the same position are different, the table entry corresponding to the position of the second sub-hollow area that is different from the hollow area data of the first sub-hollow area in the third preset position table is marked as a second mark; The hollow area coding data is obtained by encoding the hollow area data in the current frame image corresponding to the current entry identifier, the marked third preset position table, and each entry in the third preset position table marked as the second mark.

9. The method according to claim 6, characterized in that The method further comprises: Generate first update information according to all data updated in the preset database; The first update information is carried in the encoded bitstream and sent to the image decoding device.

10. An image encoding device, characterized in that: The device comprises: The acquisition module acquires window information in the current frame image when detecting that the current frame image is a frame image after scene switching, wherein the window information includes: each window identifier and window data corresponding to each window identifier; wherein the window data includes: window pixel data, window position and window size; A search module searches for a target window identifier that is identical to the window identifier from a preset database according to each window identifier; the preset database includes: a correspondence between the window identifier and the window data in the historical frame image; the window identifiers stored in the preset database are all different; The processing module encodes the target window identifier and the window data corresponding to the window identifier not included in the preset database to obtain window encoding data, and after obtaining the window encoding data, updates the preset database according to the window encoding data.