Shallow compressed video error concealment method and system based on wavelet transform

Through the shallow compressed video error hiding method based on wavelet transform, the decoding problem caused by packet loss during transmission of JPEG XS encoded signals is solved, and the video decoding is ensured successfully and the wrong pixel points are reconstructed without increasing bandwidth pressure.

CN118524223BActive Publication Date: 2025-05-13COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410689836.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-05-13
Estimated Expiration
2044-05-30

AI Technical Summary

Technical Problem

When processing JPEG XS encoded signals, network jitter causes packet loss in video streaming, affecting decoding accuracy. The existing solutions will increase bandwidth pressure when the data volume is large, making it difficult to effectively solve the decoding problem caused by packet loss.

Method used

The shallow compressed video error hiding method based on wavelet transformation is adopted. By obtaining the shallow compressed code stream, the slice index size and total number of slices are determined, slice analysis and sub-packet parsing are performed, the wrong pixel points are reconstructed, and the packet loss code stream is successfully decoded.

Benefits of technology

It realizes the decoding problem caused by packet loss without increasing the broadband burden, ensures that the video decoding is successful and the wrong pixel points are reconstructed, and improves the stability and quality of video streaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118524223B_ABST
    Figure CN118524223B_ABST
Patent Text Reader

Abstract

The present invention relates to a shallow compression video error concealment method and system based on wavelet transform, belonging to the technical field of coding and decoding, the method comprising: obtaining a shallow compression code stream; in the case where the shallow compression code stream includes multiple slices and the slices are decoded one by one, determining the index size of the current slice and the total number of slices; according to the comparison result of the index size of the current slice and the total number of slices, parsing each slice, parsing the area of ​​each slice, and parsing the sub-data packet of the area of ​​each slice to perform decoding of each slice; and obtaining the error area in the decoded image, reconstructing the pixel point based on the error pixel gradient value and target threshold of the error area. The method realizes the successful decoding of the compressed code stream with packet loss, and reconstructs the error pixel caused by the packet loss. And the decoding problem caused by packet loss is solved without increasing the broadband burden.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of coding and decoding, and in particular relates to a shallow compression video error concealment method and system based on wavelet transform. Background Art

[0002] In related technologies, JPEG XS encoded signals use the User Datagram Protocol (UDP), which provides unreliable transmission services, as the transport layer protocol. Due to network jitter, packet loss may occur during video stream transmission, affecting decoding accuracy. Currently, solutions to the JPEG XS packet loss problem are limited both domestically and internationally. The SMPTE 2022-7 standard backup stream approach is commonly used to address packet loss. Two data streams with identical content are transmitted over different links, with the primary and backup data sharing the same RTP header and RTP payload. Alternatively, a forward error correction (FEC) approach, designed for unequal error protection in JPEG XS, is employed. The redundancy level is determined based on the impact of byte loss in the JPEG XS codestream on the decoded image, and forward error correction (FEC) is used to recover lost packets. Both of these approaches place significant pressure on bandwidth given the rapidly growing volume of video data. Therefore, optimizing video decoding has become a pressing issue. Summary of the Invention

[0003] In view of the above shortcomings of the existing technology, the purpose of the present invention is to provide a wavelet transform-based shallow compression video error concealment method and system. This method can successfully decode a compressed bitstream with packet loss and reconstruct erroneous pixels caused by packet loss. It also solves the decoding problem caused by packet loss without increasing the bandwidth burden.

[0004] In a first aspect of the present invention, a wavelet transform-based shallow compression video error concealment method is proposed, comprising: S1, obtaining a shallow compression code stream; S2, when the shallow compression code stream includes multiple slices and the slices are decoded one by one, determining the index size of the current slice and the total number of slices; S3, based on a comparison result of the index size of the current slice and the total number of slices, parsing each slice, parsing the area of ​​each slice, and parsing the sub-data packet of the area of ​​each slice to perform decoding on each slice; and obtaining an error area in the decoded image, and reconstructing the pixel points based on the error pixel gradient value of the error area and a target threshold.

[0005] Furthermore, based on a comparison result of the index size of the current slice and the total number of slices, parsing is performed on each slice, including: if the index size of the current slice is less than or equal to the total number of slices, detecting whether a data error occurs in the current slice; if a data error occurs in the current slice, not decoding the current slice, and synchronizing the decoder with a slice next to the current slice to perform parsing on each slice; obtaining the shallow compressed code stream after parsing each slice; and determining whether a sequence number of a slice in the parsed shallow compressed code stream is discontinuous. If so, determining that a packet is lost in the parsed shallow compressed code stream, and replacing the lost packet data with zeros.

[0006] Furthermore, parsing the region of each slice and parsing the sub-data packets of the region of each slice includes: obtaining a shallow compressed data bare stream; when parsing each slice one by one, determining whether the expected entropy coded data length of each region in each slice is unreasonable; if unreasonable, the entire region cannot be parsed, and parsing of other regions of the slice is performed; if the expected entropy coded data length of the region is reasonable, obtaining the sub-data packets of each region, wherein each region includes multiple sub-data packets; determining the entropy coding length of the sub-data packets according to the data sub-packet size, the bit plane count sub-packet size, and the symbol sub-packet size of the sub-data packets of each region, and performing parsing of the sub-data packets based on the entropy coding length and the preset length.

[0007] Furthermore, based on the entropy coding length and the preset length, the sub-data packet is parsed, including: determining whether the entropy coding length is the preset length; if the entropy coding length is not the preset length, not parsing the sub-data packet, but parsing other sub-data packets in the corresponding area; until the parsing of the sub-data packets in the area of ​​each slice is completed.

[0008] Further, the error area in the decoded image is obtained, and the pixel points are reconstructed based on the error pixel gradient value and the target threshold of the error area, including: locating the error image line number during decoding, and obtaining the error image of the decoded image; according to , calculate the error pixel gradient value of the error area; wherein, Indicates the gradient value of the X axis, , Indicates the gradient value of the Y axis, , where P1, P2, P3, P5, P6, P7 and P8 all represent pixel values ​​of pixel points; when the gradient value of the error pixel in the error area is greater than the target threshold, the pixel point is reconstructed by using the difference based on the gradient direction; when the gradient value of the error pixel in the error area is not greater than the target threshold, the pixel point is reconstructed by using smooth interpolation.

[0009] Further, when the error pixel gradient value in the error area is greater than the target threshold, the pixel point is reconstructed by using the difference based on the gradient direction; when the error pixel gradient value in the error area is not greater than the target threshold, the pixel point is reconstructed by using smooth interpolation, including: when the error pixel gradient value in the error area is greater than the target threshold, according to , calculate the reconstructed pixel points; wherein, ,in, Represents the coordinates to be interpolated, Indicates the pixel coordinates corresponding to the gradient direction, h1 and h2 indicate the pixel values ​​of the pixels closest to the error pixel in the gradient direction; when the gradient value of the error pixel in the error area is not greater than the target threshold, , calculate the reconstructed pixel point, where f1 represents the pixel value above the pixel point to be interpolated, f2 represents the pixel value to the left of the pixel point to be interpolated, and f3 represents the correct pixel below the pixel point to be interpolated, 、 、 Both represent the distance between the pixel point and the point to be interpolated.

[0010] Furthermore, the method also includes: when the shallow compression code stream is to compress an uncompressed image, generating an RTP data packet of the shallow compression code stream, and performing encoding and decoding of the shallow compression code stream based on the RTP data packet; wherein, generating the RTP data packet of the shallow compression code stream includes: initializing the header information of the shallow compression code stream, the header information initialization includes RTP header information initialization, target header information initialization and Box header information initialization; combining the shallow compression code stream data and the initialized header information to generate the RTP data packet of the shallow compression code stream and sending it to the network.

[0011] According to a second aspect of the present invention, a shallow compression video error concealment system based on wavelet transform is proposed, which is characterized by comprising: an acquisition module for acquiring a shallow compression code stream; a determination module for determining the index size of a current slice and the total number of slices when the shallow compression code stream includes multiple slices and the slices are decoded one by one; a parsing module for parsing each slice, parsing the area of ​​each slice, and parsing the sub-data packet of the area of ​​each slice according to the comparison result of the index size of the current slice and the total number of slices, so as to perform decoding on each slice; and obtaining the error area in the decoded image, and reconstructing the pixel points based on the error pixel gradient value of the error area and the target threshold.

[0012] In a third aspect of the present invention, an electronic device is proposed, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the methods described in the first aspect of the present invention.

[0013] A fourth aspect of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute any one of the methods according to the first aspect of the present invention.

[0014] The beneficial effects of the present invention are as follows:

[0015] The present invention discloses a shallowly compressed video error concealment method and system based on wavelet transform according to an embodiment of the present invention. The method and system obtain a shallowly compressed bitstream; when the shallowly compressed bitstream includes multiple slices and each slice is decoded one by one, the index size of the current slice and the total number of slices are determined; based on the comparison result of the index size of the current slice and the total number of slices, each slice is parsed, the region of each slice is parsed, and the sub-data packets of the region of each slice are parsed to perform decoding of each slice; and the error region in the decoded image is obtained, and the pixel points are reconstructed based on the error pixel gradient value and target threshold of the error region. The method achieves the successful decoding of a compressed bitstream with packet loss and the reconstruction of the error pixels caused by the packet loss. The decoding problem caused by packet loss is solved without increasing the bandwidth burden. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are only for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference numerals represent the same components. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also obtain other drawings based on these drawings.

[0017] Figure 1 is a flow chart of a shallow compression video error concealment method based on wavelet transform according to an embodiment of the present invention;

[0018] Figure 2 is a flow chart of a shallow compression video error concealment method based on wavelet transform according to a specific embodiment of the present invention;

[0019] Figure 3 is a schematic diagram of a gradient direction according to an embodiment of the present invention;

[0020] Figure 4 is a schematic diagram of pixel points required for smooth area interpolation according to one embodiment of the present invention;

[0021] Figure 5 is a schematic diagram of an ablation experiment evaluation result with 5 data packets lost per frame according to an embodiment of the present invention;

[0022] Figure 6 is a schematic diagram comparing visual effects of an ablation experiment with 5 data packets lost per frame according to an embodiment of the present invention;

[0023] Figure 7 is a schematic diagram of shallow compression end-to-end encoding, decoding and transmission according to an embodiment of the present invention;

[0024] Figure 8 is a structural block diagram of a shallow compression video error concealment system based on wavelet transform according to an embodiment of the present invention;

[0025] Figure 9 It is a structural block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the scope of protection of the present invention.

[0027] Furthermore, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts disclosed in the present invention.

[0028] In the description of the present invention, it should be noted that, unless otherwise expressly specified and limited, the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second" and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance. The terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0029] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of methods and systems consistent with certain aspects of the present invention, as detailed in the appended claims.

[0030] The present invention proposes a wavelet transform-based shallow compression video error concealment method, system and related equipment. Specifically, the wavelet transform-based shallow compression video error concealment method, system and related equipment of the embodiment of the present invention are described below with reference to the accompanying drawings.

[0031] Figure 1 This is a flowchart of a wavelet transform-based shallow compression video error concealment method according to one embodiment of the present invention. It should be noted that the wavelet transform-based shallow compression video error concealment method according to this embodiment of the present invention can be applied to the wavelet transform-based shallow compression video error concealment system according to this embodiment of the present invention. This wavelet transform-based shallow compression video error concealment system can be configured on an electronic device or in a server. This embodiment of the present application is not limited to this.

[0032] like Figure 1 As shown, the shallow compression video error concealment method based on wavelet transform includes:

[0033] S110, obtaining a shallow compression code stream.

[0034] In the embodiment of the present invention, the description is carried out by taking the JPEG XS code stream as a shallow compression code stream as an example.

[0035] S120: When the shallow compression code stream includes multiple slices and the slices are decoded one by one, determine the index size of the current slice and the total number of slices.

[0036] In an embodiment of the present invention, when the shallow compression codestream is a JPEG XS codestream, and the JPEG XS codestream includes multiple slices, the image height (height) and slice height (slice_height) in the JPEG XS header parameters are extracted from the codestream to calculate the total number of slices (slice_num). Here, slice_num = height / slice_height.

[0037] S130, based on the comparison result of the index size of the current slice and the total number of slices, parse each slice, parse the area of ​​each slice, and parse the sub-data packet of the area of ​​each slice to perform decoding on each slice; and obtain the error area in the decoded image, and reconstruct the pixel points based on the error pixel gradient value of the error area and the target threshold.

[0038] In an embodiment of the present invention, when the index size of the current slice is less than or equal to the total number of slices, the current slice is detected for data errors. If a data error occurs in the current slice, the current slice is not decoded, and the decoder is synchronized with the slice next to the current slice to perform parsing for each slice. A shallow compressed bitstream is obtained after parsing each slice. A determination is made as to whether the sequence numbers of the slices in the parsed shallow compressed bitstream are discontinuous. If so, packet loss is determined in the parsed shallow compressed bitstream, and the lost packet data is replaced with zeros. The region of each slice is then parsed, as well as the sub-data packets within the region of each slice. For specific implementation methods, reference can be made to the subsequent embodiments.

[0039] In one embodiment of the present invention, the number of error image lines is located during decoding to obtain the error image of the decoded image; , calculate the error pixel gradient value in the error area; where, Indicates the gradient value of the X axis, , Indicates the gradient value of the Y axis, , where P1, P2, P3, P5, P6, P7, and P8 represent pixel values. If the gradient value of the error pixel in the error region is greater than the target threshold, the pixel is reconstructed using gradient direction-based interpolation. If the gradient value of the error pixel in the error region is not greater than the target threshold, the pixel is reconstructed using smooth interpolation. For specific implementation methods, please refer to the subsequent embodiments.

[0040] It should be noted that parsing each slice is performed by synchronizing the decoder with the slice immediately following the current one. This resynchronization effectively identifies error locations, clears erroneous data, and reestablishes lost synchronization to prevent error propagation. Resynchronization reestablishes synchronization between the decoder and the bitstream after a data error is detected. Typically, data between two synchronization points where data errors occur is discarded. In the case of a shallowly compressed codestream (i.e., a JPEG XS codestream) consisting of multiple slices, the JPEG XS codestream is a hierarchical structure composed of multiple slices. The decoder decodes the codestream slice by slice, and each slice can be decoded independently. The JPEG XS standard specifies that the slice header consists of SLH (FF 20), Lslh (00 04), and Yslh (00 idx). SLH is the slice header marker field, with a value of FF 20; Lslh is the segment size in bytes, with a value of 00 04; Yslh is the slice index; and idx is the index number. The present invention uses the first two parts of the slice header marker SLH and Lslh (FF 20 00 04) as resynchronization markers to regain synchronization.

[0041] It should be noted that packet loss in the parsed JPEG XS codestream is determined by determining whether the slice sequence numbers in the parsed JPEG XS codestream are discontinuous. The lost data is then replaced with zeros, allowing for parsing of each slice region and the sub-packets within each slice region. Frequency-domain error concealment maintains the integrity of the codestream, recovering as much correct data as possible in the frequency domain and providing more useful information to the decoder. Frequency-domain error concealment cannot completely restore the correct image, so pixels are reconstructed based on the gradient values ​​of the erroneous pixels in the decoded image after parsing and a target threshold. Spatial-domain error concealment, on the other hand, utilizes information from surrounding pixels to conceal the erroneous pixels after decoding, maximizing image quality.

[0042] According to an embodiment of the present invention, a shallowly compressed video error concealment method based on wavelet transform obtains a shallowly compressed bitstream; when the shallowly compressed bitstream includes multiple slices and the slices are decoded one by one, the index size of the current slice and the total number of slices are determined; based on the comparison result of the index size of the current slice and the total number of slices, each slice is parsed, the region of each slice is parsed, and the sub-data packets of the region of each slice are parsed to perform decoding of each slice; and the error region in the decoded image is obtained, and the pixel points are reconstructed based on the error pixel gradient value and target threshold of the error region. This method can successfully decode a compressed bitstream with packet loss and reconstruct the error pixels caused by the packet loss. It also solves the decoding problem caused by packet loss without increasing the bandwidth burden.

[0043] In order to make it easier for those skilled in the art to understand the present invention, Figure 2 is a shallow compression video error concealment method based on wavelet transform according to a specific embodiment of the present invention, such as Figure 2 As shown, the shallow compression video error concealment method based on wavelet transform includes:

[0044] S210, obtaining a shallow compression code stream.

[0045] S220: When the shallow compression code stream includes multiple slices and the slices are decoded one by one, determine the index size of the current slice and the total number of slices.

[0046] In the embodiment of the present invention, the implementation of steps S210 - S220 may refer to the implementation of steps S110 - S120 described above, and the present invention does not limit this.

[0047] S230 , performing parsing on each slice according to the comparison result of the index size of the current slice and the total number of slices.

[0048] In an embodiment of the present invention, when the index size of the current slice is less than or equal to the total number of slices, detecting whether a data error occurs in the current slice; when a data error occurs in the current slice, not decoding the current slice, and synchronizing the decoder with the slice next to the current slice to perform parsing of each slice; obtaining a shallow compressed code stream after parsing each slice; determining whether sequence numbers of slices in the parsed shallow compressed code stream are discontinuous; if so, determining that a packet is lost in the parsed shallow compressed code stream, and replacing the lost packet data with zeros.

[0049] The decoder synchronizes with the slice following the current slice by reading the bitstream byte by byte and storing the detected slice header bit position (slice_pos) and index (idx) in the corresponding array (slice_pos[idx]. During this phase, the resynchronization process utilizes the positions of the SLH and Lslh segments, as well as the slice index. Accordingly, when decoding slice by slice, if an error is detected, the decoder discards the data between the error point and the next synchronization point and stops decoding. The decoder's synchronization point with the bitstream is then shifted to the next slice_pos[idx] position, achieving resynchronization. This process repeats until decoding is complete. Through this process, the decoder can confine errors to the actual corrupted slice, preventing them from propagating throughout the entire bitstream. This resynchronization not only reduces the impact of errors but also resolves the issue of decoder synchronization failures due to packet loss.

[0050] The implementation of replacing lost packet data with zeros can be understood as utilizing the sequence number (seq) field in the RTP header to detect packet continuity in the received JPEG XS RTP stream. The correct seq should be incremented one by one. If seq is discontinuous, it indicates packet loss. It should be noted that packet loss detection should be performed before resynchronization. This process is explained again to more fully explain the entire frequency domain error concealment process. Both resynchronization and frequency domain error concealment are optimized in the decoding reference code. After detecting packet loss, the present invention adopts a zero interpolation strategy, padding the lost RTP packets with zeros. Since each RTP packet in the JPEG XS codestream has the same length when it is packaged, zero interpolation is effectively equivalent to replacing the lost data with zeros in the frequency domain. This process maintains the correctness of the codestream length and provides a data foundation for the subsequent decoding and error concealment stages.

[0051] For example, when a JPEG XS codestream includes slices A, B, C, and D, each slice is parsed one by one. When parsing slice A and a data error occurs in slice A, slice A is not decoded. The decoder synchronizes with slice B and begins decoding slice B, further detecting whether there are data errors in slice B until decoding of slice D is completed. A JPEG XS codestream is then obtained after parsing each slice. The parsed JPEG XS codestream now includes slices B, C, and D, where slice B corresponds to sequence number 2, slice C corresponds to sequence number 3, and slice D corresponds to sequence number 4. When sequence numbers 2, 3, and 4 are determined to be discontinuous, packet loss is determined in the codestream, i.e., packet loss sequence number 1 is determined. The data in sequence number 1, i.e., the data in slice A, is then replaced with zeros. The JPEG XS codestream now includes slices A, B, C, and D.

[0052] S240: Analyze the area of ​​each slice.

[0053] In an embodiment of the present invention, a shallow compressed data raw stream is obtained; when each slice is parsed one by one, it is determined whether the expected entropy coded data length of each region in each slice is unreasonable. If unreasonable, the entire region cannot be parsed, and parsing of other regions of the slice is executed.

[0054] The shallowly compressed raw data stream, or JPEG XS raw data stream, is a JPEG XS codestream without header information. This raw JPEG XS data stream can be understood as stripping the aforementioned header information (RTP header, BOX information, and ST2110-22 header information) added during transmission from the RTP packet, extracting the payload, and forming the original raw JPEG XS data stream. This process is called depacketization; its goal is to remove the encapsulated ST2110-22 header and RTP header, resulting in a JPEG XS codestream that begins with the start-of-codestream marker (SOC) (value FF 10) and ends with the end-of-codestream marker (EOC) (value FF 11).

[0055] The parsed slice region can be understood as each slice consisting of multiple precincts. During decoding, the expected entropy coded data length Lprc in the region header must be parsed to select the truncation position Q for all bands in the region and the region refinement R. If the field value is illegal or inconsistent with the actual read value, the information for the entire region cannot be parsed and decoded. This region is then skipped and decoding continues with the next region. If there are no errors in the region header field, each sub-packet within the region is decoded one by one. The implementation of decoding each sub-packet within the region can be found in the subsequent embodiments.

[0056] For example, each slice includes multiple areas, and each area corresponds to the expected entropy coded data length. For example, slice A includes area A1, area A2, and area A3. Area A1, area A2, and area A3 are parsed one by one. If it is determined that the expected entropy coded data length of area A1 is unreasonable, area A2 is parsed. If the expected entropy coded data length of area A2 is reasonable, the sub-data packet of area A2 is parsed to decode the sub-data packet.

[0057] S250: Parse the sub-data packets of each slice area.

[0058] In an embodiment of the present invention, when the expected entropy coded data length of the region is reasonable, a sub-data packet of each region is obtained, wherein each region includes multiple sub-data packets; the entropy coding length of the sub-data packet is determined based on the data sub-packet size, the bit plane count sub-packet size, and the symbol sub-packet size of the sub-data packet of each region, and the sub-data packet is parsed based on the entropy coding length and the preset length.

[0059] The subpackets in the parsed region can be understood as each region consisting of several subpackets. The data subpacket size Ldat[p,s], bit plane count subpacket size Lcnt[p,s], and symbol subpacket size Lsgn[p,s] in the subpacket header record the entropy-coded length of the subpacket. If the entropy-coded length actually read during decoding does not match the expected length, the subpacket cannot be decoded. The subpacket is skipped and decoding continues with the next subpacket until all data is decoded, followed by operations such as inverse quantization.

[0060] Correspondingly, in an embodiment of the present invention, based on the entropy coding length and the preset length, sub-data packets are parsed, including: determining whether the entropy coding length is the preset length; if the entropy coding length is not the preset length, not parsing the sub-data packet, but parsing other sub-data packets in the corresponding area; until the parsing of the sub-data packets in the area of ​​each slice is completed.

[0061] For example, in slice A, area A1, area A2, and area A3, where area A1 includes a first sub-data packet and a second sub-data packet, determine the entropy coding length of the first sub-data packet based on the data sub-packet size, the bit plane count sub-packet size, and the symbol sub-packet size in the first sub-data packet, and judge whether the entropy coding length of the first sub-data packet meets the preset length. If not, do not parse the first sub-data packet, and execute parsing of the second sub-data packet until the sub-data packet parsing of each slice area is completed.

[0062] S260 , obtaining an error region in the decoded image, and reconstructing pixels based on an error pixel gradient value in the error region and a target threshold.

[0063] In an embodiment of the present invention, the number of error image lines is located during decoding to obtain the error image of the decoded image; , calculate the error pixel gradient value in the error area; where, Indicates the gradient value of the X axis, , which is calculated based on the X-axis convolution kernel and the pixel values ​​of the surrounding pixels; Indicates the gradient value of the Y axis, , that is, it is calculated based on the Y-axis convolution kernel and the pixel values ​​of the surrounding pixels, where P1, P2, P3, P5, P6, P7 and P8 all represent the pixel values ​​of the pixels; when the gradient value of the error pixel in the error area is greater than the target threshold, the difference based on the gradient direction is used to reconstruct the pixel point; when the gradient value of the error pixel in the error area is not greater than the target threshold, the smooth interpolation is used to reconstruct the pixel point.

[0064] That is to say, after parsing each slice, parsing the area of ​​each slice, and parsing the sub-data packet of the area of ​​each slice to perform decoding of each slice, the number of error image rows located during decoding can be obtained, the error image of the decoded image can be obtained, and the error pixel point can be reconstructed.

[0065] In the case of an error image, the image gradient in the error image is the rate of change of the pixel along the x-axis and y-axis, reflecting the speed of change of the image pixel value. For the edge of the image, the gradient value is larger; for the smooth part, the gradient value is smaller. The gradient value of the error image is calculated using the Sobel operator, and the gradient in the horizontal, vertical and diagonal directions is considered from the x-axis and y-axis directions respectively, and a weighted sum is performed. Then the gradient values ​​of the x-axis and y-axis can be calculated. and , where the gradient value of the X axis is You can get the vertical boundary and the gradient value of the Y axis The horizontal boundary can be extracted. , calculate the error pixel gradient value in the error area.

[0066] In an embodiment of the present invention, when the gradient value of the error pixel in the error area is greater than the target threshold, , calculate the reconstructed pixel points, where ,in, Represents the coordinates to be interpolated, Indicates the pixel coordinates corresponding to the gradient direction, h1 and h2 indicate the pixel values ​​of the pixels closest to the error pixel in the gradient direction; when the gradient value of the error pixel in the error area is not greater than the target threshold, , calculate the reconstructed pixel point, where f1 represents the pixel value above the pixel to be interpolated, f2 represents the pixel value to the left of the pixel to be interpolated, and f3 represents the correct pixel below the pixel to be interpolated. 、 、 Both represent the distance between the pixel point and the point to be interpolated.

[0067] That is, if If is greater than the target threshold, the pixel point p is considered to be an edge pixel and the pixel point is reconstructed using interpolation based on the gradient direction. If the error value is not greater than the target threshold, point p is considered to be in a smooth region and smooth interpolation is used to reconstruct the pixel. The target threshold is the average of all gradients. The gradient value of the error pixel in the error region is calculated to determine whether it is an edge pixel or a pixel in the smooth region. Different pixels are processed differently. For edge pixels, spatial domain error hiding is used to operate on the decoded image in the spatial domain, fully utilizing the effective information to reconstruct the error point. This further improves image quality based on frequency domain hiding and conceals the error pixel.

[0068] The interpolation reconstruction of pixel points based on the gradient direction can be understood as follows: for edge pixel point p, the gradient angle of this point can be calculated by the formula, and the gradient angle value is between 0 and , are mapped to 9 gradient directions, the gradient directions are as follows Figure 3 As shown. The gradient angle , if the gradient angle satisfy , then the gradient direction is Directions, e.g. for When satisfied , then the gradient direction is ,Right now Direction. Once the gradient direction is determined, interpolation is performed along the gradient direction, using the valid information of the correct pixel and the surrounding restored pixels for direction-based interpolation reconstruction. The corresponding pixel in the neighborhood along the gradient direction is found, and its distance di to the point to be interpolated can be calculated using the following formula. , and then according to , to calculate the reconstructed pixel point. h1 and h2 represent the pixel values ​​of the pixel closest to the error pixel in the gradient direction. and Indicates the distance between the pixel and the point to be interpolated.

[0069] Among them, the use of smooth interpolation to reconstruct pixels can be understood as using smooth interpolation for pixels in the smooth area, using the correct or reconstructed pixel information above, left, and below the pixel to be interpolated to perform bilinear interpolation. The required pixel position is as follows: Figure 4 shown.

[0070] In this embodiment of the present invention, when a JPEG XS codestream includes multiple slices and each slice is decoded individually, the index size of the current slice and the total number of slices are determined. Based on the comparison of the current slice index size and the total number of slices, each slice is parsed, its region is parsed, and sub-packets within its region are parsed to decode each slice. Furthermore, error regions in the decoded image are obtained and pixels are reconstructed based on the gradient values ​​of the error pixels in the error regions and a target threshold. This shallow compression error concealment method based on wavelet transforms addresses the problem of the original shallow compression coding standard lacking an error-tolerant extension. This method is sensitive to packet loss and has extremely low error tolerance, which can lead to decoding failures. This method enables successful decoding of compressed codestreams with packet loss and reconstructs error pixels caused by packet loss.

[0071] In another embodiment of the present invention, to verify the effectiveness of the various components of the proposed error concealment method, ablation experiments were conducted to compare the results of the proposed method without the proposed algorithm, with resynchronization alone, with resynchronization alone and frequency-domain error concealment, and with full resynchronization, frequency-domain error concealment, and spatial-domain error concealment. Hereinafter, "without XS-EC" represents the original JPEG XS decoding method, "XS-EC-A" represents resynchronization alone, "XS-EC-AB" represents resynchronization and frequency-domain error concealment without spatial-domain error concealment, and "XS-EC-ABC" represents the full XS-EC method.

[0072] (1) Objective assessment

[0073] Evaluation Metrics: PSNR is an indicator used to evaluate image or video quality. The calculation formula for PSNR is as follows: ,in, Indicates the maximum possible pixel value (for 10-bit images, Take 1023), (Mean Squared Error) represents the mean square error. , where m and n represent the width and height of the image respectively, Indicates that the original image is at position The pixel value at It is the pixel value at the same position in the image to be measured. A higher PSNR value indicates a smaller difference between the original image and the image to be measured.

[0074] VMAF is a metric used to assess video quality. It combines multiple visual quality assessment methods, including structural similarity, mean square error, and peak signal-to-noise ratio. VMAF compares the quality differences between the original video and the image being measured, assigning a score from 0 to 100 to indicate the relative level of video quality. A higher VMAF score indicates better video quality.

[0075] Among them, the ablation experiment evaluation results of losing 5 data packets per frame are as follows: Figure 5 shown.

[0076] Judging from PSNR and VMAF, using only resynchronization (XS-EC-A) allows for successful JPEG XS decoding. However, due to inaccurate error location, excessive data is discarded, significantly degrading image quality. Using only resynchronization and frequency-domain concealment (XS-EC-AB) allows for the majority of correct data to be decoded, but frequency-domain data estimation is still inaccurate, and some errors still exist in the decoded image. The full implementation of the proposed XS-EC algorithm (XS-EC-ABC) further conceals errors using a spatial-domain concealment algorithm after frequency-domain concealment, improving image quality. This proposed error concealment scheme not only ensures successful decoding but also improves image quality to a certain extent.

[0077] (2) Visual effect comparison

[0078] The experiment played two video streams after error concealment and compared the visual effects of human eyes. Figure 6 shown.

[0079] As can be seen from the above figures, frequency domain error concealment (XS-EC-AB) greatly improves decoding quality based on resynchronization (XS-EC-A), while spatial domain error concealment (XS-EC-ABC) further hides errors after frequency domain error concealment.

[0080] S270 , when the shallow compression code stream is a compressed uncompressed image, generating an RTP data packet of the shallow compression code stream, and performing encoding and decoding on the shallow compression code stream based on the RTP data packet.

[0081] In an embodiment of the present invention, the header information of the shallow compression code stream is initialized, and the header information initialization includes RTP header information initialization, target header information initialization and Box header information initialization; the shallow compression code stream data and the initialized header information are combined to generate an RTP data packet of the shallow compression code stream and send it to the network.

[0082] Among them, the target header information is the ST2110-22 header information;

[0083] The RTP header information is initialized using the self-encoded function RtpHeaderInit(), whose parameters are the RTP data packet structure that complies with the relevant provisions of RFC3550.

[0084] The ST2110-22 header is initialized using the self-encoding function RtpXSInit(). The SMPTE 2110-22 payload header in the JPEG XS-encoded signal video IP stream is 4 bytes long. According to the RFC9134 standard, the first 4 bytes of each packet's payload header are header flags. Its parameters contain flag information consistent with the RFC9134 standard.

[0085] The Box header information is initialized using the self-encoded function BoxHeaderInit(). The Box information includes a 42-byte Video Support Box (VS Box). The LBOX (Box Length) is 4 bytes, the TBOX (Box Type) is 4 bytes, and the DBOX (Box Data) is 34 bytes. The DBOX primarily contains the Video Information Box (Video Information Box), the Profile and Level Box (Profile and Level Box), and the Color Specification Box (CS Box). The 18-byte LBOX (4 bytes), the TBOX (4 bytes), and the DBOX (10 bytes) contain flag information consistent with the first part of the JPEG XS standard. Each frame contains only one Box header, located in the first data packet of each frame.

[0086] In an embodiment of the present invention, an RTP data packet of a JPEG XS code stream is generated, and encoding of the JPEG XS code stream can be performed based on the RTP data packet. Figure 7 As shown, the RTP data packet can be packaged to achieve the sending of the RTP data packet, wherein the RTP data packet packaging includes packaging the initialized RTP header information, packaging the initialized Box header information, and packaging the initialized ST2110-22 header information.

[0087] After encoding the JPEG XS code stream, it can be decoded based on the decoding end, such as Figure 7 As shown, the specific implementation method of decoding at the decoding end is:

[0088] Step 1: Thread 1 uses Winpcap to monitor the network card and capture IP video packets. Winpcap capture is a three-step process: The first step is to construct a network device list to obtain information about all local adapters to identify the network card receiving the destination stream. The second step is to set parameters, firstly, to enable or disable promiscuous mode for packet capture on the network card. If promiscuous mode is enabled, all packets arriving at the network card are captured; if promiscuous mode is disabled, only packets destined for the network card address are captured. Larger kernel and user buffers are then set to prevent packet loss. The third step is to enable the network card adapter and use the pcap_loop function and callback function to continuously capture video stream packets and store them in blocking queue 1. The smallest unit element stored in blocking queue 1 is a packet.

[0089] Step 2: Thread 2 reads the IP video data packet. First, it initializes the data structure stored in memory based on the IP data packet. It then continuously retrieves data packets from blocking queue 1 in the order they were stored and assigns values ​​to the structure, allowing thread 3 to extract the JPEG XS codestream information from the data packet. The datastream values ​​in the data packet are primarily calculated based on the offsets of different fields in the packet, as specified by the RFC9134 encapsulation standard. If the first byte of the data packet is represented by pkt[0], then the first field of the ST2110-22 payload header, the transport mode T, can be represented by the first bit of the pkt

[54] (54 = 14 + 20 + 8 + 12) byte. Other fields are obtained by analogy using offsets. The ST 2110-22 encapsulation protocol indicates that each data packet has a fixed 4-byte header, and the first packet of each frame has an additional 60 bytes of box format description. Therefore, when calculating the offset bytes, it is important to consider whether it is the first packet.

[0090] Step 3: Thread 3 processes the data, continuously extracting JPEG XS compressed codestream information from each frame's data packet and sequentially storing it in Blocking Queue 2 (Blocking Queue 2 stores compressed codestream information) for retrieval by Thread 4. First, the first packet of the video frame is located by checking that both the SEP counter and the P counter in the ST 2110-22 payload header standard are equal to 0. If not, the data packet is read again until all are zero. Based on the video content information in the Box of the first packet, the video's frame rate, width, chroma, sampling, and other parameters are obtained, and the image pixel buffer size for storing and playing the video is set (for example, for a 4K video with a 4:2:2 sampling format, the buffer size for one frame is "height * width * 2"). Then, once the first packet of the image is identified, the data packets are read sequentially based on the sequence number to obtain all the JPEG XS compressed codestream information for the same frame. These are then sequentially spliced ​​into a single frame stream (different compression ratios result in different codestream sizes). This is done until the packet marker bit M is 1, at which point the frame streams are sequentially stored in Blocking Queue 2. Thread 4 is then started for multi-threaded decoding. The minimum storage unit of blocking queue 2 is the size of one frame of code stream.

[0091] Step 4: Thread 4 performs video decompression (i.e., video decoding), continuously retrieving frames of the JPEG XS codestream from blocking queue 2 according to the video stream frame counter, decoding the decoded image YUV data, and storing it in blocking queue 3. After decoding is complete, the Simple DirectMedia Layer (SDL) multimedia development library (SDL) determines the image display characteristics and the video stream sampling depth to determine whether downsampling to 8 bits is necessary for display. When downsampling 10-bit video, it is important to consider whether the original 10-bit pixel components are interleaved across consecutive bytes or stored in two bytes using little-endian or big-endian format. Finally, the complete frame of Y, U, and V image pixel information is stored in a packed format and passed to blocking queue 3 for retrieval by thread 5. This step is repeated, continuously storing image data in blocking queue 3. The smallest storage unit in blocking queue 3 is a complete frame of YUV data. That is, blocking queue 3 stores pixel information for a single frame.

[0092] Step 5: Thread 5 plays the video by calling the SDL multimedia library, continuously reading pixel information from blocking queue 3 for playback. The SDL video playback process consists of two steps: initialization and looping through the display. First, the SDL window, renderer, and texture are initialized in that order. Then, the complete YUV data of a frame of image data is repeatedly retrieved from blocking queue 3 as a texture parameter for texture updates, rendering, and video display.

[0093] Step 6: Thread 6 outputs monitoring information, which is divided into per-frame evaluation indicators and overall evaluation indicators. Per-frame evaluation indicators include the extraction stream duration, decoding duration, and playback duration; overall evaluation indicators include the total number of packets, number of packet losses, frame rate, packet capture time, etc.

[0094] As can be seen, when the shallow compression bitstream is a compressed uncompressed image, RTP packets for the shallow compression bitstream are generated, and encoding and decoding of the shallow compression bitstream are performed based on the RTP packets. This achieves shallow compression end-to-end encoding, decoding, and transmission. This shallow compression end-to-end encoding, decoding, and transmission mainly consists of two parts: the encoder and the decoder. The encoder can be divided into three modules: encoding, packaging, and transmission. The encoder is responsible for encoding the input YUV video source into a JPEG XS bitstream as required. The packaging module consists of three parts: the first part encapsulates the bitstream with RTP header information; the second part encapsulates the bitstream with the corresponding BOX information, as long as the first packet of each frame contains the BOX information; and the third part encapsulates the bitstream with the header information specified in ST2110-22. The sending module first splits the bitstream into a preset size, smaller than the Maximum Transmission Unit (MTU), to accommodate the payload of an RTP packet. It then combines the previously mentioned header information (RTP header, BOX information, and ST2110-22 header information) with the split payload and sends the RTP packet to the destination. The decoding module performs multi-frame parallel decoding, capturing the video stream IP packets, reading them, extracting the JPEG XS bitstream, decompressing the video, playing the video, and outputting monitoring information (including per-frame metrics and overall metrics). Per-frame metrics include extraction, bitstream duration, decoding duration, and playback duration; overall metrics include total packet count, packet loss, frame rate, and packet capture time). The program is divided into six independent multi-threaded parallel tasks, utilizing queue-buffered communication to improve program efficiency. Threads 1 and 2 share the IP stream packets in blocking queue 1 for communication and synchronization. Threads 3 and 4 share the JPEG XS compressed code stream data in blocking queue 2 for communication and synchronization. Threads 4 and 5 share the pixel data of the decoded complete frame in blocking queue 3 for communication and synchronization. Therefore, the present invention, based on shallow compression end-to-end encoding, decoding, and transmission, helps break down hardware barriers, reduce production costs, promote the popularization of ultra-high-definition technology, and promote localized R&D across the entire product chain. It also solves the problem of the high cost of integrating a complete hardware encoding, decoding, and playback system into the hardware.

[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0096] According to one aspect of the embodiments of the present invention, a shallow compression video error concealment system based on wavelet transform is also proposed. Figure 8 It is a structural block diagram of a shallow compression video error concealment system based on wavelet transform; Figure 8 Shown, including:

[0097] An acquisition module 810 is configured to acquire a shallow compressed code stream;

[0098] a determination module 820 configured to determine an index size of a current slice and a total number of slices when the shallow compressed code stream includes multiple slices and the slices are decoded one by one;

[0099] The parsing module 830 is configured to parse each slice, parse the region of each slice, and parse the sub-data packets of the region of each slice based on the comparison result of the index size of the current slice and the total number of slices, so as to decode each slice; and obtain the error region in the decoded image, and reconstruct the pixel points based on the error pixel gradient value of the error region and the target threshold.

[0100] According to an embodiment of the present invention, a shallowly compressed video error concealment system based on wavelet transform obtains a shallowly compressed code stream; when the shallowly compressed code stream includes multiple slices and the slices are decoded one by one, the index size of the current slice and the total number of slices are determined; based on the comparison result of the index size of the current slice and the total number of slices, each slice is parsed, the region of each slice is parsed, and the sub-data packets of the region of each slice are parsed to perform decoding of each slice; and the error region in the decoded image is obtained, and the pixel points are reconstructed based on the error pixel gradient value and target threshold of the error region. This enables the successful decoding of a compressed code stream with packet loss and the reconstruction of the error pixels caused by the packet loss. It also solves the decoding problem caused by packet loss without increasing the bandwidth burden.

[0101] Optionally, the parsing module 830 is specifically configured to, when the index size of the current slice is less than or equal to the total number of slices, detect whether a data error occurs in the current slice; when a data error occurs in the current slice, not decode the current slice, and synchronize the decoder with the slice next to the current slice to perform parsing of each slice; obtain the shallow compressed code stream after parsing each slice; and determine whether the serial numbers of the slices in the parsed shallow compressed code stream are discontinuous. If so, determine that packet loss occurs in the parsed shallow compressed code stream, and replace the packet loss data with zero.

[0102] Optionally, the parsing module 830 is specifically used to obtain a shallow compressed data bare stream; when parsing each of the slices one by one, it is determined whether the expected entropy coded data length of each region in each slice is unreasonable; if it is unreasonable, the entire region cannot be parsed, and parsing of other regions of the slice is performed; if the expected entropy coded data length of the region is reasonable, a sub-data packet of each region is obtained, wherein each region includes multiple sub-data packets; according to the data sub-packet size, the bit plane count sub-packet size, and the symbol sub-packet size of the sub-data packet of each region, the entropy coding length of the sub-data packet is determined, and based on the entropy coding length and the preset length, the sub-data packet is parsed.

[0103] Optionally, the parsing module 830 is specifically used to determine whether the entropy coding length is the preset length; if the entropy coding length is not the preset length, the sub-packet is not parsed, and other sub-packets in the corresponding area are parsed; until the parsing of the sub-packets in the area of ​​each slice is completed.

[0104] Optionally, the parsing module 830 is specifically used to locate the number of error image lines during decoding and obtain the error image of the decoded image; according to , calculate the error pixel gradient value of the error area; wherein, Indicates the gradient value of the X axis, , Indicates the gradient value of the Y axis, , where P1, P2, P3, P5, P6, P7 and P8 all represent pixel values ​​of pixel points; when the gradient value of the error pixel in the error area is greater than the target threshold, the pixel point is reconstructed by using the difference based on the gradient direction; when the gradient value of the error pixel in the error area is not greater than the target threshold, the pixel point is reconstructed by using smooth interpolation.

[0105] Optionally, the parsing module 830 is specifically configured to, when the error pixel gradient value in the error area is greater than the target threshold, , calculate the reconstructed pixel points; wherein, ,in, Represents the coordinates to be interpolated, Indicates the pixel coordinates corresponding to the gradient direction, h1 and h2 indicate the pixel values ​​of the pixels closest to the error pixel in the gradient direction; when the gradient value of the error pixel in the error area is not greater than the target threshold, , calculate the reconstructed pixel point, where f1 represents the pixel value above the pixel point to be interpolated, f2 represents the pixel value to the left of the pixel point to be interpolated, and f3 represents the correct pixel below the pixel point to be interpolated. 、 、 Both represent the distance between the pixel point and the point to be interpolated.

[0106] Optionally, the device also includes a generation module for generating an RTP data packet of the shallow compression code stream when the shallow compression code stream is a compression of an uncompressed image, and performing encoding and decoding of the shallow compression code stream based on the RTP data packet; wherein, generating the RTP data packet of the shallow compression code stream includes: initializing the header information of the shallow compression code stream, the header information initialization includes RTP header information initialization, target header information initialization and Box header information initialization; combining the shallow compression code stream data and the initialized header information to generate the RTP data packet of the shallow compression code stream and sending it to the network.

[0107] According to one aspect of an embodiment of the present invention, an electronic device is provided.

[0108] Figure 5 FIG. 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Figure 5 As shown, the electronic device may include one or more ( Figure 5 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a microprocessor (Microprocessor Unit, referred to as MPU) or a programmable logic device (Programmable logic device, referred to as PLD)) and a memory 104 for storing data. In an exemplary embodiment, the electronic device may further include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 5 The structure shown is only for illustration and does not limit the structure of the above terminal device. Figure 5 More or fewer components than shown, or with Figure 5 Equivalent functions or comparisons shown Figure 5 Shown are different configurations with more functionality.

[0109] Memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the level determination method in the embodiments of the present invention. Processor 102 executes the computer programs stored in memory 104 to execute various functional applications and data processing, thereby implementing the aforementioned method. Memory 104 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some embodiments, memory 104 may further include memory remotely located relative to processor 102, and such remote memory may be connected to a terminal device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0110] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a communications provider of a switching device. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0111] The present invention provides a computer-readable storage medium having a computer program stored thereon. The computer program can be loaded and executed by a processor to implement the shallow compression video error concealment method based on wavelet transform as described in the first aspect.

[0112] The applicant of the present invention has made a detailed explanation and description of the implementation examples of the present invention in conjunction with the drawings in the specification. However, those skilled in the art should understand that the above implementation examples are only preferred implementation plans of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, and is not a limitation on the scope of protection of the present invention. On the contrary, any improvements or modifications based on the inventive spirit of the present invention should fall within the scope of protection of the present invention.

[0113] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0114] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0115] Finally, it should be noted that the above embodiments are merely illustrative of the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be covered by the scope of protection of the present invention.

Claims

1. A shallow compression video error concealment method based on wavelet transform, characterized in that: include: S1, obtain shallow compression code stream; S2, when the shallow compression code stream includes multiple slices and the slices are decoded one by one, determining the index size of the current slice and the total number of slices; S3, according to the comparison result of the index size of the current slice and the total number of slices, parsing each slice, parsing the region of each slice, and parsing the sub-data packet of the region of each slice, so as to perform decoding on each slice; and executing to obtain an error region in the decoded image, and reconstructing pixel points based on an error pixel gradient value and a target threshold in the error region; The method of obtaining an error region in a decoded image and reconstructing pixels based on an error pixel gradient value and a target threshold in the error region includes: locating the error image row number during decoding and obtaining an error image of the decoded image; and reconstructing pixels based on an error pixel gradient value and a target threshold in the error region. Calculate the error pixel gradient value of the error area; wherein G x Indicates the gradient value of the X axis, G x =-p1+p3-2*p2+2*p5-p6+p8, G y Indicates the gradient value of the Y axis, G y =-p1-2*p2-p3+p6+2*p7+p8, where P1, P2, P3, P5, P6, P7 and P8 all represent pixel values ​​of pixel points; when the gradient value of the error pixel in the error area is greater than the target threshold, the pixel point is reconstructed by interpolation based on the gradient direction; when the gradient value of the error pixel in the error area is not greater than the target threshold, the pixel point is reconstructed by smooth interpolation; Wherein, when the gradient value of the error pixel in the error area is greater than the target threshold, the pixel point is reconstructed by interpolation based on the gradient direction, and when the gradient value of the error pixel in the error area is not greater than the target threshold, the pixel point is reconstructed by smooth interpolation, including: when the gradient value of the error pixel in the error area is greater than the target threshold, the reconstructed pixel point is calculated according to p=(d2*h1+d1*h2) / (d1+d2); wherein, Among them, (x, y) represents the coordinates to be interpolated, (x i ,y i ) represents the pixel coordinates corresponding to the gradient direction, h1 and h2 represent the pixel values ​​of the pixels closest to the erroneous pixels in the gradient direction; when the gradient value of the erroneous pixel in the erroneous area is not greater than the target threshold, p=(d3*f1+f2+d1*f3) / (d1+d3+1), the reconstructed pixel is calculated, wherein f1 represents the pixel value above the pixel to be interpolated, f2 represents the pixel value to the left of the pixel to be interpolated, f3 represents the correct pixel below the pixel to be interpolated, and d1, d2, and d3 all represent the distance between the pixel and the point to be interpolated.

2. The shallow compression video error concealment method based on wavelet transform according to claim 1 is characterized in that: According to the comparison result of the index size of the current slice and the total number of slices, parsing of each slice is performed, including: When the index size of the current slice is less than or equal to the total number of slices, detecting whether a data error occurs in the current slice; In the case that a data error occurs in the current slice, the current slice is not decoded, and the decoder is synchronized with a slice next to the current slice to perform parsing of each slice; Obtaining the shallow compression code stream after parsing each of the slices; It is determined whether the serial numbers of the slices in the shallow compression code stream after parsing are discontinuous. If so, it is determined that there is packet loss in the shallow compression code stream after parsing, and the packet loss data is replaced by zero.

3. The shallow compression video error concealment method based on wavelet transform according to claim 2 is characterized in that: Parsing the area of ​​each slice and parsing the sub-data packets of the area of ​​each slice, including: Get the raw stream of shallowly compressed data; When parsing each of the slices one by one, it is determined whether the expected entropy coded data length of each region in each of the slices is unreasonable. If unreasonable, the entire region cannot be parsed, and parsing of other regions of the slice is performed; If the expected entropy coded data length of the region is reasonable, obtaining a sub-data packet of each region, wherein each region includes a plurality of sub-data packets; According to the data sub-packet size, the bit plane count sub-packet size, and the symbol sub-packet size of the sub-packet in each area, the entropy coding length of the sub-packet is determined, and based on the entropy coding length and the preset length, the sub-packet is parsed.

4. The shallow compression video error concealment method based on wavelet transform according to claim 3 is characterized in that: Based on the entropy coding length and the preset length, performing parsing of the sub-data packet includes: Determining whether the entropy coding length is the preset length; When the entropy coding length is not the preset length, the sub-data packet is not parsed, and other sub-data packets in the corresponding area are parsed; Until the sub-data packets of the area of ​​each slice are parsed.

5. The shallow compression video error concealment method based on wavelet transform according to claim 1, characterized in that: The method further comprises: In the case where the shallow compression code stream is a compression of an uncompressed image, an RTP data packet of the shallow compression code stream is generated, and encoding and decoding of the shallow compression code stream is performed based on the RTP data packet; The RTP data packet of the shallow compression code stream is generated, including: Initializing the header information of the shallow compression code stream, wherein the header information initialization includes RTP header information initialization, target header information initialization and Box header information initialization; The shallow compression code stream data and the initialized header information are combined to generate an RTP data packet of the shallow compression code stream and send it to the network.

6. A shallow compression video error concealment system based on wavelet transform, characterized in that: include: An acquisition module is used to obtain a shallow compression code stream; A determination module, configured to determine an index size of a current slice and a total number of slices when the shallow compression code stream includes multiple slices and the slices are decoded one by one; a parsing module, configured to parse each slice, parse a region of each slice, and parse a sub-data packet of the region of each slice according to a comparison result of the index size of the current slice and the total number of slices, so as to decode each slice; and executing to obtain an error region in the decoded image, and reconstructing pixel points based on an error pixel gradient value and a target threshold in the error region; The method of obtaining an error region in a decoded image and reconstructing pixels based on an error pixel gradient value and a target threshold in the error region includes: locating the error image row number during decoding and obtaining an error image of the decoded image; and reconstructing pixels based on an error pixel gradient value and a target threshold in the error region. Calculate the error pixel gradient value of the error area; wherein G x Indicates the gradient value of the X axis, G x =-p1+p3-2*p2+2*p5-p6+p8, G y Indicates the gradient value of the Y axis, G y =-p1-2*p2-p3+p6+2*p7+p8, where P1, P2, P3, P5, P6, P7 and P8 all represent pixel values ​​of pixel points; when the gradient value of the error pixel in the error area is greater than the target threshold, the pixel point is reconstructed by interpolation based on the gradient direction; when the gradient value of the error pixel in the error area is not greater than the target threshold, the pixel point is reconstructed by smooth interpolation; Wherein, when the gradient value of the error pixel in the error area is greater than the target threshold, the pixel point is reconstructed by interpolation based on the gradient direction, and when the gradient value of the error pixel in the error area is not greater than the target threshold, the pixel point is reconstructed by smooth interpolation, including: when the gradient value of the error pixel in the error area is greater than the target threshold, the reconstructed pixel point is calculated according to p=(d2*h1+d1*h2) / (d1+d2); wherein, Among them, (x, y) represents the coordinates to be interpolated, (x i ,y i ) represents the pixel coordinates corresponding to the gradient direction, h1 and h2 represent the pixel values ​​of the pixels closest to the erroneous pixels in the gradient direction; when the gradient value of the erroneous pixel in the erroneous area is not greater than the target threshold, p=(d3*f1+f2+d1*f3) / (d1+d3+1), the reconstructed pixel is calculated, wherein f1 represents the pixel value above the pixel to be interpolated, f2 represents the pixel value to the left of the pixel to be interpolated, f3 represents the correct pixel below the pixel to be interpolated, and d1, d2, and d3 all represent the distance between the pixel and the point to be interpolated.

7. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Mixed video fault tolerance method based on multiple description encoding and error hiding

    CN101175216A

  • Image processing device, video display device with same, and error concealment method

    CN101998032A