Loss information calculation method and apparatus, and related device

By calculating the frequency band component loss information of the output image and the target image of the neural network, the problem of insufficient specificity of the neural network loss information is solved, thereby improving the learning effect and image processing performance.

WO2026007784A1PCT designated stage Publication Date: 2026-01-08VIVO MOBILE COMM CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/103711
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-06-26
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

The loss information calculation of neural networks is not targeted enough, resulting in poor learning performance.

Method used

By acquiring the output image of the neural network and the target frequency band components of the target image, the loss information between the two is calculated, including the loss values ​​of high frequency, low frequency and mid frequency components. The loss information is calculated using loss functions such as L1, L2 or SSIM.

Benefits of technology

It increases the neural network's focus on the target frequency band, improves the learning effect and image processing performance of the neural network, and improves image quality, especially in super-resolution and image enhancement tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025103711_08012026_PF_FP_ABST
    Figure CN2025103711_08012026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and discloses a loss information calculation method and apparatus, and a related device. The loss information calculation method in embodiments of the present application comprises: acquiring a target frequency band component of an output image of a neural network; acquiring a target frequency band component of a target image corresponding to the output image; and calculating loss information of the neural network, the loss information comprising loss information between the target frequency band component of the output image and the target frequency band component of the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Loss information calculation method and device and related equipment

[0001] Cross-reference to Related Applications

[0002] The present application claims priority to Chinese Patent Application No. 202410886308.4, filed on July 3, 2024, the contents of which are incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present application belongs to the technical field of computers, and specifically relates to a loss information calculation method and device and related equipment. BACKGROUND

[0004] Neural networks mainly learn based on loss information. In some related technologies, the loss information of an image-related neural network is calculated equally for the loss of all image information of all pixels in an entire image. Therefore, the loss information of the neural network is less targeted, which leads to poor learning effect of the neural network. SUMMARY

[0005] The embodiments of the present application provide a loss information calculation method and device and related equipment, which can solve the problem of poor learning effect of the neural network caused by poor targeting of the loss information of the neural network.

[0006] In a first aspect, a loss information calculation method is provided, comprising:

[0007] obtaining a target frequency band component of an output image of a neural network;

[0008] obtaining a target frequency band component of a target image corresponding to the output image;

[0009] calculating loss information of the neural network, the loss information comprising loss information between the target frequency band component of the output image and the target frequency band component of the target image.

[0010] In a second aspect, a loss information calculation device is provided, comprising:

[0011] a first obtaining module configured to obtain a target frequency band component of an output image of a neural network;

[0012] a second obtaining module configured to obtain a target frequency band component of a target image corresponding to the output image;

[0013] a calculating module configured to calculate loss information of the neural network, the loss information comprising loss information between the target frequency band component of the output image and the target frequency band component of the target image.

[0014] In a third aspect, a loss information calculation apparatus is provided, which is configured to perform the steps of the loss information calculation method provided in the embodiments of the present application.

[0015] In a fourth aspect, an electronic device is provided, which includes a processor and a memory, the memory storing a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the loss information calculation method provided in the embodiments of the present application.

[0016] In a fifth aspect, an electronic device is provided, which includes a processor and a communication interface, and the processor is configured to obtain a target frequency band component of an output image of a neural network, obtain a target frequency band component of a target image corresponding to the output image, and calculate loss information of the neural network, the loss information including loss information between the target frequency band component of the output image and the target frequency band component of the target image.

[0017] In a sixth aspect, an electronic device is provided, which includes a memory configured to store video data, and a processing circuit configured to implement the steps of the loss information calculation method provided in the embodiments of the present application.

[0018] In a seventh aspect, a readable storage medium is provided, which stores a program or instructions, and the program or instructions, when executed by a processor, implement the steps of the loss information calculation method provided in the embodiments of the present application.

[0019] In an eighth aspect, a chip is provided, which includes a processor and a communication interface, the communication interface and the processor being coupled, and the processor is configured to run a program or instructions to implement the steps of the loss information calculation method provided in the embodiments of the present application.

[0020] In a ninth aspect, a computer program / program product is provided, which is stored in a storage medium, and the computer program / program product is executed by at least one processor to implement the steps of the loss information calculation method provided in the embodiments of the present application.

[0021] In the embodiments of the present application, a target frequency band component of an output image of a neural network is obtained, a target frequency band component of a target image corresponding to the output image is obtained, and loss information of the neural network is calculated, the loss information including loss information between the target frequency band component of the output image and the target frequency band component of the target image. Since the loss information includes loss information between the target frequency band component of the output image and the target frequency band component of the target image, loss information for the target frequency band can be obtained, so that the loss information of the neural network is more targeted, and the learning effect of the neural network is improved. BRIEF DESCRIPTION OF DRAWINGS

[0022] Fig. 1 is a schematic diagram of a coding system according to an embodiment of the present application;

[0023] Fig. 2 is a schematic diagram of an encoder according to an embodiment of the present application;

[0024] Fig. 3 is a schematic diagram of a decoder according to an embodiment of the present application;

[0025] Fig. 4 is a flowchart of a method for calculating loss information according to an embodiment of the present application;

[0026] Fig. 5 is a schematic diagram of a frequency domain transform according to an embodiment of the present application;

[0027] Fig. 6 is a schematic diagram of loss information according to an embodiment of the present application;

[0028] Fig. 7 is a schematic diagram of a device for calculating loss information according to an embodiment of the present application;

[0029] Fig. 8 is a schematic diagram of an electronic device according to an embodiment of the present application;

[0030] Fig. 9 is a schematic diagram of a terminal according to an embodiment of the present application. DETAILED DESCRIPTION

[0031] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of the present application.

[0032] The terms "first", "second", etc. in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second" are usually a category, and are not limited to the number of objects, for example, the first object can be one or more. In addition, "or" in the present application means at least one of the connected objects. For example, the protection scope of "A or B" at least covers three schemes, namely, scheme one: including A and not including B; scheme two: including B and not including A; scheme three: including A and B. In addition, the terms "A and / or B", "at least one of A and B", "at least one of A or B" also at least cover the above three schemes respectively. The character " / " generally represents that the objects before and after are in an "or" relationship.

[0033] FIG. 1 is a schematic diagram of a coding system 10 according to an embodiment of the present application. The technical solutions of the embodiments of the present application relate to coding (CODEC) of video data, including encoding or decoding. The video data includes original uncoded video, coded video, decoded (e.g., reconstructed) video, or syntax elements, etc.

[0034] As shown in FIG. 1, the coding system 10 includes a source device 100 that provides encoded video data to be decoded and displayed by a destination device 110. In particular, the source device 100 provides the video data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 can comprise any of one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile telephone, a wearable device (e.g., a smart watch or a wearable camera), a television, a camera, a display device, a vehicle head unit, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video gaming console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an airplane, a robot, a satellite, etc.

[0035] In the example of FIG. 1, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of a video encoding device, while the destination device 110 represents an example of a video decoding device. In other examples, the source device 100 and the destination device 110 can not include some of the components of FIG. 1, or can include other components not shown in FIG. 1. For example, the source device 100 can receive video data from an external data source, such as an external camera. Also, the destination device 110 can interface with an external display device, rather than include an integrated display device. For another example, the memory 102, the memory 113 can be external memories.

[0036] Although FIG. 1 illustrates the source device 100 and the destination device 110 as separate devices, in some examples, the source device 100 and the destination device 110 can be integrated in one device. In such embodiments, the corresponding functions of the source device 100 and the destination device 110 can be implemented using the same hardware or software, or using separate hardware or software, or any combination thereof.

[0037] In some examples, the source device 100 and the destination device 110 can engage in one-way video transmission or two-way video transmission. If two-way video transmission, the source device 100 and the destination device 110 can operate in a substantially symmetrical manner, i.e., each of the source device 100 and the destination device 110 includes an encoder and a decoder.

[0038] The data source 101 represents a source of video data (i.e., raw, uncoded video data) and provides the encoder 200 with successive pictures containing the video data that the encoder 200 encodes. The data source 101 of the source device 100 can include a video capture device, such as a video camera, a video archive containing previously captured raw video, or a video feed interface to receive video from a video content provider. As another alternative, the data source 101 can generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In these cases, the encoder 200 encodes the captured, pre-captured, or computer-generated video data. The encoder 200 can rearrange the pictures from the received order (sometimes referred to as "display order") into the encoding order. The encoder 200 can generate a bitstream including encoded video data. The source device 100 can then output the encoded video data via the output interface 104 onto the communication medium 120 for reception or retrieval by, e.g., the input interface 111 of the destination device 110.

[0039] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general purpose memories. In some examples, the memory 102 can store raw video data from the data source 101, and the memory 113 can store decoded video data from the decoder 300. Additionally or alternatively, the memories 102, 113 can store software instructions that are executable by, e.g., the encoder 200 and the decoder 300, respectively. Although the memory 102 and the memory 113 are shown separately from the encoder 200 and the decoder 300 in this example, it should be understood that the encoder 200 and the decoder 300 can also include internal memories for functionally similar or equivalent purposes. If the encoder 200 and the decoder 300 are deployed on the same hardware device, the memory 102 and the memory 113 can be one and the same memory. Moreover, the memories 102, 113 can store encoded video data that is output from the encoder 200 and input to the decoder 300, for example. In some examples, portions of the memories 102, 113 can be allocated as one or more video buffers, e.g., to store raw, decoded, or encoded video data.

[0040] In some examples, source device 100 can output encoded data from output interface 104 to storage 113. Similarly, destination device 110 can access encoded data from storage 113 via input interface 111. Storage 113 or storage 102 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, Digital Versatile Discs (DVDs), Compact Discs Read-Only Memories (CD-ROMs), flash drives, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.

[0041] Output interface 104 can include any type of medium or device capable of sending encoded video data from source device 100 to destination device 110. For example, output interface 104 can include a transmitter or a transceiver, e.g., an antenna, configured to transmit encoded video data from source device 100 directly to destination device 110 in real-time. The encoded video data can be modulated according to a communication standard and transmitted to destination device 110.

[0042] Communication medium 120 can include transient media, such as wireless broadcasts or wired networks. For example, communication medium 120 can include radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cable). Communication medium 120 can form a portion of a packet-based network, such as a local area network, a wide area network, or a global network, such as the Internet. Communication medium 120 can also be in a form of storage media, such as a hard drive, flash drive, compact disc, digital video disc, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.

[0043] In some embodiments, communication medium 120 can include routers, switches, base stations, or any other equipment that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) can receive the encoded video from source device 100 and provide the encoded video data to destination device 110, such as via a network transmission. The server can include a web server (e.g., for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached storage (NAS) device, etc. The server can implement one or more HTTP streaming protocols, such as Moving Picture Experts Group (MPEG) Media Transport (MMT) protocol, Dynamic Adaptive Streaming over HTTP (DASH) protocol, HTTP Live Streaming (HLS) protocol, or Real Time Streaming Protocol (RTSP), etc.

[0044] Destination device 110 can access the encoded video data from a server, such as through a wireless channel (e.g., a Wi-Fi connection) or a wired connection (e.g., a Digital subscriber line (DSL), a cable modem, etc.) to a network that accesses the encoded video data stored on the server.

[0045] Output interface 104 and input interface 111 can represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards or the IEEE 802.15 standards (e.g., ZigBee™), the Bluetooth standard, etc., or other physical components. In examples where output interface 104 and input interface 111 comprise wireless components, output interface 104 and input interface 111 can be configured to transfer data, such as encoded video data, according to a WIFI, Ethernet, cellular (such as 4th Generation Mobile Communication Technology (4G), LTE (Long-Term Evolution), LTE-Advanced, 5th Generation Mobile Communication Technology (5G), 6th Generation Mobile Communication Technology (6G), etc.), or other protocol.

[0046] The techniques of this disclosure can be applied to video coding in support of one or more multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions, digital video that is encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0047] Input interface 111 of destination device 110 receives an encoded video bitstream from communication medium 120. The encoded video bitstream can include syntax elements and coded data units (e.g., slices, pictures, groups of pictures, or other units) that, when decoded, reproduce the video data. Display device 114 displays the decoded video data to a user. Display device 114 can comprise a Cathode ray tube (CRT), a liquid-crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other type of display device.

[0048] The encoder 200 and the decoder 300 can be implemented as one or more of various processing circuitry, which can include one or more microprocessors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), discrete logic circuitry, hardware, or any combinations thereof. When the techniques are implemented partially in software, a device can store instructions for the software in a suitable, non- transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure.

[0049] The encoder 200 and the decoder 300 can process based on the following video coding standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding (HEVC)), H.266 (also known as Versatile Video Coding (VVC)), Moving Picture Experts Group 2 (MPEG-2), MPEG-4, VP8, VP9, Alliance for Open Media Video 1 (AV1), Audio Video coding Standard 1 (AVS1), AVS2, AVS3, or a next generation video standard protocol, which are not limited in the embodiments of the present application.

[0050] Generally, the encoder 200 and the decoder 300 can perform block-based coding of pictures. The term “block” generally refers to a structure including data to be processed (e.g., encoded, decoded, or otherwise used in the encoding or decoding process). For example, a block can include a two-dimensional matrix of samples of luma or chroma data. For example, the encoder 200 and the decoder 300 can code video data represented in a YUV format.

[0051] Referring to FIG. 2, which is a structural diagram of an encoder 200 provided by an embodiment of the present application, the encoder 200 can be the encoder 200 in FIG. 1. In the example of FIG. 2, the encoder 200 includes a memory 201, a coding parameter determination unit 210, a residual generation unit 202, a transform processing unit 203, a quantization unit 204, a dequantization unit 205, an inverse transform processing unit 206, a reconstruction unit 207, a filter unit 208, a decoded picture buffer (DPB) 209, and an entropy encoding unit 220.

[0052] The memory 201 can store video data to be encoded, for example, the encoder 200 can receive video data from the data source 101 shown in FIG. 1 and store it. In some examples, the memory 201 can be on the same chip as other components of the encoder 200 (as shown in FIG. 2), or it can be independent of the chip on which the components are located.

[0053] The coding parameter determination unit 210 includes a mode selection unit 211, an inter prediction unit 212, and an intra prediction unit 213. The inter prediction unit 212 is configured to obtain a first prediction block of the current block using an inter prediction mode, the intra prediction unit 213 is configured to obtain a second prediction block of the current block using an intra prediction mode, and the mode selection unit 211 is configured to obtain a target prediction block according to the first prediction block and the second prediction block, and determine a final prediction mode. In addition, the coding parameter determination unit 210 can also include other functional units, such as a functional unit for determining the division manner of a coding unit (CU), a functional unit for determining the transform type of the residual data of the CU, or a functional unit for determining the quantization parameter of the residual data of the CU, etc.

[0054] For ease of description and understanding, the CU to be processed in the current picture is referred to as the current CU, and the image block to be processed in the current CU is referred to as the current block or the image block to be processed, for example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.

[0055] The inter prediction unit 212 can include a motion estimation unit and a motion compensation unit. For inter prediction of the current block, the motion estimation unit can perform motion search to identify one or more matching reference blocks in one or more reference pictures (for example, one or more previously coded pictures stored in the DPB 209).

[0056] The motion estimation unit can form one or more motion vectors (MVs) of the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit can obtain a predicted value of the precision indicated by the motion vector by interpolation.

[0057] The encoding parameter determination unit 210 can provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives original uncoded video data of the current block from the memory 201 and computes a residual between the current block and the target prediction block to obtain a residual block. In some examples, the functions of the residual generation unit 202 can be implemented using one or more subtractor circuits that perform binary subtraction.

[0058] As an example, the encoding parameter determination unit 210 can provide syntax elements representing the encoding parameters to the entropy encoding unit 220 for encoding. The encoding parameters include one or more of a partitioning mode of the CU, a final prediction mode, a transform type of the residual data of the CU, or a quantization parameter of the residual data of the CU.

[0059] The transform processing unit 203 can perform a transform on the residual block output by the residual generation unit 202 to obtain a transform coefficient block, which can include a discrete cosine transform (DCT), an integer transform, a directional transform, or a Karhunen-Loeve transform, among others. In some examples, the encoder 200 can not include the transform processing unit 203.

[0060] The quantization unit 204 can quantize the transform coefficients in the transform coefficient block according to a quantization parameter (QP) value associated with the current block to produce a quantized transform coefficient block.

[0061] The inverse quantization unit 205 and the inverse transform processing unit 206 can perform inverse quantization and inverse transform, respectively, on the transform coefficient block to obtain a reconstructed residual block. The reconstruction unit 207 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the encoding parameter determination unit 210.

[0062] The filter unit 208 can perform one or more filter operations on the reconstructed block. For example, the filter unit 208 can be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, among others. In some examples, the encoder 200 can not include the filter unit 208.

[0063] The encoder 200 stores the reconstructed picture resulting from the reconstructed block in the DPB 209. For example, in examples where the operation of the filter unit 208 is not needed, the reconstruction unit 207 can store the reconstructed block to the DPB 209. In examples where the operation of the filter unit 208 is needed, the filter unit 208 can store the filtered reconstructed block to the DPB 209. The inter prediction unit 212 obtains the reconstructed picture from the DPB 209 to perform inter prediction for blocks of a subsequent picture to be encoded. In some examples, the DPB 209 can be replaced by other types of memory.

[0064] The entropy encoding unit 220 can entropy encode syntax elements of other components in the encoder 200 to output encoded video data. For example, the entropy encoding unit 220 can entropy encode quantized transform coefficient blocks from the quantization unit 204. As another example, the entropy encoding unit 220 can entropy encode syntax elements from the coding parameter determination unit 210, such as motion information for inter prediction or intra mode information for intra prediction.

[0065] It can be appreciated that the components of the encoder 200 described in FIG. 2 are only an example and do not constitute a limitation to the embodiments of the present application.

[0066] FIG. 3 is a structure schematic diagram of a decoder 300 according to an embodiment of the present application. The decoder 300 can be the decoder 300 described in FIG. 1. In the example of FIG. 3, the decoder 300 includes a coded picture buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.

[0067] The entropy decoding unit 302 can receive encoded video data from the CPB 301 and entropy decode the video data to obtain syntax elements indicative of coding parameters including one or more of a partitioning of a CU, a final prediction mode, a transform type of residual data of the CU, or a quantization parameter of residual data of the CU, etc.

[0068] In the case where the syntax elements include the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter prediction mode, a prediction block of the current CU can be obtained by an inter prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra prediction mode, a prediction block of the current CU can be obtained by an intra prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 can further include units for performing prediction functions according to other prediction modes.

[0069] CPB 301 can obtain encoded video data from a communication medium 120 as illustrated in FIG. 1 and store it. The DPB 307 is used to store decoded pictures. The CPB 301 and the DPB 307 can also be replaced by other types of memory in some examples, the application does not make specific limitations. In some examples, the CPB 301 can be on the same chip as other components of the decoder 300 (as illustrated), or it can be independent of the chip on which the components are located.

[0070] The decoder 300 can perform the reconstruction operation separately for each block. The entropy decoding unit 302 can entropy decode syntax elements of quantized transform coefficients and transform information (e.g., QP or transform mode indication) to obtain quantized transform coefficients. The quantized transform coefficients are dequantized by the inverse quantization unit 303 to obtain a transform coefficient block including transform coefficients. The transform coefficient block is inverse transformed by the inverse transform processing unit 304 to generate a residual block corresponding to the current block, which is the inverse operation of the above-mentioned transform.

[0071] The reconstruction unit 305 can reconstruct the current block from the prediction block and the residual block. For example, the reconstruction unit 305 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the current block.

[0072] The filter unit 306 can perform one or more filter operations on the reconstructed block. For example, the filter unit 306 can be of the types referred to with respect to the filter unit 208, which are not repeated here. In some examples, the operation of the filter unit 306 can be skipped.

[0073] The decoder 300 can store the reconstructed picture resulting from the reconstructed block in the DPB 307. For example, in examples in which the operation of the filter unit 306 is not performed, the reconstruction unit 305 can store the reconstructed block to the DPB 307. In examples in which the operation of the filter unit 306 is performed, the filter unit 306 can store the filtered reconstructed block to the DPB 307. The decoder 300 can output decoded pictures (e.g., decoded video) from the DPB 307 for subsequent presentation to a display device, such as the display device 114 of FIG. 1.

[0074] The loss information calculation method provided by the embodiments of the application will be described below with reference to the accompanying drawings. The loss information calculation method provided by the embodiments of the application can be executed by an encoding end, for example, the encoder 200 shown in FIG. 1 or FIG. 2. The loss information calculation method provided by the embodiments of the application can be executed by a decoding end, for example, the decoder 300 described in FIG. 1 or FIG. 3. Wherein, the encoding end and the decoding end can be realized by software, hardware or a combination thereof, when it is realized by hardware, the encoding end can be referred to as an encoding end device or a video encoding device, and the decoding end can be referred to as a decoding end device or a video decoding device.

[0075] Please refer to FIG. 4, which is a flow chart of a loss information calculation method provided in the embodiments of the present application. As shown in FIG. 4, the method comprises the following steps:

[0076] Step 401: Obtain a target frequency band component of an output image of a neural network.

[0077] The neural network can be a convolutional neural network, a deep convolutional neural network, a feedforward neural network, or a recurrent neural network.

[0078] The neural network is an image-related neural network, for example, a neural network used for image-related tasks in the field of video coding, such as a neural network used for super-resolution (SR), a neural network used for image enhancement, or a neural network used for loop filtering.

[0079] The output image is an image output by the neural network, i.e., an image obtained after processing by the neural network.

[0080] It should be noted that the type or structure of the neural network is not limited in the embodiments of the present application, and the function of the neural network is also not limited. Specifically, it can be a neural network related to image or video coding.

[0081] In the embodiments of the present application, the target frequency band component is a component of a target frequency band, which can be a low-frequency component, a high-frequency component, or a medium-frequency component.

[0082] Step 402: Obtain a target frequency band component of a target image corresponding to the output image.

[0083] The target image corresponding to the output image can be an original image, which can be an image targeted by the neural network, i.e., the neural network learns to target the image, or the target image can also be referred to as a positive image sample or a real image sample in the learning process of the neural network. The target image and the output image have the same size.

[0084] It should be noted that the execution order of steps 401 and 402 is not limited in the embodiments of the present application. Step 401 can be executed first, followed by step 402, or step 402 can be executed first, followed by step 401, or steps 401 and 402 can be executed simultaneously.

[0085] Step 403: Calculate the loss information of the neural network, which includes the loss information between the target frequency band component of the output image and the target frequency band component of the target image.

[0086] The loss information of the neural network can be loss information between a target frequency band component of the output image and a target frequency band component of the target image.

[0087] In some embodiments, the loss function for calculating the loss information can be a loss function L1 (i.e., an absolute error loss function), a loss function L2 (i.e., a mean square error loss function), or a structural similarity index (SSIM) loss function, etc.

[0088] It should be noted that the loss function used to calculate the loss information in the embodiments of the present application is not limited.

[0089] In the embodiments of the present application, since the loss information includes loss information between a target frequency band component of the output image and a target frequency band component of the target image, loss information for the target frequency band can be obtained, and the loss information of the neural network is more targeted, which is beneficial to improving the learning effect of the neural network.

[0090] In the embodiments of the present application, the loss information calculation method described above can be executed by the encoding end, i.e., the encoding end executes each step described above. In some embodiments, the loss information calculation method described above can also be executed by other electronic devices other than the encoding end, such as each step in the above method being executed by an electronic device that executes neural network learning. The electronic device can be a computer, a server, a mobile phone, or the like.

[0091] In the embodiments of the present application, steps 401 to 403 can be configured to be executed in the neural network described above, or can be executed using an algorithm or program other than the neural network described above, and no limitation is made in this regard.

[0092] As an optional embodiment, the above method further includes the following steps:

[0093] The neural network is learned based on the loss information.

[0094] In the embodiments of the present application, the learning of the neural network can also be referred to as the training of the neural network.

[0095] Through the learning described above, the attention degree of the neural network to the target frequency band can be improved, and the processing performance of the neural network for a task associated with the target frequency band can be improved. For example, if the target frequency band is high frequency, the attention degree of the neural network to high frequency information can be improved, and the high frequency component in the image can be directly paid more attention to. Thus, for a super-resolution task, the conversion effect from low resolution to high resolution can be improved.

[0096] It should be noted that the embodiments of the present application are not limited to performing the above steps, for example, in some embodiments, after obtaining the loss information, the loss information is sent to other devices, and other devices learn the neural network based on the loss information.

[0097] In some embodiments, the method can further include performing a video coding related image task such as SR or image enhancement using the learned neural network to improve video coding performance.

[0098] As an optional implementation, the component of the target frequency band of the output image of the neural network includes:

[0099] The output image of the neural network is subjected to frequency domain transformation to obtain frequency domain information of the output image, and a target frequency band component is obtained from the frequency domain information of the output image.

[0100] The frequency domain transformation of the output image of the neural network can be frequency domain transformation of the output image by using a spatial-to-frequency domain transformation method to obtain the frequency domain information of the output image. The frequency domain information of the output image includes components of each frequency band in the output image.

[0101] The target frequency band component can be directly extracted from the frequency domain information of the output image.

[0102] The target frequency band component can be understood as information of the target frequency band in the frequency domain information of the image.

[0103] In this embodiment, the target frequency band component of the output image can be obtained by frequency domain transformation, so that the loss information corresponding to the target frequency band can be obtained for the neural network whose output image does not include frequency domain information, so as to improve the performance of the neural network.

[0104] It should be noted that the embodiments of the present application are not limited to obtaining the target frequency band component by the above frequency domain transformation, for example, in some embodiments, the output image of the neural network itself carries frequency domain information, so that the target frequency band component can be directly extracted from the output image.

[0105] As an optional implementation, the component of the target frequency band of the output image of the neural network includes:

[0106] The output image of the neural network is subjected to frequency domain transformation to obtain frequency domain information of the output image, and a target frequency band component is obtained from the frequency domain information of the output image.

[0107] The frequency domain transformation of the target image can be a frequency domain transformation of the target image by using a space domain to frequency domain transformation method to obtain frequency domain information of the target image.

[0108] The target frequency band component can be directly extracted from the frequency domain information of the target image.

[0109] In this embodiment, the target frequency band component of the target image can be obtained by frequency domain transformation, so that the loss information corresponding to the target frequency band can also be obtained for the target image without frequency domain information, thereby reducing the limitations of loss information calculation.

[0110] It should be noted that the target frequency band component of the target image obtained by the above frequency domain transformation is not limited in the embodiments of the present application. For example, in some embodiments, the frequency information of the target image can be directly obtained from other devices.

[0111] Optionally, the frequency domain transformation includes one of the following:

[0112] Discrete wavelet transform (DWT), discrete cosine transform (DCT), discrete Fourier transform (DFT).

[0113] In this embodiment, multiple frequency domain transformations are supported to improve the flexibility of loss calculation.

[0114] In some embodiments, the order of the frequency domain transformation can be set according to the complexity requirement. For example, the first order transformation of DWT is used, as shown in FIG. 5. The digital image represents the output image or the target image, and the high frequency component and the low frequency component obtained by row decomposition, column decomposition, column reconstruction and row reconstruction are represented by H and L in FIG. 5.

[0115] As an optional embodiment, the target frequency band component includes at least one of the following:

[0116] High frequency component, medium frequency component, low frequency component.

[0117] In this embodiment, the loss information associated with at least one of the high frequency component, the medium frequency component and the low frequency component can be calculated, so that the neural network can pay more attention to the high frequency component, the medium frequency component or the low frequency component, thereby improving the processing details of the neural network in the image processing process, and further improving the processing performance of the neural network.

[0118] For example, by calculating the loss information of the high-frequency component, the neural network can pay more attention to the information of the high-frequency component, thereby improving the ability of the neural network to learn the information of the high-frequency component from the data, and enabling the network to supplement more high-frequency detail information in the high-resolution reconstruction process, so as to improve the recovery ability of the image high-frequency information.

[0119] In some embodiments, in the case that the target frequency band component includes one of a high-frequency component, a medium-frequency component and a low-frequency component, more targeted loss information can be obtained. For example, in the case that the high-frequency component and the loss function LI, the above loss information can include loss information calculated by the following formula:

[0120] wherein the L1_DWT_loss represents the loss information between the high-frequency component of the output image and the high-frequency component of the target image, n is the number of sample points of the high-frequency component of the output image and the target image, Y i is the high-frequency component of the target image, f(x i ) is the high-frequency component of the output image.

[0121] A specific diagram can be as shown in FIG. 6, wherein FIG. 6 takes the neural network as an example of the SR network, and SR in FIG. 6 represents the above output image, and GT represents the above target image.

[0122] Optionally, the calculating the loss information of the neural network comprises at least one of the following:

[0123] calculating a first loss value between the high-frequency component of the output image and the high-frequency component of the target image;

[0124] calculating a second loss value between the medium-frequency component of the output image and the medium-frequency component of the target image;

[0125] calculating a third loss value between the low-frequency component of the output image and the low-frequency component of the target image.

[0126] wherein the first loss value, the second loss value and the third loss value can be calculated by using the same or different loss functions.

[0127] In this embodiment, the loss value of any one of the high-frequency component, the medium-frequency component and the low-frequency component, or the loss value of multiple components can be obtained. In the case of calculating one loss value, the neural network can pay more attention to the high-frequency component, the medium-frequency component or the low-frequency component, so as to improve the processing details of the neural network in the image processing process, and further improve the processing performance of the neural network. In the case of calculating the above-mentioned multiple loss values, the neural network can pay attention to multiple frequency bands, improve the ability of the neural network to learn multiple frequency band information from data, and further improve the performance of the neural network.

[0128] Optionally, in the case where the target frequency band component includes at least two of the high-frequency component, the medium-frequency component and the low-frequency component, the loss information between the target frequency band component of the output image and the target frequency band component of the target image includes a weighted average loss value of at least two of the first loss value, the second loss value and the third loss value.

[0129] The weighted average loss value of at least two of the first loss value, the second loss value and the third loss value can be a weight value of each frequency band configured or agreed in advance, the loss value of each frequency band is multiplied by the corresponding weight, and then averaged to obtain the weighted average loss value. For example, the weights corresponding to the high-frequency component, the medium-frequency component and the low-frequency component are weight 1, weight 2 and weight 3, respectively, and the weighted average loss value can be equal to one of the following:

[0130] (first loss value x weight 1 + second loss value x weight 2) / 2;

[0131] (first loss value x weight 1 + third loss value x weight 3) / 2;

[0132] (second loss value x weight 2 + third loss value x weight 3) / 2;

[0133] (first loss value x weight 1 + second loss value x weight 2 + third loss value x weight 3) / 2.

[0134] In this embodiment, the loss information of the neural network can be obtained based on the loss value information of multiple frequency bands, so that the neural network can pay attention to multiple frequency bands, improve the ability of the neural network to learn multiple frequency band information from data, and further improve the performance of the neural network. Moreover, since it is a weighted average loss value, the weights of each frequency band are considered, so that the neural network can pay more attention to multiple frequency bands, and further improve the ability of the neural network to learn multiple frequency band information from data and improve the performance of the neural network.

[0135] It should be noted that in the case that the target frequency band component includes at least two of a high frequency component, a medium frequency component and a low frequency component, the loss information in the embodiment of the present application is not limited to be obtained by weighted average, for example, the loss average of at least two can also be directly calculated.

[0136] As an optional implementation of the above, the output image includes:

[0137] The up-sampling image output by the neural network.

[0138] In this implementation, since the output image includes the up-sampling image output by the neural network, in the case that the neural network is deployed at the decoding end, the image compressed at the encoding end can be supported to be the image after down-sampling, thereby saving the transmission bandwidth. In addition, for SR, the decoding end decodes to obtain the reconstructed low-resolution image, and outputs a high-quality up-sampling image, i.e., an original resolution image, through the neural network, so as to improve the image processing effect.

[0139] It should be noted that the output image in the embodiment of the present application is not limited to be the up-sampling image, for example, in some implementations, the image can also be not sampled or down-sampled.

[0140] In the embodiment of the present application, the target frequency band component of the output image of the neural network is obtained, the target frequency band component of the target image corresponding to the output image is obtained, and the loss information of the neural network is calculated, the loss information including the loss information between the target frequency band component of the output image and the target frequency band component of the target image. Since the loss information includes the loss information between the target frequency band component of the output image and the target frequency band component of the target image, the loss information for the target frequency band can be obtained, so that the loss information of the neural network is more targeted, which is beneficial to improving the learning effect of the neural network.

[0141] The loss information calculation method provided in the embodiment of the present application can be executed by a loss information calculation device. As an example, the device can be an electronic device, or a component in the electronic device, such as a chip, a circuit, etc. In the embodiment of the present application, the loss information calculation device is taken as an example to illustrate the loss information calculation device provided in the embodiment of the present application.

[0142] Please refer to FIG. 7, which is a structure diagram of a loss information calculation device provided in the embodiment of the present application, as shown in FIG. 7, the loss information calculation device 700 includes:

[0143] The first obtaining module 701 is configured to obtain the target frequency band component of the output image of the neural network.

[0144] The second obtaining module 702 is configured to obtain the target frequency band component of the target image corresponding to the output image.

[0145] The computing module 703 is configured to calculate loss information of the neural network, the loss information including loss information between target frequency band components of the output image and target frequency band components of the target image.

[0146] Optionally, the first obtaining module 701 is configured to perform frequency domain transformation on the output image of the neural network to obtain frequency domain information of the output image, and obtain the target frequency band components from the frequency domain information of the output image.

[0147] Optionally, the second obtaining module 702 is configured to perform frequency domain transformation on the target image corresponding to the output image to obtain frequency domain information of the target image, and obtain the target frequency band components from the frequency domain information of the target image.

[0148] Optionally, the frequency domain transformation includes one of the following:

[0149] Discrete wavelet transform (DWT), discrete cosine transform (DCT), and discrete Fourier transform (DFT).

[0150] Optionally, the target frequency band components include at least one of the following:

[0151] High frequency components, medium frequency components, and low frequency components.

[0152] Optionally, the computing module 703 is configured to perform at least one of the following:

[0153] Calculate a first loss value between high frequency components of the output image and high frequency components of the target image;

[0154] Calculate a second loss value between medium frequency components of the output image and medium frequency components of the target image;

[0155] Calculate a third loss value between low frequency components of the output image and low frequency components of the target image.

[0156] Optionally, in a case where the target frequency band components include at least two of the high frequency components, the medium frequency components, and the low frequency components, the loss information between the target frequency band components of the output image and the target frequency band components of the target image includes a weighted average loss value of at least two of the first loss value, the second loss value, and the third loss value.

[0157] Optionally, the output image includes:

[0158] An up-sampled image output by the neural network.

[0159] The loss information calculation apparatus is advantageous in improving learning effect of the neural network.

[0160] The loss information calculation apparatus 700 provided by the embodiments of the present application can realize each process of the method embodiments of FIG. 4 and achieve the same technical effects. To avoid repetition, details are not described herein.

[0161] As shown in FIG. 8, the embodiments of the present application further provide an electronic device 800, which includes a processor 801 and a memory 802. The memory 802 stores programs or instructions executable on the processor 801. For example, when the electronic device 800 is an encoding end device, the programs or instructions, when executed by the processor 801, realize each step of the loss information calculation method embodiments described above and achieve the same technical effects. When the electronic device 800 is a decoding end device, the programs or instructions, when executed by the processor 801, realize each step of the loss information calculation method embodiments described above and achieve the same technical effects. To avoid repetition, details are not described herein. Optionally, the memory 802 can be the memory 102 or the memory 113 in the embodiments of FIG. 1, and the processor 801 can realize the functions of the encoder 200 or the decoder 300 in the embodiments of FIGS. 1-3.

[0162] The embodiments of the present application further provide an electronic device, which includes a memory configured to store video data, and a processing circuit configured to realize each step of the loss information calculation method embodiments described above. Optionally, the memory can be the memory 102 or the memory 113 in the embodiments of FIG. 1, and the processing circuit can realize the functions of the encoder 200 or the decoder 300 in the embodiments of FIGS. 1-3.

[0163] The embodiments of the present application further provide an electronic device, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to realize the steps in the method embodiments of FIG. 4. The device embodiments correspond to the method embodiments described above. Each implementation process and implementation manner of the method embodiments described above can be applicable to the terminal embodiments and achieve the same technical effects.

[0164] The processor or processing circuit of the embodiments of the present application can include a general-purpose processor, a special-purpose processor, etc., such as a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), an artificial intelligent (AI) processor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a network processor (NP), a field programmable gate array (FPGA) or other programmable logic devices, a gate circuit, a transistor, a discrete hardware component, etc. The communication interface of the embodiments of the present application can include a transceiver, a pin, a circuit, a bus, etc.

[0165] The electronic device described above can be a terminal, or can be other devices other than the terminal, such as a server, a network attached storage (NAS), etc.

[0166] The terminal can be a mobile phone, a tablet personal computer, a laptop computer, a notebook computer, a personal digital assistant (PDA), a palm computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile Internet device (MID), an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, a robot, a wearable device, a flight vehicle, a vehicle user equipment (VUE), a shipboard device, a pedestrian user equipment (PUE), a smart home (a home device with a wireless communication function, such as a refrigerator, a television, a washing machine, or furniture), a game console, a personal computer (PC), a teller machine, or a self-service machine, and the like. The wearable device includes a smart watch, a smart bracelet, a smart earphone, smart glasses, smart jewelry (a smart bracelet, a smart necklace, a smart ring, a smart necklace, a smart anklet, a smart necklace, and the like), a smart wristband, smart clothing, and the like. The vehicle-mounted device can also be referred to as a vehicle-mounted terminal, a vehicle-mounted controller, a vehicle-mounted module, a vehicle-mounted component, a vehicle-mounted chip, or a vehicle-mounted unit, and the like. It should be noted that the specific type of the terminal is not limited in the embodiments of the present application.

[0167] The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server. The cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.

[0168] For example, the electronic device can include, but is not limited to, a source device 100 or a destination device 110 as shown in FIG. 1.

[0169] Taking the electronic device as an example, FIG. 9 is a schematic diagram of a hardware structure of a terminal according to an embodiment of the present application.

[0170] The terminal 900 includes, but is not limited to, at least part of components such as a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909, and a processor 910.

[0171] Those skilled in the art can understand that the terminal 900 can further include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 910 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The terminal structure shown in FIG. 9 does not constitute a limitation on the terminal, and the terminal can include more or fewer components than those shown, or combine certain components, or different component arrangements, which are not described here.

[0172] It should be understood that in the embodiments of the present application, the input unit 904 can include a graphics processor 9041 and a microphone 9042. The graphics processor 9041 processes image data of a still picture or a video obtained by an image acquisition device (such as a camera) in a video acquisition mode or an image acquisition mode, or can process obtained point cloud data. The display unit 906 can include a display panel 9061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 can include two parts of a touch detection device and a touch controller. The other input devices 9072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), trackballs, mice, joysticks, etc., which are not described here.

[0173] In the embodiments of the present application, the radio frequency unit 901 can transmit downlink data from the network side device to the processor 910 for processing after receiving the downlink data. In addition, the radio frequency unit 901 can send uplink data to the network side device. Generally, the radio frequency unit 901 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc.

[0174] The memory 909 can be used to store software programs or instructions and various data. The memory 909 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 909 can include a volatile memory or a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 909 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0175] The processor 910 can include one or more processing units; optionally, the processor 910 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 910.

[0176] The processor 910 is configured to obtain a target frequency band component of an output image of a neural network; obtain a target frequency band component of a target image corresponding to the output image; and calculate loss information of the neural network, the loss information including loss information between the target frequency band component of the output image and the target frequency band component of the target image.

[0177] Optionally, the obtaining the target frequency band component of the output image of the neural network comprises:

[0178] The output image of the neural network is subjected to a frequency domain transformation to obtain frequency domain information of the output image, and a target frequency band component is obtained from the frequency domain information of the output image.

[0179] Optionally, the target frequency band component of the target image corresponding to the output image is obtained by:

[0180] The target image corresponding to the output image is subjected to a frequency domain transformation to obtain frequency domain information of the target image, and a target frequency band component is obtained from the frequency domain information of the target image.

[0181] Optionally, the frequency domain transformation includes one of:

[0182] Discrete wavelet transform (DWT), discrete cosine transform (DCT), and Fourier transform (DFT).

[0183] Optionally, the target frequency band component includes at least one of:

[0184] High frequency component, medium frequency component, and low frequency component.

[0185] Optionally, the loss information of the neural network is calculated by at least one of:

[0186] A first loss value between a high frequency component of the output image and a high frequency component of the target image is calculated.

[0187] A second loss value between a medium frequency component of the output image and a medium frequency component of the target image is calculated.

[0188] A third loss value between a low frequency component of the output image and a low frequency component of the target image is calculated.

[0189] Optionally, when the target frequency band component includes at least two of the high frequency component, the medium frequency component, and the low frequency component, the loss information between the target frequency band component of the output image and the target frequency band component of the target image includes a weighted average loss value of at least two of the first loss value, the second loss value, and the third loss value.

[0190] Optionally, the output image includes:

[0191] An up-sampled image output by the neural network.

[0192] The terminal described above is beneficial to improving the learning effect of the neural network.

[0193] It can be understood that the implementation process of each implementation manner mentioned in the embodiment can refer to the related description of the loss information calculation method of the method embodiment, and achieve the same or corresponding technical effects. To avoid repetition, it will not be described here.

[0194] The embodiment of the present application further provides a readable storage medium, which stores a program or instructions, and the program or instructions are executed by a processor to realize the processes of the loss information calculation method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0195] The processor is the processor in the terminal in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a ROM, a RAM, a magnetic disk, or an optical disk. In some examples, the readable storage medium can be a non-transitory readable storage medium.

[0196] The embodiment of the present application further provides a chip, which includes a processor and a communication interface, the communication interface is coupled with the processor, and the processor is used to run a program or instructions to realize the processes of the loss information calculation method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0197] It should be understood that the chip mentioned in the embodiment of the present application can include a system on chip (SOC), and can also include a standalone display chip, etc.

[0198] The embodiment of the present application further provides a computer program / program product, which is stored in a storage medium and is executed by at least one processor to realize the processes of the loss information calculation method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0199] It should be noted that, in this document, the terms “comprises”, “comprising”, or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements does not only include those elements, but also includes other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the statement “comprises a” does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element. In addition, it should be pointed out that the scope of the methods and apparatus in the embodiments of the present application is not limited to the order of performing the functions shown or discussed, and can also include performing the functions in a substantially simultaneous manner or in a reverse order, for example, the described method can be performed in an order different from the described order, and various steps can be added, omitted, or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0200] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of a computer software product and a general hardware platform as necessary, and of course can also be realized by hardware. The computer software product is stored in a storage medium (such as a ROM, a RAM, a magnetic disc, an optical disc, etc.), and includes a plurality of instructions for enabling a terminal or a network side device to execute the method described in each embodiment of the present application.

[0201] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative rather than limiting. Those skilled in the art can make many forms of embodiments under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims, and these embodiments all belong to the protection of the present application.

Claims

1. A loss information calculation method, comprising: obtaining target frequency band components of an output image of a neural network; obtaining target frequency band components of a target image corresponding to the output image; calculating loss information of the neural network, the loss information comprising loss information between the target frequency band components of the output image and the target frequency band components of the target image.

2. The method of claim 1, wherein, The obtaining of the target frequency band components of the output image of the neural network comprises: performing frequency domain transformation on the output image of the neural network to obtain frequency domain information of the output image, and obtaining the target frequency band components from the frequency domain information of the output image.

3. The method of claim 1 or 2, wherein, The obtaining of the target frequency band components of the target image corresponding to the output image comprises: performing frequency domain transformation on the target image corresponding to the output image to obtain frequency domain information of the target image, and obtaining the target frequency band components from the frequency domain information of the target image.

4. The method of claim 2 or 3, wherein, The frequency domain transformation comprises one of: discrete wavelet transform (DWT), discrete cosine transform (DCT), and Fourier transform (DFT).

5. The method according to any one of claims 1 to 4, characterized in that, The target frequency band components comprise at least one of: high frequency components, medium frequency components, and low frequency components.

6. The method of claim 5, wherein, The calculation of the loss information of the neural network comprises at least one of: calculating a first loss value between high frequency components of the output image and high frequency components of the target image; calculating a second loss value between medium frequency components of the output image and medium frequency components of the target image; calculating a third loss value between low frequency components of the output image and low frequency components of the target image.

7. The method of claim 6, wherein, In a case where the target frequency band components comprise at least two of the high frequency components, the medium frequency components, and the low frequency components, the loss information between the target frequency band components of the output image and the target frequency band components of the target image comprises a weighted average loss value of at least two of the first loss value, the second loss value, and the third loss value.

8. The method of any one of claims 1 to 7, wherein, The output image comprises: an up-sampled image output by the neural network.

9. A loss information calculation apparatus, comprising: a first obtaining module configured to obtain target frequency band components of an output image of a neural network; a second obtaining module configured to obtain target frequency band components of a target image corresponding to the output image; a calculation module configured to calculate loss information of the neural network, the loss information comprising loss information between the target frequency band components of the output image and the target frequency band components of the target image.

10. The apparatus of claim 9, wherein, The first obtaining module is configured to perform frequency domain transformation on the output image of the neural network to obtain frequency domain information of the output image, and obtain the target frequency band components from the frequency domain information of the output image.

11. The apparatus of claim 9 or 10, wherein, The second obtaining module is configured to perform frequency domain transformation on the target image corresponding to the output image to obtain frequency domain information of the target image, and obtain the target frequency band components from the frequency domain information of the target image.

12. The apparatus of any one of claims 9-11, wherein, The target frequency band components comprise at least one of: high frequency components, medium frequency components, and low frequency components.

13. The apparatus of claim 12, wherein, The calculation module is configured to perform at least one of: calculating a first loss value between high frequency components of the output image and high frequency components of the target image; calculating a second loss value between medium frequency components of the output image and medium frequency components of the target image; and calculating a third loss value between low frequency components of the output image and low frequency components of the target image. calculating a second loss value between a mid-frequency component of the output image and a mid-frequency component of the target image; calculating a third loss value between a low-frequency component of the output image and a low-frequency component of the target image.

14. An electronic device comprising a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implement the steps of the loss information calculation method according to any one of claims 1 to 8.

15. A readable storage medium, the readable storage medium storing programs or instructions, the programs or instructions, when executed by a processor, implement the steps of the loss information calculation method according to any one of claims 1 to 8.

16. A computer program product stored in a storage medium, the computer program product, when executed by at least one processor, implement the steps of the loss information calculation method according to any one of claims 1 to 8.

17. A chip comprising a processor and a communication interface, the communication interface and the processor being coupled, the processor being configured to run programs or instructions, implement the steps of the loss information calculation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Binocular image super-resolution reconstruction method based on multistage intensified attention mechanism

    CN116797461A

  • Training method of image reconstruction model, main control equipment and image reconstruction method

    CN116993852A

  • Image reconstruction method and system based on self-encoding neural network

    CN117218149A

  • Method and apparatus for training video generation model, storage medium, and computer device

    US20240212252A1